Thirty-six entries in, my error log stopped adding up
I started keeping a log of every time the AI gets something wrong in my work. What it got wrong, what caught it, and whether the catch was mine or the system’s. My sponsor asked me to keep it after I brought him two of my own mistakes, and I think he was right that writing them down is more useful than remembering them.
For a while the log was just a list. Then it got long enough to count, and counting is where it got uncomfortable.
Eighteen of the thirty-six entries are recorded as caught by me asking a sceptical question. That looked like a good number. I was the check, the check was working. But a sceptical question is also an invitation, and models are agreeable. Anthropic published a paper on this in 2023, “Towards Understanding Sycophancy in Language Models,” which measured how often a model abandons a correct answer once a user pushes back on it. I had read it. It did not occur to me to apply it to my own log.
So I do not actually know how many of my eighteen catches were catches. Some of them were probably the model folding.
I split the categories after that, which helps a little. A catch where I named the specific error is different from a catch where I only expressed doubt. But splitting the categories does not give me a control. To know the real number I would have to ask the same question in both directions, sometimes doubting a correct answer, and see how often the model changes its mind when it should not.
I have not built that yet. It is the next thing.
What I keep thinking about is that the log only became useful when it got long enough to contradict me. For the first fifteen or twenty entries it was a diary. It confirmed what I already believed, which is that I was being careful. It took thirty-six before the arithmetic stopped closing.