The question almost always comes in the break, never in the seminar room.
In the room, people ask about tools. Which software, which version, what it costs. In the break, cup in hand, they ask about something else.
That day, someone asked:
"Your AI agents, the ones that work for you. Do they make mistakes?"
Do my AI agents make mistakes?
"Yes," I said. "And no."
There are two kinds of mistake, and only one of them belongs to the agents. The first happens because the agent is missing context that a person inside the company would have automatically. The second happens because my instruction did not say what I meant. The first kind is obvious straight away. The second looks like success.
Which mistakes do the agents make themselves?
An agent only knows what is in its field of view.
It sees the file I hand it. It doesn't see the conversation I had yesterday, or the reason something is set up one way and not another. It is missing the context a colleague picks up automatically after two years inside the company. I described how far that goes when my agents had no memory.
These mistakes are annoying, but they are easy to spot. Something comes back that doesn't fit, and I notice straight away. I add the missing context, and the second attempt is right.
And which mistakes are mine?
The other kind doesn't look like a mistake. It looks like success.
Recently I wanted a particular seminar to appear on my home page. I told my website agent, and it did it. It wrote to tell me how it had done it.
Half an hour later, a different seminar was missing from the page. My October course at tecTrain (page in German), a full day in Vienna. Nobody had deleted it. It was still in the list, in the right place, with date, location and booking link.
It just wasn't visible any more.
What I asked for had a side effect I had not thought about. The agent did not cause it. It carried it out, cleanly and quickly and without a moment's hesitation.
An instruction to a machine is not a command with exactly one meaning. It is more like a lasso: it catches everything within its reach, including the things I had not thought about.
Why is the second kind the expensive one?
The first kind announces itself. Something looks wrong, so I go and look.
The second kind never announces itself. Everything looks right, because everything was carried out correctly. The only way to notice it is to check somewhere you have no reason to check.
A colleague carrying out my instruction is a safeguard at this point. Not because they are cleverer, but because they sit in the same room and have the same list in front of them. They say: "Hang on, that drops the October date, is that what you want?"
You stop hearing that sentence the moment execution runs cleanly.
How I fixed it
I didn't simply put it back. That would have looked fine for a day and left the cause in place.
Instead I got rid of the point where I could choose anything at all. Today the date decides which events come first, and nobody else. That is less flexible. It is also less wrong.
The question I have handed back ever since
These days, when someone asks me in the break whether my agents make mistakes, I still say: yes and no.
And then I ask a question back.
Not: do you trust the machine. But: when did you last check whether what you tell it is what you actually mean?
Because it will not check it for you. It will do it.
Common questions
Do AI agents make mistakes? Yes, but not only their own. Some of the mistakes start with the agent, because it is missing context that a person inside the company brings automatically. The rest start with whoever gave the instruction, because it says something other than what was meant.
How do I tell one kind from the other? Agent mistakes look wrong: something comes back that does not fit. Instruction mistakes look right, because they were carried out correctly. They only surface when somebody checks somewhere they have no reason to check.
What helps against the second kind? Remove the point where things are picked by hand. As long as an exception is possible, sooner or later one gets built and then forgotten. A fixed rule is less flexible and less wrong.
So does an AI agent replace an employee? In execution, often yes. In asking back, no. A colleague carrying out an instruction says "Hang on, that drops the other one, is that what you want?" That question back is a safeguard, and it disappears the moment execution runs cleanly.

