Your First AI Agent Should Be Boring
The agents that earn their keep are narrow, reversible and dull. Here's the line between a job an agent can finish and one it should never be handed.
Every agent demo you have seen was built to be impressive. It takes a vague instruction, thinks visibly for a while, touches six systems, and produces something that looks like a person did it.
The agents I build for clients are not like that, and the ones that survive contact with a real business are never like that. They do one job, they do it the same way every time, and the interesting part is not what they can do — it’s where they have been told to stop.
That distinction is worth more than any model comparison, so let me be specific about it.
The line between answering and acting
I’ve written before about where AI genuinely pays off, and the strongest case on that list is still reading documents. The invoice extraction system I built for a finance team is a good example of the category: 900 supplier invoices a month, previously hand-keyed, now extracted automatically with a person confirming a pre-filled form in about ten seconds.
That system is not an agent. It answers a question — what does this document say? — and hands the answer to a human, who decides what happens next.
An agent is what you get when nobody wants to be the one clicking confirm 900 times. It takes the job, not the question: get this invoice into the ledger. It reads the document, checks the supplier exists, matches it against the purchase order, notices the total does not agree, and either resolves that or escalates it to a person with the discrepancy already explained.
The model doing the reading is the same in both cases. What changed is that something is now allowed to do things on the back of what it read. That is the entire step change, and every difficulty that follows comes from it.
Four things a job needs before an agent can hold it
I turn down more agent projects than I take, and it is almost always because the job fails one of these.
1. There is a definition of done you could check yourself.
“Get the invoice into the ledger” passes: either it’s there, matched and correct, or it isn’t. “Improve our supplier relationships” does not. If you cannot write the check that proves the job was finished, you cannot build the agent, because you have no way to know when it has failed. You have only a feeling.
2. The work is reversible, or the agent stops before the part that isn’t.
A drafted invoice can be deleted. A sent invoice can be credited, awkwardly. A payment cannot be unmade, and a rude email to your largest customer cannot be unsent.
This is the single most useful design tool available. You do not have to choose between “the agent does nothing” and “the agent has your credit card.” You put the stop immediately before the irreversible step. In practice almost every agent I build ends its run by preparing something for a human to release — and the value is already banked by then, because the preparation was the work.
3. Its systems can tell it whether it worked.
An agent that writes to a system and gets no meaningful confirmation back is flying blind, and a blind agent repeats itself. If your CRM’s API returns success for a write that silently failed validation, the agent will cheerfully do it four more times.
This sounds like a technical footnote and it is actually the most common reason a promising project dies in week two. Before scoping any of this, the real question is not “how smart is the model” but “what can this system tell me about its own state.”
4. A named person owns it being wrong.
Not a team. A person. When an agent makes a bad call — and it will — someone has to see it, judge it and adjust the rules. Agents that belong to everybody get switched off within a month, because the first time one embarrasses somebody, nobody has the standing to defend it or the authority to fix it.
What a real one looks like
Take the job that shows up in nearly every business I work with: chasing money you are owed.
The manual version is someone opening the accounting system on a Monday, working out which invoices are past due, deciding which customers can be chased and which need handling gently, writing the emails, and remembering to look again next week. It takes a couple of hours and it slips the moment anything else is on fire.
The agent version runs every morning and does this:
- Pulls the overdue list from the accounting system.
- Drops anyone flagged in dispute, on a payment plan, or contacted in the last five days.
- Works out which chase this is — first, second, or final — from the contact history.
- Drafts the right message for that stage, in your wording, with the invoice attached.
- Puts the drafts in front of a person, in one batch, with anything unusual flagged at the top.
Step five is the whole design. It is not a limitation to be removed in phase two — it is the thing that makes the other four safe to run unattended. The two hours are gone either way; what’s kept is the judgment about which customers get chased and how hard, which was never the boring part.
Notice how little of that is the model. Steps 1, 2 and 5 are ordinary code. The AI does step 3’s messier reading and step 4’s drafting. That ratio is normal. Most of a good agent is plumbing, and the projects that go wrong are the ones that got that backwards.
The ones I talk people out of
The agent that runs the business. One system that knows everything and handles anything. It demos beautifully and cannot be evaluated, which means it can never be trusted, which means it gets used twice and abandoned. Narrow agents can be measured. Measured things get trusted.
The agent that needed a script. If every run does the same five steps in the same order with no judgment involved, you do not need a model in the loop. You need a scheduled job, which is cheaper, faster and cannot surprise you. About a third of the “agent” enquiries I get are this, and the honest version of that conversation takes ten minutes.
The agent bolted onto a process nobody has written down. If the humans doing this job disagree about how it works, an agent will pick one of their answers and apply it consistently to everybody. Automating a process you have not mapped is the classic way to make a bad process faster rather than better.
The agent whose failures are quiet. Any agent I build says something when it breaks. Silent failure is worse than no automation at all, because people stop checking, and the damage accumulates behind the assumption that it is handled.
Where to start
Pick the job somebody on your team does every week that they would describe as tedious rather than difficult. Write down the checklist they actually follow. Then mark the steps that need a judgment call, and count them.
If it’s zero, you want automation, not an agent, and it will cost you less.
If it’s one or two, that’s a genuinely good first agent — narrow enough to measure, valuable enough to notice.
If it’s most of the steps, you have found something worth doing carefully, and probably not first.
The whole point of starting boring is that the second one is much easier to argue for once the first has been quietly working for three months.
Wondering whether the tedious job in your week is an agent, an automation, or neither? Book a free call and describe it — I’ll tell you which, including when the answer is “a scheduled script and twenty minutes.” More on how I approach AI projects.