Turning Customer Requests Into Completed Work
Requests arrive in a hundred different shapes, and the real cost is everything a person does before they can answer one. Here is how we architect an agent to do that part.
Most businesses do not have a problem answering customer requests. They have a problem getting to the point where the request can be answered.
This is the architecture we use when that is the bottleneck. The worked example below is a request-handling agent, because that is the shape the problem takes most often — but the structure is the same whether the requests are customer emails, internal tickets or supplier queries. What changes per business is the tools and the thresholds, not the design.
The work is not the answer
A request lands. Some are trivial:
Can you update my details?
Others need someone to go and find out:
My order hasn’t arrived. Can you check what’s happening?
Nobody on the team struggles to answer the second one. What costs them is the sequence they have to run first:
- Work out what the customer is actually asking.
- Identify the customer.
- Find their account.
- Read the previous conversations.
- Check the orders or transactions.
- Search the internal documentation for the policy.
- Decide what should happen.
- Do it, in another system.
- Write the reply.
- Update the record.
One request, that is fine. Several hundred a week and it is a full-time role — and it is a role staffed by your most experienced people, because steps 6 and 7 need judgment. They spend most of their day on the nine steps that do not.
Why I did not just automate it
The obvious first move is a workflow. New request, create ticket, notify, assign:
New request
↓
Create ticket
↓
Send notification
↓
Assign employee
That works beautifully when the process is predictable, and this one is not. Two customers describe the same problem in completely different words. One request arrives with the order number, the next has “the thing I bought last month”. Some need three systems checked, some need one, some need a policy read first. Some can be finished automatically and some must never be.
There is no fixed path to encode, because the path depends on what you find when you start looking. A conventional workflow can only branch on conditions you knew about when you wrote it. This problem needs something that can decide what to check next based on what the last check returned.
That is the actual test for whether a job needs an agent, and it is worth being strict about it — most jobs do not pass, and the ones that do not are cheaper and more reliable as ordinary automation.
Give it the job, not the question
I was not trying to build “ask our AI anything”. A chatbot answers:
Your order was shipped on Tuesday.
An agent takes the whole task:
Find out why this order has not arrived and resolve it if you are allowed to.
That second sentence contains multiple steps, an investigation whose shape is not known in advance, a decision, and an action. The system has to understand the goal, gather what it needs, judge what it found, and do something about it — stopping at the edge of its authority.
The overall shape:
Customer request
│
▼
Understand
the request
│
▼
What do I need
to know?
│
┌───┴────┐
▼ ▼
Business Internal
systems docs
│ │
└───┬────┘
▼
Investigate
│
▼
Decide next step
│
┌───┴─────────┐
▼ ▼
Permitted Needs
action approval
│ │
▼ ▼
Execute Human
│ decides
└──────┬──────┘
▼
Reply and
update record
The important thing is that the model is not generating a reply. It is operating inside a workflow that happens to have judgment in the middle of it.
Step 1 — Turn the request into a task
The customer writes:
I’ve been waiting ages for this. Tracking hasn’t moved since Monday. Can you check what’s going on?
There is no keyword to match here. “Tracking hasn’t moved” is a delivery problem; so is “where is my stuff”; so is a forwarded carrier email with no text at all. The first job is to convert an unstructured message into something the rest of the system can act on:
intent: delivery_issue
customer: identified from sender address
order: not stated — needs lookup
sentiment: frustrated
requires_lookup: true
This is the one step where a language model is unambiguously the right tool. Everything after it is a program that happens to consult a model at specific points.
Step 2 — Let it gather its own context
The agent does not guess and it does not work from what happens to be in its prompt. It is given tools and decides which it needs:
get_customer()
get_orders()
get_shipping_status()
get_previous_conversations()
search_knowledge_base()
This is the substantive difference between an agent and a chatbot. The model can reach the systems where your business data actually lives, so it is reasoning about this customer’s real order rather than about the general concept of a late delivery.
It also means the quality ceiling is set by your systems, not by the model. If your order system cannot tell you whether a delivery was attempted, no agent can work that out either. In practice this is where most of the scoping conversation goes.
Step 3 — Investigate
Suppose it finds:
Order #18392
Status Shipped
Last scan Monday
Due Wednesday
Today Friday
Two days past due with no movement. That is unusual, and it is not yet an answer. Depending on what it sees, the next thing worth checking might be the carrier’s status, a previous failed delivery, the address on the order, an internal note, or whether four other customers on the same route have reported the same thing this morning.
Nobody wrote that decision tree. The agent works out what it needs to know next from what it has just learned — which is the entire reason this job justifies an agent rather than a flowchart.
Step 4 — Decide, inside boundaries I set
Once it has enough, it picks from outcomes I have defined in advance:
- Still within the delivery window — explain the status, no action needed.
- Delayed but moving — reply, and set a follow-up to check again.
- Apparently lost — escalate, with the investigation attached.
- Entitled to a replacement under policy — prepare the replacement request.
The agent is choosing between known outcomes. It is not inventing a resolution, and it cannot decide that this customer deserves something no policy covers. That constraint is not a limitation I plan to remove later; it is what makes the system testable.
Step 5 — Act, without holding the authority
This is where an agent becomes worth building, and where most of the engineering care goes.
create_followup()
update_ticket()
send_notification()
create_replacement_request()
schedule_callback()
These are real calls into real systems. But note where the permission lives. The agent decides:
I need to create a replacement request.
The application decides:
This agent may create a replacement request, for this customer, under this value, in this state.
The model proposes and the application authorises. Those are two different pieces of software, and keeping them apart is what lets you reason about the system at all. It means a bad decision by the model is a rejected call rather than an incident — and it means the answer to “what can this thing do?” is a list you can read, not a prompt you have to trust.
Every call is logged with the reasoning that led to it. When something goes wrong at 3pm on a Tuesday, you need to be able to reconstruct why, and “the model decided to” is not an answer anyone can act on.
Step 6 — Know exactly where to stop
Say the investigation concludes the customer is owed a refund above the threshold the agent is allowed to authorise. It stops — but it stops usefully:
Agent investigates
↓
Agent prepares recommendation
↓
Human reviews and decides
↓
Agent carries out the decision
The person does not redo the investigation. They read a summary of what was found, and make the one call that genuinely required them. That is the shape of every human-in-the-loop design worth having: the human keeps the decision, not the legwork.
What actually changes for the person
Before, the whole sequence belonged to them. After, it looks like this:
Request
│
▼
Agent investigates and prepares
│
├──────────────┐
▼ ▼
Resolved Decision
automatically for a human
│ │
└──────┬───────┘
▼
Completed
The employee is still in the process. What has gone is the part of it that never needed them.
This is the same pattern as the invoice extraction system I built for a finance team, one level up. There, the model read a document and a person confirmed it. Here, the system reads a request, works out what it means, goes and finds the context, and presents a decision that is ready to make.
Why it is not one enormous prompt
The tempting version is to put everything in the context window — all the documents, the customer data, the business rules, the API list — and ask it to sort the problem out.
That demos well and fails in production, because a system people depend on needs things a prompt cannot give it: specific tools, permissions that are enforced somewhere other than the instructions, structured outputs you can validate, error handling for when a system is down, retries that are safe to repeat, an audit trail, and a defined escalation path.
Most of that is ordinary software engineering. The AI is a component inside it, with a job description — read these things, search these sources, use these tools, take these low-risk actions, and hand everything else to a person.
Rules, automation, AI, agents, humans
I deliberately did not make everything an agent. The architecture ends up layered, and each layer is chosen for a reason:
| Layer | When it is the right answer |
|---|---|
| Rules | The logic is fixed and you can write it down |
| Automation | The steps are known and always run in the same order |
| AI | Something has to be read, classified or drafted |
| Agent | The next step depends on what the last step found |
| Human | The decision is consequential, contested or irreversible |
Reaching for the top of that list first is the most expensive mistake in this field. A scheduled script that always works beats an agent that usually does, and a lot of “we need AI” problems are really process problems wearing a costume.
What to measure
The point was never an impressive demo, so the evidence has to be operational. Before building anything, we agree what would count as it having worked. Usually:
- The share of requests closed without a person investigating at all.
- Average time from request arriving to it being resolved.
- Hours a month returned to the team.
- The escalation rate, and whether it is falling as the thresholds are tuned.
- How often a human overturned the agent’s recommendation — the number that tells you whether to widen its authority or narrow it.
That last one matters most and is the one people forget to instrument. An agent whose recommendations are accepted 98% of the time should be given more room. One running at 70% has a policy problem, not a model problem, and no amount of prompt tuning will fix it.
We run the whole thing alongside the existing process first, deciding nothing and touching nothing live, until those numbers exist. Then the go-live is an evidence-based decision instead of a hopeful one.
From AI that answers to AI that acts
A chatbot can tell you what to do. An assistant can help you do it. An agent can be handed the task and take it as far as its boundaries allow — which is a genuinely different way to think about what software is for.
The goal is not to put AI everywhere. It is to find the work that does not need to consume a person’s day, and build something that can safely take it off them.
The goal isn’t more AI. It’s less work.
If your team spends more time getting ready to answer than answering, that is the interesting problem — not “how do we use AI”. Book a free call and describe the sequence they run before they can reply. More on how we approach AI systems.