Agents
What Is an AI Agent?
An agent is a system that decides what to do next. That single property is what separates it from the automation you already run and from the chatbot you already have — and it is also what makes it harder to build.
Drafted with AI assistance and edited by the Bralak engineering team. No client examples, performance figures or benchmark comparisons appear in these pieces — every technical claim is one a reader can check independently.
The word has been stretched to cover almost anything with a language model attached to it, which makes it close to useless in a procurement conversation. So here is a definition narrow enough to exclude things: an AI agent is a system that is given a goal rather than a procedure, and decides its own next step.
Everything else people list as a defining feature — tool use, memory, autonomy, multi-step behaviour — follows from that one property or is optional decoration on top of it. If a system cannot choose its next action, it is not an agent, however sophisticated the model behind it.
Automation, chatbot, agent
The clearest way to see the boundary is to hold the same task still and change only the system doing it. Take: a supplier invoice arrives by email and needs to reach the right place in the ledger.
You wrote the procedure. Extract these fields from this PDF layout, match on this supplier code, post to this account, and if anything is missing, stop and raise a task. It is fast, cheap and completely predictable. It handles the invoice you designed it for, and it fails visibly the moment a supplier changes their template.
It can read the invoice and answer questions about it. Ask what the VAT total is and you will get the right number. Ask it to post the invoice and nothing happens, because it has nowhere to post to. It produces text; the world is unchanged.
It is given the goal — get this invoice correctly into the ledger — and access to the systems that would let it. It reads the document, notices the supplier code does not match anything on file, searches the vendor master for a near match, finds the account was renamed last quarter, posts against it, and flags the rename for a human to confirm. Nobody wrote that sequence down.
That third description is the appeal and the whole problem in one paragraph. Nobody wrote the sequence down, which means nobody can read the sequence to check it before it runs.
An automation does what you specified. An agent does what it decided. The engineering is almost entirely in the second sentence.
The loop
Under the branding, every agent architecture is the same four-stage cycle, repeated until the goal is met or a limit is hit. It is worth knowing because every failure mode and every control point maps onto one of the four stages.
- InputRequestA goal arrives
- RetrievalContext retrievalGrounding in your data
- ReasoningPlan formationDecompose and plan
- SystemTool callOperate your systems
- DecisionVerificationCheck the result
- ActionResponse
- Human escalationApproval gateHuman confirms anything irreversible
Gather the current state — the request, the retrieved documents, the result of the last action, the contents of a record. This is where most quality problems originate, and they are invisible later: an agent reasoning correctly over the wrong context produces a confident, well-argued, wrong answer.
Decide what to do next given the goal and the state. In practice this is a model call whose output is a structured choice — which tool, with which arguments — rather than prose.
Call the tool. Query the database, send the email, create the ticket, post the ledger entry. This is the only stage that touches anything outside the process, and therefore the only stage where a mistake has consequences that outlive the request.
Check the result against the goal before continuing. Did the write succeed? Does the returned record look like what was asked for? Is the loop making progress, or has it tried three variations of the same failing call? This stage is the one most often left out, and leaving it out is what turns a wrong step into a run of twenty wrong steps.
A system with a rich model and no verify stage is not a cautious agent. It is an open loop with a budget.
What tools and memory actually add
Two capabilities get discussed as if they were features to be switched on. Both are better understood as decisions about blast radius.
Tools are the agent’s hands
A tool is a function the model is allowed to call — search_orders(customer_id), issue_refund(order_id, amount). The model does not execute anything itself; it emits a request to call one, and your code decides whether to honour it. That gap is where every meaningful control lives: argument validation, permission checks, spend limits, approval gates on the actions that are hard to reverse.
The practical consequence is that the tool list is the security boundary, not the prompt. A prompt that says never issue a refund over a certain amount is a request. A tool that rejects the call is a rule. Design at the second level and the first becomes a convenience rather than a defence.
Memory is a retrieval problem wearing a friendlier name
Agents are usually described as having short-term memory — the current conversation — and long-term memory across sessions. The second one is not a special faculty. It is a store you write to and retrieve from, with all the ordinary questions that implies: what gets written, who can read it back, how it expires, and what happens when what it contains is out of date.
Where agents fail
These are the four that account for most of what goes wrong in production. None is exotic, and all four are addressable — but only if the system was designed expecting them.
Each step takes the previous step’s output as fact. A small misreading at step two is load-bearing by step six, and the reasoning after it is internally consistent, which is exactly why it reads as trustworthy. Short loops with verification between steps contain this; long unattended chains do not.
The model has no calibrated sense of when it is out of its depth, so the failure arrives in the same register as the success. Any output a human is meant to check must carry its sources, or the check is theatre.
Given an under-specified goal, the agent optimises for something adjacent to what you meant — closing the ticket rather than resolving the problem. Constrain the definition of done in the tools and the verify stage, not in adjectives.
The agent reads a document, a web page or an email; that text contains instructions; the agent follows them. This is not a bug to be patched but a structural property of systems where retrieved content and instructions share a channel. The defence is that untrusted content can never expand what the agent is permitted to do — which brings you back to the tool list.
When an agent is the right choice
Agents are the more expensive option. They cost more per run, they are harder to test, and they need monitoring that deterministic software does not. That expense buys exactly one thing: the ability to handle variation you could not enumerate in advance.
| If the work is… | Build… | Because… |
|---|---|---|
| The same shape every time | An automation | A decision engine adds cost and non-determinism to a problem that had neither |
| Answering questions from documents | Retrieval with citations | Nothing needs to be decided or done — see RAG |
| Variable, judgement-heavy, spread across systems | An agent | This is the case where enumerating the procedure is the thing that is impossible |
| Consequential and irreversible | An agent with an approval gate | Autonomy and reversibility trade against each other; put the human where the damage would be |
The useful test before committing to one: could a capable new colleague do this task from the goal alone, given access to the same systems and a week to learn? If yes, an agent is a reasonable fit. If they would also need judgement you cannot articulate, or authority nobody would give a new starter, the honest answer is a narrower system with a person still in it.
Most projects that stall did not stall on model quality. They stalled because a process that was never written down turned out to be the actual deliverable, and writing it down was the work nobody had scheduled.
This piece supports our AI Agent Development page, which covers how we design, build and run these systems in production.
Your Next Intelligent System Starts Here.
Tell us what you’re trying to improve, automate or build. We’ll help you identify the right AI strategy and engineering path.