Bralak AI

Agents

When an AI Agent Is the Wrong Answer

Most of the systems we are asked to build as agents should not be agents. The autonomy that makes an agent impressive is a cost, it is paid on every run, and there is a specific test for whether you are getting anything back for it.

, Founder7 min read

Drafted with AI assistance and edited by the Bralak engineering team. No client examples, performance figures or benchmark comparisons appear in these pieces — every technical claim is one a reader can check independently.

We build agents. It is the first service on this site, and it is the thing most enquiries ask for by name. It is also, more often than not, the wrong answer to the problem described in the enquiry — and saying so at the first call rather than the last is the more useful thing we can do with the conversation.

This is not a caution about AI. It is a caution about one architecture. The work is usually real; the autonomy is usually the part that is not.

An agent decides the sequence of steps at run time. If you already know the sequence, you are paying for a decision you had already made.

What autonomy actually costs

The appeal of an agent is that you do not have to specify the path. You describe the goal, give it tools, and it works out the order. That is genuinely powerful, and it is not free. The bill arrives in five places, and it arrives on every run rather than once at build time.

  1. Latency

    A planned sequence is one round trip per step. An agentic loop is a model call to decide the step, the step, and often a model call to evaluate the result. The same work takes several times longer, and the difference is felt by whoever is waiting.

  2. Cost

    Reasoning tokens are spent choosing between actions you could have chosen once, at design time, for nothing. At low volume this is invisible. At the volume that made the workflow worth automating, it is the line item.

  3. Testability

    A deterministic workflow has a fixed set of paths and you can test them. An agent has a distribution of paths, so correctness becomes a statistical claim requiring an evaluation set, a threshold and a regression suite — real engineering that a fixed sequence simply does not need.

  4. Debuggability

    When a workflow fails you look at the step that failed. When an agent fails you reconstruct why it chose the branch it chose, which requires tracing you have to have built beforehand and reasoning you can only infer.

  5. Blast radius

    Giving a system the discretion to choose actions means giving it the discretion to choose the wrong one. Every tool becomes a permission question, and the answers become approval gates, which become queues that people have to staff.

None of these is an argument against agents. They are the price of the thing agents are for. The question is only ever whether the problem is buying anything with it.

The test: is the path knowable in advance?

One question separates the two architectures, and it is answerable in a meeting without building anything.

The common failure is drawing the diagram, finding two or three branches, and concluding that the branching justifies an agent. It does not. Two branches are an if-statement. Ten known branches are a state machine, and a state machine is a better engineering artefact than an agent in almost every respect — it is faster, cheaper, exhaustively testable, and it fails in ways you can enumerate before it ships.

The genuine agent case is when the number of branches is not knowable, because the work depends on what the tools return: a research task where each source suggests the next, a diagnosis where the second check depends on the first result, a task spanning systems where the relevant ones are not known until you look.

Four shapes that are almost never agents

These come up repeatedly in enquiries, and in each of them the agent framing is doing work the framing itself invented.

Asked forUsually isWhat the agent framing costs
An agent that answers customer questions from our documentationA retrieval system with a good refusal thresholdTool-calling machinery and approval design around a system whose only action is to reply. All of the overhead, none of the consequence
An agent that processes incoming documentsAn extraction pipeline with a review queue for low-confidence fieldsNon-deterministic ordering over a process that must be auditable per document, and a much harder story when a regulator asks why one was handled differently
An agent that keeps our CRM up to dateAn integration with idempotent writes and a reconciliation passReasoning tokens spent per record to decide something a field mapping already decides, and a system that can be creative with your data
An agent that triages and routes ticketsA classifier, plus rules you already have written downA slower, less measurable version of a step that a small model does in a hundred milliseconds and that you can evaluate on a fixed test set
The request, what it usually is, and what it costs to build it as an agent

In every row, the model is still doing the hard part — reading unstructured input and forming a judgement. What is removed is the discretion over what happens next, which was never the difficult bit and is where all the risk lives.

The shape that is usually right

Most production systems that work look the same, and it is not the architecture in the diagram on the vendor’s homepage: a deterministic workflow, with model calls at the specific steps that need interpretation, and one bounded agentic region — if any — around the part that genuinely varies.

  • The frame is code. Triggers, sequencing, retries, idempotency, the audit record. These are solved problems and they should be solved the solved way.
  • The model reads and judges. Classify the email, extract the fields, decide whether the case is in policy, draft the reply. Each is a scoped call with a testable output.
  • Agency is scoped to the unknown part. Where investigation is genuinely open-ended, an agent runs inside a step budget, with a small tool set and a defined terminal state — and the workflow around it stays deterministic.
  • The person stands where the action is irreversible. Not everywhere, which produces a queue nobody reads, and not nowhere.

This is less exciting to describe and considerably easier to operate. It is also cheaper to change: when the process moves, you edit a step rather than re-tuning a prompt and re-running an evaluation suite to find out what else moved with it.

When an agent genuinely is the answer

It would be dishonest to write all of the above and leave the impression we think agents are a fashion. They are the right architecture in a real and growing set of cases, and the case has a recognisable shape.

  • The task requires investigation — the second action depends on what the first returned, and the space of first results is large.
  • The work spans systems whose relevance is not known in advance, so the tool set cannot be reduced to a fixed call sequence.
  • The volume of distinct cases is high enough that enumerating them as rules is a maintenance burden that will not be kept up.
  • The cost of a wrong action is bounded and reversible, or the irreversible actions are few enough to sit behind gates without creating a full-time review job.

Notice that the last one is a constraint on the business process rather than on the model. It is the one most often skipped in scoping, and it is the one that decides whether the finished system can actually be left running.

If there is one thing to take from this: decide the architecture from the shape of the process, not from the shape of the market. The process rarely changes as fast as the vocabulary does.

This piece supports our AI Agent Development page, which covers how we design, build and run these systems in production.

Your Next Intelligent System Starts Here.

Tell us what you’re trying to improve, automate or build. We’ll help you identify the right AI strategy and engineering path.