# What Breaks Between a Prototype and Production

The prototype worked. The production system is late, and nobody can say why the estimate was wrong. It was not wrong about the model — it was wrong about which parts of the problem the prototype was allowed to skip.

- Cluster: Production
- Published: 2 September 2026 (2026-09-02)
- Reading time: ~7 min
- By Lakshay Chauhan, Founder — https://www.bralakai.com/authors/lakshay-chauhan

> Drafted with AI assistance and edited by the Bralak engineering team. No client examples, performance figures or benchmark comparisons appear in these pieces — every technical claim is one a reader can check independently.

There is a pattern that repeats often enough to be worth naming. A team builds a prototype in a fortnight. It is genuinely impressive: it answers questions about the company’s own data, it drafts the reply, it books the appointment. Everyone agrees it should ship. Four months later it has not shipped, and the honest internal summary is that nobody is quite sure where the time went.

The time went somewhere specific. A prototype is not a small production system — it is a production system with a set of assumptions granted to it, and every one of those assumptions is withdrawn at the same moment. The estimate was not wrong about how hard the model was. It was wrong about what the prototype had been allowed to skip.

**A demo needs one path to work. Production needs every other path to fail in a way somebody can act on.**

## The seven assumptions a prototype is granted

Not shortcuts, and not bad practice — a prototype exists to test feasibility, and testing feasibility means deferring everything that is not feasibility. The mistake is not making the assumptions. It is estimating the production build as though they were still true.

- **One user, no concurrency** — It was run by the person who built it. Nothing was contended, nothing was queued, no rate limit was reached, and no two runs touched the same record.
- **Clean input** — The test documents were the readable ones. The scanned fax, the password-protected PDF, the spreadsheet with merged header cells and the email with the entire thread quoted underneath were not in the sample.
- **Nothing fails** — Every API answered. Nothing timed out mid-transaction, no token expired, no supplier deployed a schema change on a Tuesday morning.
- **Everyone is allowed to see everything** — Retrieval ran as an administrator. Nobody asked what happens when the person asking may only see three of the five documents that answer their question.
- **Correctness by inspection** — The builder read the outputs and judged them good. There was no held-out set, no threshold, and no way to tell whether tomorrow’s prompt change made it better or only different.
- **Nothing needs an audit** — No one had to reconstruct why a particular run did what it did, six weeks later, for someone who was not in the room.
- **Change is free** — The prompt was edited fifty times in an afternoon with no review, no versioning, and no risk that an improvement to one case silently broke nine others.

Withdraw all seven at once and the remaining work is not a polish pass. It is most of the system, and it is engineering rather than AI — which is precisely why it is missed by an estimate produced while looking at a working demo.

## Where the four months actually go

Three areas absorb most of it, and none of them is model quality.

### Integration, in both directions

The prototype read from a copy of the data. Production reads from the live system, and writes back to it. Writing back is where the work is: idempotency so that a retry does not duplicate the action, reconciliation so that a partial failure across two systems is detected rather than assumed away, and a defined answer to what happens when the third system in the chain is down. A read-only prototype has not touched any of this, which is why it finished in a fortnight.

### The permission model

This is the one that most often forces a redesign rather than an addition. If retrieval was built without entitlements, adding them is not a filter on the output — filtering the answer is a leak waiting to be summarised. The user’s permissions have to become part of the query, which means the index has to carry the access metadata, which means re-ingesting the corpus. Discovering that in month three is expensive; it is knowable in week one.

### Evaluation, and the confidence to change anything

Without a held-out set, every subsequent change is a guess with a demo attached. Teams feel this as a slowdown they cannot explain: each fix takes longer than the last, because each one requires manually re-checking the cases the fix might have broken. Building the evaluation set is not overhead on the way to shipping. It is the thing that makes shipping repeatable, and it is cheapest to build while someone still remembers what the right answers were.

## The trade-off nobody states out loud

There is a real tension here and it is worth being honest about it, because the advice "build it properly from the start" is both correct and frequently wrong.

A prototype built with production discipline takes three or four times as long, and its entire purpose is to answer a question that might return *no*. Spending a quarter to discover that the data quality does not support the use case is a worse outcome than spending a fortnight to discover the same thing. The assumptions are not a failure of rigour; they are what makes a cheap answer possible.

> **The resolution: prototype the risk, not the demo** — A prototype should attack the assumption most likely to be fatal — usually data quality, integration access, or whether permissions can be expressed in the index at all — rather than the part that presents best in a meeting. Skip concurrency, skip the audit trail, skip the review queue. Do not skip the thing that could kill the project, because a prototype that avoided it has proved nothing and has cost a fortnight to prove it.

This is why the permission question belongs in week one even though permissions are a production concern. It is not that entitlements must be *built* early. It is that whether they can be expressed at all in the retrieval design determines whether the design survives, and finding out is cheap only while nothing has been built on top of it.

## The failure mode: shipping the prototype

The worst outcome is not the project that overruns. It is the one that does not — where commercial pressure and a convincing demo combine, the assumptions are never withdrawn, and the prototype goes live.

These systems fail in a characteristic way. They work for weeks, because the assumptions hold most of the time; the input mostly is clean, the APIs mostly do answer, and most users mostly are allowed to see most things. Then a document arrives in an unexpected format, or two runs hit the same record, or someone sees a passage they should not have. The incident is not a model failure and cannot be fixed by improving the prompt, and by then the system is load-bearing, so it is fixed under pressure by the people who now depend on it.

The tell is worth learning to hear. When someone says "it works, we just need to productionise it", the word *just* is carrying the second half of the project.

## Six questions to ask a working prototype

Before the meeting where everyone agrees it should ship, these six turn an impression into an estimate. None requires a technical background to ask, and the answers are usually known immediately.

1. **Whose permissions did it run as, and what happens when the asker has fewer?**
2. **What did it do when the input was malformed?** If nobody tried, that is the answer.
3. **Which of its actions writes to a real system, and which of those can be undone?**
4. **How would we know tomorrow whether a change made it better?** A name for the evaluation set, or there is not one.
5. **What did we deliberately not test, and why?** A team that cannot answer this has not been prototyping — it has been building slowly.
6. **If this made a wrong decision six weeks ago, could we reconstruct why?**

Six unfavourable answers do not mean the prototype failed. They mean it did its job: it proved the idea works and it left the engineering visible instead of hidden. That is a good position, and it is a considerably better one than a schedule built on the belief that the remaining work is a polish pass.

## Where this fits

This piece supports [How We Work](https://www.bralakai.com/how-we-work).

## Related reading

- [How to Identify AI Opportunities in Your Business](https://www.bralakai.com/insights/how-to-identify-ai-opportunities) — Most AI programmes do not fail on technology. They fail because the first project was chosen for how interesting it sounded rather than for whether it could be finished. Here is an audit you can run yourself, before anybody is asked for a budget.
- [What Is RAG?](https://www.bralakai.com/insights/what-is-rag) — Retrieval-augmented generation closes the gap between what a model was trained on and what your organisation actually knows. The interesting part is not the generation. It is that a well-built RAG system changes what happens when it does not know.

## Your Next Intelligent System Starts Here.

Tell us what you’re trying to improve, automate or build. We’ll help you identify the right AI strategy and engineering path.

- [Book an AI Strategy Call](https://www.bralakai.com/contact)
- [Start a Project](https://www.bralakai.com/contact)

---

*Bralak AI — Building Intelligent Solutions · Automating the Future.* Bralak AI Pvt. Ltd. — Noida, UP, India.

- Canonical page: https://www.bralakai.com/insights/prototype-to-production
- Agent index: https://www.bralakai.com/llms.txt · full text: https://www.bralakai.com/llms-full.txt
- Contact: info@bralakai.com · https://www.bralakai.com/contact