How we work
Five phases, each ending in a decision you make.
Every phase produces something you own and can act on without us. Stopping at the end of any of them is a legitimate outcome, and discovery is deliberately standalone.
- InputDiscover
- ReasoningArchitect
- DecisionPrototype
- SystemProduction
- ActionOptimize
Phase 01
Discover
Understand the workflow before designing anything. Most AI initiatives fail on selection, not execution, so this phase exists to make sure the right problem was chosen.
Typically 1–3 weeks
Activities
- Sessions with the people who actually run the process, not only the people who own it
- Map the current workflow end to end, including the exceptions nobody documented
- Assess data availability, quality and access, and what integration would require
- Score candidate workflows on volume, repetition, definability and measurability
Deliverables
- Prioritised opportunity shortlist with complexity honestly assessed
- Current-state workflow map
- Recommended starting point, with the reasoning
- Written note of anything we found that is a process or data problem rather than an AI one
Decision point
You have a recommendation you could hand to any competent team, including one that is not us. Proceeding is a separate decision.
Phase 02
Architect
Design the system, the data flow and the safety model together. Deciding what an agent may do unsupervised is an architectural question, not a configuration one.
Typically 1–2 weeks
Activities
- Target architecture: models, retrieval, tools, memory and orchestration
- Integration design against your real systems, including authentication and failure paths
- Classify every action by reversibility and set where approval is required
- Security and data-flow review, with your security team where you have one
- Define the evaluation set that will decide whether this is working
Deliverables
- Architecture document with the decisions and their trade-offs recorded
- Integration and data-flow specification
- Action classification and approval policy
- Evaluation plan with the questions and expected answers agreed up front
Decision point
The architecture is specific enough to quote implementation against. Quoting before this point is guesswork.
Phase 03
Prototype
Build the risky part first, against real data. A prototype that avoids the hard assumption has proved nothing.
Typically 2–6 weeks, depending on integration depth
Activities
- Implement the core loop against your actual data, not a sample
- Connect to at least one real system, read-only where write access is not yet approved
- Run the evaluation set and record where it fails
- Review sessions with the people who will use it
Deliverables
- Working prototype you can operate yourself
- Evaluation results, including the failures
- Revised scope where the prototype changed our mind
Decision point
Feasibility is now known rather than assumed. This is the cheapest place to stop, and stopping here is a legitimate outcome.
Phase 04
Production
The difference between the prototype and production is the failure paths, the observability and the approval workflow — which is most of the remaining work.
Typically 4–12 weeks, depending on how many systems it touches
Activities
- Harden integrations: idempotency, retry with backoff, dead-letter handling, reconciliation
- Implement approval gates, guardrails and input validation
- Instrument tracing so every run is attributable to a step afterwards
- Regression suite wired to run on every change
- Deployment, environment separation, secrets handling and access control
- Handover: runbook, documentation and a working session with the team who owns it
Deliverables
- Production system with monitoring and alerting
- Runbook and operational documentation
- Regression and evaluation suites in CI
- Named owner on your side who can operate it without us
Decision point
The system runs, is observable, and someone on your team can run it. Support scope is agreed separately.
Phase 05
Optimize
Widen the scope from what the logs show, not from what was hoped for. This is where approval gates get relaxed — if the observed error rate justifies it.
Ongoing, reviewed on a cadence you set
Activities
- Review throughput, exception rate and where cases are actually routing
- Close documentation gaps the system logged when it could not answer
- Re-tune retrieval and prompts against evaluation results, not intuition
- Recommend which approval gates can safely be relaxed, and which cannot
Deliverables
- Operating review against the evaluation set
- Prioritised list of next workflows, with evidence from this one
- Updated approval policy
Decision point
Each expansion is its own decision, made from measured behaviour rather than momentum.
Your Next Intelligent System Starts Here.
Tell us what you’re trying to improve, automate or build. We’ll help you identify the right AI strategy and engineering path.