Evidence first. Then a baseline.
Five stages, run in order, on a sprint of two weeks or a program of six months. The last stage feeds the next engagement's second stage, which is why we draw it as a loop rather than a line.
STAGE 1
Evidence
What is actually happening — in the data, in the contracts, and in what people do when nobody is presenting. We interview, sample and read before proposing anything.
A written picture of the current state, with its gaps named.
STAGE 2
Baseline
Measure it before we change it: time, cost, error rate, review effort. A claim of improvement is worthless without the number it improved on.
A baseline measurement both sides agree on.
STAGE 3
Design
Decide the approach and where a human has to be. Retrieval, routing, review points, privacy and IP positions — signed off by the people who own the risk.
A design your architecture and risk leads have approved.
STAGE 4
Ship
Build and release in slices, against an evaluation set, with guardrails and a way back. Something real is in use before the engagement ends.
Working software or a live process, in production.
STAGE 5
Prove
Report against the baseline, including the parts that did not improve. Then decide what is worth doing next, on evidence rather than enthusiasm.
A result measured against the baseline, and an instrument you keep.
Prove feeds the next Baseline.
Three rules that don't move
Nothing ships without a human-review design
Before anything reaches a user, we write down who checks the output, on what basis, and what happens when it is wrong. If we cannot answer that, it does not ship.
You own all prompts, code and documentation
Everything we write is yours, in your repositories, with no license back to us. You can end the engagement and keep working.
Every engagement leaves an instrument behind
A measure, a test set, a scorecard, a register — something that keeps working after we leave and lets you check the next claim yourself.
"AI programmes fail on data readiness and evaluation — not on modelling. So we gate both before a single agent reaches a user."
For a firm handling privileged material, that discipline isn't optional. We build AI-native legal applications with Claude Code and the Azure AI stack — evaluated on a held-out set you control, human-in-the-loop, and auditable end to end.
PHASE 1
Use-case & value case
Score candidates on value, feasibility, data availability and risk. Sponsor approves the top three and how success is measured.
- GATE
PHASE 2
Data & platform readiness
Readiness scorecard, lineage, and access controls. Below the agreed threshold, no build starts.
PHASE 3
Responsible-AI design
Retrieval strategy, model routing, human-in-the-loop, plus bias, privacy and IP review — signed by your architecture and risk leads.
- GATE
PHASE 4
Build & evaluate
Iterative build against an evaluation harness — accuracy, groundedness, latency, cost — and red-teamed for prompt injection and data leakage.
PHASE 5
Deploy, operate & scale
Guardrails, observability, cost caps and rollback. Drift monitoring and version control keep it honest in production.
What you receive
- Ranked use-case backlog with a business case per candidate
- Data-readiness scorecard and remediation plan
- The evaluation harness and its results — on a held-out set you control
- Operating runbook, monitoring dashboard and benefit tracker
Responsible-AI controls
- No client data to public models without consent
- Governed LLM access — model routing and allow-lists
- Prompt and output logging, with training-data exclusion by contract
- IP in all outputs assigned to you
We report on
- Evaluation pass rate against a held-out set
- Cost per transaction
- Share of outputs requiring human correction
- Time to first production use case, and adoption
- Surface AIProcess discovery
- AI GatewayGoverned LLM access
- AgamiKnowledge management
Delivery accelerators included at no licence cost.
Start with one sprint
Two to three weeks, one approver, a deliverable you keep. Tell us the problem and we'll come back within one business day with a scope, a date and a fee.
Scope a sprint →