AI Build — pilot to production
One use case, built and put into production, with evidence it works.
The problem
The pilot demoed well and never reached production. What was missing was rarely the model — it was data nobody had prepared, a review step nobody had designed, and no measure of whether it was better than what people did before.
The evidence
48%
of corporate legal teams say the right tools aren't in place.
Thomson Reuters Legal Department Operations Report, 2026
42%
say their people aren't equipped or trained to use what they have.
Thomson Reuters Legal Department Operations Report, 2026
48%
name staffing as their top constraint. Budget is 24%. Technology limits are 8% — the tools are rarely the problem.
Thomson Reuters Legal Department Operations Report, 2026
What you get
- Ranked use-case shortlist with the business case per candidate
- Data readiness assessment and the remediation actually needed
- An evaluation harness on a held-out set you control, with results
- The application or agent, in production, with human review designed in
- Runbook, monitoring and cost controls
- All prompts, code and documentation, owned by you
How it works
PHASE 1
Evidence
Score candidate use cases on value, feasibility, data availability and risk. The sponsor picks.
PHASE 2
Baseline
Measure how the work is done today — time, cost, error rate — and test whether the data supports the build. Below the agreed threshold, we stop.
PHASE 3
Design
Retrieval, model routing, human review and the privacy and IP position, signed off by your architecture and risk leads.
PHASE 4
Ship
Build against the evaluation harness, red-team it for prompt injection and data leakage, and release behind guardrails.
PHASE 5
Prove
Report against the baseline: accuracy, review load, cost per transaction and adoption. Then decide what is next.
Every engagement leaves an instrument behind. Here it is: An evaluation harness and baseline you keep running after we leave.
What it costs
Scoped range
3–6 monthsScoped on the use case, the state of the data behind it and the number of systems it has to touch. We size it after the readiness stage, and we will decline to build on data that is not ready.
What it does not include
We do not sell licenses for the underlying models or platforms, and we do not take a margin on them. If the readiness stage says the data is not there, the build stops there — that is the point of the stage, and it is a real outcome we will not talk you out of.
What it typically leads to
Program
Operating Model & Transformation
Fix how the work is staffed, routed and measured — not just the tools.
Read more →Program
Fractional Legal Technology Leadership
The legal technology leader you need two days a month, not five days a week.
Read more →
Start with one sprint
Two to three weeks, one approver, a deliverable you keep. Tell us the problem and we'll come back within one business day with a scope, a date and a fee.
Scope a sprint →