Everyone has an opinion, nobody has a shortlist
Forty use-case ideas across six departments, no shared scoring method, and no way to tell an eight-week build from an eighteen-month one.
The Agent Proof Sprint ends with a working agent on your own data, its accuracy measured against your own documents, and a production cost model at your volume. Fixed fee, credited in full against your pilot — and refunded in full if no prototype ships.
Three days on site · Accuracy measured on your data · Facilitated by the CTO, not a sales engineer
By the time most mid-market teams reach us, they've run two or three GenAI experiments that looked promising in a demo and died on contact with real data, real permissions, or real procurement. These are the patterns we see most.
Forty use-case ideas across six departments, no shared scoring method, and no way to tell an eight-week build from an eighteen-month one.
Vendor POCs run on ten clean sample documents. Your actual estate is scanned faxes, three EHR exports and a shared drive nobody has audited since 2019.
PHI, PII and residency questions surface at security review, not at design. The build stalls while legal reverse-engineers what was already shipped.
Token cost, retrieval infrastructure, human review and ongoing evaluation never make it into the business case, so the ROI number collapses under scrutiny.
A single enthusiastic engineer holds all the context. When they take leave or leave outright, the initiative quietly stops.
Leadership has committed to an AI position publicly, and the engineering answer is still “it depends” — which reads as no answer at all.
Each format is self-contained — you can stop after any one of them and still have something usable. Most teams start at 02. If you're not sure, say so on the call and we'll tell you which one fits, including if the answer is none of them.
A shared, non-technical vocabulary for the people who approve budget and carry the risk. No model architecture, no prompt tricks — what these systems can and cannot do, where the liability actually sits, and what “we use AI” commits you to in front of a regulator or an enterprise customer's security questionnaire.
We map your actual workflows, not a generic industry value chain, and score every candidate use case on the two axes that decide whether it ships: how ready the underlying data is, and how much a wrong answer costs. Most ideas die here. That's the point — killing eight bad use cases in a day is worth more than starting a ninth.
Two days, one use case, running code. We build on our Rapid Agent Builder framework against a sample of your real data — under NDA, in your environment where possible — so what you see on day two is behaviour on your documents, not a curated demo. You keep the prototype and the evaluation harness whether or not you engage us further.
The same build discipline as Workshop 03, run inside a regulated workflow — pre-denial claim validation, prior-auth document review, AML alert triage, or sanctions screening. We work against de-identified or synthetic data unless a BAA is in place, and the compliance constraint is a design input from hour one rather than a review gate at the end.
A stalled pilot rarely fails loudly. It absorbs two engineers for nine months, produces something nobody trusts enough to put in front of a customer, and quietly ends. Before you spend that, it's worth knowing what the workflow is actually costing you today.
Everything below leaves with you at the end of the sprint. There is no clawback and no obligation to engage us afterwards.
A running agent on your own data, source included. Not a demo environment — something you can put in front of the people who'd use it.
Your definition of “good” expressed in numbers, with accuracy measured against it on your documents rather than a public benchmark.
Inference, retrieval, human review and ongoing evaluation, projected at your volume. The figure your CFO will ask for and most vendors can't produce.
Two or three of your high-volume processes mapped end to end, including the manual steps nobody ever documented.
Every candidate scored on data readiness and cost-of-error — and the rejections written down with reasons, so the debate doesn't restart next quarter.
Everything standing between this prototype and something you'd trust in front of a customer, sized and sequenced.
A written recommendation with the reasoning attached, signed by the CTO. Board-ready as written, including when the answer is no.
Written into the engagement terms rather than negotiated at the end — whether or not you work with us afterwards.
Scope your sprint →Every engagement is a fixed fee agreed before we start — no hourly billing, no change orders, no separately invoiced discovery phase. What it costs depends on the depth, the workflow and whether regulated data is in scope, which is a ten-minute conversation rather than a form.
Right for you ifthe gap is board-level understanding and there's no engineering question on the table yet. The fee credits in full against a Proof Sprint.
Right for you ifyou want a build answer rather than a strategy answer. The fee is credited in full against a 90-day pilot — proceed, and the sprint costs you nothing.
Right for you if the workflow touches PHI or sits under an auditor, and the compliance answer has to hold up alongside the technical one.
Three guarantees, written into the engagement terms. None is conditional on you buying anything afterwards.
If you do not leave day three with a working prototype operating on your own data, you pay nothing. Not a partial refund, not a credit note — the full fee returned.
Start a 90-day pilot with us within 60 days of the sprint and 100% of the fee is credited against it. If you proceed, the sprint cost you nothing.
If we conclude you should not build it, we say so in writing and you keep every artefact. We would rather lose the pilot than sell you a build that fails in month four.
The same four movements run through every format — only the depth changes. It's the sequence we use on paid engagements, compressed into the time available.
Start from the business problem and the cost of getting it wrong — not from a capability looking for somewhere to land.
Test the idea against the data you actually hold. Most use cases fail here, and finding that out in a day is the cheapest failure available.
Put something running in front of the people who'd use it. Opinions about AI converge fast once there's a real output on screen.
Decide the review points, audit trail and escalation path while the design is still cheap to change.
The prototype, the evaluation harness and the source leave with you, whether you engage us afterwards or not. It's written into the engagement terms rather than negotiated at the end.
Most vendor workshops are co-funded by a cloud provider, which is why the recommendation always lands on that provider's stack. Ours aren't, so model and platform choice stays an open engineering question.
Thomas runs these sessions himself. There is no handoff from a workshop team to a delivery team, and no incentive to scope something that sounds good and delivers badly.
A significant share of use cases that reach us shouldn't be built — the data isn't there, the volume doesn't justify it, or a rules engine does the job for a fraction of the cost. Saying so is part of the deliverable.
There are good reasons to choose someone else. Here's the shape of the market as we see it, so you can rule us out quickly if we're the wrong fit.
| Option | Typical length | What you walk out with | Choose them when |
|---|---|---|---|
| Global SI workshop programme | 4 hours – 1.5 days | Prioritised opportunity map, maturity score, follow-on proposal | You need a globally consistent programme across many business units and a brand name your board already recognises. |
| Hyperscaler-funded workshop | 2 hours – 1 day | Use-case shortlist scored against that provider's services | You've already standardised on one cloud and want the fastest possible route to a funded POC on it. |
| Boutique executive session | Half day | Leadership alignment, risk framing, discussion guide | The gap is purely board-level understanding and there's no engineering question yet. |
| Internal, run by your own team | Whatever you can spare | Whatever your team already knows, made explicit | You have in-house people who've shipped production GenAI before. If you do, run it yourself — it'll be better and cheaper. |
| 10decoders | 3 days | Running prototype on your data, evaluation harness, cost model, written go / no-go — fee credited against your pilot | You want a build answer rather than a strategy answer, you're mid-market, and you'd rather be told no early than sold a programme. |
Categories describe common market patterns, not any single named provider. Durations reflect publicly published offerings as of August 2026.
The differentiator is that our CTO facilitates every sprint personally, and the engineers in the room are the ones who would do the build. That is also the constraint.
Any more and either the sprint or the delivery work behind it degrades. We would rather turn one away than run four badly.
The first six book at the rates on this page. Rates are reviewed after that, and the founding cohort keeps theirs.
Anything booked and scheduled before 31 December 2026 is honoured at these figures, whenever it actually runs.
Regulated Proof requires a BAA in place before day one — typically two to three weeks. Start that conversation early if the workflow touches PHI.
No 50-page proposals. We'll tell you which level fits your situation, what a realistic engagement looks like, and what it would cost — in one direct meeting.

Thomas leads 10decoders' AI engineering practice and sits in on the scoping call himself — so the person mapping your engagement is the one who has shipped it before. His teams build and deploy agents for mid-market healthcare and fintech companies, with enterprise grade build experience for clients like IBM, Dedalus and Harris Healthcare. He'll be straight with you about what's worth doing and what isn't.
Love what we're doing? Want to partner and sell our products or services?
Explore partner programs →Three fields. We'll reply within one business day.