Buyer tool / Healthcare AI
Score your healthcare AI vendors on the twelve things that decide whether the system ever runs.
Most healthcare AI evaluations compare feature lists and pilot demos. Those aren't what fails. What fails is PHI handling nobody stress-tested, an integration that stalls in the EHR queue, and a delivery model that hands you a canvas and expects your team to build the worker.
This scorecard weights the twelve criteria that actually separate vendors after signature. We've published our own scores — including the three where we don't score full marks — and left the second column open for whoever else is on your shortlist.
Scored against the default weighting below. Every criterion is backed by evidence a buyer can independently verify — certificate numbers, reference calls, contract clauses.
The distinction the scorecard is built around
A platform is not a solution. Someone still has to build the worker.
The healthcare AI market sells two very different things under one word. Confusing them is the single most expensive mistake in a mid-market evaluation, because the cost of the second model shows up in your headcount plan, not the vendor's invoice.
You buy the tooling. You supply the build team.
Licence covers the platform, the connectors and the studio. Getting a prior-authorisation workflow into production requires solution architects, integration engineers, a clinical SME and someone to own it at 2am — sourced by you, or from the vendor's partner network at day rates.
Where the cost hides: the internal team you didn't budget for, and the 9–14 months before the first workflow touches a real claim.
You buy the working system. We build and operate it.
We scope one healthcare workflow, build the digital worker on CheiAI, integrate it into your record and claims systems, validate it against your own historical data, and then run it — with named engineers, monitoring and an escalation path that belongs to us.
What you own: a workflow in production and the outputs it produces. Not a backlog of platform configuration.
The framework
Twelve criteria, four categories, weighted to reflect what breaks deployments.
Each criterion scores 0–5. The category weights below are our defaults; you can change them in the scorecard to match your own risk profile. Under each criterion is the evidence to demand — from us, and from everyone else you're evaluating.
Delivery & production ownership
Default weight — 30 / 100Whether anything reaches production, and who is accountable when it does.
01Post-go-live operation ownership
After launch, does the vendor run the system, or hand you a runbook? Ask specifically who receives the alert at 2am and whose payroll they are on.
Evidence to demand: the support model in contract language, on-call rota structure, and the name of the escalation owner.02Time to first workflow in production
Not time to demo, and not time to pilot. Time until one real workflow processes real records under real volume. Count from contract signature.
Evidence to demand: two references who can state the date they signed and the date the first workflow went live.03Delivery bench depth and continuity
Can the vendor absorb a scope increase or an engineer leaving without restarting your project? Boutiques stall; large SIs rotate staff off the account.
Evidence to demand: engineer headcount, delivery locations, and the named team that stays on your account after go-live.Healthcare data & clinical fit
Default weight — 28 / 100Whether the system survives contact with your actual records.
04PHI handling and HIPAA posture
Where PHI sits, who can see it, whether it leaves your tenancy, and what happens to it inside model calls. A signed BAA is table stakes, not a differentiator.
Evidence to demand: BAA, data-flow diagram including model inference, retention policy, and de-identification approach.05Record, claims and document integration
Healthcare data arrives as faxes, scanned PDFs, HL7 messages and free-text notes. Ask what the vendor does with the 30% that is malformed — not the clean 70% in the demo.
Evidence to demand: a run against a sample of your own worst documents, scored before any contract is signed.06Accuracy validation method
How accuracy is measured, on whose data, and what the published failure modes are. A vendor that cannot describe where its system fails has not measured it.
Evidence to demand: validation methodology, gold-set construction, and a written list of known failure modes.Security, compliance & auditability
Default weight — 24 / 100What your security review and your auditor will actually ask for.
07Independent certification
Certifications held by the delivering entity, current and verifiable by number — not a parent company's badge reused in a deck.
Evidence to demand: certificate numbers, issuing body, scope statement and expiry date.08Zero-trust architecture and access control
Least-privilege access for engineers, secrets management, network segmentation, and whether offshore delivery staff can reach production PHI at all.
Evidence to demand: access-control matrix by role and geography, plus the last penetration test summary.09Audit trail, traceability and explainability
For any automated decision, can you reconstruct which inputs produced it, which model version ran, and who reviewed it? This is what a payer audit or a clinical incident review will require.
Evidence to demand: a live trace of one decision, end to end, in the vendor's own environment.Commercial risk
Default weight — 18 / 100What it costs in total, and what happens if you want to leave.
10Contractual accountability for outcomes
Whether performance claims survive into the contract as service levels with consequences, or evaporate between the deck and the MSA.
Evidence to demand: draft SLA clauses with remedies attached, provided before commercial negotiation closes.11True cost including your build team
Licence plus implementation plus the internal headcount the model requires. Score this on the three-year total, not year-one list price.
Evidence to demand: a three-year TCO including the FTEs the vendor expects you to assign.12Portability and exit path
If you leave in year three, what do you keep? Prompts, evaluation sets, integration code and fine-tuned artefacts should be yours in a usable form.
Evidence to demand: a written exit clause listing every artefact handed back and in what format.The scorecard
Score us against anyone on your shortlist. The second column is yours.
Our column is fixed — those are the scores we're willing to defend on a reference call. Set the other column using the evidence you've collected, adjust the category weights to your priorities, and print the result for your evaluation file.
Where we don't score full marks
Three criteria where a different vendor may beat us — and how to tell.
A scorecard where the publisher wins every row is marketing, not an evaluation instrument. These are the rows where we'd expect a well-prepared competitor to score higher, and the questions that will reveal whether they actually do.
Portability and exit path
Digital workers we build run on CheiAI, our orchestration layer. You keep the integration code, prompts, evaluation sets and data artefacts, and we document the exit path — but rebuilding the orchestration elsewhere is real work. A vendor delivering purely on your own cloud primitives has a cleaner answer here.
US federal compliance regimes
We hold ISO 27001 and ISO 9001 and operate to HIPAA requirements. We do not hold FedRAMP authorisation. If your workload falls under federal programme requirements, that constraint is real and you should weight criterion 07 accordingly.
Analyst presence and brand cover
We are not in the major analyst quadrants. If your board requires an analyst-recognised name to sign off the decision, that is a legitimate constraint and we would rather you raise it in the first conversation than the fourth.
Behind our scores
What each of our scores is backed by.
The products behind the delivery
DocuFindr — document intelligence for records, claims and correspondence, including the malformed scanned material that breaks generic extraction.
CheiAI — the orchestration layer digital workers run on, taking a defined workflow from specification through to production with the tests attached.
VuFindr — computer vision and video analytics, where the workflow involves physical environments rather than documents.
Healthcare and regulated-sector clients
Our healthcare work runs across provider and health-technology organisations, alongside regulated delivery in financial services where the same audit and data-handling standards apply.
Named with client permission. Reference calls arranged on request during evaluation.
References
Don't take our word for any of this. Take theirs.
Every score in the column above is a claim until someone who has already lived through a delivery with us confirms it. So we'd rather you stopped reading this page and spoke to them. Three reference calls, arranged by us, run without us — and each one mapped to the criteria they can actually speak to.
A client running a digital worker in production today
Someone past go-live, where the system is processing real volume and our engineers are still on the escalation path.
- 01Who picks up when it breaks at 2am — and how fast?
- 02What was the gap between signing and the first live workflow?
- 11How many of your own people did you have to assign?
A client whose delivery didn't go smoothly
Every delivery organisation has these. We'd rather hand you one than have you find it later. Ask what went wrong and what we did about it.
- 03Did the team change under you, and did that cost time?
- 06Where did accuracy fall short of what was promised?
- 10Did the contract protect you, or did you have to negotiate?
The security or compliance reviewer who cleared us
Not the sponsor — the person whose job was to find a reason to say no. Usually the most useful 20 minutes in an evaluation.
- 04What did our PHI data flow look like under scrutiny?
- 08Could offshore engineers reach production data?
- 09Could you reconstruct a single automated decision end to end?
No account manager on the line, no note-taker. You get the contact and you run it in your own time.
We tell them who you are and which workflow you're evaluating. Nothing beyond that.
If all three calls are glowing, you've learned nothing you couldn't have read in a case study.
If there's a specific client on our page you want to speak to, ask. We'll go and request it.
