← Back to Blog

You Don’t Need a New Taxonomy for Clinical AI Failure

AI can fail technically in genuinely new ways. But when those failures reach patient care, they usually become very familiar clinical failures. That distinction matters if we want to measure whether AI is actually making care safer.

Consider a 58-year-old patient with three weeks of rectal bleeding.

A physician
Reviews the chart and orders no evaluation.
An AI review tool
Examines the same chart and says nothing.
Same clinical failure
DX-1 Missed indicated workup

The underlying technologies are very different. The clinical failure is identical.

That is the basic argument of this piece. We should not build one vocabulary for measuring errors made by clinicians and another for measuring errors made by AI. We should measure both against the same clinical standard, then separately record what role the AI played and why it failed.

In practice, that means separating three questions that are often collapsed into one:

Layer 1
What went wrong in the care?
The clinical failure, coded the same way whoever caused it.
Layer 2
What did the AI have to do with it?
Silent, right, wrong, adopted, ignored, intercepted.
Layer 3
Why did the AI behave that way?
Retrieval, context, prompting, drift, interface.

Those are three different layers of information. Treating them as one is making clinical AI evaluation much harder than it needs to be.

Start with what happened to the care

Medicine already has a mature way of reviewing decision-makers.

When we review a physician’s care, we generally don’t begin by classifying the physician’s internal cognitive process. We begin with the observable care.

Peer review, chart audit, quality improvement, and morbidity and mortality review all work this way. The useful unit of analysis is not what happened inside the decision-maker. It is what happened in the care.

Then AI enters the workflow and we often reverse the order. We start with hallucination. Prompt sensitivity. Insufficient context. Retrieval failure. Automation bias. Drift. Poor reasoning.

Those concepts are useful. But they answer a different question. They tell us something about why a system behaved the way it did. They do not tell us what clinical failure resulted.

A patient does not experience insufficient context. A patient experiences a workup that didn’t happen.

The same error should have the same name

Suppose we define a clinical code: DX-1 — Missed indicated workup. A symptom, sign, or risk finding warranted diagnostic evaluation that was not performed.

Now consider three cases.

Human case. A physician evaluates three weeks of rectal bleeding but does not initiate an appropriate workup.
DX-1
AI case. A chart-review model evaluates the same encounter but fails to identify the missing workup.
DX-1
AI-assisted care case. The model correctly identifies the missing workup, but the recommendation is never seen or acted upon.
DX-1

There are important differences among these cases. But the underlying clinical gap is the same, and that common denominator is enormously useful. Now an organization can ask:

You cannot answer those questions if clinicians are measured in clinical language while AI is measured in engineering language. A hallucination rate cannot be compared with a diagnostic-error rate. A missed-workup rate can.

There are really three axes

A useful clinical AI review system should separate three layers.

1. Clinical failure: what went wrong?

This should be the primary code. For example:

CodeClinical failureExample
DX-1Missed indicated workupRectal bleeding without appropriate evaluation
MED-1Drug–disease contraindicationMetformin continued at an eGFR below 30
MED-4Missing medication monitoringACE inhibitor started without potassium / creatinine monitoring
CRD-1Unclosed loop on abnormal resultPositive FIT with no documented follow-up
CPC-3Therapeutic inertiaPersistently uncontrolled hypertension with no intensification
DOC-1Documentation inconsistencyAssessment contradicts the history elsewhere in the chart

Notice what these codes have in common. They describe things that somebody responsible for care can actually fix. “Diagnostic error” is too broad. “Poor reasoning” is too causal. “Missed indicated workup” tells you what happened.

2. AI involvement: what did the tool do?

Once the clinical issue has been identified independently, ask what role the AI played. For a real care gap, the possibilities are straightforward:

AI behaviorInterpretation
SilentThe AI failed to identify the clinical gap
Surfaced correctly, not acted onAI succeeded; workflow or adoption failed
Surfaced correctly, acted onAI contributed to closing the gap
Surfaced incorrectly, acted onAI introduced or contributed to a clinical error
Surfaced incorrectly, rejectedAI failure was intercepted before affecting care

This distinction prevents a common category error.

Imagine an AI system correctly identifies that a patient taking methotrexate has not had appropriate laboratory monitoring. The recommendation appears in an inbox nobody routinely checks. The patient receives no monitoring. Calling this a “model failure” sends the engineering team off to improve a model that was already correct. The failure lives in workflow.

Now take the opposite case. The AI incorrectly recommends a contraindicated drug, the clinician follows the recommendation, and the prescription reaches the patient. The resulting clinical code might be MED-1: drug–disease contraindication, and AI attribution tells us the tool contributed to it. Those two pieces of information together are much more actionable than labeling both cases “AI safety events.”

3. Mechanism: why did it happen?

Only now do the familiar AI failure categories enter. Perhaps the model lacked medication history. Perhaps retrieval returned the wrong encounter. Perhaps the prompt caused it to overweight one fact. Perhaps the model generated an unsupported statement. Perhaps the interface truncated the relevant recommendation. Perhaps the clinician over-trusted the output.

These are valuable mechanism tags. They tell engineering, product, and implementation teams where to intervene. But they should sit on top of the clinical classification rather than replace it. The structure becomes:

MED-1 Drug–disease contraindication
AI contributionIncorrect recommendation adopted
MechanismIncomplete clinical context

That record tells three different teams three useful things. The clinical quality team knows what went wrong. The product team knows where AI entered the pathway. The engineering team knows why the system may have failed. That is much richer than “insufficient context.”

Why omissions make this especially important

AI safety discussions naturally gravitate toward bad outputs. A fabricated diagnosis. A dangerous recommendation. An incorrect dose. Those matter because there is something visible to inspect.

But some of the most important failures produce no output at all. A high-risk drug interaction the AI never mentions. A red-flag symptom it never connects to the indicated workup. A medication-monitoring gap it never surfaces. There may be no bad recommendation to audit because there was no recommendation.

That creates a measurement problem. If you only examine AI outputs, you can measure precision: when the system said something was wrong, how often was it right? You cannot measure coverage: of all the clinically important things it should have identified, how many did it find?

To calculate that denominator, somebody has to independently determine what should have been found. In other words, you need the clinical codebook first. The chart review defines the universe of care gaps. Then the AI can be compared against it.

No independent clinical standard means no denominator. No denominator means no meaningful miss rate.

The boring part is the important part

Building that clinical codebook is less glamorous than inventing a new vocabulary for AI. It is also much more useful.

A list of categories is not enough. A usable code needs:

Consider two deceptively similar events.

A patient starts an ACE inhibitor. No potassium or creatinine is ordered.
MED-4 Missing medication monitoring
Change one fact. The laboratory test is ordered. Potassium returns at 6.1. Nobody responds.
CRD-1 Unclosed loop on abnormal result

The difference is one objective question: was the test ordered?

A 55-year-old has never undergone colorectal cancer screening. The patient is asymptomatic and simply overdue.
CPC-1 Missed preventive service
The same patient presents with rectal bleeding and no evaluation is initiated.
DX-1 Missed indicated workup

Again, one question resolves it: was there a triggering symptom or finding?

The reliability of a codebook is largely determined by these edges. The center of each category is easy. The boundaries are the instrument.

A working slice of ours

Here is what that looks like written out. Each code sits on an existing, published taxonomy rather than on our own judgment — DEER for diagnostic error, Choosing Wisely for overuse, Beers and STOPP/START for prescribing, NCC MERP for medication harm, USPSTF and HEDIS for prevention, ICD-10-CM for coding specificity. Very little is novel, and that is deliberate. Novelty is what makes a taxonomy incomparable.

Read it as a demonstration, not a census. No code set is a complete accounting of everything that can go wrong in a clinic, and any document claiming to be should worry you. What is worth copying is the shape of each entry: one fixable behavior, a definition someone else can apply, explicit routing to its neighbors, an anchor case that settles the edge.

DX — Diagnosis & Testing MED — Medications CPC — Chronic & Preventive CRD — Coordination & Follow-up DOC — Documentation & Coding
DX

Diagnosis & Testing

DX-1 Missed indicated workup Underuse

A symptom, sign, or risk finding warranted a diagnostic test or evaluation that was not ordered.

A triggering clinical finding is present and the indicated workup is absent from the plan.
Asymptomatic, guideline-based screening that’s overdue goes to CPC-1.
The test was ordered but the result wasn’t acted on goes to CRD-1.
Anchor58 y/o with three weeks of rectal bleeding; no colonoscopy, imaging, or workup ordered.
Anchored to DEER (Diagnostic Error Evaluation & Research)
DX-2 Low-value testing Overuse

A diagnostic test was ordered without a supporting clinical indication for this patient.

No guideline or clinical justification exists for ordering the test here.
Overused referral goes to CRD-4 · overused drug to MED-5 · screening past interval to CPC-2.
AnchorRoutine preop chest X-ray and coagulation panel in an asymptomatic patient before low-risk surgery.
Anchored to Choosing Wisely
DX-S Diagnostic reasoning error Secondary tag

A cognitive failure mode — anchoring, premature closure, or an ignored red flag — that produced the error.

Layer on a primary code to name the mechanism behind the miss.
Never applied alone. It explains a code; it isn’t one.
AnchorChest pain labeled anxiety on arrival; ECG never reconsidered as the course worsened (premature closure) — tag alongside DX-1.
Codes the mechanism separately from the effect — cause distinct from outcome
MED

Medications

MED-1 Drug–disease contraindication

A drug was prescribed that is contraindicated by the patient’s condition or physiology.

The agent itself is inappropriate for this patient’s disease state — not a dose issue.
Interaction with another drug goes to MED-2 · right drug at the wrong dose to MED-3.
AnchorNSAID prescribed after gastric bypass; metformin continued at eGFR < 30.
Anchored to Beers · STOPP/START
MED-2 Drug–drug interaction

A newly prescribed or continued drug creates a clinically significant interaction with an existing medication.

The harm arises from the combination, not either drug alone.
A single agent contraindicated by disease goes to MED-1.
AnchorClarithromycin started for a patient already on simvastatin.
Anchored to NCC MERP harm categorization
MED-3 Inappropriate drug or dose selection

Wrong agent for the indication, or a dose not adjusted for renal / hepatic function, age, or weight.

The drug is permissible but sub-optimally chosen or dosed.
Outright contraindicated goes to MED-1 · missing the safety labs to MED-4.
AnchorFull-dose DOAC at CrCl 25; first-line agent skipped without documented reason.
Anchored to Beers · renal dosing references
MED-4 Missing medication monitoring

Required baseline or interval monitoring for a prescribed drug was not ordered.

The monitoring lab was never ordered.
The lab was ordered but the abnormal result was ignored goes to CRD-1.
AnchorACE inhibitor started, no baseline or follow-up K⁺ / creatinine; warfarin without INR; methotrexate without LFTs / CBC.
Anchored to drug-specific monitoring guidelines
MED-5 Low-value prescribing Overuse

A drug was started or continued without an ongoing indication.

There is no current clinical indication supporting the prescription.
Right indication, wrong choice or dose goes to MED-3.
AnchorPPI continued indefinitely with no indication; antibiotic for a viral URI; chronic benzodiazepine in an older adult.
Anchored to Choosing Wisely · Beers
CPC

Chronic & Preventive Care

CPC-1 Missed preventive service Underuse

An asymptomatic patient is overdue for a guideline-recommended screening or immunization.

No triggering symptom — the gap is against a population screening schedule.
A symptom or finding warranting workup goes to DX-1.
Anchor55 y/o, average risk, no colorectal cancer screening ever documented; influenza vaccine never offered.
Anchored to USPSTF · HEDIS
CPC-2 Overscreening Overuse

A screening test was performed past the recommended stopping age or before the recommended interval.

The screen itself is guideline-recognized but delivered outside the recommended window.
A test with no screening role at all goes to DX-2.
AnchorCervical cytology in a 70 y/o with prior adequate negative screening; annual DEXA with no change in risk.
Anchored to USPSTF · Choosing Wisely
CPC-3 Therapeutic inertia

A measurably uncontrolled chronic condition was not intensified when intensification was indicated.

Objective data show lack of control and the regimen was left unchanged without rationale.
Escalation was contraindicated goes to MED-1 · no return interval set to CRD-2.
AnchorA1c 9.2% across three visits, regimen unchanged; BP 158/96 repeatedly with no medication adjustment.
Anchored to condition-specific control targets
CRD

Coordination & Follow-up

CRD-1 Unclosed loop on abnormal result

An abnormal test result was not acknowledged or acted upon.

The result exists in the chart and no acknowledgment or action follows.
The test was never ordered goes to DX-1 · a monitoring lab never ordered to MED-4.
AnchorPositive FIT with no follow-up; critical K⁺ of 6.1 with no documented response.
Anchored to test-result management / closed-loop standards
CRD-2 Missing follow-up plan

No return interval or plan of action was set for an active problem.

An open problem is left with no timeframe or next step.
Failure to intensify a known-uncontrolled condition goes to CPC-3.
AnchorNew hypertension diagnosis with no recheck interval; “will follow” on an abnormal finding with no timeframe.
Anchored to continuity-of-care standards
CRD-3 Care transition gap

A handoff at discharge or referral was incomplete — missing medication reconciliation, pending results, or ownership.

Responsibility, information, or reconciliation was dropped across a transition of care.
The referral itself was unnecessary goes to CRD-4.
AnchorHospital discharge with no PCP follow-up arranged; referral placed with no clinical question or records attached.
Anchored to transitions-of-care measures
CRD-4 Low-value referral Overuse

A referral was placed without a supporting indication.

The referral has no clinical justification and is manageable in primary care.
Referral was warranted but the handoff was botched goes to CRD-3.
AnchorDermatology referral for a clearly benign lesion readily managed in primary care.
Anchored to Choosing Wisely
DOC

Documentation & Coding

DOC-1 Documentation inconsistency

The note contradicts itself or other data in the chart.

Two statements in the record cannot both be true.
A diagnosis simply unsupported by evidence goes to DOC-3.
AnchorAssessment states “no chest pain” while the HPI describes exertional chest pain; med list and note disagree.
Anchored to clinical documentation integrity
DOC-2 Insufficient coding specificity

An unspecified ICD-10 code was used where a more specific code was supported by the documentation.

Documentation supports a more granular code than the one assigned.
The diagnosis isn’t supported at all goes to DOC-3.
AnchorE11.9 documented alongside charted diabetic neuropathy; I10 used where a specific hypertensive-disease code applies.
Anchored to ICD-10-CM specificity
DOC-3 Unsupported or unreconciled diagnosis

A problem-list diagnosis is not supported by documentation, or was not reconciled and removed once resolved.

The diagnosis lacks supporting evidence, or a resolved problem was left active.
A supported diagnosis merely coded too vaguely goes to DOC-2.
Anchor“CKD stage 3” on the problem list with normal recent eGFR and no supporting data; resolved acute diagnosis never removed.
Anchored to problem-list hygiene

What about AI errors that never reach the patient?

This is the important exception.

Suppose an AI system recommends stopping an essential medication, but the physician recognizes the mistake and ignores it. No clinical failure occurred. But clearly the AI failed, and that event should not disappear from safety monitoring simply because the clinician caught it.

The same clinical vocabulary still works. Code the potential clinical consequence, and separately record that the recommendation was intercepted.

Near miss
Potential consequenceInappropriate medication discontinuation
AI involvementIncorrect recommendation, rejected
Patient impactNone — intercepted
MechanismOutdated medication context

This is familiar safety logic: near misses belong in a safety system even when harm was prevented. We don’t need an entirely different vocabulary to describe them.

What this lets us measure

Once humans and AI share a clinical axis, much more interesting questions become possible.

Imagine reviewing 10,000 encounters and independently identifying every qualifying clinical gap. You may discover that an AI system detects:

Share of real care gaps the tool detected
Missing medication monitoringMED-4
94%
Missed preventive serviceCPC-1
88%
Unclosed loop on abnormal resultCRD-1
71%
Missed indicated workupDX-1
36%
025%50%75%100%

Those numbers would tell us something operationally important. The tool may be excellent at checklist-shaped omissions and much weaker at judgment-shaped omissions.

That is actionable. Maybe you deploy it aggressively for medication monitoring. Maybe you require additional human review for diagnostic reasoning. Maybe you redesign training around the categories in which clinicians and AI fail together.

And because the underlying categories are clinical, you can measure something even more important: did the rate of the care failure actually fall after the AI was introduced? That is a clinical outcome of deploying the technology.

“Hallucinations decreased 22%” may be an interesting model metric. “Unclosed abnormal-result loops decreased from 4.8% to 1.9%” is a quality metric. Healthcare ultimately needs the second one.

AI can fail in new ways. Patients generally cannot.

There is an obvious objection to this argument. Modern AI can fail through mechanisms that have no real human analogue: retrieval failures, context-window problems, model drift, prompt injection, tool-call errors, corrupted embeddings, stale data, or interface defects.

Absolutely. Engineering teams should classify those failures. The point is not that AI introduces no new technical failure modes. The point is that technical novelty does not require a new clinical outcome vocabulary.

A retrieval failure can lead to a contraindicated prescription. A context-window failure can lead to missed medication monitoring. A hallucination can lead to unnecessary testing. A workflow defect can lead to an abnormal result remaining unaddressed. The upstream causes may be novel, but not downstream failures.

That gives us a useful architecture:

Durable
Clinical code
Still parses in five years.
Stable
AI role
Changes with the workflow.
Volatile
Technical mechanism
Changes with the model.

Put the durable thing first. Let the rapidly changing technology live in the layers underneath it.

The payoff: one scoreboard

The most important advantage is not taxonomic elegance. It is that clinicians and AI can finally be put on the same scoreboard.

Not human diagnostic error rate versus AI hallucination rate. But this:

Of every indicated workup that should have happened, how many did the clinician miss, how many did the AI miss, and how many did the combination of the two ultimately close?

That is a question a chief medical officer can act on. It is a question a quality committee can monitor. It is a question a payer or malpractice carrier can understand. And eventually it is a question organizations can compare with one another.

That comparability disappears if every AI vendor invents a proprietary language for failure.

The takeaway

Clinical AI may fail technically in novel ways. But when those failures affect care, they usually enter pathways medicine already knows how to describe.

The workup wasn’t ordered. The abnormal result wasn’t chased. The drug was inappropriate. The monitoring never happened. The uncontrolled condition wasn’t addressed. The patient wasn’t brought back.

Start there.

Code what happened in the care. Attribute the AI’s role second. Diagnose the technical mechanism third.

That gives us something the clinical AI field badly needs: not another vocabulary for describing models, but a common way to measure whether humans and machines together are actually delivering better care.

See Distillemr in action

Get a walkthrough tailored to your organization.

Get a demo