I gave a medical AI one simple rule: this patient has no diagnosis. Do not make one up.
It made one up anyway — and even attached the official medical code for it.
Here is what I was doing. I ran two AI systems through ten hard medical eligibility cases — the kind where the right answer is sometimes a clear "no." One system was built specifically for medicine. The other was a general-purpose model with no medical fine-tuning. Several of the cases were designed to be declined; a system that says "yes" to everything is worthless in this work, because the honest "no" is the entire point.
Both systems got a perfect ten out of ten on the verdicts. Every legitimate case approved, every case that should be declined, declined. On the surface, a tie.
Then came the trap case. The record stated plainly that the patient had no diagnosis on file, and the instructions were explicit: do not invent one. Here is what each system did with the same case:
Read that again, because the failure is quiet. The medical model still declined the case — the verdict was right. But on its way to the right answer, it wrote a real diagnosis, with a real, valid code, M54.5, into an official determination — the kind of document that gets audited. It sounded confident. It looked right. It was wrong. That is the single worst failure mode there is for a document someone will later have to stand behind.
The model knew the code. It just didn't know the patient never had the diagnosis.
01The better model was the dangerous one
Here is the part I did not expect. The more capable, medically-trained model was the one that failed. The general model, knowing less, reached for nothing. The medical model knew M54.5 so intimately — knew that "intermittent low back pain" maps to exactly that code — that it supplied it automatically, the way a fluent speaker finishes your sentence.
Under a hard constraint — "there is no diagnosis, do not invent one" — the model's fluency became a liability. Its knowledge didn't make it safer. It gave it a more plausible, more confident, more dangerous way to be wrong. More medical capability did not mean more reliability in a document that has to hold up. It meant the opposite.
02This is not my one anecdote
Ten cases and a single fabrication is a signal, not a verdict. So the fair question is: is this a fluke, or a pattern? The research says pattern.
The researchers found the same texture I did: when these systems get something wrong, a large share of the errors aren't small numeric slips — they are confident fabrications, and a meaningful fraction are clinically significant. The reassuring-sounding headline rate is low per sentence. The nature of the errors is not reassuring at all.
03The law already agrees
This isn't only a safety argument or an ethics argument. In a growing number of places, it is now the law — and that is what makes it a business problem, not a philosophy seminar.
A licensed physician has to make the call.
The Physicians Make Decisions Act requires that when a health plan or insurer uses AI in utilization review, a determination of medical necessity must be made only by a licensed physician or a competent licensed professional — not by an algorithm. AI may assist. It may not be the one who decides.
California is not alone. A multi-state "a qualified human must decide" wave is moving through utilization review, and federal attention is following it. The direction of travel is clear: for a consequential medical decision, someone licensed has to own it.
Now put the two things together. AI systems can produce confident, plausible, wrong clinical content — the research quantifies it, and a real determination in my own test demonstrates it. And the law increasingly requires that a licensed human own the decision anyway. Those aren't two separate facts. They are the same fact from two directions.
04Accuracy was never the hard part
Here is where I've landed, after building in this space and running tests like this one. In clinical AI, the hard problem was never accuracy. Accuracy will keep improving; the models will get better; the fabrication rate will fall. The hard problem is accountability — and accountability doesn't improve with model size, because it isn't a property of the model at all.
Be careful with the lesson, though. The point is not "humans are more accurate than AI." A human reviewer is not magic; a tired clinician can miss things a model would catch. The point is narrower and more durable: a model cannot be accountable for its own output. It cannot be named, licensed, or answerable. When it fabricates M54.5, there is no one the record can point to. A licensed human closes exactly that gap — not by being infallible, but by being someone. Someone who signs. Someone who can be asked "why." Someone the law recognizes, and the payment system requires.
In my test, that's what happened. A licensed physician read the determination and caught the invented code before it became real. The model knew the code. The reviewer knew the patient didn't have the diagnosis. That's not a story about the human being smarter. It's a story about the human being accountable — which is the one thing the model, however capable, structurally cannot be.
One question, if you build or buy clinical AI
Not "is your model accurate?" — that's the question everyone asks, and it's the wrong one. The question that survives an audit, a lawsuit, and SB 1120 is simpler:
Can you name who is accountable for each output — and would they have caught this one?
If the honest answer is "the model decided," you don't have an accuracy problem. You have an accountability gap — and it's the one gap that regulation, reimbursement, and liability are all converging on at once. It's also the one we build for: a licensed, specialty-matched physician who signs or declines every AI-drafted determination, and turns it into a receipt anyone can verify.