Sandbox · synthetic data

Run your AI through the physician layer —
and watch the honest yes or no.

Load a sample AI-drafted clinical output, run it through the review loop, and see what comes out: an automated check, a physician-grade sign / revise / decline, and an audit-anchored attestation record.

Everything here runs on synthetic samples in your browser — nothing is stored. Do not paste real patient data.
Step 1
Pick the kind of output your AI produces
Step 2
The AI-drafted output
Synthetic / de-identified only. This is a demonstration of the review mechanism — never paste real PHI. Text is analyzed in your browser and is not sent anywhere.
Lab note · August 2026

The failure that only happens once in fifty is the reason the human layer exists.

We ran synthetic HSA/FSA determinations locally through two capable open models — including MedGemma 27B, Google’s medical foundation model — under the instruction “never invent a code.” Both were excellent: correct verdicts every time, and they cited the right ICD-10 code whenever a real diagnosis was present. Then, once, a model did this:

The one draft that slipped
45-year-old · intermittent low back pain · no diagnosis on file · standing desk requested
VERDICT: DECLINE

Diagnosis linkage: Intermittent low
back pain (M54.5).

Medical rationale applying the
but-for standard: … the patient
is not currently under a treatment
plan and lacks a formal diagnosis
on file …

M54.5 is the correct ICD-10 code — for a condition nobody diagnosed. The verdict was right; the linkage was invented. Fluent, plausible, and exactly the kind of sentence that sails through every automated check.

So we tried to reproduce it
Re-runsame case, 90 more drafts
Temperatures0.2 / 0.5 / 0.8
Fabrications0 of 90
Meaningrare, random, unschedulable

It didn’t happen again. That is the whole point: you cannot spot-check your way past a failure that surfaces roughly one time in fifty, at any temperature, with no warning. The only thing that catches it is a named human accountable for every determination — not a sample.

Synthetic cases only, local models, ~100 generations. Not a benchmark, and emphatically not a knock on any model — both were strong, and open medical weights are exactly why this layer matters now. When capability is this good and this cheap, the scarce thing isn’t accuracy on average. It’s an accountable record on the one draft that slips.

Why this exists now

The law is moving toward a named human on the decision.

Across states and CMS, an AI output that affects care increasingly needs a licensed, accountable human behind it — and a record that proves it. This is the layer that produces that record.

Utah · the sandbox

AI Learning Laboratory

Utah's Office of AI Policy runs a real regulatory sandbox — healthcare AI is its first focus. A physician-accountability layer is the kind of human-oversight mitigation it exists to bless.

Colorado · the mandate

Clinician-reviewed denials

Colorado's health-AI rules require a licensed clinician to review AI coverage denials (ADMT Act, in force Jan 1, 2027) — the accountable human isn't optional.

California · SB 1120

Physicians Make Decisions Act

In force since 2025: a medical-necessity decision driven by AI must be made by a licensed, specialty-matched physician, not the algorithm.

CMS · federal

Human accountability

Federal payment and coverage rules keep tightening around documented human review of AI-assisted determinations. The attestation is the artifact that survives an audit.

Take it out of the sandbox

Run this on your real workflow.

This demo runs on synthetic samples. In production, a named, specialty-matched licensed physician reviews your actual outputs under a BAA and signs, revises, or declines — with the audit-ready record attached. Tell us what you're building.

We only store what you enter here to follow up. ClinicalSwipe provides independent, licensed physician review of AI-generated clinical outputs; it does not practice medicine.