The broken ruler: Why the healthcare industry is measuring AI notes the wrong way, and what they should be doing instead
August 19, 2026
Ambient AI scribes have moved fast. They now draft clinical notes in roughly a third of physician practices, and the results people care about — less time in the chart, lower burnout — are real. But having sat on both sides of this, as the clinician writing the note and as the administrator responsible for quality and safety, I keep coming back to a question that is easy to ask and surprisingly hard to answer: how do we actually know the notes are good, and that they are safe? Despite the success of AI scribes, this is still the question that keeps clinical leaders up at night.
Most programs answer with a quality score. A trained reviewer reads the note, rates it on a handful of dimensions, and the numbers get averaged into a single grade. When a new version of the AI scores higher, it ships. It feels rigorous. Increasingly, it isn't. Here at Suki, our team conducted our own analysis to find the real answer that clinical leaders are looking for.
The score everyone trusts was built for a different era
PDQI-9, the rating scale behind most of these programs, was created in 2012 — before ambient AI existed. It was designed to give clinicians feedback on their own writing and to spot the problems of that era: copy-pasted templates, outdated information, notes that rambled. It was validated on a small set of hospital charts, reviewed by clinicians who had read the patient's full record.
Modern AI writes nothing like a rushed clinician. It produces notes that are clean, organized, and easy to read — and that can still invent a medication dose, change a number, or quietly leave out something the patient said. The very polish these tools are good at is what makes their mistakes hard to catch.
A note can read beautifully and still be wrong in a way that matters.
What we found when we put it to the test
At Suki, we ran the standard scoring approach across four specialties, comparing two versions of our AI on real note samples, with trained physician reviewers. We weren't testing the AI. We were testing the test.
Clinical AI Informaticist Karl Swanson, DO, MS, and Machine Learning Engineer Nikita Gupta deployed PDQI-9 across Internal Medicine, Oncology, Pediatric, Orthopedic Surgery) with two production-candidate LLM models on 84 paired notes total (23 IM, 20 Onc, 20 Peds, 21 Ortho).
The finding that mattered most: When two qualified reviewers scored the same note, they agreed no more often than they would have by flipping a coin — and on the single dimension that matters most for patient safety, catching a made-up or incorrect fact, they didn't agree at all.
We set out to prove whether the test itself was trustworthy, and when we saw two reviewers rate the exact same note with a fabricated detail completely differently, that told us more than any average score ever could.
Nikita Gupta
Suki Machine Learning Engineer
In other words, the same note could earn very different grades depending on who happened to review it. Two versions of the AI could trade places in the rankings on reviewer opinion alone. If a scoring system can't tell you reliably whether a note is accurate, then a higher average score is not evidence that one version is better — or safer — than another.
A lot of Clinical AI folks think about evaluations with a more clinical lens, which makes sense; however, it can cause blind spots in more robust methodologies and appropriate ML literature. There are emergent ambiguities and uncertainties in medicine that PDQI9 and other holistic Likert approaches fail to capture. By breaking text down into units of facts, we can examine what each fact actually entails and whether it was justified or not from the inputs to the model, which helps reduce ambiguity and uncertainty.
Karl Swanson DO, MS
Suki Clinical AI Informaticist
Why this should matter to you
Almost every meaningful decision in a health system sits at the intersection of quality, safety, cost, regulatory exposure, and long-term strategy, and leaders rarely get to optimize one of these without answering to the others. A shaky quality score touches all of them at once:
- It undermines the case leaders make to their own stakeholders. When a vendor choice or an AI safety review has to be defended to a board, a quality committee, or a regulator, the measurement underneath that decision has to hold up to scrutiny. A score that changes with the reviewer will not.
- It lets the highest-risk errors hide in plain sight. A single fabricated detail buried in an otherwise excellent note barely moves an overall grade, and yet that is precisely the kind of error that can reach a patient and cause real harm.
- It weakens the year-over-year story. If the underlying measurement is not consistent, then “quality improved this year” becomes a claim that can neither be trusted nor honestly challenged — a difficult position for anyone accountable for the program.
- It does not help teams improve. A grade of 3.2 versus 4.1 tells you reviewers disagreed. It does not tell you whether the issue was a made-up fact, a missing detail, or awkward phrasing, so it gives the people running the program nothing concrete to act on.
Better questions to ask your vendor
You don't need to be a statistician to raise the standard. A few plain questions separate marketing from measurement:
- Do you measure specific types of error — made-up facts, missing information, misplaced details — not just an overall grade?
- Do independent reviewers actually agree when they score the same note? Can you show it?
- When you release a new version, is there a real test behind the decision — or just a higher average?
How Suki thinks about it
We published Karl and Nikita's findings — including where the industry-standard score falls short — because we think the bar for measuring clinical AI should be higher, and because we're building to meet it. Suki's approach focuses on the specific errors that carry clinical risk, measured in a way that holds up when different reviewers look at the same note, so the claims we make about quality and safety can actually be defended.
When two qualified reviewers scored the same note, they agreed no more often than they would have by flipping a coin — and on the single dimension that matters most for patient safety, catching a made-up or incorrect fact, they didn't agree at all.
Bobby Garg
Suki Medical Director
Suki is committed to rigorously evaluating the infrastructure governing AI scribe deployment in healthcare. That's the same rigor we bring to our product itself: an Ambient Clinical Intelligence layer embedded at the point of care, used across major EHRs, that on average lifts practice satisfaction by 81%, delivers $1,688 in incremental monthly revenue per user, and returns time to clinicians — without asking you to take “it scored well” on faith.
If you're evaluating ambient AI for your organization, we'd welcome a conversation about how we measure quality and safety, and a look at Suki in your own workflows. Get in touch today.



