Ambient AI is past the pilot stage. Most health systems are still running the pilot-stage scorecard.
Adoption, satisfaction, documentation time, after-hours work — those were the right metrics when the open question was whether clinicians would use the technology at all. Most systems have their answer on that narrow point. That’s not the same as knowing whether the technology is creating value, and the scorecard should reflect the difference.
The question worth asking is what happens once physicians actually get that time back.
A University of Edinburgh-led review of 27 publications on ambient AI scribes raised questions about whether these tools miss nonverbal cues, shift what patients are willing to disclose, or change how clinicians reason through a note. Nothing in the review shows ambient AI hurting outcomes — the evidence isn’t there yet either way. But it’s a fair prompt to ask whether “physicians like it and it saves time” is a complete success metric for an enterprise deployment. I don’t think it is.
To be clear, this isn’t an argument for pulling back. Plenty of health systems are already reporting real-time savings from these tools, and that’s a legitimate reason to keep scaling them. The point is narrower: time saved is the first data point, not the last one.
I’d want to know if patterns are showing up at scale. Are clinicians fixing the same section of every note? Is the same category of information getting dropped? Does accuracy hold across specialties, or only in the ones where it was piloted? Does anyone still read the note closely once they trust the tool?
Apply the same standard to the money.
Time saved is not value created. Does after-hours work actually go down, or just shift? Does the system gain real capacity, or just finish notes faster? Does rework fall, or move to someone else’s desk? Do staffing models change at all, or does everyone keep their same schedule and just leave earlier?
In my view, retention, recruiting, and lower burnout are legitimate returns — they don’t have to show up as an extra patient visit or a headcount cut. But name the outcome you’re underwriting before you buy the tool, then go check if it happened.
Where CFOs should push back
In every deployment I’ve been part of, the subscription fee ends up being the smallest number in the business case. Integration, training, clinical informatics, security review, analytics, support, monitoring, governance — that’s where the budget actually goes.
A good share of that cost never appears on an invoice. CMIO staff, physician champions, legal, compliance, cybersecurity, operations, revenue cycle, and finance all spend hours on this that nobody’s billing for.
Track where the work goes, not just where it disappears. If physicians spend less time writing notes but more time correcting them — or coding and billing teams inherit new exception work — you haven’t removed work from the system. You’ve relocated it.
That’s why you should be skeptical of any ROI math that converts minutes saved directly into dollars. A physician who gets ten minutes back a day doesn’t automatically produce another appointment. The clinic still needs demand, nursing coverage, an exam room, and an open slot to use that time. The investment can still be the right call. The return just isn’t the one on the vendor’s slide.
I don’t think this calls for a new governance committee. It calls for a short list of measures and one person accountable for each. For an enterprise deployment, I’d track four things:
Is clinician burden actually going down?
Is the record holding up, or are the same errors recurring?
What operational workflow changed because clinicians have time back?
What financial or strategic value is the organization actually capturing?
We spent the first phase asking whether ambient AI could cut documentation burden. Most systems now have a working answer to that. The harder question is what the time back is worth—better workflows, a physician who sticks around, real scheduling capacity, or a number the CFO can defend. Proving the tool saves time was the easy part. Proving what that time is worth is next.


