technology

The Five-Minute Wall

AI agents won't replace your doctor. But they might give your doctor back the five minutes they lost to a keyboard.

Sathyan··26 min read
A physician's hands at a keyboard, separated from an empty patient chair by a wall of floating EHR panels — patient summary, clinical notes, allergy alerts, drug interactions, orders, a click counter reading 236, and a clock at 5:00

I was in Chennai last year when I needed my medical records from a hospital in eastern India. Same hospital chain. Same brand. Same patient portal login. I called the help desk, gave them my patient ID, and waited.

They couldn't pull my records.

The hospital I was standing in and the hospital that had treated me were part of the same private network — one of India's largest — and their systems didn't talk to each other. The person on the phone was apologetic. "Sir, we don't have access to that branch's data. You'll need to contact them directly and request a transfer."

A transfer. Of my own health records. Between two buildings that share a logo.

I work in healthcare data. I've built EHRs end to end — each of them acquired more than once since — and I've spent two decades building systems that move, process, and adjudicate enrollment files, claims, and eligibility transactions: the operational plumbing that keeps insurance working. I know exactly why that phone call went the way it did. The systems were built for the hospital, not for me. The data exists. The bridge was never built because the patient was never the customer.

That phone call is where this essay starts. Because when people talk about AI agents in healthcare, they talk about replacing doctors. And that conversation is so far from the actual problem that it would be funny if people's lives weren't tangled up in it.

Two Minutes

A doctor in a government hospital outpatient department in India sees between two hundred and three hundred patients a day. That's not a typo. Two hundred to three hundred.

The average consultation time, measured across a systematic review of sixty-seven countries and twenty-eight million consultations: two minutes.

Two minutes to listen, examine, diagnose, prescribe, and document. Two minutes in which a human being walks in with a problem and walks out with — hopefully — a direction.

In rural community health centres, the situation is worse. Eighty percent of specialist positions are unfilled. Not understaffed — unfilled. No surgeon. No paediatrician. No obstetrician. In four out of five centres, the position simply doesn't exist as a filled role. Seventy-four percent of India's doctors practise in urban areas, serving thirty percent of the population. The other seventy percent gets what's left.

A gastroenterology OPD at a government teaching hospital in Bihar. Two hundred patients on the list. One senior consultant. The residents do the pre-screening — history, vitals, a quick exam — and the consultant sees each patient for the final assessment. The residents are good. They're also exhausted. They're also documenting everything by hand, or on a system that crashes twice a day, and they know that anything they miss will not be caught by a second pair of eyes because there is no second pair of eyes. The consultant trusts them because there is no alternative to trusting them.

India shows the extreme end of a capacity problem that wears different clothes in every country.

In the UK, 7.27 million people are on NHS waiting lists. The constitutional target — ninety-two percent of patients treated within eighteen weeks of referral — hasn't been met since 2016. The median wait is nearly twelve weeks. A hundred and six thousand people have been waiting over a year. In some specialties — ENT, oral surgery — barely half are seen within eighteen weeks.

In the US, the average wait for a new patient appointment across the fifteen largest metro areas is thirty-one days. The longest it's ever been measured. In Boston, it's sixty-five days. One dermatology office in Portland, Oregon logged a two-hundred-and-ninety-one-day wait. A projected shortage of up to eighty-six thousand physicians by 2036 means this gets worse, not better.

The physician is the bottleneck. And we're filling that bottleneck with paperwork.

The Click Tax

A physician in the United States spends forty-nine percent of their working time on electronic health records and desk work. Twenty-seven percent on direct patient care. For every hour spent with a patient, nearly two hours go to the EHR.

Read that again. The system designed to support clinical care now consumes twice as much physician time as clinical care itself.

An emergency physician makes approximately four thousand mouse clicks during a ten-hour shift. Ordering an aspirin: six clicks. Documenting a back pain physical exam: forty-seven clicks. Completing an entire admitted chest pain encounter: a hundred and eighty-seven clicks. The EHR has become a tax on care, collected one click at a time. (Hill, Sears & Melanson, "4000 Clicks," American Journal of Emergency Medicine, 2013)

And the tax doesn't end when the shift does. Family physicians spend an average of eighty-six minutes every night doing EHR work at home — what the industry grimly calls "pajama time." Nearly a quarter of all physicians spend more than eight hours a week on after-hours documentation. A third of senior family medicine residents report three or more hours a day of after-hours ambulatory EHR work.

The burnout numbers follow predictably. Forty-two percent of American physicians report burnout. Sixty-two percent cite bureaucratic tasks as the primary cause — not difficult patients, not long hours, not emotional toll. Paperwork. Among physicians who walk away from clinical practice early, the average age of departure has dropped from fifty-seven to forty-eight since 2008 — nearly a decade sooner than the generation before them. The top reasons they give researchers are the "hassle factor" and stress, which is a survey's polite vocabulary for not having gone to medical school to become a data entry clerk.

Nurses are leaving too. A hundred and thirty-eight thousand have left the US workforce since 2022. Forty percent intend to leave by 2029. Documentation workload is cited by forty-three percent as a top contributor. The WHO reports that at least a quarter of health workers globally showed symptoms of anxiety, depression, or burnout between 2020 and 2022, with no significant improvement since.

The system is haemorrhaging the people it cannot afford to lose, and the wound is administrative.

The Alert That Cried Wolf

Every EHR worth deploying has drug-drug interaction alerts, allergy flags, dosage warnings. This is table stakes. The technology exists, and it has existed for years.

It also doesn't work.

A meta-analysis across sixteen studies found that ninety percent of all clinical decision support alerts are overridden by physicians. Ninety percent. Drug-drug interaction alerts are overridden up to ninety-six percent of the time. Drug-allergy alerts up to ninety-five percent. In one survey, fifty-five percent of physicians admitted to dismissing alerts without reading them at all.

The reason is straightforward: the alerts are context-blind. A physician prescribes ibuprofen for a patient on warfarin and the system screams. Every time. Regardless of whether this is a one-time dose for a headache or chronic therapy in a patient with a history of GI bleeds. The system doesn't know. It fires the same red pop-up for a theoretical risk and a genuine danger, and after seeing a thousand theoretical risks, the physician stops reading.

In 2013, a patient received a dose thirty-eight times higher than intended — a 3,800% overdose — after every human in the chain clicked past the warnings. The alerts had fired. The technology had done its job. But it had fired so many times before, on so many non-issues, that the humans on the other end had been trained by the system itself to stop paying attention.

A primary care physician faces roughly fifty-six alerts a day, spending forty-nine minutes processing them. In high-volume settings, it's a hundred to two hundred. Over eighty percent are false or non-actionable. The system has turned its own safety feature into noise.

This is the gap that an agent — a real one, not a pop-up with a better skin — could actually close.

What an Agent Actually Changes

An agent that listens to a doctor-patient conversation and structures the documentation in real time does something subtler than saving five minutes per visit — it gives back the five minutes the doctor was never supposed to lose.

Think about what clinical documentation looked like before EHRs. A physician dictated notes. A medical transcriptionist — often working from a BPO in another city, or another country — listened to the recording and produced a structured document. The physician reviewed and signed. The process was slow, expensive, and depended on a human being who understood medical terminology well enough to turn a stream of speech into a coherent record.

An agent does that transcription faster, cheaper, and with the ability to structure the output against a template in real time. But that's the obvious part. The interesting part is what happens when the agent has context.

Pre-screening that actually prepares. A physician assistant or a junior resident walks into a room with a patient they've never met. Today, they open the EHR, scroll through a wall of data — past visits, medications, allergies, lab results — and try to piece together what matters in the two minutes before the patient starts talking. An agent could assemble that brief before the door opens. Not a data dump — a one-page summary. Here's what's relevant today. Here's what changed since the last visit. Here's what to ask about.

Contextual alerts, not pop-ups. The current alert system fires on theoretical interactions regardless of context. An agent that knows this patient's INR history, their actual bleed risk, whether the ibuprofen is a one-time dose or long-term therapy — that agent can decide whether an interaction is worth interrupting the doctor for. And when it does interrupt, it speaks in clinical weight: "Given her current warfarin dose and last INR of 3.2, this is a real interaction, not a textbook one. Published data shows GI bleeds in three of the last two hundred patients with similar profiles." That's a colleague picking their moment — not a pop-up that's been crying wolf for a decade.

Junior doctors who learn faster. A resident doing a pre-screening doesn't have fifteen years of pattern recognition. An agent that listens to the conversation and quietly surfaces relevant differentials, flags screening guidelines for this patient's demographics, reminds them of drug interactions they haven't memorised yet — that builds clinical judgment faster instead of replacing it. The senior physician's time is the most expensive, scarcest resource in the system. Everything that makes a PA or junior doctor more effective is a multiplier on that scarce resource.

Records that follow the patient. My hospital records couldn't cross a city within the same network. An agent backed by a canonical patient model — pulling from enrollment history, claims, lab results, pharmacy data, prior visits — could stitch those fragments into a coherent picture. The data exists. What's missing is the intelligence to find the fragments, match them to the right person, and assemble them before the doctor needs them.

None of this replaces a doctor's judgment. Every one of these use cases puts the physician in the loop — reviewing, overriding, deciding. An agent that drafts a summary is producing a draft. An agent that flags an interaction is flagging, not prescribing. The moment an agent makes a clinical decision without a human checkpoint, it has crossed a line that no amount of accuracy justifies crossing.

The Products That Should Exist

Modern healthcare IT products — EHRs, HIMS, practice management systems — are sold to hospitals and payers, not to patients or clinicians. The buyer is an administrator. The product is optimised for billing, compliance, and reporting. The clinician is a user, not a customer. And users get what customers pay for.

I've built these products, and I'll say the part the sales deck hides: the software is the easy half. A large hospital system buying a new EHR will spend twelve to eighteen months in configuration before the first real record moves — templates, workflows, order sets, formulary integrations, interface builds, testing. All of it manual, all of it expensive, all of it fragile. The product assumes someone else will handle the standing-up. That someone is usually a consulting firm billing by the hour, and the meter runs for a fiscal year. Payer systems carry the same weight on the other side: every major health plan runs a claims adjudication engine configured over decades — thousands of rules, hundreds of exceptions, documentation that lives in someone's head because the wiki was last updated in 2019.

This is the most underrated surface for agents in all of healthcare IT — bigger than documentation, because implementation pain is the reason the better system never gets bought. A hospital that knows the migration costs two years of institutional courage stays on the system everyone hates. Here is what the product that absorbs its own implementation looks like:

It reads before it asks. Point it at the legacy system's exports — the templates, the order sets, the interface specs, the adjudication rules — and it proposes the new configuration instead of handing you a blank workbook. The clinical rules still need a clinician's eye and the payment rules still need an analyst's sign-off. The three thousand mechanical mappings underneath them need neither.

It proves parity before cutover. Run the old system and the new one side by side on the same live inputs for a month. An agent compares every output — every claim priced, every eligibility answer, every order routed — and the discrepancies, ranked by clinical and financial weight, become the punch list. Cutover stops being an act of faith and becomes the closing of a list.

It interviews the institution. The rules nobody documented are the ones that break migrations. An agent that reads the configuration, the change history, and the exception queues, then writes the spec nobody ever wrote — asking a human only where the evidence contradicts itself — recovers the institution's memory before the person carrying it retires.

It makes implementation a number. Time to first live transaction should be printed on the box, the way range is printed on an EV. The buyer who demands that number will change this industry faster than any mandate. Ask two vendors how long until you're live. One says "six weeks." The other says "we'll staff four consultants and see." The first gave you a commitment. The second gave you a meter that's already running.

The products that should exist do more than optimise a workflow — they absorb the pain of standing themselves up. The first vendor in any category to make switching cheap resets what buyers believe is possible — and wins every deal along the way.

The Alert You Don't Show

There's one proposal above that deserves harder scrutiny than I gave it, because I wrote it and it reads clean: the agent that decides which interactions are worth interrupting for.

Firing an alert is a safe act. The system pointed at a risk, a human dismissed it, and the record shows diligence. Suppressing an alert is a clinical judgment made by omission — and the paragraph I wrote above never says who signs for it. If the agent decides the ibuprofen interaction is theoretical for this patient and stays silent, and the patient bleeds, the physician never saw a warning, the vendor's logs show the model behaved as designed, and the question of whose judgment failed has no name attached to it. An interruption engine needs what every other clinical actor has: an owner. A named clinician who approves the suppression policy, an audit trail of every alert withheld, and a measured false-negative rate — how often the agent stayed quiet about an interaction that mattered — before it ever touches a live prescriber.

The regulatory ground is just as unsettled. The FDA's 2022 guidance on clinical decision support turns partly on whether the clinician can independently review the basis for a software recommendation — that criterion is much of what keeps advisory software on the "not a medical device" side of the line. An alert that never fires presents no basis to review. Whether an engine that withholds warnings keeps that carve-out is a question I can't answer, and as far as I can tell, neither can anyone currently selling one.

Who Keeps the Five Minutes

Everything above makes a promise, and the title of this essay repeats it: the agent documents the visit, and the doctor gets five minutes back. Assume the technology works perfectly — the note writes itself, the clicks disappear, pajama time goes to zero.

Now ask the question the vendors never answer. Who keeps the minutes?

In a system with a thirty-one-day queue for a new appointment, recovered time is inventory. A scheduler looking at a physician who finishes visits five minutes faster has an obvious, rational, defensible move: book more visits. The administrator who approved the tool's budget needs the ROI slide to say something, and "the doctor breathes now" has never survived a budget review. Throughput has. The five minutes get harvested, the treadmill speeds up, and the burnout the tool was bought to fix continues — at a higher frame rate.

There's a second thing the pitch never explains, and I've spent two decades on the side of the wall where the explanation lives. The forty-seven clicks behind a back pain exam were never a UX accident. Each one is evidence for someone downstream — a medical-necessity reviewer, an auditor working three years in arrears, a coding level that has to survive challenge, a prior authorization desk. The clinical note stopped being written for the next clinician a long time ago; its real reader is a reviewer the physician will never meet, employed by the company that pays the claim.

Which means there are two ways to give a doctor time back, and they are not the same product. One makes compliance faster — the ambient scribe, the documentation agent, everything in the sections above. The other makes compliance smaller — shrinks what the system demands in the first place. In 2021, the US simplified the billing rules for office visits: physicians could code by medical decision-making or total time, and history and exam elements stopped counting toward the bill. No model was deployed; someone upstream simply stopped asking. The honest coda is that measured EHR time fell only modestly — one rule changed while templates, habits, and every other downstream demand stayed put. The lever works exactly as far as the demand actually shrinks, which is an argument for pulling it harder, not for putting it down.

The demand side belongs to my industry. A payer that answers prior authorization by API in seconds, against published criteria. A payer that gold-cards its physicians — Texas has required this by law since 2021: a ninety percent approval rate over six months earns an exemption from prior authorization on those services, because at that point the paperwork was never catching anything. A payer whose adjudication trusts structured data enough to stop demanding prose as proof. Every one of those removes clicks that the best ambient scribe can only accelerate.

So here is the test I'd apply to any tool that promises the five minutes back — including the ones I've proposed in this essay. Ask who buys it, and what number they show their board. Ask what mechanism — a scheduling policy, a contract term, a panel-size commitment — converts saved minutes into longer visits or a shorter queue. If nobody can name the mechanism, the minutes already have an owner. Saved time flows to whoever controls the schedule, and the schedule has never belonged to the doctor.

What Has to Be True

Later in this essay I measure vendors against the evidence and find seven studies, one with patients. Fairness demands the same instrument pointed at my own proposals. Every "an agent could" above is a hope until its preconditions are named — so here they are, falsifiable, the way a spec would state them.

ProposalWhat has to be true before it's real
Ambient documentationPhysician editing time is measured and lower than the writing time it replaced — a wrong draft can cost more than a blank page. After-hours EHR time drops at your site, not the vendor's reference site. A fixed fraction of generated notes gets audited by a human every week, with a named owner for that number.
Contextual alertsThe patient's INR history and problem list are queryable at the moment of prescribing. Every suppressed alert is logged and reviewable. A false-negative rate exists — measured against a retrospective corpus, not asserted. A named clinician signs the suppression policy. The regulatory question from "The Alert You Don't Show" has an answer in writing.
Pre-visit briefsThe longitudinal record actually exists — my Chennai records problem is solved before the summary is attempted, or the brief silently omits whatever the agent never saw. Every line in the brief carries provenance a doctor can check in one tap. The wrong-patient rate is measured and reviewed.
Records that follow the patientIdentity resolution runs with a conservative bias — a wrong merge is worse than a missed merge, because it attaches one person's history to another. Consent rails exist for every source. Lineage traces every fact back to the file it came from.

One category has already cleared several of these bars, and it deserves to be named rather than gestured at. Kaiser Permanente's medical group ran ambient AI scribes with over seven thousand physicians across roughly 2.5 million patient encounters between October 2023 and December 2024, and published the results: more than fifteen thousand hours of documentation time saved compared with non-users, with measured reductions in after-hours EHR time (NEJM Catalyst, 2025). That's the standard — a named deployment, a real denominator, a delta someone else can check. Ambient documentation is the one proposal in this essay with precedent behind it. The others are still hopes with preconditions, and now at least the preconditions are written down.

And one precondition sits underneath all of them. Every proposal above assumes an EHR to draft into, a longitudinal record to summarise, a claims history to mine. The gastroenterology OPD in Patna from the top of this essay has none of those. For most of India's public health system, the fixes in this essay arrive a decade after the pain unless they're designed for that world from the start — different rails, different workers, different assumptions. That world deserves its own essay, and it will get one.

The Discovery Problem

Before a patient ever reaches the five minutes, they have to find the doctor — and the tools for finding one, Practo in India, Zocdoc in the US, run on gamed star ratings instead of anything clinical. A doctor with a good front desk climbs the rankings while a genuinely excellent one who won't play the review game sits invisible. The patient can't see conditions treated, volumes, wait times, or whether their insurance is accepted. They see stars. An agent could answer the question the patient actually has: who can see me soon, treats what I have, and takes my insurance?

The first version of this essay said, right here, "the data exists." I have to retract that before someone quotes it back at me. It half-exists. CMS audits of Medicare Advantage provider directories have repeatedly found errors in roughly half of all listings — wrong location, wrong phone number, a false "accepting new patients" flag. A Senate Finance Committee secret-shopper study of mental health provider listings managed to book an appointment eighteen percent of the time. The industry has a name for this: ghost networks. So the agent's real job in discovery is harder and more interesting than assembly — reconciling sources that contradict each other, weighting a claims feed that says a physician is active against a directory that says they left two years ago, and answering "unverified" when the evidence is thin. An agent that inherits the directory's confidence inherits its errors.

What's Broken in How This Is Being Sold

The pitch from most vendors goes like this: "Our AI agent automates clinical workflows, reduces documentation burden by forty percent, and improves patient outcomes." The demo is polished. The slides have case studies. The ROI projections are confident.

The evidence says something different.

A 2026 scoping review in npj Digital Medicine searched five databases for studies on agentic AI in healthcare. They found seven. Seven eligible studies, spanning emergency medicine, oncology, radiology, and rehabilitation. Most were exploratory. Most lacked robust clinical validation. Only one involved actual patients.

Seven studies. One with patients. And the industry is selling certainty.

That review's scope is agentic AI — systems that decide and act — and the ambient documentation results earlier in this essay are the exception that sharpens the point: the best-evidenced use of AI in clinical work today is the one closest to transcription. Gartner projects that around forty percent of agentic AI projects will be cancelled by the end of 2027 — over unclear value, governance problems, and cost. The most mature deployments today are in patient access, clinical documentation, prior authorisation, and revenue cycle operations. The administrative layer. Where the results hold up, they come from hybrid models — agents paired with trained human teams, quality oversight, and explicit escalation paths.

The "AI will replace doctors" headline is the loudest and the most wrong. Agents don't replace doctors. At their best, they replace the forty-seven clicks it takes to document a back pain exam. They replace the fifty-six alerts a day that a physician has been trained to ignore. They replace the eighty-six minutes of pajama time that burns out a family physician before they hit fifty. They replace the administrative weight that is driving clinicians out of the profession a decade before their time.

The vendor that understands this builds a tool that gives time back — and a contract that says who keeps it. The vendor that doesn't builds a demo that looks impressive and collapses the first time it encounters a companion guide variation it hasn't seen before.

Agents Fail Differently

Software you're used to crashes. It throws an error, stops, tells you something went wrong. A crashed system is inconvenient. A crashed system is also honest.

Agents fail fluently. They produce a confident, well-formatted, entirely wrong answer. A hallucinated drug interaction. A plausible-sounding dosage that's off by a factor of ten. A patient summary that merges two different patients' histories into one coherent — and completely fictional — narrative.

In most software, a confident wrong answer is a bug. In medicine, it's a patient safety event.

There's a second failure mode, quieter than the clinical one, and my side of the industry will meet it first. The entire audit apparatus of healthcare billing treats documentation as costly evidence — a thorough note is credible partly because it took effort to produce. When notes cost nothing to generate, that assumption dies. An agent tuned to produce "complete" documentation will drift toward notes that support higher billing codes, fluently, plausibly, and at scale — upcoding laundered through eloquence, with nobody having typed a dishonest word. The payer's traditional defense, auditing the notes, was built for a world where no clinic could produce ten thousand immaculate-looking records a day. Someone now can. Every payment-integrity team in the country is about to learn what fluent failure means.

This is why the "AI proposes, rules dispose" architecture matters. An agent that drafts, summarises, or suggests is useful. An agent that acts without a human checkpoint is dangerous. The verification step is not optional. The deterministic gate — schema checks, dosage range validation, drug interaction rules, identity matching — is not a nice-to-have. It's the difference between a tool and a liability.

Clinical responsibility does not transfer to a tool. Anything that reaches a patient record, a referral, or a prescription is the physician's output and their accountability, however it was generated. An agent that drafts a note does not carry a medical licence. The doctor who signs it does.

Anyone building agents for clinical use needs to internalise this: the failure mode is not "the system crashed." The failure mode is "the system sounded right and no one checked." That is more dangerous in medicine than any system outage.

Internalising it means operational answers, not architectural diagrams. What fraction of generated notes does a human audit each week, and who owns that number? Which signal is monitored for drift, and who gets paged when it moves? Where is the kill switch, and who is authorised to pull it without a meeting? Unglamorous questions — and a deployment that can't answer all four is an experiment running on patients without a consent form.

The Opinion That Matters Most

The argument of this essay is easy to state and hard to hold onto, because every vendor deck is built to make you drop it. The technology is real. The evidence is thin. The potential is enormous. And the people who should be shaping it — the clinicians, the ASHA workers, the care gap analysts, the enrollment specialists who know why the queue breaks — are being told to learn a framework before they've earned an opinion.

They already hold the only opinion that matters.

The most valuable thing a clinician brings to this field is knowing what actually goes wrong in a clinic at four o'clock on a Friday, why a workflow that looks inefficient on paper exists for a good reason, and what a safe failure looks like — none of which is code.

The most valuable thing an operations analyst brings is knowing which queue breaks first when volume spikes, why a rule that seems redundant exists because of a payer who went bankrupt in 2014, and what "resolved" actually means versus what the system reports.

Engineers cannot supply that. Learn enough of the technology to be a credible partner in the conversation, and your judgment stays the scarce ingredient. The self-paced path that used to be an appendix here — mental models to healthcare IT foundations, none of it requiring a line of code to start — now has its own room: The Scarce Ingredient.

Enjoyed this?

Get new articles delivered to your inbox. No spam, unsubscribe anytime.

Related Articles

More from Narchol