Dr. Desmond Daley
The 3 AM Test for Healthcare AI | Dr. Desmond Daley
Why healthcare AI pilots fail at scale. A practical guide to building AI that survives real clinical workflows, 3 AM chaos, and the human limits of attention.
All articles
At 3:07 AM, the unit is loud in the way only hospitals can be. Monitors chirp. The hallway fluorescents hum. A nurse is juggling a new admit, a family asking for updates, and a pager that will not stop vibrating. The resident is trying to reconcile meds with one eye on the clock and one eye on the patient. Somewhere in that chaos, an AI alert pops up.
It is technically correct. It might even be clinically useful.
It is also ignored.
Not because clinicians are stubborn. Not because they "hate change." Not because the model is bad.
It is ignored because it arrives at the wrong moment, in the wrong place, with the wrong burden. It asks for attention when attention is the scarcest resource in the building. It adds a step when the workflow is already collapsing under the weight of ten other steps. It offers a score without ownership, without context, without a clear "what happens next."
That is what healthcare AI feels like in real life.
Now walk upstairs to the ballroom. The vibe flips. Smooth slides. Clean charts. Confident claims. AI, machine learning, predictive analytics, autonomous agents. You would think the future is already here.
That disconnect is why healthcare AI keeps stalling out.
We keep building for the ballroom and then acting surprised when the unit rejects it.
## The Uncomfortable Truth
Here is the uncomfortable truth: AI does not transform healthcare. People do. Systems do. Workflow does. Accountability does. The model is the easy part compared to embedding intelligence into a machine designed for reliability, liability, and human limits.
## A Pilot Is Proof. Scale Is Commitment.
Pilots succeed in controlled conditions. Data is curated. Edge cases are trimmed. Users are hand-picked champions who want the thing to work. Extra support exists off the books, like quiet heroism you never see on a dashboard. Manual checks. Shadow spreadsheets. Late-night fixes. A temporary layer of humans holding the system together.
Scale removes the scaffolding. That is the point of scale. The system has to stand on its own, across shifts, roles, staffing levels, and chaos. It has to survive Tuesday afternoon and Friday night and 3 AM. It has to keep working when the champions rotate off service and the project team moves to the next shiny pilot.
This is where the mythology breaks. We treat healthcare AI like a technology purchase when it is actually a transformation program. Technology purchases do not require an adoption owner. Transformation programs do. Technology purchases do not require Day-2 operations. Transformation programs do. Technology purchases do not demand a redesigned workflow. Transformation programs do.
Most healthcare AI fails for the same reason, over and over: we prove intelligence and skip operational reality.
## The First Collapse: Infrastructure
The first collapse is infrastructure, because pilots are allowed to be fragile. In early phases, teams tolerate brittle pipelines and one-off integrations. The solution "works," but only because someone is quietly babysitting it. Then rollout hits real IT. Legacy systems. Identity and access control. Latency. Downtime windows. Security reviews. Monitoring. Logging. Failure modes. Suddenly every dependency becomes a risk and every exception becomes a liability. The model may still perform, but the organization cannot absorb it.
## The Second Collapse: Data
The second collapse is data, because healthcare data is not one thing. It is a quilt made of billing claims, labs, vitals, imaging, notes, and half-truths trapped in free text. It lives in disconnected systems recorded in inconsistent ways. Fields get mapped and flattened until the clinical meaning changes. Updates arrive at different frequencies, breaking sequence and context. Workflows evolve, guidelines update, populations shift, and the model slowly becomes less relevant.
Models often do not fail dramatically. They fade. Alerts feel less helpful. Clinicians stop believing. The system stays "on," which makes leaders think it is fine, while the people doing the work treat it like background noise.
## The Third Collapse: Workflow
The third collapse is workflow, because friction compounds under pressure. A product team might see an extra click as trivial. A clinician sees it as an interruption at the worst time. An AI recommendation that lives on another screen is not decision support, it is a new chore. Insight that arrives too early or too late is not insight, it is clutter.
This is what "resistance" looks like in healthcare. It is not ideological opposition. It is rational adaptation. People route around what does not serve them. They use it only in low-risk moments. They bypass it when the unit is slammed. Adoption becomes selective. Trust becomes conditional. That is the beginning of the quiet fade-out.
## The Burnout Layer
Then comes burnout, the limit nobody wants to talk about because it does not fit neatly in a KPI. AI can be accurate and still fail at the human layer. If the system demands sustained attention, vigilance fatigue wins. If alerts are conservative to avoid misses, volumes climb. Clinicians triage notifications while also triaging patients. Over time the brain protects itself. The alert is no longer a signal. It becomes noise.
This is where many leaders misread the situation. They see a deployed system and assume usage. The unit sees a deployed system and feels another drain on cognitive bandwidth. Accuracy does not matter if attention is unavailable.
## The Ownership Problem
Now add the most lethal ingredient: unclear ownership.
During pilots, everyone "owns" the AI in a fuzzy, collaborative way. The innovation team. The vendor. The data team. A few champions. It works until it does not. Then a conflict happens. An output clashes with clinical judgment. The workflow breaks. The data pipeline shifts. A clinician asks, "Should I trust this today?"
Who decides?
If there is no answer, the system is not owned. It is installed.
Without clear decision rights and escalation paths, risk gets pushed downward. Individuals carry the liability. Individuals make the call. Individuals learn to ignore the system to protect themselves. The AI becomes a suggestion box with no accountable adult in the room.
## Governance and Compliance
Governance often arrives late as a reaction to pain. Something goes wrong, or almost goes wrong, and suddenly everyone wants policies. Committees appear. Approvals multiply. Oversight becomes synonymous with delay because it was bolted on after the fact. That is backwards. Governance is supposed to be designed in, so the system can move safely with speed.
Compliance fails in the same way. If compliance lives in training slides, reminders, and manual reviews, it will not scale. People will forget. Workarounds will happen. Exception handling will be underspecified, and the system will silently rely on humans to fill gaps under pressure. That is not compliance. That is hope.
## The Economics Break
Then the economics break, because pilot math is a lie by omission. Early deployments are subsidized by hidden labor and narrow scope. At enterprise scale, costs surface everywhere: integration, monitoring, retraining, security, governance, support, change management. Costs grow faster than usage. Value often accrues slowly, diffusely, and in the wrong budget line. IT pays while a different department benefits. Leadership looks at the spreadsheet and sees a widening gap.
## Measurement Fails
Finally, measurement fails, because we keep worshipping accuracy as if it is the point. It is not. Accuracy is table stakes. Impact is the point.
A model can have a beautiful AUC and still deliver zero meaningful change. If it does not save time, reduce cognitive load, change decisions at the right moment, or prevent harm in a measurable way, it will not survive budgeting season. Not because leaders are evil, but because resources are finite and healthcare is always on fire.
This is why so many initiatives end the same way: not a dramatic shutdown, just a slow drift into "we will revisit next quarter."
## The Autonomous Agent Era
Now, in 2026, the industry is sprinting into the autonomous agent era. Agents are not just chatbots that answer questions. They execute tasks. They draft notes, place orders, route messages, chase prior authorizations, coordinate scheduling, monitor inboxes, and attempt to orchestrate workflows across systems.
This is powerful. It is also dangerous in the specific way that power is always dangerous.
The more an agent can do, the more your organization must be explicit about what it is allowed to do, when, under what supervision, and with what fallback. Agents do not forgive ambiguity. They operationalize it. They escalate it. They make the blast radius larger because they touch more systems and move faster than humans can compensate.
If your workflows are brittle, agents will not fix that. They will stress test it at scale.
## Where Do We Go From Here?
So where do we go from here?
We stop chasing hype and start building for reality.
Pick one workflow. Not a platform. Not an "AI strategy." One workflow where time, trust, and safety are on the line. ED disposition. Discharge planning. Prior auth. Inbasket triage. Medication reconciliation. Nursing documentation. Choose a problem you can describe in one sentence and measure in one metric.
Then run the 3 AM test. Ask what happens when the unit is slammed, staffing is thin, and people are operating on fumes. Does the AI reduce steps or add them? Does it land inside the workflow or on another screen? Does it provide clarity or demand interpretation? Does it come with a clear escalation path? Do users know what to do when it is wrong?
Name an adoption owner. A real person, with authority, not a committee and not a vague "shared responsibility." Someone who owns training, workflow fit, metrics, and the right to pause the system if it becomes harmful.
Design Day-2 operations before go-live. Monitoring. Drift detection. Retraining cadence. Incident response. Change control. Auditability. If you cannot support it like a system, do not ship it like a system.
Measure impact, not applause. Time saved. Cognitive load reduced. Decisions changed. Harm avoided. If you cannot quantify the translation from prediction to action, you are not scaling intelligence. You are scaling theater.
This is not slower. This is faster than wasting a year.
## The Bottom Line
Healthcare does not need more AI slogans. It needs operational courage. The unglamorous discipline of building systems that respect human limits, clinical liability, and the reality that trust is earned through lived experience, not training slides.
If it cannot survive 3 AM, it does not scale. If it cannot reduce burden, it does not stick. If nobody owns it after launch, it fades.
The future of healthcare AI will not be decided by the next breakthrough model. It will be decided by whether we are willing to do the work that turns intelligence into infrastructure.