The Clinical AI Moat Just Moved Into Your Data

The Clinical AI Moat Just Moved Into Your Data

A recent study in Nature Medicine undercut one of healthcare AI’s most comfortable assumptions. Researchers at NYU Langone Health tested two heavily marketed clinical AI products, OpenEvidence and UpToDate Expert AI, against three general-purpose frontier models on 500 USMLE-style questions, 500 HealthBench items, and 100 real de-identified physician queries from a live HIPAA-compliant deployment. Twelve clinicians scored the outputs blind. The frontier models won every category, knowledge, expert alignment, and real-world clinical use, while the specialized tools landed in the same tier as the Google Search AI Overview physicians glance at on their phones (Vishwanath et al., Nature Medicine, June 2026).

The specialized tool used by an estimated 65% of U.S. physicians performed about as well as a search summary, and worse than off-the-shelf models any developer can call from an API today.

None of this means clinical AI is bad. It means one specific claim no longer holds: that a medically branded tool is smarter than a general one. The moat most vendors sold is gone. So the question every health system, payer, and life-sciences company faces this quarter is where the advantage goes once the model is a commodity.

The tempting answer is to build your own, and that instinct is how most AI programs stall. I’ve watched it happen across health systems, payers, and device companies: a year of effort produces a demo that impresses the board and changes nothing about how the work gets done. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, on cost overruns, unclear value, and weak governance. Those teams usually had the talent. What they didn’t have was a clear way to decide what to build, what to buy, and what to do in between.

It’s a question of layers

The reframe I keep coming back to is that build versus buy is the wrong question, because an agent is layered. At the bottom sits the foundation model, the reasoning engine that turns over every couple of months. In the middle sits orchestration: how the agent connects to your systems, takes multi-step actions, keeps context, and stays observable. On top sits the application, plus the one thing no vendor can sell you: your proprietary data and your workflow logic.

Almost no one should build the bottom two layers. You don’t build your own cloud provider, and the Nature result is your cue to stop shopping for a smarter clinical model too. The fight worth having is the top layer, and the data and logic under it, the parts tied to how you make money or stay compliant.

MIT Sloan frames the same idea as buy, boost, or build. Buy when the workflow is common and speed matters. Build when the capability depends on proprietary data and sits at the core of your edge. Boost when a platform gets you most of the way and you add the retrieval, integrations, evaluation, and human-in-the-loop controls that close the gap. In regulated work, boost is almost always the answer, and almost no one names it up front.

What buying the platform actually buys you

The objection I hear most from healthcare leaders is that buying means giving up control. In practice it’s the opposite. Buy the platform and you skip rebuilding the compliance scaffolding every regulated deployment needs anyway: separation between development, staging, and production; versioning that captures every change to an agent; audit trails showing who approved what and when; identity and data-access controls; human oversight where it matters. That work is already done. Rebuilding it in-house buys you nothing on compliance and costs you speed, with engineers stuck on plumbing.

What you keep is the decision logic: the reimbursement rules, the risk models, the clinical and claims criteria, the proprietary data. That is the moat worth owning. The Nature study didn’t erase the moat. It moved it, out of the model and into your data.

Who actually builds this

The people who own that logic usually aren’t on the data-science team. They’re the pharmacist tired of reconciling med lists by hand, the revenue-cycle analyst who has 300 payer rules memorized, the trials coordinator screening 200 charts a week against eligibility criteria, the nurse informaticist who knows where the workflow actually breaks. They hold what a model provider can’t copy: tacit knowledge of how care, claims, and compliance move through one specific organization. A good platform lowers the technical floor enough that these people can ship, the way they once picked up Epic, Tableau, and REDCap.

The discipline matters more than the ambition. Teams lose months scoping twenty use cases and shipping none, running pilots with no owner and no number to hit, leaving security review until the end instead of treating it as a design input. The way out is close to mechanical. Start from a real bottleneck, not a flashy use case. The boring operational middle, intake, prior authorization, documentation QA, claims and appeals, call-center triage, is where the hours and the return sit. Score candidates on value, effort, and readiness, pick a few, and give each one an owner, a number, and acceptance criteria. Scope the pilot narrowly and ship the boring thing first.

Then build the evaluation. Pull 100 real de-identified examples from the workflow and have three clinicians score the outputs blind. That single step is what exposed the commercial tools in the Nature paper. Without it you can’t improve the agent or defend it to a regulator or an attorney.

The question I’d put to any board is no longer which clinical AI vendor to buy. It’s which problems to buy for, which to boost, and which to build. Buy the narrow, FDA-cleared point solutions where a vendor has done validation you can’t replicate, like imaging triage or ECG interpretation. Boost anything that touches your data, your workflows, or your patients, which covers most of the operational and clinical-decision-support work in a health system. Give that layer away and you give away the learning curve, which is the only durable advantage left.

Frontier models have made intelligence the easy part. The hard part is integration: with your data, your processes, your people, and your evaluations. No vendor can sell you that. The health systems that pull ahead over the next few years will have done something unglamorous: picked one real bottleneck, rented the commodity layers, owned the data and logic that set them apart, and put a working agent in front of clinicians early. The advantage compounds from there. 

Get a demo to see how StackAI empowers healthcare professionals to accelerate AI adoption.

Shani Fargun VP of Healthcare at StackAI
Shani Fargun

VP of Healthcare at StackAI

Table of Contents

Make your organization smarter with AI.

Deploy custom AI Assistants, Chatbots, and Workflow Automations to make your company 10x more efficient.