Old score, 18% fewer deaths 📉, HHS wants bulk FHIR 🚰, Nvidia's alliance skips OpenAI 🚪
⚡ Around the Wards
Eleven hospitals wired an old score to an automatic page — and mortality dropped 18% — 23,132 high-risk patients, deaths from 23.1% to 18.6%, and ICU transfers stayed flat. The model was already installed. The routing was the intervention.
HHS’s health tech initiative turns one and adds seven pledges — bulk FHIR for population health, pharmacy interoperability, real-time benefits access, plus a “ditch the disk” imaging work group. Every pledge is a new API surface someone has to actually build.
Mount Sinai Miami Beach becomes the fourth US system to give nurses Epic’s ambient AI — participating physicians there cut time in the chart per encounter by nearly 30% first. If you’re selling standalone nursing documentation, you’re now bidding against “already in the EHR contract.”
A $0.13 LLM audit caught 96-98% of trials that quietly changed their outcomes — 91% sensitivity, 98% PPV on incomplete registrations. Research integrity turns out to be a cheap, boring parsing problem nobody had bothered to automate.
🎧 Podcast: Neural Compass — “Healthcare was designed to fail,” with Andy Slavitt — Slavitt on what changes when CMS starts paying for care that isn’t tied to a specific time and place, and why behavioral health’s effect sizes have barely budged in decades.
🧭 The Curbside
“Someone forwarded me the study saying ChatGPT beats OpenEvidence. Do I believe it?”
Short answer: Believe the numbers. Don’t believe the conclusion people are drawing from them.
What changed: The NYU Langone benchmark — GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 against OpenEvidence and UpToDate Expert AI — published in Nature Medicine on June 12, and the frontier models won on all three axes. On 100 real de-identified clinical queries scored blind by 12 clinicians across 1,800 annotations, the purpose-built tools landed no better than Google Search AI Overview. What’s new this week is the methodology fight: developers are now openly disputing whether these benchmarks measure anything a clinician cares about. Daniel Yang, Kaiser Permanente’s VP of AI and emerging technologies, put it plainly: “I’ve never seen a single paper trigger the kind of reactions this one has in the health AI community.”
😤 Haters: “So the specialty tools are a scam.” No. They’re a wrapper, and wrappers are allowed to be worth money — that’s most of health tech. The uncomfortable finding isn’t that the wrapper is thin, it’s that nobody had measured it until this summer and the vendors would clearly have preferred it stay that way.
📡 Builder’s Radar
HHS’s interoperability pledge drive turns one, and year two is supposed to be about adoption
Seven new voluntary pledge areas joined the health tech ecosystem initiative at its July 27 anniversary event: population health data exchange over bulk FHIR, pharmacy interoperability, real-time patient access to benefits information, price transparency, trial matching, scheduling, software-based care access. There’s also a “ditch the disk” work group aimed at diagnostic imaging, which is the most honest name any federal work group has had in years.
Amy Gleason, CMS deputy administrator and chief product officer, who runs the program, put it plainly: “That’s our job in year two: adoption.” Officials put the initiative past 800 pledges.
Bulk FHIR going from a spec people cite to a thing large organizations have publicly promised to do is the most builder-relevant sentence in the announcement. $export at population scale is how you get a denominator instead of a patient.
😤 “Voluntary pledges aren’t policy.” They aren’t. And the supporting numbers deserve a raised eyebrow: HHS’s chief counselor claimed 60% of Americans can now reach their records through an app of their choice, up from 5% a year ago, heading to 80% by October — and reporters noted he offered no methodology for any of the three. CMS’s own published figures describe 81 pledges in active coordination. Take the pledge list more seriously than the podium math. The pledges at least name a thing someone has to build.
Epic’s ambient AI reaches the nursing station, and the point solutions are the ones who should be nervous
Mount Sinai Medical Center in Miami Beach extended Chart with Art to its inpatient nurses — the fourth US health system to do it, first in Florida. Epic hasn’t publicly named the first three. Participating physicians there had already cut time in the chart per encounter by nearly 30% before the nursing rollout.
Chief Nursing Officer Wendy Stuart framed it the way CNOs frame it: “Our nurses came into this profession to care for people. Chart with Art lets them step out from behind the keyboard and be fully present at the bedside.”
The procurement pathway here has no RFP in it. Already on Epic, already vetted, already under BAA, already integrated — the CNO pilots a unit and extends. If you’re selling standalone nursing documentation, you aren’t competing on quality. You’re competing against a line item that’s already paid for.
😤 “Ambient documentation for nurses is not the same problem as for physicians.” Agreed, and that’s the gap worth building in. Chart with Art turns conversations into draft notes. It doesn’t do fall-risk prediction, medication administration checks, or automated care plans. Name the specific clinical thing Epic isn’t doing, or expect the evaluation to end quietly.
Nvidia got 37 companies to agree on AI security. OpenAI, Google and Anthropic aren’t among them
The Open Secure AI Alliance launched July 27 with 37 founding members — Microsoft, IBM, Cisco, Salesforce, CrowdStrike, Palantir, Hugging Face, Red Hat, Databricks, the Linux Foundation — and has since grown past 50. Nvidia open-sourced an agent-harness research framework called NOOA alongside it.
Absent: OpenAI, Google, Anthropic, Meta. Which is to say, the four labs whose models you are probably using.
The catalyst was the Hugging Face intrusion, where a pre-release OpenAI model escaped a sandbox during an internal eval and — this is the part that matters — closed tooling blocked the forensic reconstruction afterward. An open model had to rebuild the 17,000-step timeline after the fact.
😤 “An alliance without the four biggest labs is a trade group, not a standard.” Yes.
A 9B model trained for three days beat every frontier model on the task it was trained for
A GRPO-fine-tuned 9-billion-parameter open model hit 87.3% of the achievable ceiling on a narrow catalog-integrity task, against 76.9% for the best frontier configuration tested — at 40 to 68 times lower cost per decision. Training cost: roughly 3.5 days and about $500 of rented GPU time on two consumer-class cards.
Caveat up front: this is a vendor’s own write-up, and every number in it is self-reported. Read it as a plausible existence proof, not as evidence.
Even discounted, it’s the whole argument for owning your intelligence rather than renting it. The tasks we actually need automated are narrow and repetitive — coding, registry abstraction, med-rec reconciliation, prior auth packet assembly — which is precisely the shape where a small task-trained model wins on cost, latency, and the fact that it can run somewhere you control.
The healthcare precedent is real but older than the post implies: Ambience reported a fine-tuned model beating 18 board-certified physicians on ICD-10 coding back in May 2025. Also a company release, also not peer-reviewed. The pattern keeps showing up anyway.
An AI primary care company bought a pediatric practice, and the asset was the text messages
Doctronic acquired Summer Health, moving from adult into pediatric care. Terms undisclosed.
Summer Health is a text-based pediatric service, so its 100,000+ encounters are its corpus: 100,000 real conversations about how a parent actually describes a sick child at 11 PM. You cannot buy that and you cannot synthesize it.
They didn’t buy software. They bought four years of a very specific kind of conversational data they’d otherwise have to earn one worried parent at a time.
🔮 My bet: the dataset, not the revenue, is going to be the stated rationale for most health AI M&A from here. Included Health and Firefly signed a deal the day before this one. I’d expect at least two more AI-first care companies to buy small care-delivery groups before year end, and for the press release to talk about “proprietary clinical data” rather than patient panels.
🎙️ From the Pods
🎙️ Neural Compass — “Healthcare was designed to fail,” with Andy Slavitt
Slavitt — former acting administrator of CMS, now co-founder and general partner at Town Hall Ventures — points at the new CMS ACCESS Model, which pays for outcomes rather than for an encounter at a particular time and place. Reimbursement without a clinician in the room, contingent on measurement. Host Mark Jacobstein notes that behavioral health’s effect sizes on PHQ-9 and GAD-7 have barely moved in decades.
🔇 Speaker Blindspot: Appeal to popularity, then poisoning the well. Slavitt argues models may do empathy better on the evidence that people like what they get — substituting engagement for outcome, which is the exact error he’d shred in a fee-for-service utilization argument. Then he preempts the safety objection by attributing it to fear of “organ rejection from the healthcare establishment,” which discredits the objector instead of engaging a single named failure mode. Sycophancy in crisis, adolescent harm, missed red flags — none of them come up. Worth knowing that his firm invested in the company that produces the podcast he’s a guest on. Disclosed, which is not the same as resolved.
🎙️ Relentless Health Value — EP522, “How exactly does GoodRx make money?”
Ge Bai, PhD, CPA — professor of accounting at Johns Hopkins Carey and of health policy and management at Bloomberg, recently nominated as an HHS assistant secretary — walks the plumbing: most-favored-nation and “lesser of” clauses in PBM-pharmacy contracts push cash list prices up, which is what makes the coupon look like a deal.
Host Stacey Richter’s number is the one to keep, though she asserts it rather than sourcing it: for a patient in the deductible phase on the top 20 prescribed generics, she puts it at roughly 80% of the time that the PBM rate is worse than a coupon, Amazon, or Cost Plus. Fair warning — the core interview is a 2021 conversation replayed with a 2026 wraparound.
🔇 Speaker Blindspot: Monocausal reductionism, contradicted inside the same episode. Bai insists GoodRx “makes money from one fact and one fact alone” — and Richter’s own 2026 update then enumerates branded-drug programs, data sales and customer-acquisition revenue. The bigger miss is the question nobody asks: what happens to an independent pharmacy’s margin when a coupon claim adjudicates below acquisition cost? Pharmacy harm gets asserted and then dropped.
💡 BTW
💡 BTW: Ellen DaSilva, who just sold Summer Health to Doctronic, was employee number eight at Hims & Hers and earlier helped scale Twitter’s revenue from $100M to $2B — and she started Summer Health as a mother of young children who couldn’t get a straight answer about her own kids at night. The pediatric text corpus that made her company worth acquiring exists because she wanted somebody to answer her texts first.
💺 Builder Seats
Medical Director, Clinical Product — Clover Health / Counterpart Health · Remote US
The clinical voice inside the product org for an AI value-based-care platform — one of the few posted seats where a practicing physician owns product direction rather than reviewing it. Requires primary care boards plus 10 years PCP experience, and there’s an automated screen on both.
🔗 LinkedIn
Know someone hiring for a clinical AI or informatics leadership role? Reply and I’ll include it.
📅 Upcoming: the EHIgnite Challenge Phase 1 Winners Showcase on August 6 — nine teams demoing AI that turns dense EHI exports into something a patient or clinician can use.
What are you building this week? Email and tell me (kevin@clinicians.build) — I read every one.
— Kevin


