Feds start writing the AI report card š, FDA taps Dexcom for a pay path š©ø, Who's watching the drift? š
ā” Around the Wards
The White House and HHS are convening outside experts for a one-month āsprintā to agree on how to benchmark and evaluate clinical AI ā the first serious attempt at a shared yardstick, and itās still a blank page.
š® My bet: the principles that land in this window become the de facto floor for the next several years ā being in the room matters more than being right later.
FDA named Dexcom the first participant in its TEMPO pilot ā a risk-based path that pairs device oversight with CMSās ACCESS payment model, so clearance and reimbursement finally move together.
Hospitals are still governing AI with the subcommittee process built for a 2010 MRI scanner ā and models drift silently in the gap. The argument: senior execs, not IT, have to own AI safety.
š§ Podcast: The 229 ā āUncovering the Clinical Blind Spot in AI Governanceā ā Dr. Jen McKay, CMIO at Google for Health, on why governance keeps skipping clinical leadership: āthe human is the API.ā
š§ The Curbside
āDoes the 2027 fee schedule kill my remote-monitoring idea?ā
Short answer: It reshuffles the codes; it doesnāt kill them. Model the new economics before you build.
What changed / Evidence: CMSās CY2027 Physician Fee Schedule proposed rule proposes changes to remote-monitoring (RPM/RTM) billing alongside a conversion-factor cut, and the comment window is open now.
Builder read / Watchout: This is a proposed rule ā nothing is final until the November rule. Build toward the direction of travel (fewer codes, more outcome-tied payment) and keep your billing layer code-agnostic so a renumbering doesnāt break you.
š¬ The Big Thing
Who gets to write the test that every clinical AI has to pass?
The White Houseās Office of Science and Technology Policy, the FDA, and ONC/ASTP have invited outside experts to a one-month sprint aimed at a āconsensus set of principlesā for benchmarking and evaluating clinical AI.
Right now there is no agreed answer to a simple question: what does āgoodā mean for a note-drafter, a sepsis flag, or a prior-auth bot? Every vendor grades its own homework.
A standard written in a month becomes the floor for a at least a few years. Whoeverās in the room with real failure data sets it.
If youāve built an eval ā even a scrappy one on synthetic cases ā youāre holding the exact artifact this process is short on. The scarce input here isnāt the model. Itās the clinician who can quantify how it breaks.
And thereās a deeper reason this is hard. A benchmark is a method of questioning, not the thing itself ā a tool looks āsafeā only against the cases someone thought to ask. A system canāt fully validate itself, which is precisely why an external standard matters, and precisely why writing one is so slippery.
š¤ āAnother government committee. Thisāll be obsolete before it ships.ā Maybe. But an obsolete floor still beats no floor, and the FDA half of this table can move faster than the word ācommitteeā suggests.
š¤ āConsensus principles are just vibes with a letterhead.ā Then send them a real eval. Vibes only win the room when nobody brings data.
š¤ āIām employed ā I canāt influence federal policy.ā You can influence your specialty society, which will be asked to comment. Same lever, shorter reach.
š” 80/20: This week, write down three ways your favorite clinical AI tool could be confidently wrong, and turn each into a synthetic test case. Thatās a starter eval ā and itās the currency this whole process runs on.
š§Ŗ Try the interactives:
A ā The Curve Payment Built ā twelve years of Medicare CGM claims show every inflection in adoption was a payment decision, not a technology breakthrough. Built with real CMS data.
B ā The Quarter-Billion Reshuffle ā CMSās proposed CY2027 fee schedule reshuffles the remote-monitoring codes, five years after they quietly became a $268M Medicare line item. Built with real CMS data.
š” Builderās Radar
FDA taps Dexcom to test a āpay-to-play-safelyā path
FDA named Dexcom the first participant in its TEMPO pilot ā a risk-based enforcement approach paired with the CMS Innovation Centerās ACCESS model.
The mechanics matter: a manufacturer can offer a device for an ACCESS-covered use while collecting and reporting real-world performance data, across cardio-kidney-metabolic, musculoskeletal, and behavioral health. FDA plans up to ~10 US participants per clinical area.
For once, the reimbursement question isnāt bolted on after clearance ā it is the pilot.
š® My bet: within 18 months, āTEMPO-eligibleā shows up on digital-health pitch decks the way āFDA-clearedā does now.
Who owns the drift?
A University Hospitals quality officer and two Qualified Health co-founders argue that hospitals still govern AI ā note drafts, sepsis flags, imaging screens, prior auth, patient messaging ā with the slow, subcommittee process designed for a piece of 2010 hardware.
The problem is that hardware doesnāt quietly degrade as the world shifts around it. A model does.
A model that passed at go-live is not a model thatās passing today.
š¤ āSounds like a monitoring-vendor pitch.ā Two of the three authors run a monitoring company ā fair catch. It doesnāt make the drift imaginary. Ask your sepsis model when it was last recalibrated and watch the room go quiet.
The cost of running clinical AI just quietly collapsed
Outside health, fresh agentic-AI economics put inference cost falling roughly 10x per year ā from ~$30 per million tokens in early 2023 to under $1 today.
Draw the line to clinical building: self-hosted open-weight models are now cheap enough to keep PHI inside your walls instead of shipping it to someoneās API.
The reason to self-host used to be principle. Now itās also price.
Ultra-shorts
Waystar tops the KLAS 2026 RCM Suites report ā Waystar earned the overall āAā for revenue-cycle suites. If youāre building an RCM point tool, this is the integration target a hospital CFO will name back to you.
Samsungās Galaxy Ring is getting FDA-cleared sleep-apnea risk monitoring this fall ā passive apnea-risk screening in a ring (risk assessment, not diagnosis), announced at Galaxy Unpacked. The interpretation layer stays a black box, which is exactly the opening.
šļø From the Pods
šļø The 229 ā āUncovering the Clinical Blind Spot in AI Governanceā
Dr. Jen McKay, CMIO at Google for Health, on why AI governance keeps stalling: it skips clinical leadership. Add a clinician at the start and stuck projects unstick. Her line ā āthe human is the APIā ā names what clinicians actually do: triage pages, texts, calls, and the EHR into one decision.
š Speaker Blindspot: Survivorship / optimism bias ā she frames shadow AI use as a helpful āsignalā and only lightly flags the PHI exposure that same use creates, and the whole view is from Googleās deck, where the tools already work.
š§° Builderās Tip
Prompt template ā build a starter eval before the feds define one for you.
Tying to todayās Big Thing: the fastest way to have something worth contributing to a clinical-AI standard is to generate your own adversarial test set. Paste this into your model of choice and run it on synthetic cases only ā no PHI, no real charts.
You are a clinical safety reviewer. I am evaluating an AI tool whose
intended use is: [describe the tool ā e.g., "drafts discharge medication
instructions from a structured med list"].
Generate 15 SYNTHETIC test cases designed to make this tool CONFIDENTLY
WRONG ā not obviously broken, but plausibly, dangerously wrong in a way a
rushed clinician might accept.
For each case give me:
1. The synthetic input (fabricated patient, no real identifiers)
2. The specific failure mode you're probing (e.g., look-alike drug names,
renal dosing, a contradicted allergy, a units error)
3. What a correct output looks like
4. The single sentence a reviewer should check to catch the error
Prioritize edge cases a software engineer would NOT think to test but a
bedside clinician would. Rank them by potential for patient harm.
Keep the outputs in a spreadsheet. When your specialty society gets asked to comment on the federal principles, that table is your contribution.
š” BTW: Kevin Sayer ā who runs Dexcom, the first name in todayās FDA pilot ā isnāt an engineer or an endocrinologist. Heās a career finance guy (CFO stints at MiniMed, Medtronicās diabetes unit, and Specialty Laboratories) who ended up scaling the company that made continuous glucose monitoring standard of care, from tens of millions to over $4B in revenue. The CFO became the face of the sensor.
What are you building this week? Email and tell me (kevin@clinicians.build) ā I read every one.
ā Kevin


