VA signs a $1.6B agent license đď¸, Hopkins won't scale one yet đ§Ş, Evals are the new PRDs đ
⥠Around the Wards
The VA awarded up to $1.6B under an âAgentic Enterprise License Agreementâ â a contract category nobody was using last week. Full treatment below.
Johns Hopkins is benchmarking agents before it deploys them â and published the five things it measures. Steal the list.
A patient brought two years of her own data and nothing could receive it â nine of fifteen visit minutes went to rebuilding it off a shared screen.
Machinify is on the block â reportedly ~$750M revenue at ~38% EBITDA margins for deciding whether your claim gets paid.
đ§ Podcast: Lennyâs Podcast â Dianne Penn, Anthropicâs first technical PM â âevals are the new PRDs.â Early-days âClaude doesnât follow instructionsâ feedback decomposed to roughly 80% âClaude would not write the right JSONâ; 30â40 examples fixed it.
đ§ The Curbside
âCan I finally point a foundation model at my registry spreadsheet?â
Short answer: Closer than you think, and this is the model class nobody in health tech is talking about.
What changed: Not a single release â a category that quietly matured. Tabular foundation models treat table prediction as in-context learning: no training run, no ML engineer. TabFM from Google Research is open and headed into BigQuery; TabPFN (Prior Labs, Nature-published) ships an API and an open package; TabICL came out of Inria. None of these landed this week â Iâm flagging the category, not an announcement.
Builder read / Watchout: Everything you actually have is a table â the flowsheet export, the claims extract, the registry CSV, the QI pull. Prototype and cold-start are where these win. They do not win at millisecond-latency production scoring, and nothing about zero-shot removes your obligation to check calibration on your own population. Also: the hosted ones are cloud services. Synthetic or de-identified data only until someone shows you a BAA.
đ¤ âSo now every department gets to build its own unvalidated risk model.â Yes. That was already true â the difference is it used to take six weeks of a data scientistâs time and now it takes an afternoon, which means the governance conversation you were deferring is due. The tool didnât create the problem. It removed your excuse for not having an answer.
âFDA just codified a new device class for a diabetes app. Does that open a lane for me?â
Short answer: It opens the regulatory lane. It does nothing for the payment lane, and those are two different pipes.
What changed: On July 24 FDA published the final order codifying âdiabetes digital behavioral therapeutic deviceâ into Class II at 21 CFR 880.5735 â prescription software, special controls, clinical data required, four mandatory limiting statements in the labeling. Be precise about the date: the classification itself has been applicable since July 7, 2023 (Better Therapeuticsâ BT-001 De Novo). What happened Friday is codification into the CFR â the tidy, citable predicate lane, spelled out in regulation text you can read in ten minutes.
Builder read / Watchout: Read the special controls before you read anything else â âthe device is not intended for use as a standalone therapyâ and âshould not be used by people with unstable psychiatric disordersâ are label constraints that shape your product, not footnotes. And separate covered from paid in your deck today: CMSâs FY2027 IPPS proposed rule (CMS-1849-P, published April 14) would repeal the Breakthrough-Device fast lane to new-technology add-on payment starting FY2028; comments closed June 9, and IPPS final rules customarily land in late July or early August. Coverage is permission to bill. The add-on is the money.
đŹ The Big Thing
The VA didnât buy software. It bought a license to agents â and nobody has priced that SKU before.
On Friday the Department of Veterans Affairs awarded Salesforce up to $1.6 billion under what the announcement calls an Agentic Enterprise License Agreement â one base year with two one-year options.
Agents go into the round-the-clock contact center, care coordination, benefits verification, and scheduling.
The number that matters isnât the $1.6B. Itâs this: the stated goal is to bring average appointment scheduling from 28 days to minutes.
Context worth holding. This is the largest integrated health system in the United States â more than 9 million veterans enrolled in its health care services, across 170 medical centers and more than 1,100 outpatient clinics, consolidating several IT contracts into one platform relationship with a CRM company.
Three days earlier, the same reporting notes, the White House folded the VA into the Genesis Mission â with the VAâs electronic health records and Million Veteran Program genomic data going toward the Energy Departmentâs model training.
The contract vehicle is the product. I canât find a prior award under the label âAgentic Enterprise License Agreementâ â and now it has a $1.6 billion comparable.
That matters more to you than the vendor name does. Once agents are a licensed line item rather than a feature inside a clinical application, procurement stops asking âdoes this tool workâ and starts asking âhow many agents, at what tier.â
The clinical question doesnât disappear. It just stops being a purchasing question.
Which is exactly the opening. Nothing in a license agreement specifies what a triage-routing agent should do when the callerâs chief complaint is chest pain and the scheduling logic says next Tuesday.
Thatâs not in the statement of work. Itâs in the head of someone who has been on the other end of that call.
đ¤ âA CRM company running care coordination for nine million veterans. What could go wrong.â Be careful which objection youâre actually making. The scheduling contact center genuinely is a customer-service problem, and Salesforce is not obviously the wrong vendor for it. The place to be nervous is the seam where routing decisions start carrying clinical weight and nobody redrew the line.
đ¤ ââHIPAA-readyâ isnât a certification. Itâs an adjective.â Correct â and itâs the announcementâs own word. Ask what the BAA covers instead.
đ¤ â28 days to minutes is a press release number and you know it.â I do. Every one of these is. But itâs a falsifiable press release number, which is more than most vendors give you, and the VA publishes wait-time data. Twelve months from now somebody can check. Write the date down.
đĄ 80/20: Go read the chi-bench and write the paragraph nobody in that VA contract wrote: for one workflow you actually know, name the single failure the agent is not allowed to make, and how you would detect it from the outside. That paragraph is worth more than the architecture diagram.
đ§Ş Try the interactives:
A â Thirteen Lanes, Fifteen Cars â FDA has opened 13 digital-behavioral-therapeutic device lanes; only 15 follow-on 510(k)s have ever driven through them, and 5 lanes â including the diabetes lane codified Friday â are empty. Built with real FDA data.
B â The Predicate Lane Economy â every FDA 510(k) product-code lane with meaningful traffic since 2016, 846 lanes and 18,335 clearances, as one brushable scatter of traffic vs. FDA review speed, with the 13 digital-therapeutic lanes highlighted. Built with real FDA data.
đĄ Builderâs Radar
Johns Hopkins wonât scale an administrative agent until it can measure five things
Health system leaders at Hopkins laid out their pre-deployment process for agentic AI in administrative operations, led by Dr. T.Y. Alvin Liu, inaugural director of the Gills AI Innovation Center.
The five measures, in their words: first-pass completion versus repeated completion rates, consistency during handoffs, reasoning failures, tool-use failures and unsafe completions. Theyâre evaluating and governing through a platform called actAVA, and theyâre explicit that reliability comes before anyone gets to talk about financial return.
đĄ 80/20: Copy those five categories verbatim into the eval spec for whatever youâre building. âConsistency during handoffsâ is the one most solo builders skip, and itâs the one that breaks first when your agent hands work to a human and takes it back.
âShe bought her own data. For two years, nothing could receive it.â
Adam Carewe, MD describes a patient who arrived with ninety days of continuous glucose readings, a year of overnight heart-rate-variability and temperature trends from her ring, and a lab panel she ordered and paid for herself â and no ingestion path except sharing her screen. This had been the shape of every visit for two years.
Nine of the fifteen available minutes went to manually reconstructing what she already owned.
Careweâs argument is that ingestion is now the commodity and the moat has moved to specifying the clinical output â and that AI synthesis clean enough to look finished is exactly what invites a clinician to rubber-stamp it.
đ¤ âThis is a workflow problem, not a technology problem.â Theyâre the same problem when the workflow is fifteen minutes long. A tool that saves nine of them is not a convenience feature, itâs the whole visit.
Payment integrity is a $750M-revenue, 38%-margin business, and itâs for sale
New Mountain Capital is shopping Machinify â the payment-integrity platform assembled from Rawlings, Apixioâs payment integrity unit, and VARIS. A JPMorgan deck reportedly shows roughly $750M in revenue, 21% CAGR, and about 38% EBITDA margins.
Those are software margins on the business of deciding whether your claim gets paid.
đŽ My bet: the buyer is a payer or a payer-adjacent platform, not a PE secondary â and within two years âpayment integrityâ and âprior authorization AIâ stop being sold as separate products, because theyâre the same model pointed at two ends of the same claim.
Where the agent stops: 98% resolution turned out to be 39 of 52 tickets
Adjacent possible. Nate Jones opens with a support-agent success story and immediately guts it â the 98% ticket-resolution figure was 39 of 52 tickets, all on one repeated issue. Then the real case: a Gumroad support agent autonomously diagnosed and fixed a charting bug in production code, and got a design detail wrong that only a human reviewer caught.
Translate that directly. Your AI intake-triage dashboard is going to look spectacular, because three-quarters of your volume is one problem.
đ¤ âSupport tickets arenât patients.â No. Thatâs the point â go look at your denominator before you quote your number to a CMIO.
Quick hit
Missing data is a modeling choice, not an accident. A new npj Digital Medicine paper predicts major adverse cardiovascular events from incomplete clinical data using interpretable multimodal AI.
Almost every published clinical model assumes a complete feature vector. Thatâs why almost every published clinical model breaks on a real EHR extract.
đď¸ From the Pods
đď¸ Lennyâs Podcast â âAnthropicâs first technical PM on token maxing, the jagged edge, and living in the future | Dianne Pennâ
Penn joined as the first technical PM when there were five product engineers and one engineer covering the entire API business, and sheâs shipped every model since. Her operating claim: the workflow is no longer write-a-PRD-then-build, itâs read the failed transcripts, classify the failure mode, and encode it as an eval. Her example â early-days âClaude was not good at following instructionsâ feedback decomposed until roughly 80% of it turned out to be âClaude would not write the right JSONâ; 30â40 examples became the eval set, and it now runs âalways 100% or like 99.9.â
âHere, you have to sweat the tokens as much as you sweat the pixels.â
đĄ Builder take: âThe scribe hallucinatedâ is not a bug report. Decompose it until you have a reproducible case with a right answer, and youâll usually find one failure mode is most of your volume.
đ Speaker Blindspot: Motte-and-bailey on evals. The motte is âmeasure what you ship.â The bailey is âevals replace product requirementsâ â and the JSON example only works because correctness is decidable; thereâs a golden answer. She never addresses domains with no golden answer, which is precisely clinical judgment. Secondary: survivorship bias. Every mechanism she credits is narrated backward from the wins, and when she says âthereâs a lot of bets that we end up turning down or turning off,â she names zero of them.
đ¤ Selling It
The Reference Customer Ask â building the answer to âwho else is using this?â before youâre asked
The clinical pitch went well, governance is solid, pricing is forecastable, and then the CFO asks the five words that kill more deals than any objection: who else is using this?
A customer is someone who bought your product. A reference customer is someone who will take a fifteen-minute call from a CFO theyâve never met and say we deployed this, hereâs the number, Iâd do it again. CFOs know the difference.
The stack has four layers, and you build it before you need it:
A peer system â same size, same payer mix, same EHR footprint.
Outcome data from their own numbers, not their enthusiasm.
A named executive who has already agreed to take the call. Ask at 90 days post-go-live, while theyâre still happy and the result is fresh â not in the moment when you need it.
A failure story. The CFO who hears only wins doesnât believe any of them.
The version of this that works looks like a peer CMIO with a matching payer mix and EHR footprint, on the phone inside 24 hours, saying one specific number about excess days. That call closes it â not the pitch, not the ROI slide.
đ
Upcoming Health IT Events
Mostly free, mostly virtual, mostly recorded â attend live or catch the transcript later.
Wed Jul 29 â Preparing Health Plans for the 2027 CMS Prior Authorization Rule (AHIP / EviCore)
2:00 PM ET ¡ Virtual ¡ Free
Walks through the CMS PA rule and the API automation stack (CRD, DTR, PAS) â a preview of the plumbing clinician-builders will be plugging into.
Thu Aug 6 â EHIgnite Challenge Phase 1 Winners Showcase (ASTP/ONC)
3:00 PM ET ¡ Virtual ¡ Free
Nine winning teams demo AI that turns dense EHI exports into plain-language, actionable info for patients and clinicians â the closest thing to a live build-in-public session this week.
Tue Aug 11 â Application Rationalization: From Infrastructure Clarity to Portfolio Control (CHIME)
12:00 PM ET ¡ Virtual ¡ Free
CIO-level playbook for clearing technical debt as a precondition for AI adoption â good background if youâre pitching a clinical-AI tool into a health systemâs existing stack.
Thu Aug 13 â Building the Consent Trust Stack (HL7 FAST, SHIFT & The Sequoia Project)
11:30 AM ET ¡ Virtual ¡ Free
Launches a scalable consent management implementation guide â worth a look if your build touches patient-directed data sharing or granular consent.
đĄ BTW: T.Y. Alvin Liu â the Hopkins physician now writing the reliability tests agents have to pass â is a vitreoretinal surgeon, and since 2020 his institution has been running FDA-approved autonomous AI in primary care offices: diabetic retinopathy screening that returns a result with no physician reading the image. The person insisting on proof before deployment already helped ship the machine that doesnât ask.
What are you building this week? Email and tell me (kevin@clinicians.build) â I read every one.
â Kevin & AI (GLM5.2/Sonnet 5/Opus 5 - Fabel too expensive)


