Blue Cross says hospital AI coding added $942M without changing care. About 70% of it came from secondary diagnoses, and many of those can be pulled from a single lab value. More below.
Basalt Health raised $20M to replace the 70-page post-acute fax. At Lifepoint, median referral processing fell from 8.5 minutes to 1.2. Itâs heading to 111 markets by year end.
Oracle put ID.me identity proofing into its EHR. Controlled-substance e-prescribing credentialing is live now. Patient intake is still ârolling out.â
North Carolina launched a shared AI evaluation network for rural hospitals. UNC and Duke will share âframeworks, vendor evaluations, lessons learned.â The funding is $4.4M over three years.
đ§ Podcast: The 229 â Derek De Young on Epic Agent Factory. Epicâs Agent Factory lead says itâs âstill in what weâre calling the development preview.â Early adopters open at the end of October, and general availability is a Q1/Q2 goal.
đ§ The Curbside
âOpenAI just released MentalHealthBench. Can I use it to test our behavioral health bot?â
Short answer: Yes, as a floor. Donât treat it as a pass.
What changed: Itâs an open benchmark. More than 80 licensed clinicians wrote the rubrics, and it covers adults, teens, caregivers and clinicians at three acuity levels, up to emergencies. The conversations are synthetic, so you donât need PHI to run it.
Watchout: The benchmark came from a company that sells a chatbot, and it grades the kind of conversation that companyâs product has. đ¤ Haters: âSo the vendor wrote the exam.â Yes. Thatâs why itâs useful as a floor and not as proof. Your own ugliest ten conversations are still the real test.
đŹ The Big Thing
If the code went up and the transfusion didnât, what does the chart actually prove?
The Blue Cross Blue Shield Association published its analysis Thursday. From early 2023 to late 2025, the share of inpatient claims coded as medically complex rose from 37% to 40%, with no matching change in patient acuity.
The breakdown works out to $942 million, and $653 million of it came from 55,000+ extra secondary-diagnosis cases. In major bowel procedures, the share coded at the highest complexity level rose from 10.2% to 22.7%.
The day before, CMS Administrator Mehmet Oz told Oracleâs summit that âshort term, AI is going to be inflationary.â
The payer doesnât have to prove the code was wrong. It only has to show that the treatment didnât move with the code.
Think about a hemoglobin of 9.8 on post-op day two. Thatâs a lab value. Whether itâs âacute blood loss anemiaâ is a clinical judgment. AI is very good at finding the lab value. It canât tell you whether anyone thought about it.
Thatâs the whole fight. Blue Crossâs example is more anemia diagnoses without more transfusions. BCBSAâs Razia Hashmi, MD, put it bluntly:
âIf it was worth coding, there should have been something done.â â Razia Hashmi, MD, BCBSA
Hospitals tell the mirror-image story. HCAâs CFO says hospitals are âbehind the payers.â NYC Health + Hospitalsâ revenue chief calls it a ârock âem, sock âem robotâ fight that nobody wins.
đ¤ âDocumenting a real diagnosis isnât upcoding.â True. Even BCBSAâs Luke Chalker calls some of it âthe stuff that hospitals should bill for.â The fight is over conditions that nobody treated. The methodology also has a limit: itâs claims-only, and BCBSA says so.
đ¤ âPayers use AI to deny. Why canât hospitals use it to code?â They can. Both sides are running the same arms race.
đ¤ âThis is a payer PR document.â It is. Itâs also a preview of the audit letter.
â Every one of these disputes turns on one question: did the diagnosis change what the team did? I think thereâs a product in linking each coded condition to the order, med or note that shows it was acted on. It wouldnât be a coding tool or a CDI tool. It would be a âtreatment-concordanceâ layer. Would anyone buy that before the first payer clawback, or only after?
đĄ Builderâs Radar
The payerâs rulebook is becoming an API
Ensemble Health Partners is embedding Penelope Healthâs policy engine into its revenue cycle platform. (Both are Thoreau companies: Thoreau agreed to buy Ensemble in June and put $100M into Penelope on Sept. 18.) Penelope turns payer coverage rules into structured data: 15,000+ procedure and drug codes, and policies covering 200M+ insured Americans.
The point is to check medical necessity before the encounter, not after the denial.
Once the rules are machine-readable, the fight moves upstream, to the moment the order gets placed.
SNOMED has 350,000 concepts. Your chest pain still maps to R07.9.
John Lee, MD, argues that coded terminologies squeeze out clinical meaning. His example is a patient with known coronary disease whose new burning chest pain gets the same âunspecifiedâ code as the vague discomfort she had before. LLMs can read the nuance the code throws away.
Read that next to the Big Thing. A machine that reads nuance well will also find more to code.
đŽ Where this lands: The next fight wonât be over structured versus unstructured data. Itâll be over who gets to cite the free text in a dispute, and by 2027 payers will be asking for the note, not just the code.
North Carolina is trying to build a shared evaluation bench
NC CHAIN links UNC Health, Duke Health, the Sheps Center and Duke-Margolis to help rural and critical access hospitals choose and govern AI tools.
$4.4M over three years works out to about $1.5M a year for a whole state. Thatâs a few staff, not an evaluation lab.
đ¤ âItâll produce a framework and a webinar series.â It might. The test is whether a published vendor evaluation ever comes out of it. If one does, itâll be the most-read document in rural health IT.
⥠Quick hits
Anthropic and OpenEvidence are offering free clinical AI in about 100 countries. Itâs a localized version of OpenEvidence for clinicians in low- and middle-income countries, including Uganda, Haiti and Mongolia. The US product doesnât change, and neither company disclosed financial terms.
Memorial Health in Savannah is holding off on about 50 Epic AI features and plans to turn them on a few at a time, with close monitoring. Having a feature available isnât the same as deciding to use it.
đď¸ From the Pods
đď¸ Stanford AIMI Grand Rounds â Maya Yiadom, MD, âAI in the Loopâ
Yiadomâs move is to treat the modelâs threshold like shared decision-making, except the health system is the patient. Simulate the options on site data, show the trade-offs, then have governance sign off on the parameter, the acceptable and unacceptable trade-offs, and the conditions that trigger a reassessment. In her words, âitâs not just the algorithm alone.â
đĄ Builder take: Write the threshold authorization before you write the alert.
đ Speaker Blindspot: Generalizing from a resource-rich site. Her governance table has about ten roles in it, from ECG techs to health-equity experts. She allows it could be just a few stakeholders and says automating governance lowers the bar, but thatâs still a bench a critical access hospital doesnât have, which is why NC CHAIN exists.
đď¸ The Nocturnists â âThe Lost Auraâ with John Lantos, MD
Lantos gave a chatbot his own cardiac data and asked for the pessimistic take, then the optimistic one. His point is that patients can now âdesign our own doctorsâ and get as many opinions as they want.
đĄ Builder take: A âshow me the other readâ button may do more for shared decision-making than another summary.
đ Speaker Blindspot: An illusion of independence. Two framings from the same model on the same data arenât two opinions. Theyâre one opinion wearing two outfits.
đĄ BTW: John Lantos was asking this question almost thirty years before ChatGPT. His 1997 book was titled Do We Still Need Doctors?
You have a unique combination of skills, experience and values. So do great things! ⌠and tell me about them at kevin@clinicians.build.
â Kevin & AI
(please verify content for yourself, partially AI generated and may contain errors)


