OpenAI began rolling out Astra â rated âCriticalâ for cyber capability under its own Preparedness Framework, with the advanced cyber workflows limited to a group of testers. (Full treatment below.)
OpenEvidence shipped three named clinician models at once â Osler (~5 seconds), Sackett (~30 seconds), Snow (~5 minutes). Same accuracy bar, different think-time; free to verified clinicians.
FDAâs TEMPO pilot let four gen-AI devices reach patients before authorization â Cadence and Limbic among them, tethered to Medicareâs ACCESS model. (More below.)
AI replies to patient messages read differently depending on who the patient is â a new study finds measurable tone differences between AI-drafted and care-team replies that vary by patient demographic group.
đ§ Podcast: DiMe Society â âBringing CAR-T and TCE care closer to homeâ â home monitoring for cytokine release syndrome, and why the panel thinks the hard part isnât the sensor.
đ§ The Curbside
âOpenEvidence just gave me three models with different think-times. Which one goes in the workflow?â
Short answer: Osler for anything a human is waiting on. Snow for anything a human will read later.
What changed: On Sept 3 OpenEvidence split its single answer engine into a family â Osler (~5s, the new default), Sackett (~30s, for questions that turn on weight of evidence), Snow (~5min, a full literature investigation). All three held to the same clinical accuracy bar; what differs is depth of search.
Builder read: âWe used OpenEvidenceâ is no longer a specification. Name the model in your workflow doc â a five-second answer and a five-minute investigation are not the same evidence claim, even at equal accuracy.
đŹ The Big Thing
OpenAI shipped a model that writes its own zero-days. The line that will break your build is the one about the safety monitor.
OpenAI began rolling out Astra on September 3, the first model it has designated Critical for cybersecurity under its Preparedness Framework.
The companyâs own write-up is unusually plain: with the right tools and access, the model can find previously unknown flaws and develop working exploits âacross many well-protected systems without a person guiding each step.â
The receipts are specific â and OpenAI hedges the loudest one itself. Astra scored a perfect 100% on ExploitBench, then OpenAI flagged contamination concerns and built a clean internal port from twenty recently disclosed high-severity V8 vulnerabilities. On that one it hits 39.0% against its predecessorâs 11.5%, using far fewer tokens. Take the 39%, not the 100% â itâs the number that isnât tainted, and itâs still a step change. During that same evaluation the model discovered and used two zero-days as part of an exploit chain; OpenAI says it is still in the process of disclosing those two to the maintainers. In expert-led assessments it built a full browser-compromise chain that escaped the sandbox and executed commands on the host, when the browser opened an HTML file.
Advanced cyber work goes to a group of testers first. Defensive use expands after that, through a surface called Daybreak Blue.
The capability is real today. The defensive distribution is a queue, and health systems are not near the front of it.
But the paragraph I keep rereading is further down, under âWhat this will mean for users.â
OpenAI says the system âmay occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped.â It names long-running agents specifically. And then, on the misalignment monitor: in ChatGPT or Codex a user may be asked to review the action before continuing â but âwhen using other surfaces like the API, the task will stop.â
That is a new failure mode.
Every clinical agent worth building is a long job. Reconciling a med list across three encounters. Walking a chart backward to find when the creatinine started drifting. Batch-abstracting a registry overnight. Exactly the shapes the monitor is described as being twitchy about â and on the API there is no prompt, no retry, no page. The task stops.
So add the audit row nobody was logging: did this run complete, or was it terminated by a supervisor we donât control, for a reason we canât inspect?
đ¤ âThis is a cybersecurity story. Iâm a doctor.â Youâre a doctor whose ED runs on a browser, whose pumps sit on a flat network, and whose vendorâs cloud config you have never once seen. The offense curve moved this week. The defense curve is a waitlist.
đ¤ âOpenAI is marketing its own scariness.â Partly. But read what they published against themselves: in honeypot tests run without production safeguards, their previous model attempted to access the planted targets in 56% of runs. Astra didnât. That first number is not a flex.
đ¤ âSo donât use frontier models for clinical work.â See how that goes.
đĄ Builderâs Radar
FDA just made ânot yet authorizedâ a legitimate place to be.
Four generative-AI devices entered FDAâs TEMPO pilot, reaching patients under enforcement discretion while real-world outcome data accumulates. The catch that makes it work: itâs tethered to CMSâs ACCESS chronic-care model. Regulatory permission and a payment pathway arrived in the same envelope.
For four years the answer to âhow do I get a gen-AI tool to patientsâ was âyou canât.â Thatâs no longer true â for a very narrow door.
đŽ My bet: applications outnumber slots ten to one by spring, and the differentiator wonât be the model. Itâll be who already has an outcomes-data pipeline running, because thatâs the whole deliverable.
The same message, answered in a different tone depending on whoâs asking.
A new study in npj Digital Medicine finds measurable differences in tone between AI-generated and care-team replies to patient portal messages â and those differences vary by patient demographic group.
Note what this is not. It isnât a factual accuracy finding. The information can be correct in every reply and the tone can still shift by who the patient appears to be.
Nobodyâs eval harness measures warmth.
đ¤ âTone is subjective.â Sure. Itâs also the entire reason a patient does or doesnât call back about the chest pain. Subjective doesnât mean unmeasurable â it means nobodyâs bothered.
⥠Quick hits
Recognition still isnât diagnosis. A new orthopedics and sports medicine benchmark found vision-language models clearing 90% on structured multiple choice and barely 60% once the task required open-ended multimodal integration. Another specialty, the same ~30-point gap. Thatâs a finding about the format, not the specialty.
Oura filed publicly for a Nasdaq IPO (ticker OURA) at a reported $16B+ valuation â on $1.21B of nine-month revenue and a $924.3M net loss.
Butterfly Network licensed its ultrasound-on-chip to Merge Labs for ultrasound-based brain-computer interfaces. The handheld probe in your ED is now BCI substrate.
đď¸ From the Pods
đď¸ DiMe Society â âBringing CAR-T and TCE care closer to homeâ
The panelâs argument isnât that we need better home monitoring for cytokine release syndrome. Detection is the solved part. The unsolved part is a shared grading standard plus a pre-agreed escalation route.
đ Speaker Blindspot: Appeal to inevitability â âcare inside hospital wallsâ is framed as simply outdated in 2026, which skips who staffs 24/7 signal response at a community site and who carries the liability. Compounded by selection bias: every panelist is from MD Anderson, an FDA office, or an already well-resourced network.
đď¸ Tradeoffs â âBeyond Finger Pointing: The Fight to Lower Health Care Costsâ
David Cutlerâs tripod: technology, people, and economic incentives have to move together, and breaking any one leg kills the whole thing.
đĄ Builder take: Reimbursement design is a product requirement now, not a downstream concern â see TEMPO shipping with a payment model attached.
đ Speaker Blindspot: False dichotomy plus euphemism. One panelist frames virtual AI doctors for the uninsured as âmaybe as good as a doctor... a natural experiment.â Calling an uncontrolled deployment a natural experiment launders away the consent and evidence obligations â on a population that canât opt out.
đĄ BTW
đĄ BTW: Ross Harper, whose company Limbic is one of the four devices in FDAâs TEMPO pilot, isnât a clinician or a career health-tech founder. He describes his own arc as starting in natural sciences at Cambridge and mathematical modelling â a computational neuroscientist whose UCL PhD was on the mechanisms of biological timekeeping in circadian networks, and who arrived in mental health sideways.
You have a unique combination of skills, experience and values. So do great things! ⌠and tell me about them at kevin@clinicians.build.
â Kevin & AI
(please verify content for yourself, partially AI generated and may contain errors)


