AI-generated code needs an AI-powered reviewer â Redesign Health built an internal system around one observation: AI makes telltale mistakes, not random ones. (The Big Thing, below.)
Microsoft announced Project Zenith â Windows for 64GB+ developer machines, built to run 30B+ parameter models locally with no usage metering. The word that matters is metering.
UW Medicine is going back to a base Epic build â customer number 30, live since 1996, stripping thirty years of customization to adopt new features faster.
đ¤ Haters: âThatâs a budget cut wearing a strategy costume.â Sure. Itâs also the reason that happens to be true.
đ§ The Curbside
âShould I fine-tune a small model instead of paying frontier prices?â
Short answer: Only once the workload is boring â high volume, bounded, measurable.
Evidence: Ben Dicksonâs weekend breakdown of Shopifyâs pipeline: a 0.8B specialist beat a frontier model on one narrow task, cutting serving cost from ~$27M/year to ~$1M/year. The hard part wasnât training. It was building the judge.
Builder read: Shopifyâs production data held no examples of a correct refusal, because refusals never looked like successes. Yours wonât either.
đŽ My bet: the first clinical workloads distilled this way arenât diagnostic â theyâre de-identification, chart summarization and prior-auth drafting, inside health systems, before any vendor sells it.
đŹ The Big Thing
Your AI wrote the code. Who reads it back?
Aron Szanto, head of technology at Redesign Health, spent Fridayâs HIMSSCast on the part of the vibe-coding story health tech coverage keeps skipping: what happens after the model writes the software.
His team built an internal reviewer called Argus, aimed at AI-produced code â code that can look completely functional while carrying problems nobody spots.
Then he said the thing I canât stop turning over. âWhen an AI makes a mistake in coding, theyâre not uniformly random ... They make telltale mistakes.â
Patterned failure is a gift. Random failure you can only catch by reading everything. Patterned failure you can build a detector for.
That distinction is load-bearing if youâre a clinician who builds things.
Every argument against clinician-built software ends in the same place: youâre not an engineer, your code will be unsafe, no security review will pass it. Fair, when the only alternative was a senior reviewer you couldnât afford and couldnât recruit.
Much weaker when the reviewer is a specialist that runs on every commit and already knows the seventeen ways an LLM gets authorization wrong.
None of which makes the regulatory surface disappear. A BAA is still a contract, SOC 2 is still an audit with a human on the other end, and no reviewer catches a design decision that should never have been made. But annual and continuous are different regimes, and continuous is the one agentic review makes cheap. On that single axis, a clinician-built tool with a harness on every commit is in better shape than the enterprise vendor with a pen test each October.
đ¤ âRedesign Health has an interest in saying this.â Theyâre a venture studio, not a code-review vendor. Argus is internal tooling built because the bugs were real. Building for your own portfolio is a better signal than launching a product, not a worse one.
đ¤ âYou still canât read what it wrote.â Neither can the staff engineer at whatever vendor your system just signed, at the volume theyâre shipping. Pick your unread code.
â If coding failures are patterned and enumerable, the same should hold for clinical failures â per model, per task, per specialty. Hallucinated citations. Anchoring on the chief complaint. Silent unit conversions. Nobody is shipping the Argus for clinical output, and I suspect the taxonomy is the product rather than the detector. I canât see the shape of it yet.
đ§Ş Try the interactive: Telltale, Not Random â two views of the same idea, built with real FDA MAUDE adverse-event data.
A â Telltale, Not Random â every device class the FDA heard about in 2024, placed in a triangle by how it fails: malfunction, injury, death. The points pile into the corners. Press one button and see where they would sit if failure were random.
B â The Shape of a Failure â 560 device classes and 2.47 million adverse event reports plotted by volume against harm. 61% of classes are more than 90% a single event type. Random would be 2%.
đĄ Builderâs Radar
A Windows box that runs a 30-billion-parameter model and never phones home
Microsoft announced Project Zenith Friday: a preconfigured Windows developer setup for machines with 64GB+ unified memory and 250GB/s+ bandwidth, aimed at running 30B+ parameter models locally without usage metering. Lenovoâs ThinkCentre X Ultra starts at $3,699 in November.
The interesting number isnât the price. Itâs the metering.
Everything you canât put on a metered API becomes possible on a machine in a room you control.
đĄ 80/20: If youâve been running Ollama locally and hitting the wall around 14B, this is the hardware class that unblocks you.
Epicâs 30th customer is ripping out its own customizations
UW Medicine CFO Jon Alford â customer 30, first go-live 1996 â says the system is heading back to a base Epic platform. âOrganizations need to get back to a base platform to be able to implement new technology faster as it starts to come at us.â
Thirty years of legacy build is thirty years of somebodyâs very good idea, each one now a reason a new feature wonât turn on.
đŽ Where this lands: by the end of 2027, âAI readinessâ is the stated rationale for de-customization at most large Epic shops. The unstated one is that nobody can staff the legacy build.
đď¸ From the Pods
đď¸ 229 â âChatGPT Did Not Come Through Epicâs Front Doorâ
Bill Russell takes the ChatGPT/UCSF integration apart: read-only, the USCDI subset rather than the full record, through the standard Cures Act FHIR path â not a privileged Epic partnership. OpenAIâs own terms say not to use it for diagnosis, which is the exact thing everyone is excited about.
đĄ Builder take: âConnected to Epicâ now means at least four different things. Ask which one before you build a roadmap on a press release.
đ Speaker Blindspot: Composition fallacy â arguing this changes nothing because systems âalready have this through our ambient vendorâ treats an enterprise contract as if it were a capability. That access is a business arrangement between two companies. It isnât something the physician on shift can invoke, or a clinician-builder can build against.
đď¸ HLTH â âHow Health Plans Can Close the Gap Between AI Ambition and Data Realityâ
A survey of 103 health plan executives on the distance between AI ambition and provider-data quality. The line that landed: âAI doesnât fix a bad foundation. It actually scales it.â A manual error costs one bad phone call; the same error learned by a model repeats at machine speed across adequacy filings, credentialing, member search and contracting.
đĄ Builder take: Before you price the model, price a wrong record. If you canât put a number on what one bad row costs downstream, your ROI math runs backwards.
đ Speaker Blindspot: Sponsored-survey framing â the research was commissioned by a provider-data company, and the finding is that provider data is the bottleneck. That doesnât make it false. It does mean the ranking of causes belongs to the sponsor.
đ§° Builderâs Tip
Tool spotlight: query the FDAâs device failure database before you design around a device.
openFDA exposes the MAUDE adverse event database as a free JSON API. No key for light use, no PHI, no BAA, thirty seconds from your laptop:
curl "https://api.fda.gov/device/event.json?search=device.generic_name:\"infusion+pump\"&count=event_type.exact"
That returns the distribution of event types â malfunction, injury, death â for a device class. Swap in pulse+oximeter, ventilator, glucose+monitor, or whatever your tool takes a feed from.
MAUDEâs failure taxonomy is the closest public thing to a spec for what your error handling has to survive. We design the happy path because the unhappy path is invisible. This makes it visible, free.
đ Upcoming: Using AI Intelligently to Solve Pressing Healthcare Challenges (AHIP/Engagys), Tue Sep 8, 2 PM ET, free.
đĄ BTW
đĄ BTW: Aron Szanto, who built the AI code reviewer in todayâs Big Thing, studied computational economics and mathematical philosophy at Harvard. The person who decided AI errors are enumerable came out of the field that asks what a formal system can and cannot prove about itself. Redesign Health
You have a unique combination of skills, experience and values. So do great things! ⌠and tell me about them at kevin@clinicians.build.
â Kevin & AI
(please verify content for yourself, partially AI generated and may contain errors)



