Video Preview
🎧 Part I Podcast free on Apple Podcasts and Spotify.
🎧 Part II Podcast episode for paid subscribers only. Also available on Apple Podcasts and Spotify.
To listen to paid episodes in Apple or Spotify, link your Substack subscription via the show settings on those platforms (instructions inside the Substack app under Subscriptions → Podcast).
Table of Contents
What actually shipped on September first
The plumbing, or how patient data probably gets from Epic into a chatbot
Epic has been playing this game for years, mostly with OpenAI’s own models
About that 99.1 percent
The startup blast radius
The governance hangover
What to watch between now and the next HIMSS
Abstract
OpenAI announced an Epic integration for ChatGPT for Healthcare, letting clinicians pull authorized patient context into ChatGPT and, in some deployments, embed ChatGPT inside the EHR layout itself
A new Healthcare Public Data plugin connects nine official sources including PubMed, ClinicalTrials.gov, DailyMed, RxNorm, and CMS Coverage as structured, queryable connectors rather than generic web search
OpenAI is leaning hard on physician eval numbers: 700k plus reviewed responses, 27 clinical use cases, 4,363 ratings, 99.1 percent rated safe, and 93 percent plus accuracy on connected data sources
Launch partners include UCSF, HCA, Cedars-Sinai, Memorial Sloan Kettering, Boston Children’s, Baylor Scott and White, and AdventHealth, which is a serious roster for a v1
The essay covers the likely technical architecture, Epic’s own AI position and its complicated Microsoft relationship, what the evals do and do not prove, which startups should be nervous, and why the real work is governance, not integration
What actually shipped on September first
OpenAI dropped a healthcare announcement the day after Labor Day, which is either a scheduling accident or a very deliberate way to ruin the first week back for every CMIO in the country. Two things shipped. First, an Epic integration for ChatGPT for Healthcare that brings authorized patient context out of the chart and into the chat window. A clinician can ask what changed since the last visit, which labs need review before this afternoon, whether meds got touched, whether a specialist wrote something important that got buried on page forty of a scanned fax. ChatGPT assembles the answer from the record and points back to the supporting chart data. Second, a Healthcare Public Data plugin that gives structured access to nine official sources, including ClinicalTrials.gov, CMS Coverage, RxNorm, DailyMed, and PubMed. Not search over these sources. Actual connectors, with the ability to work against specific records, identifiers, fields, and policy versions.
The Epic piece comes in two flavors. One is EHR context flowing into the ChatGPT workspace, so a clinician preps for clinic in ChatGPT instead of clicking through Chart Review tabs. The other is ChatGPT embedded directly into an EHR layout in supported deployments, meaning the model shows up inside the workflow instead of asking the clinician to leave it. That second flavor is the one that matters commercially, because the entire history of clinical software teaches one lesson over and over: whatever lives inside the Epic frame wins, and whatever requires a second browser tab dies a slow death measured in abandoned pilots.
Everything sits inside ChatGPT for Healthcare, a governed workspace with role based access, SSO, audit logs, and an applicable BAA, which makes OpenAI a business associate under HIPAA for these deployments. Individual clinicians on the eligible clinician tier can get the public data plugin, but the EHR integration is enterprise only, which is the correct call and also conveniently a great enterprise sales motion. Launch partners include UCSF, HCA, Cedars-Sinai, Memorial Sloan Kettering, Boston Children’s, Baylor Scott and White, and AdventHealth. That is not a list of gullible pilot shops. HCA alone runs around 190 hospitals and one of the most industrialized IT operations in the industry. When they show up in the launch graphic, procurement conversations at every other system get shorter.
The plumbing, or how patient data probably gets from Epic into a chatbot
The announcement is light on architecture, which is normal for a launch post and maddening for anyone who has ever actually stood up an Epic integration. So some informed reading between the lines. The phrase authorized patient context is doing heavy lifting. The plausible path is the standard one: SMART on FHIR app launch for the embedded experience, FHIR R4 reads against the Epic instance for chart data, scoped to the patient in context and the user’s security class, with the health system controlling which resources are exposed. Epic’s APIs cover the USCDI data classes plus a long tail of Epic specific endpoints, and any vendor who has been through the open.epic and Vendor Services gauntlet knows the difference between what the brochure says is available and what your specific customer’s version, build, and security team will actually let you touch. The embedded layout language maps to Epic’s web embedding patterns, where an outside app rides inside Hyperspace or Hyperdrive with a launch token carrying user and patient context. None of this is exotic. What is new is who is on the other end of the pipe.
The parts the post does not answer are the parts every security review will ask in the first ten minutes. Is patient data retained by OpenAI, and for how long, and is any of it eligible for model training under any configuration, and what does the audit log actually capture, prompt text included or just metadata. Zero data retention is table stakes language in healthcare AI contracting now, and the post says nothing either way, which means the answer lives in the enterprise agreement, which means it is negotiable, which means everyone should negotiate it. Minimum necessary is another fun one. HIPAA’s minimum necessary standard was written for humans requesting records, and it gets philosophically weird when the whole value proposition is a model that reads broadly across the chart so it can tell you what matters. A model that pulls the entire longitudinal record to answer whether the potassium is trending is doing something a records clerk would get written up for. Nobody has a clean answer here, and OCR guidance is not exactly sprinting to provide one.
Then there is the question of whose permission this all runs on. Epic does not have a public quote in the announcement, which is worth noticing. The integration presumably runs through the standard vendor pathways, and Epic controls those pathways, and Epic has demonstrated repeatedly that it will use that control when it feels like it. Ask Particle Health, currently litigating an antitrust case after Epic cut off portions of its data access in a dispute nominally about treatment purpose under Carequality. Ask CureIS, which filed its own suit alleging Epic squeezes third parties that touch its customers. The point is not that Epic will kneecap OpenAI. The point is that the world’s most valuable AI company just built a flagship healthcare product whose critical dependency is the goodwill of a privately held company in Verona, Wisconsin that has never once lost a staring contest.
Epic has been playing this game for years, mostly with OpenAI’s own models
Here is the part that makes the whole thing delicious. Epic has been shipping generative AI features since 2023, over a hundred of them at this point across In Basket draft replies, note summarization, coding suggestions, patient message drafting, and its Ask ART style chart questioning tools. And the models under most of that hood have historically been OpenAI models, delivered through Azure OpenAI, because Epic’s strategic AI partner is Microsoft, and Microsoft’s strategic AI partner is, or at least was in the uncomplicated days, OpenAI. So for a couple of years the arrangement was tidy. OpenAI made the models, Microsoft wrapped them in Azure compliance clothing, Epic embedded them in workflow, health systems paid Epic and Microsoft, and OpenAI got paid at the bottom of the stack without ever holding a BAA or sitting in a hospital security review.
This announcement is OpenAI deciding the bottom of the stack is a bad place to live. Going direct to health systems with its own branded workspace, its own BAA, its own embedded Epic experience, and its own enterprise sales team puts OpenAI in competition with the feature roadmap of its distribution partner’s most important healthcare relationship. Epic now gets to decide how enthusiastically to support an integration that competes with Ask ART and friends, Microsoft gets to decide how it feels about its model supplier disintermediating an Azure revenue stream, and health system CIOs get to sit through three different vendors pitching what is functionally the same chart summarization demo, two of which run on the same underlying model family. Somewhere a strategy consultant is billing four hundred an hour to draw this as a triangle.
Epic’s counterweight is real, though. Cosmos sits on de-identified records covering roughly 300 million patients, and Epic’s pitch is increasingly that the interesting AI is the AI trained and validated on that corpus, delivered natively, with no third party BAA and no new vendor risk assessment. Epic also holds a bit under half the US acute care hospital market by facilities and more than half of beds, and its customers skew toward exactly the large academic and multi-state systems on OpenAI’s launch partner list. Which means the launch partners are, almost by definition, Epic shops that decided the native roadmap was not moving fast enough or was not general enough. That is the actual market signal in this announcement. Seven brand name systems just said out loud that they want a general purpose reasoning layer over the chart, not another point feature per workflow, and they are willing to onboard a new business associate to get it.
About that 99.1 percent
OpenAI came armed with numbers, and to their credit the numbers are more specific than the usual AI healthcare vapor. A network of hundreds of physicians across 60 countries, 49 languages, and 26 specialties has reviewed more than 700,000 model responses to date. For the EHR context work specifically, physicians rated responses across 27 use cases, things like pre-visit review, medication reconciliation, clinical timelines, and handoff summaries, with 4,363 ratings and 99.1 percent of responses judged safe. A separate two round eval on nuanced questions over large US healthcare datasets found more than 93 percent of responses rated good or better on accuracy for each of five connected sources.
Now do what this audience always does and squint at the denominators. Safe is a floor, not a ceiling. A response can be rated safe and still be unhelpful, incomplete, or subtly wrong in ways that do not trip a harm flag. 99.1 percent safe across 4,363 ratings means roughly 39 responses were rated something other than safe, and the interesting document is the one describing those 39, which use cases they clustered in, what the failure modes looked like, and what inter-rater agreement was. Handoff summaries and med rec are precisely the use cases where the cost of an omission is asymmetric, where the thing the model fails to surface matters more than anything it says. Rating a summary safe requires the rater to know what was in the chart that the summary skipped, which is a much harder evaluation than reading the output and vibing. Same story on the 93 percent good or better accuracy figure. Good or better is a bar that includes good, and one wrong answer out of fifteen on coverage policy versions or trial eligibility criteria is a number that would get a human analyst put on a performance plan.
None of this is a dunk. Publishing use case level physician evals at all puts OpenAI ahead of most of the market, and 700k reviewed responses is a genuinely large human feedback operation. The gripe is that these are vendor conducted, vendor summarized evals with no public methodology, no confusion matrices, and no per use case breakdown, at exactly the moment health systems are being told to treat AI procurement like device procurement. HTI-1 already forces certified EHR developers to publish source attribute transparency for predictive decision support. The cultural expectation is drifting toward model cards with actual statistics. The vendor who publishes the ugly table first, failure modes and all, is going to win a surprising amount of trust from clinical informatics people who are professionally allergic to marketing percentages.
The startup blast radius
Every platform announcement in healthcare AI triggers the same ritual, in which founders post that this validates the space while their investors quietly reopen the competitive slide. So, honestly, who gets hurt. The most exposed category is the pure chart summarization and pre-visit prep startups, the companies whose entire product is a FHIR pipe, a prompt library, and a nice UI for asking questions of the record. That was always the thinnest wrapper in the industry, and it now competes with the model vendor itself, embedded in Epic, holding a BAA, with a physician eval program bigger than most startups’ user counts. Those companies have maybe a year to become a workflow or become a memory.
OpenEvidence is the more interesting case. It built a very large clinician user base and a rich valuation on being the trusted medical answer engine, with licensed content partnerships and a distribution flywheel through individual doctors. The Healthcare Public Data plugin walks directly at the evidence retrieval part of that, and everyone in the industry can see the next move the LinkedIn commentariat already called: content partnerships with the likes of UpToDate or other medical knowledge publishers, at which point ChatGPT for Healthcare is an answer engine with the chart attached. OpenEvidence’s defense is specialization, physician trust, and the fact that Wolters Kluwer and Elsevier get to choose their partners carefully, since licensing your crown jewel content to the platform that might eat you is a decision publishers have gotten burned on before. Watch where the UpToDate license lands. It is the most consequential unsigned contract in clinical AI.
The ambient documentation crowd, Abridge and Ambience and Suki and Microsoft’s Dragon Copilot, are safer for now because ambient capture is a genuinely hard audio and workflow problem, and Abridge in particular has spent its multi-billion valuation building deep Epic integration and health system relationships. But the strategic pattern should bother them: OpenAI went from model supplier to application vendor in one announcement, and there is no law of nature saying documentation is not next. Meanwhile revenue cycle and coverage tooling people should look hard at the CMS Coverage connector. Structured, version-aware access to NCDs and LCDs inside a general reasoning workspace is quietly a prior auth and denials research tool, and the number of RCM point solutions that are essentially a coverage policy lookup with an interface is larger than anyone in RCM wants to admit. And a note for the interop and data infrastructure layer broadly: every model vendor that goes direct to the chart increases the value of clean, permissioned, well governed pipes and the legal scaffolding around them. Records do not move themselves, and the compliance surface area of AI reading charts at scale is about to make everyone rediscover why release of information is a regulated discipline and not a file transfer.
The governance hangover
The sharpest early commentary on this launch came from clinical informatics people, not investors, and the concern was consistent. When an organization builds its own AI tools on top of the EHR, it controls the guardrails, the logging, the scope of what the model can see, and the eval loop. When users connect the org’s Epic instance to a general purpose chat product, those guardrails now live wherever OpenAI decided to put them, and the org’s visibility depends entirely on what monitoring and evaluation tooling OpenAI exposes to workspace admins. If that tooling is thin, health systems will be flying blind on how the feature is actually used, which prompts are being run against which patients, and what the failure rate looks like in their own population rather than in OpenAI’s eval set. Audit logs that satisfy HIPAA are not the same thing as observability that satisfies a quality committee.
To be fair to the doomers’ critics, the barriers to chaos are real. Nobody’s average attending is connecting the Epic prod environment to ChatGPT from the parking lot. This requires a BAA, workspace configuration, IT enablement, security review, and all the usual enterprise ceremony, and the EHR integration is not available on individual accounts at all. The uncontrolled shadow AI problem, clinicians pasting chart text into consumer chatbots, predates this launch and is arguably reduced by giving people a sanctioned, logged, BAA-covered place to do the same thing. The honest framing is that this launch converts a shadow governance problem into an explicit governance workload, which is progress, but workload nonetheless.
And the workload is genuinely large. Someone has to decide which roles get access, which use cases are blessed, what the escalation path is when the model whiffs on a med list, how output gets labeled, and how any of this squares with the regulatory perimeter. The FDA’s clinical decision support guidance keeps non-device status for tools where the clinician can independently review the basis for the recommendation, which is why everyone’s marketing copy leans so hard on pointing back to supporting chart information, and why the moment these tools start triaging or prioritizing patients rather than summarizing them, the device question gets loud. States are piling on too. California already requires disclaimers on generative AI clinical communications to patients, Colorado’s AI act looms over high risk automated systems, and the patchwork only grows. Add the malpractice question nobody has case law for yet, where a clinician relied on an AI pre-visit summary that omitted the one thing that mattered, and the governance committee’s agenda writes itself. Governance is not the boring part of this launch. Governance is the product surface where this succeeds or dies, and the vendors who treat monitoring, evals, and admin controls as first class features rather than compliance homework are going to take the enterprise market.
What to watch between now and the next HIMSS
A few tells will reveal how this actually goes. Watch whether Epic says anything at all, and in what tone, because a warm co-marketing motion versus pointed silence versus a new Vendor Services pricing memo are three very different futures. Watch whether Cerner, sorry, Oracle Health, shows up in the supported EHR list next, since the announcement’s careful phrasing about supported EHRs suggests Epic is the first pipe and not the last, and Oracle is desperate for a differentiator. Watch the UpToDate and broader content licensing chessboard, because whoever assembles chart context plus licensed evidence plus public data in one governed workspace has functionally rebuilt the clinician’s entire information environment. Watch whether OpenAI publishes real eval methodology, per use case numbers, and admin-facing monitoring tools, because that is the difference between a healthcare product and a healthcare-flavored product. And watch the pilot partners’ podium talks in about a year, because UCSF and HCA will have actual utilization data, and health systems are constitutionally incapable of not presenting a slide about it.
The meta point is simple and worth saying plainly. Foundation model vendors have decided healthcare distribution is worth owning directly, the EHR vendors have decided AI is a native feature and not a partner category, the content publishers hold licensing leverage they have not fully spent, and the startups in between are about to find out which of them were companies and which were features. Patient data plumbing, governance infrastructure, and clinical evidence licensing just became the three most strategic assets in the stack. Everything else is a prompt away from being commodity. Plan the next two quarters accordingly, and maybe send the AI governance committee some coffee. They are going to need it
.


