Video Preview
🎧 Part I Podcast free on Apple Podcasts and Spotify.
🎧 Part II Podcast episode for paid subscribers only. Also available on Apple Podcasts and Spotify.
To listen to paid episodes in Apple or Spotify, link your Substack subscription via the show settings on those platforms (instructions inside the Substack app under Subscriptions → Podcast).
Table of Contents
The Intranet Was Supposed to Win
Private Blockchains and the Consortium Graveyard
Encarta, Britannica, and the Encyclopedia That Anyone Could Edit
Why Open Wins the Market and Loses the P&L
Open Source AI Is the Whole Ballgame Right Now
Healthcare’s Long, Weird Relationship With Open
Metriport and the Patient Context Layer
How Metriport Actually Makes Money
What Could Go Wrong
Abstract
Open systems tend to beat closed ones on adoption, durability, and total surplus created, but the surplus mostly lands with users and the ecosystem rather than the originator, which is why capital keeps funding closed things instead.
Three case studies: corporate intranets and walled-garden online services vs the public internet, permissioned enterprise blockchains vs public chains, and Encarta and Britannica vs Wikipedia. Same pattern each time.
Open source AI matters more than any prior open vs closed fight because the thing being closed is the reasoning layer for everything else, and healthcare cannot audit, govern, or trust a model it cannot inspect.
Healthcare has a history of open standards (HL7, FHIR, VistA, Mirth) getting captured, abandoned, or relicensed, and a current market structure where the record itself sits behind a handful of gatekeepers.
Metriport raised $26M (Series A, led by TJ Parker at Matrix with ARTIS and YC, $28.4M total) to build open source infrastructure that pulls fragmented longitudinal records into structured patient context for apps, clinicians, and AI agents. Customers cited include Amazon One Medical, Sollis Health, and Color Health.
Monetization paths that actually work for open core in this specific market: hosted network access and compliance as the paid layer, usage-based pricing on the retrieval and normalization pipeline, proprietary quality and dedup layers on top of an open substrate, and eventually per-agent or per-query context fees from AI vendors who need a longitudinal record and have no legal or technical way to build one themselves.
Risks: relicensing pressure, network gatekeepers cutting access, EHR vendors building their own context layer, and the classic open core problem of the free tier being good enough.
The Intranet Was Supposed to Win
Rewind to the mid nineties and the smart money in corporate IT was not on the internet. It was on the intranet. Lotus Notes, Novell, proprietary groupware, and a whole generation of consultants selling the idea that your company’s information should live inside a controlled network with controlled clients and controlled protocols. On the consumer side the same logic ran through CompuServe, Prodigy, and AOL, each of which was a curated online service with its own content, its own email, its own forums, and its own idea of what you should be allowed to see. AOL peaked at something like 26.5 million subscribers around 2002. That is a lot of people paying monthly for a walled garden, and it looked like a fantastic business right up until it wasn’t.
The thing that beat all of it was TCP/IP plus HTTP plus HTML, none of which anyone owned. Tim Berners-Lee’s decision to put the web protocols into the public domain in 1993 is one of the highest leverage licensing choices in history, and it produced roughly zero dollars for CERN. Netscape briefly looked like it could own the browser layer, then Microsoft crushed it by giving IE away, then Netscape open sourced Mozilla out of desperation, and two decades later the dominant browser engine on the planet is Chromium, which is open source and maintained mostly by Google because owning the substrate is worth more to Google than charging for it. Meanwhile the intranet vendors either pivoted into cloud collaboration tools that ride on open protocols or quietly died.
Notice the shape of that story because it repeats. The closed thing had a better business model on paper. The open thing had no business model at all. The open thing won anyway because every incremental developer, user, and dollar of investment on the open side made the open side more valuable to everyone else, while every incremental dollar on the closed side made it more valuable mostly to the owner. Users are not dumb. Given the choice they pick the network where the value accrues to them.
The people who got rich off the open internet were not the people who built the protocols. They were the people who built on top of them and figured out how to charge for something scarce that the open layer created: attention, distribution, hosting, trust, search. That distinction, between the layer that is open and the layer where the money is, is the entire subject of this essay.
Private Blockchains and the Consortium Graveyard
The blockchain version of this is more recent and more embarrassing because a lot of very serious institutions bought into it with very serious budgets. Around 2015 to 2018 the enterprise pitch was that public chains like Bitcoin and Ethereum were fine for speculators but real businesses needed permissioned ledgers: known participants, no tokens, no miners, governance by a consortium of incumbents. Hyperledger Fabric out of IBM and the Linux Foundation, R3 Corda backed by a consortium of banks, JPMorgan’s Quorum, and dozens of industry specific consortia in shipping, insurance, and yes, healthcare.
The results were not great. IBM and Maersk’s TradeLens, which was supposed to put global container shipping on a permissioned ledger, shut down at the end of 2022 after five years because it could not get enough carriers and ports to join a network that Maersk effectively controlled. The Australian Securities Exchange spent seven years trying to replace its CHESS settlement system with a blockchain built on Digital Asset’s technology, then scrapped it in late 2022 and wrote down roughly A$250 million. The healthcare consortia, the ones that were going to put provider credentialing or claims adjudication on a shared ledger between competing payers, mostly produced press releases and a pilot or two. B3i, the insurance industry consortium, folded in 2022.
The underlying problem was structural, not technical. A permissioned ledger among competitors requires those competitors to agree on governance, and the whole point of a blockchain was supposed to be that you did not need to trust a governing party. A private chain gives you the worst of both worlds: the cost and complexity of distributed consensus plus the political overhead of a consortium, with none of the credible neutrality that makes a public chain worth building on. Nobody wants to build their business on a database their competitor can turn off.
Meanwhile Ethereum, which anyone can read, write to, fork, or build on, processed more real economic activity in a random month of 2024 than every enterprise consortium chain combined did over their entire lifespan, and stablecoins on public chains settle more volume annually than Visa. Whatever your view on crypto as an asset class, the market spoke clearly on open vs permissioned infrastructure. Open won, again, and again the people who built the base layer did not capture most of the value. Ethereum’s founders did fine, to be fair, but the exchanges, the stablecoin issuers, and the L2s captured the recurring revenue. The pattern holds.
Encarta, Britannica, and the Encyclopedia That Anyone Could Edit
Microsoft Encarta launched in 1993 on CD ROM, cost around a hundred bucks, and was for a while a genuinely great product with licensed Funk and Wagnalls content, multimedia, and a real editorial staff. Britannica had been the gold standard for two centuries and a full print set ran well over a thousand dollars. Both were closed in every sense: closed editorial process, closed licensing, closed distribution. Both had professional writers, professional fact checkers, and brand equity money cannot buy.
Wikipedia launched in January 2001 as a side project to Nupedia, which was itself supposed to be a free but expert reviewed encyclopedia with a seven step editorial process. Nupedia produced about two dozen finished articles in its first year. Wikipedia, which threw out the review process and let anyone edit, hit twenty thousand articles in the same period. By 2005 Nature ran a comparison of science entries and found Wikipedia’s error rate was in the same neighborhood as Britannica’s, which caused an enormous fight and which Britannica disputed, but the fact that the comparison was even plausible was the whole story. Microsoft killed Encarta in 2009. Britannica stopped printing in 2012 after 244 years. English Wikipedia sits at nearly seven million articles and the project runs in over three hundred languages on a budget that would not cover the marketing line for a mid sized SaaS company.
What is instructive for healthcare people is why the closed encyclopedias lost. It was not that Wikipedia’s articles were better one at a time. Early on they were often worse. It was that Wikipedia’s marginal cost of a new article was effectively zero and its marginal cost of a correction was effectively zero, so coverage and freshness compounded in a way that no editorial staff could match. A closed system has to decide what is worth covering. An open system covers everything and lets the traffic sort it out. The niche entry on a rare disease, an obscure billing code, or a defunct medical device company exists on Wikipedia because someone who cared wrote it, not because an editor calculated it would justify its cost.
And once more, the Wikimedia Foundation is a nonprofit that runs on donations. The value went to every student, every clinician doing a quick lookup, every LLM that trained on the corpus. Google reportedly gets a huge share of its knowledge panel content from Wikipedia and pays essentially nothing for it. The open thing became infrastructure and infrastructure is notoriously hard to bill for.
Why Open Wins the Market and Loses the P&L
So the pattern is consistent enough to state as a rule. Open systems win on adoption because the cost of joining is low and the value of joining rises with every other participant. They win on durability because no single owner can kill them, relicense them, or price them out of reach. They win on innovation because anyone can build on them without permission. And they win on trust because you can inspect what you are relying on. In a world of adversarial vendors, that last one matters more every year.
Open systems lose on the income statement because the same properties that make them win make them hard to charge for. If the code is free and the protocol is public, the thing you are selling has to be something adjacent: hosting, support, compliance, convenience, proprietary extensions, or the network position that comes from being the default implementer. The history of commercial open source is the history of companies trying to find that adjacent thing and either succeeding wildly or getting eaten.
The successes are real and they are big. Red Hat sold to IBM for $34 billion in 2019 on a business that was essentially support subscriptions for software anyone could download. MongoDB is a public company worth tens of billions on the strength of Atlas, its hosted service, after the self hosted product got commoditized. GitLab, HashiCorp, Elastic, Confluent, and Databricks all built real revenue on top of open cores. Android is open source and runs on roughly seventy percent of the world’s smartphones, and Google monetizes it through Play and search defaults rather than the OS itself.
The failures and the near failures are instructive too. Elastic, MongoDB, Redis, and HashiCorp all relicensed away from OSI approved licenses because AWS and other hyperscalers were hosting their software and capturing the managed service revenue without contributing back. Each relicensing triggered a community fork: OpenSearch from Elasticsearch, Valkey from Redis, OpenTofu from Terraform. HashiCorp’s move to the Business Source License in 2023 preceded its sale to IBM for $6.4 billion, which you can read as a win or as a company that could not make the math work independently. The lesson is not that open core fails. The lesson is that the moat cannot be the code. It has to be something the code produces that a hyperscaler or a competitor cannot easily replicate: the network, the data, the compliance posture, the operational know how, or the brand that customers trust when it matters.
There is also a simpler, dumber reason we do not see more open source: the founders and investors who could build it look at the cap table math and choose closed. A closed product with a clean license and a per seat price is easy to model and easy to defend in a board meeting. An open product requires a story about where the money comes from that most Series A partners are not built to evaluate. Capital flows to legibility. Open is illegible at the seed stage and obvious in hindsight, which is a terrible combination for fundraising.
Open Source AI Is the Whole Ballgame Right Now
Every prior open vs closed fight was about a layer of the stack. This one is about the layer that is starting to do the thinking for all the other layers, which is why it is the most important version of the argument the world has faced and why the stakes inside healthcare are higher than in almost any other sector.
The state of play as of this writing: the frontier of closed models is held by a handful of labs charging per token through APIs, and the frontier of open weights sits somewhere between six and eighteen months behind depending on the benchmark and who is measuring. Meta’s Llama series, Mistral, Alibaba’s Qwen, DeepSeek, and a rotating cast of others have kept the gap from widening into a chasm. DeepSeek R1 in early 2025 was the moment the market noticed that a lab with a fraction of the compute budget could ship a reasoning model with open weights that was competitive with the best closed models, and the resulting one day drawdown in Nvidia’s market cap was something like six hundred billion dollars, which is a fun way to learn that open source has macro consequences. Even OpenAI, which had spent years arguing that open weights were dangerous, shipped a set of open weight models in 2025, which tells you the competitive pressure was real.
Why does this matter more than the browser wars or the encyclopedia wars? Because a closed model is a closed reasoning process. You cannot audit it, you cannot reproduce it, you cannot run it inside your own security perimeter, you cannot guarantee it will still exist or behave the same way next quarter, and you cannot tell a regulator what it does beyond what the vendor tells you. For a consumer chatbot that is a nuisance. For a clinical decision support tool, a prior authorization engine, a coding model, or a PBM’s formulary logic, it is a governance problem that most compliance departments have not yet fully priced. FDA’s approach to AI enabled devices leans heavily on predetermined change control plans and transparency about the model, which is a lot easier when you can actually inspect the model.
Then there is the data question, which healthcare people understand better than anyone. If the reasoning layer is closed and the context layer that feeds it is also closed, the vendor that owns both effectively owns the patient. That is the outcome the industry spent two decades fighting through interoperability rules, information blocking penalties, and TEFCA, and it would be a shame to lose it in a year because a model vendor bundled the record with the brain. Open weights plus open context infrastructure is the only configuration where the health system, the patient, and the regulator retain any leverage. Everything else is a rerun of the EHR lock in story with a bigger budget.
Healthcare’s Long, Weird Relationship With Open
Healthcare is not a stranger to open. It just keeps getting the ending wrong.
HL7 v2 is technically an open standard and it runs an enormous share of the world’s clinical interfaces, but the way it was implemented turned every integration into a bespoke consulting engagement. FHIR was the do over: RESTful, JSON, modern, genuinely open, and it worked well enough that CMS and ONC mandated FHIR APIs for payers and certified EHRs. FHIR is arguably the biggest open standard win the sector has had, and it took a federal mandate to force adoption because the vendors had no commercial reason to make their data easy to get.
VistA, the VA’s electronic health record, was built by government employees starting in the late seventies and released into the public domain. It was for a long time the most widely deployed EHR in the country, it consistently scored at or near the top on clinician satisfaction surveys, and it spawned a small ecosystem of commercial vendors like Medsphere and DSS. The VA then decided to replace it with a commercial Cerner system on a contract originally valued around $10 billion that has since grown past $16 billion, been paused, restarted, and generated a stack of oversight reports about patient safety incidents. You can argue the decision either way, but as a natural experiment in open vs closed it was not kind to closed.
Mirth Connect was the workhorse open source integration engine for a generation of health IT shops. It was acquired by Quality Systems, which became NextGen, and in 2024 NextGen moved Mirth to a commercial license and ended the open source distribution. A community fork appeared almost immediately, because of course it did, and thousands of hospitals and vendors who had built on Mirth got a lesson in what it means to depend on an open project with a single corporate owner. OpenMRS, OpenEMR, and a few other open EHRs persist mostly in global health and small practice settings where the commercial vendors never bothered to compete.
The current market structure for the record itself is the part that should keep people up at night. The two big national exchange frameworks, Carequality and CommonWell, plus the eHealth Exchange and now TEFCA’s QHINs, are the plumbing through which longitudinal records actually move between organizations. Access to that plumbing is gated by implementer agreements, purpose of use rules, and the discretion of the major EHR vendors whose customers hold most of the data. In 2024 Epic cut off Particle Health’s access to Carequality over a dispute about whether Particle’s customers were using data for treatment, and Particle responded with an antitrust suit.
Carequality later released a resolution to the dispute that both Epic and Particle accepted, and Particle’s antitrust lawsuit against Epic has continued separately. Epic has denied the claims and sought dismissal, but the key point for this article remains the same: a company whose entire product depends on network access is one policy decision away from losing its product.
Metriport and the Patient Context Layer
Metriport announced a $26 million Series A on August 27, led by TJ Parker, general partner at Matrix, with participation from ARTIS and Y Combinator, bringing total funding to $28.4 million. The company came out of YC and has been building in the open for a few years: a medical records API, connectivity to the major exchange networks, a FHIR based data model, and tooling to consolidate records from many sources into a single longitudinal view. The code is on GitHub and the company has been fairly loud about the open source positioning, which in this sector is unusual enough to be a brand. Customers cited include Amazon One Medical, Sollis Health, and Color Health.
The pitch is simple and the timing is the interesting part. An AI agent that sees one encounter or one EHR database is not particularly useful for anything beyond drafting a note. To do real clinical or administrative work it needs the patient’s history across every system that has touched them, normalized into a consistent structure, deduplicated, reconciled, and retrievable at the moment the agent needs it. The pipeline looks roughly like this: network connectivity to HIEs, QHINs, and the exchange frameworks, then normalization into a consistent model, then assembly into a longitudinal record, then retrieval and context construction, then finally the agent. Metriport is positioning itself across the middle three steps. Not the network, which it does not own, and not the agent, which everyone and their cousin is building, but the layer that turns network access into something an agent can actually use.
That layer is genuinely hard and mostly unglamorous. Records arrive as CCDs that are technically valid and practically useless, as FHIR bundles from a dozen different implementations that each interpret the spec slightly differently, as PDFs of scanned faxes, and as duplicate encounters from three organizations that all billed the same visit. Reconciling that into a clean medication list or a problem list that a clinician would trust is a data engineering problem with a long tail of edge cases. It is also exactly the kind of problem where open source compounds, because every customer that hits a weird edge case and contributes a fix makes the pipeline better for everyone, and where a closed vendor has to pay engineers to find those cases one at a time.
There is going to be a fight over who becomes the data and context API for healthcare AI. The big EHR vendors will argue it should be them, since they hold the source data and already have the network relationships. Cloud providers will argue it should be them, since they host everything anyway. Independent aggregators like Health Gorilla, Zus, 1upHealth, and Particle will each make a case. The model labs would love it to be them. Metriport’s argument is that the context layer should be open because every other option puts a toll booth between the patient’s history and the systems that need it, and healthcare has spent twenty years and several acts of Congress trying to remove exactly those toll booths.
How Metriport Actually Makes Money
Here is where the open source purists usually get quiet and the investors start squinting, so it is worth being concrete. A $26 million Series A implies a plan to get to something like tens of millions in ARR within a few years. The code being free does not preclude that. The following is how a company in this exact position gets paid, roughly in order of how quickly the revenue shows up.
The first and most obvious revenue line is hosted network access. Connecting to Carequality, CommonWell, eHealth Exchange, and the QHINs is not something you do by cloning a repo. It requires implementer agreements, legal review, security attestations, ongoing compliance, certificates, and relationships with the network operators that take months to establish and can be revoked. A customer can self host the entire Metriport stack and still needs someone to be the implementer of record on the networks, and that someone charges for it. This is the moat that is not the code, and it is a good one. The hyperscaler that wants to fork the repo and host it also has to go become a Carequality implementer and take on the purpose of use liability, which is a lot less attractive than spinning up a managed Elasticsearch cluster. The pricing model here is likely a platform fee plus per patient or per query charges for network retrievals, and the unit economics work because the network costs are largely fixed while the query volume scales with the customer.
The second line is usage based pricing on the normalization and consolidation pipeline. Even a customer with their own network access will happily pay per document or per patient for a service that turns a pile of CCDs into a clean FHIR bundle with deduplicated meds, reconciled problems, and structured labs. This is the MongoDB Atlas move: the open source software is the on ramp and the hosted, scaled, monitored, SLA backed version of it is the product. Self hosting is real for the big customers who want control and a hedge, and it costs them engineers, so most of the market picks the hosted tier and pays for the privilege of not running it. The self hosted option is not lost revenue. It is the sales pitch for the hosted option, because the customer knows they are never locked in.
The third line is the proprietary quality layer, which is where open core gets its name and where the license boundary matters. The base pipeline can be open. The models that do clinical entity extraction from unstructured text, the reconciliation logic that decides which of three conflicting medication records is current, the confidence scoring, the specialty specific views, and the retrieval tooling that packages context for an LLM with the right chunking, the right citations back to source documents, and the right filtering for purpose of use can all sit behind a commercial license. This is the part where healthcare’s data mess is a feature. The open substrate handles the eighty percent of records that are well formed. The paid layer handles the twenty percent that determine whether a clinician trusts the output, and that twenty percent is where all the value is.
The fourth line, and the one that turns this from a nice infrastructure business into a strategically important one, is the AI vendor channel. Every company building a clinical or administrative agent needs longitudinal patient context and almost none of them have a legal or technical path to build it themselves. They are not covered entities in the right way, they do not have network access, and they are not going to spend two years and a compliance team getting it. A context API priced per agent call or per patient context assembled, sold to the ambient scribe vendors, the prior auth automation shops, the coding companies, the care management platforms, and the model labs themselves, is a revenue line that scales with the entire AI in healthcare category rather than with any one customer’s patient panel. This is the Stripe analogy that every infra company reaches for, and for once it is roughly right: be the thing every app calls, take a small cut of every call, and let the apps compete with each other.
There are supporting lines too. Enterprise support and services for the health systems and payers that self host, the Red Hat model, which is unsexy and prints money. Compliance packaging: the BAA, the SOC 2, the HITRUST, the audit logs, and the purpose of use enforcement, sold as a bundle to customers whose compliance team will not sign off on a bare repo. Cloud marketplace listings that let a health system buy through their existing AWS or Azure commit, which shortens sales cycles considerably. And eventually, if the company gets big enough, a governed data network play where customers who contribute to the consolidated record can query aggregate insights, though that runs into privacy and consent walls quickly and is a later chapter.
The honest summary is that the code is a customer acquisition strategy and a trust signal, the network access and compliance posture are the moat, the hosted pipeline is the recurring revenue, the proprietary quality layer is the margin, and the AI vendor channel is the upside. None of that requires closing the core. All of it requires being disciplined about what the core is.
What Could Go Wrong
Plenty, and it is worth naming the failure modes because they are the same ones that have caught every open core company before.
The relicensing trap is the obvious one. At some point a hyperscaler or a well funded competitor will fork the repo, and the board will have a conversation about moving to a source available license. Elastic, Mongo, Redis, and HashiCorp all had that conversation and all took the hit to community trust that followed. In healthcare the community is smaller and more risk averse, and the customers who chose Metriport specifically because it was open would notice. The defense is to make sure the moat was never the code in the first place, which is the argument above, but the temptation will be real when the first big fork shows up.
Network gatekeeping is the risk specific to this sector. The Particle episode is the template: a network operator or a dominant EHR vendor decides that an aggregator’s customers are not using data for a permitted purpose, or that the aggregator’s AI use cases fall outside the exchange agreements, and access gets restricted. TEFCA was supposed to reduce this risk by putting exchange rules under a common agreement with a recognized coordinating entity, and it may, but the QHINs are still largely operated by the same incumbents and the purpose of use rules for AI context retrieval are not settled. An open context layer with lots of AI customers is exactly the kind of participant that a cautious network operator looks at sideways.
The EHR vendors could also just build it. Epic has the source data for a large fraction of the country, has its own network, has Cosmos as a research dataset, and has made no secret of its AI ambitions. A context API from the vendor that holds most of the records would be closed, expensive, and probably good enough for its own customers, and it would arrive with a sales force that is already in every large health system in the country. The counter is that the vendor’s context layer would only ever cover the vendor’s data well, and the whole point is the longitudinal record across systems, but that argument requires customers to care about the cross system view more than they care about a single throat to choke.
And there is the plain old open core problem: the free tier might be good enough. If the self hosted stack handles the customer’s needs and the customer has the engineers to run it, the hosted tier has to be meaningfully better, not just more convenient. That is a product bar the company has to keep clearing every quarter, and it gets harder as the open version improves, which it will because that is what open does.
None of that is a reason not to do it. Every case in this essay had a moment where the closed alternative looked safer, better funded, and more obviously monetizable, and in every case the open thing ended up as the infrastructure and the closed thing ended up as a case study. Healthcare is late to that pattern because its data has been locked up longer and more thoroughly than almost any other sector’s, and because the people who lock it up have been very well paid to do so. The AI wave changes the calculus, because for the first time the cost of a closed context layer is not just inconvenience or a bad interface. It is a closed reasoning process operating on a closed record, accountable to no one the patient can see. A $26 million bet that the context layer should be open, and that there is money in being the one who runs it well, is a bet that healthcare will eventually follow the same script as everything else. History says that is a pretty good bet. The P&L says it is a hard one. Both things are true, and that is the whole reason it is worth watching.


