Data Is the Moat
But the Moat has to be at the National level
A blue-chip venture firm just underwrote the biological data-infrastructure thesis. They are right about the premise, and they are reading it through the lens of drug development. Widen the lens, and you find where national capability lives.
A biological data factory is not a volume play. It is a specification for making the right variables measurable at the right organizational layer. This is a fundamental argument about information density, not about corporate formation — and it sits squarely upstream of every business case currently being stacked on top of the stack.
BVP just underwrote a neighboring conclusion from the opposite shore. In “Building biology-native data infrastructure for the AI era” (Hedin, Jalbut, Dai, April 13, 2026), they argue that the next generation of biotech winners will be defined by three pillars: biology-native data generated at scale, agentic AI embedded in R&D workflows, and closed-loop automation. It is a coherent articulation of the stack that venture capital is prepared to fund. Bonus feature: they have provided the market map to back it up - always love that.
I want to take this seriously because it is the most coherent public articulation I have seen of a thesis I have been developing within DARPA as a national capability to prevent and create strategic surprise. When a leading venture firm underwrites “data infrastructure beats model architecture,” it is a strong signal of capital-formation events. It tells founders and LPs that generating purpose-built biological data is an investable business, which changes the build-versus-buy math. It also changes the calculus for every program we think about, where the automatic assumption is/was that Uncle Sam had to seed the data layer alone.
So this piece is not a rebuttal, nor a full-throated endorsement (we don’t do that as USG anyway…); it is, however, a calibration. BVP is smart, and the thesis is good. They are also looking at it through the one lens that’s most comfortable and financially durable: drug development, which is the correct lens for a venture firm and a narrowing one for the capability as a whole. I agree with a lot of what’s being said; it matters that they are the ones saying it, and there are a few things that come into view when you widen the lens as we have to do here. These are the pieces I’d like to focus on because they are the most important parts of this infrastructure that will not be built by venture capital alone.
The premise is correct, and the evidence is clean
The core BVP claim is that the Cambrian explosion of AI biology models has already happened, and that it did not solve the problem. I would not fight a single number in their case.
By the Epoch AI count, more than 380 biology models shipped in 2025, up from fewer than ten a decade earlier. Roughly 63% of them trained on protein sequence and structure data from UniProt and the Protein Data Bank.
That last number is the whole argument compressed into one statistic.
The PDB is magnificent, and it is biased as they clearly point out: it is overwhelmingly populated by proteins that were stable enough, soluble enough, and crystallizable enough to make it through a structural pipeline designed decades ago for entirely different reasons. We trained a generation of models on a survivorship-biased sample of the proteome and then expressed surprise that they generalized poorly to the disordered, membrane-bound, and transient.
This is the part BVP nails: the bottleneck is not architecture. It is a substrate. You can stack transformers on the PDB until the heat death of the universe, and you CANNOT learn the parts of biology that the PDB never contained, which are MOST of biology.
Much value is creatable already - great. Insilico Medicine nominated a preclinical candidate for idiopathic pulmonary fibrosis after synthesizing and testing only 78 molecules, compared with the usual 1e4-6 candidates, in roughly 18 months. The fact is that worked so far as that molecule has since reported Phase IIa proof-of-concept in Nature Medicine (June 2025). That is the first time an AI-discovered target and an AI-designed molecule have produced a human clinical signal. It is the clearest demonstration we have that a tighter data-design loop compresses discovery.
And the market is now pricing the thesis directly. GSK committed roughly $50M to NOETIK’s virtual-cell models in January. Lilly reportedly pays Chai Discovery a mid-eight-figure annual fee for access to its biologics design platform.
Well-known other fact: In April, Anthropic acquired Coefficient Bio at around $400M (well done, both team and Dimension Cap).
So we agree on the premise. Where we part company is on what the premise is for.
Widening the aperture, take one: drug discovery is the application; grokking biology is the capability
BVP’s frame is pharma R&D. The stakes they cite are clinical failure rates, the cost of an approved therapy, and candidate-nomination speed. Inside that frame, every argument they make is correct. It is simply one frame.
A venture thesis is supposed to be narrow. It should point at the slice of a capability with the clearest near-term return and the most defensible moat. Drug discovery is that slice. It has paying customers, measurable milestones, and a willingness to license data at eight- and nine-figure levels. If I were deploying a fund, I would scope it exactly the way they did.
A national capability cannot be scoped that way. The reason to build a biological data factory is not to nominate drug candidates faster. It is to give machine intelligence the data to model biology at every scale at which biology matters: host-pathogen dynamics in a developing outbreak, tissue-level response to perturbation, the process control of a biomanufacturing run, and the ecology of an environment under stress, to name a few that traverse scale (more on scale here →
Drug discovery is one application of a general capability. It is the application the market will overweight precisely because it pays, meaning it is also the one most likely to crowd out everything else.
This is the first place where the venture frame and the national frame genuinely diverge. Not because BVP is wrong about drugs. Because drugs are not the mission.
Widening the aperture, take two: “more, better, multi-modal” is a stop along the way, not the destination
BVP’s prescription is more data, better data, multi-modal data (I don’t hate that!). They enumerate the modalities: genomics, transcriptomics, proteomics, pathology, and clinical outcomes. The implicit claim is that if you assemble enough modalities at enough scale, the model will find the structure.
That phrase, “multi-modal data”, is doing a great deal of cognitive compression. A modality is a projection of an underlying biological state, not the state itself. Single-cell RNA-seq provides a single snapshot of a cell. It does not give you the cell. The right question is never simply “more data.” It is: more of which variables, measured at which organizational layer, at what time resolution, under what perturbation, against what notion of state?
Phil Anderson’s 1972 essay “More Is Different” is the load-bearing idea here. Each organizational layer of nature introduces new variables that cannot be reduced to those of the layer below. The useful state variables of a tissue are not the concatenation of the state variables of its cells. This means you cannot assume that piling up modalities at the molecular layer will teach a model anything about the dynamics that govern the tissue layer. You can build an enormous, beautifully curated, fully multi-modal dataset that measures the wrong projections and teaches a model nothing about the layer you actually care about.
The BVP map stops one level short of “the right projection.” It enumerates instruments and modalities usefully, but does not yet ask whether those modalities are the right basis for the state they aim to predict.
That is the conceptual difference I’d like to hammer on, and why we call this concept a data factory. A factory is designed backward from the variables a model needs in order to be useful at a given layer, and you build the infrastructure to produce exactly those variables. Nobody is funding that (yet), because it is harder to define and slower to monetize than “we generate proprietary multi-modal data.”
Widening the aperture, take three: the commons they stand on
There is a quiet point worth surfacing in the BVP map. The foundations they rightly credit as the bedrock of computational biology (PDB, HGP, ChEMBL) were all publicly funded, pre-competitive, exquisite, and EXPENSIVE. No venture fund built them. No venture fund could have, because their returns were nowhere to be found at first, then accumulated slowly, and still remain impossible to capture.
This new layer of foundational data has the same challenging economics. Standardized perturbation-response atlases, cross-layer reference measurements, and benchmarks that let one company’s proprietary model be trusted and compared against another’s, to name just a few, are all expensive. So these high-cost, non-excludable goods are textbook examples of public goods, and that is the textbook reason the private market underinvests in them.
Venture capital will fund the proprietary slice, deep maybe, but never wide (prove me wrong), and they won’t be public (again happy to think about PPPs and so on). That is exactly what the BVP map shows, and it is exactly the right thing for venture to fund. What venture will not fund is the reference layer that makes all those proprietary slices interoperable, comparable, and trustworthy. Markets do not build commons. They build ON them. The pre-competitive data layer is not a gap in the BVP thesis by accident. It is a gap created by economic necessity, and it is the bedrock where a non-market mandate must operate.
The strongest argument against all of this
I owe the reader the best version of the counter-argument, including the version that wounds my own position rather than BVP’s.
The strongest objection is that the binding constraint is neither models nor discovery-stage data. It is the clinical wall. Nearly 90% of drugs that enter clinical trials fail — and that figure is a clinical-stage number, not an all-pipeline one. The failures focus on efficacy and toxicity that only manifest in humans. Insilico has reached Phase IIa, not approval. If most of the cost and most of the failure live in animal and human trials, then a faster, richer data loop compresses the cheap front third of the funnel while leaving the expensive back two-thirds untouched. On this view, both “better models” and “better data” optimize the part of the process that is neither the source of value nor the cause of failure.
This is a serious critique, and I take it seriously. My response is that it wounds the drug-discovery framing far more than it wounds the data-factory thesis, and that it is, in fact, an argument for the broader scope rather than against it.
If the clinical wall is the problem, the right response is to generate data that predict clinical failure earlier: human-relevant immunogenicity, ADMET, and perturbation responses in systems that actually resemble the patient. Those are harder projections, at higher organizational layers, exactly the ones the current data ignores. And it is an argument for all the applications of biological understanding that have no clinical trial at all (surveillance, manufacturing, environmental biology), where the clinical wall does not exist, and the data gap is just as real.
If purely computational discovery is real, an AI-native drug, from an AI-selected target, an AI-designed molecule, through an AI-designed clinical trial, likely will clear Phase III in the next few years. If by 2029 none have, we haven’t learned from all our data. It might be that no entity can generate enough private data to crack that nut.
None of that weakens the point for national-level Data Factories.
What happens next, and what would I build?
Three predictions, offered with a confidence modifier.
First, the frontier labs will NOT vertically integrate into data generation. This makes the commons less likely to be built privately, not more, because the most capable players like OAI and Anthropic will be acquiring data rather than generating it themselves. That increases the value for the public case; it does not lower it.
Second, the proprietary-data licensing market BVP describes will get VERY hot and then get (more) crowded, and the differentiator will migrate from “we have proprietary data” to “we measure the right projections of the right layer for your model (maybe it’s already there, though I don’t think so).”
The companies that internalize the Anderson point will pull away from the ones still counting modalities.
Third, the pre-competitive layer will not build itself, and if it is going to exist this decade, it will require the same kind of public, mission-driven investment that built the PDB and the Human Genome Project.
That is the part I am working on, and it is the part the venture thesis structurally cannot and should not reach.
So I will end one layer out from where the BVP thesis lands.
The data is the moat.
They are right about that. But the deepest, most valuable, and least defensible part of that moat is exactly the part the market will not fund. That part is a national capability, not a venture return.
Right thesis, through the drug-development lens. The wider scope is not a correction to their work. It is a proper division of labor.


I love this view on AI and data. There is a big public-private conversation needed to support this. Regulators are blocking this accidentally from happening on the manufacturing and we are missing a common infrastructure to manage this. It would leap us forward massively if we had this infrastructure and approach to data.
Love this point, hits home: “If the clinical wall is the problem, the right response is to generate data that predict clinical failure earlier: human-relevant immunogenicity, ADMET, and perturbation responses in systems that actually resemble the patient. Those are harder projections, at higher organizational layers, exactly the ones the current data ignores. And it is an argument for all the applications of biological understanding that have no clinical trial at all (surveillance, manufacturing, environmental biology), where the clinical wall does not exist, and the data gap is just as real.”