<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Michael Koeris]]></title><description><![CDATA[Director of the Biological Technologies Office @DARPA. BOD @addgene. Serial founder Vulcan Biologics, GDMC, Sample6 (acq.), Corvium (acq.); frm. Prof @KeckGrad!]]></description><link>https://michaelkoeris.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!tBNY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364420d5-dc01-434d-a135-febb6c24dff2_4000x4000.jpeg</url><title>Michael Koeris</title><link>https://michaelkoeris.substack.com</link></image><generator>Substack</generator><lastBuildDate>Wed, 12 Aug 2026 05:04:26 GMT</lastBuildDate><atom:link href="https://michaelkoeris.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Michael Koeris]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[michaelkoeris@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[michaelkoeris@substack.com]]></itunes:email><itunes:name><![CDATA[Michael Koeris]]></itunes:name></itunes:owner><itunes:author><![CDATA[Michael Koeris]]></itunes:author><googleplay:owner><![CDATA[michaelkoeris@substack.com]]></googleplay:owner><googleplay:email><![CDATA[michaelkoeris@substack.com]]></googleplay:email><googleplay:author><![CDATA[Michael Koeris]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[On Biological Data Generation 2/n: The Economics of Metrology]]></title><description><![CDATA[What to measure, what a measurement preserves, and what it costs to make an observation identifiable]]></description><link>https://michaelkoeris.substack.com/p/on-biological-data-generation-2n</link><guid isPermaLink="false">https://michaelkoeris.substack.com/p/on-biological-data-generation-2n</guid><dc:creator><![CDATA[Michael Koeris]]></dc:creator><pubDate>Wed, 29 Jul 2026 16:13:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EURW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In <code>1/n</code>, I used Philip Anderson&#8217;s &#8220;More Is Different&#8221; as a starting point for thinking about the generation of biological data. Anderson accepted that higher-order systems obey the laws governing their constituent parts. His harder point was that those laws do not automatically supply the effective variables needed to describe the organized system above them. A cell obeys molecular physics, but a useful cellular description does not emerge simply by listing every molecule. As the organization changes, the observables required to explain the behavior change with it. <a href="https://doi.org/10.1126/science.177.4047.393?utm_source=chatgpt.com">DOI.org</a></p><p>Michael Bronstein and Luca Naef extend this argument in a complementary direction. Their &#8220;black-box data&#8221; framework questions whether biological measurements should continue to be designed mainly for human interpretation without aid, or if some data sources should be optimized for machine learning instead: higher throughput, less immediate interpretability, and more algorithmic processing between raw measurement and the desired outcome. Importantly, the black box refers to the inverse problem between measurement and useful output, not to any relaxation of experimental rigor. The measurement still needs to contain enough signal, metadata, calibration, and structure for that inverse problem to be solvable. <a href="https://pubs.rsc.org/ka/content/articlelanding/2026/sc/d6sc01189f?utm_source=chatgpt.com">Royal Society of Chemistry Publications</a></p><p>That brings the argument into metrology&#8212;the science of measurement and its application&#8212;and then one step further into the <strong>economics of metrology</strong>. Metrology conventionally includes calibration, reference standards, uncertainty, quality assurance, and traceability. The economic layer accounts for what a measurement resolves, what it averages, which structures survive the measurement, and how much downstream work is required before the resulting observation becomes useful. This is a useful conceptual extension without straying from the technical core of the discipline. <a href="https://www.nist.gov/metrology?utm_source=chatgpt.com">NIST</a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://michaelkoeris.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://michaelkoeris.substack.com/subscribe?"><span>Subscribe now</span></a></p><p>Most assay comparisons begin with cost per sample because it is visible, procurement-friendly, and easy to place in a spreadsheet. Reagents, sample preparation, instrument time, sequencing, imaging, labor, compute, and storage can all be assigned numbers. That ledger is necessary, but incomplete. The differentiators are at a higher level of abstraction. An assay can win the spreadsheet by removing the structure that made the underlying question identifiable. Therefore the cost then reemerges downstream as extra modalities, more disturbances, more replicates, stronger modeling assumptions, or a model that fails outside the conditions in which it was trained.</p><p>Bulk sequencing is the canonical example. It was transformative because it scaled. DNA was an unusually favorable substrate: stable, linear, amplifiable [Nobel][https://www.nobelprize.org/prizes/chemistry/1993/summary/], and readable [Nobel][https://www.nobelprize.org/prizes/chemistry/1980/summary/]. But bulk measurement also averages; it collapses cell identity, spatial position, local exchange, neighborhood, topology, and time. That was very useful early on and has given us many, many advances, but we tend to forget that we&#8217;re looking at aggregate measures/statistical ensembles. It&#8217;s much less useful when the questions to be answered concern organization, communication, perturbation response, or decision-making.</p><p>The unit I want to reason about is therefore the <strong>cost per task-relevant, identifiable observation</strong>. An observation is identifiable when it preserves enough of the relevant state and context that the latent variable, transition, or mechanism can be inferred without a cascade of heroic assumptions. In current machine-learning parlance, such an observation may serve as a useful training example, but actually, we don&#8217;t want to be constrained to today&#8217;s training paradigm. Rather, we recognize and celebrate that the scientific object is broader: it may be useful to constrain a mechanistic model, a causal graph, a Bayesian inference procedure, a dynamical simulator, or a control policy. So let&#8217;s get into it!</p><p>A first-pass cost accounting might look like:</p><pre><code><code>C_identifiable &#8776; C_total / N_identifiable
</code></code></pre><p>where:</p><pre><code><code>C_total &#8776; C_acquisition
        + C_sample preparation
        + C_calibration
        + C_quality control
        + C_perturbation
        + C_replication
        + C_analysis
        + C_experimental attrition
</code></code></pre><p><code>N_identifiable</code> is the number of observations that retain enough biological structure to answer the task, i.e., a molecular-abundance question, a localization question, and a cell-fate question require different observations even when they begin with the same sample.</p><p>This accounting also forces a cleaner separation between entities and variables. A receptor names an entity. Its measurable coordinates include surface density, occupancy fraction, activation state, clustering, internalization rate, and turnover. A glycan names a molecular feature class; motif abundance, branching index, linkage composition, and spatial distribution are variables. A cytokine can be measured by concentration, secretion flux, diffusion field, receptor occupancy, clearance rate, and response trajectory. The cost of metrology appears when those variables must be calibrated, compared across runs, resolved in time, or preserved in their spatial context.</p><h2><strong>From coordinate count to identifiable observation</strong></h2><p>Snapshot dimensionality remains useful as an upstream accounting device. A rough way to keep the coordinate count explicit is:</p><pre><code><code>D_snapshot &#8776; N_units &#183; d_unit + E_edges &#183; d_edge + V_fields &#183; d_field
</code></code></pre><p>where <code>N_units</code> is the number of represented units, <code>d_unit</code> is the number of scalar coordinates per unit, <code>E_edges</code> counts represented interactions or adjacencies, <code>d_edge</code> is the number of coordinates per edge, <code>V_fields</code> is the number of spatial voxels or field samples, and <code>d_field</code> is the number of coordinates associated with each field location.</p><p>This formula only counts the chosen representation. Consider - in the biological context (we&#8217;ll get to chemistry - promise) - a high-content screen containing <code>1e6</code> cells, each represented by a 200-coordinate morphological embedding:</p><pre><code><code>1e6 cells &#215; 200 coordinates/cell &#8776; 2e8 scalar coordinates
</code></code></pre><p>That is already a large measurement object before adding doses, timepoints, perturbations, replicates, batches, or adjacency relationships. The number does not reveal how many independent biological states are present, whether the coordinates are comparable across plates, or whether they retain the variable required for the task. It describes the ambient measurement space before calibration, constraints, correlation, compression, and identifiability enter.</p><p>The relationship among the relevant quantities is then:</p><pre><code><code>raw coordinates
&#8594; quality-controlled observations
&#8594; calibrated and comparable observations
&#8594; task-relevant identifiable observations
</code></code></pre><p>Raw data density can be enormous at every layer. Sequencing produces reads and bases. LC-MS produces chromatographic and spectral features. Imaging produces pixels, voxels, segmented objects, and trajectories. Single-cell assays produce sparse matrices with millions or billions of entries. The economics depend increasingly on the conversion between these stages rather than on raw coordinate production alone.</p><p>For the first reduction to practice, I am narrowing the hierarchy to three layers: molecular, subcellular, and cellular. These are the first layers where the data-factory concept can be specified with enough precision to discuss instruments, variables, quality controls, perturbations, and model updates.</p><p></p><p>Read this as a first-pass ledger for three different measurement burdens; it&#8217;s not exhaustive, and I didn&#8217;t even get into the surface receptors/interactors as a cell-2-cell comms layer. The molecular factory pays to convert a physical signal into a calibrated molecular state. The subcellular factory pays to preserve organization while measuring it. The cellular factory pays to connect baseline state, perturbation, trajectory, and outcome strongly enough that a response becomes identifiable.</p><h2><strong>A small-molecule perturbation as a metrology problem</strong></h2><p>Suppose a small molecule produces a heterogeneous response in an apparently uniform cell population. Some cells recover, some enter a durable arrest, and some commit to apoptosis. The experimental task is to determine which physical interaction initiated the response, how that interaction propagated through intracellular organization, and which state transition separated the eventual outcomes.</p><p>A molecular experiment might begin with target engagement, bulk phosphoproteomics, metabolomics, or time-resolved transcriptomics. These methods can produce extremely dense data. They may identify candidate targets, pathway activation, modification states, and metabolic consequences. Yet population averaging can obscure the distinction between responder and non-responder cells. If only a minority commits to apoptosis, the aggregate profile mixes commitment, compensation, resistance, and recovery into one measurement. The dataset can be molecularly rich while leaving the transition of interest weakly identified.</p><p>Single-cell molecular profiling restores heterogeneity. Measurements taken before treatment and at selected later times can expose distinct cell-state distributions and identify populations associated with stress, adaptation, arrest, or death. Because most such assays destroy the cells, temporal trajectories are reconstructed across different cells rather than observed continuously within the same cell. The reconstruction may be excellent, but it depends on assumptions about correspondence, timing, and the geometry of the state space.</p><p>Live-cell imaging preserves a different part of the problem. Individual cells can be followed through mitochondrial depolarization, protein translocation, organelle remodeling, nuclear changes, migration, division, arrest, or death. Timing and trajectory remain visible. Molecular breadth narrows, and the measurement itself can perturb the system through labels, illumination, or environmental constraints. For the question of which cells commit, when they commit, and what organization-state precedes the transition, however, the imaging observation can be more identifiable than a much denser molecular profile.</p><p>The strongest design may combine these measurement regimes rather than maximize any one of them: the combined experiment will be economically far superior because it reduces the number of rescue experiments required to reconstruct what any single modality discarded.</p><p>The mechanism is:</p><pre><code><code>measurement choice
&#8594; structure preserved or averaged
&#8594; target state or transition identifiable or confounded
&#8594; downstream experimental burden
</code></code></pre><p>This is the core of the economics of metrology. Acquisition cost enters at the beginning. The larger bill often depends on what has to be reconstructed after the measurement.</p><p>A related precedent is the Connectivity Map&#8217;s L1000 assay. Rather than measure a full transcriptome directly for every perturbation, L1000 measures a reduced set of 978 landmark transcripts and computationally infers much of the remaining expression state. That choice enabled a dataset of approximately 1.3 million perturbational profiles and represented more than a thousand-fold scale-up of the original Connectivity Map. The representation was intentionally sparse; its value depended on preserving enough task-relevant signal for mechanism-of-action and perturbation comparisons. <a href="https://www.sciencedirect.com/science/article/pii/S0092867417313090?utm_source=chatgpt.com">ScienceDirect</a></p><p>Ok so we see the utility of combining three layers - not news. But how to do it will become interesting. Let&#8217;s briefly get through the more precise description and definition of the three layers here.</p><h2><strong>Molecular layer: validated physical state</strong></h2><p>At the molecular layer, the factory has to make physical and chemical states comparable across samples, conditions, runs, and eventually sites. Molecular identity matters, but identity alone is a thin description. Depending on the task, the relevant coordinates may include abundance, concentration, charge state, protonation state, conformation, modification occupancy, binding occupancy, reaction rate, diffusion coefficient, degradation half-life, local environment, and interaction kinetics.</p><p>Proteins add sequence, fold, complex membership, post-translational modification, localization, turnover, and activity state. Metabolites and small molecules add solubility, stability, transport, reaction participation, target engagement, off-target interaction, and cellular effect. These variables are coupled: abundance can change without activity changing; binding affinity can remain constant while residence time changes; a modification can matter only in a particular compartment; the same metabolite can carry different functional significance depending on local concentration and flux.</p><p>DNA sequencing was a special case because the mapping from physical substrate to symbol sequence became unusually direct. The rest of molecular metrology contains harder inverse problems. LC-MS can generate dense chromatographic and spectral objects, but metabolite identity, peptide assignment, modification localization, cross-run alignment, quantitative comparability, and missingness require additional inference. A binding or kinetic assay may generate fewer raw coordinates, but each observation contains controlled concentration and temporal context. High-throughput proxy measurements may sacrifice direct interpretability while preserving latent signal that a model can exploit&#8212;the regime Bronstein and Naef describe as black-box data. <a href="https://pubs.rsc.org/en/content/articlehtml/2026/sc/d6sc01189f?utm_source=chatgpt.com">Royal Society of Chemistry Publications</a></p><p>The burden at this layer therefore concentrates in the conversion from physical signal to validated state. Sample preparation, separation, standards, calibration, provenance, dynamic range, replication, failure handling, and quality control are part of the measurement. A mass spectrum without reliable annotation may still be valuable latent data, but its relationship to the target quantity must be learned or reconstructed. A concentration estimate without traceability may be difficult to compare across instruments or sites. A kinetic parameter without a controlled input distribution may not transfer to the environment in which the molecule acts.</p><p>The molecular factory is an engineering loop already:</p><pre><code><code>make or obtain the material
&#8594; prepare and separate
&#8594; detect and characterize
&#8594; calibrate and quality-control
&#8594; associate signal with physical state
&#8594; update the model
&#8594; select the next molecule, condition, or measurement
</code></code></pre><p>Its unit of progress is a validated physical-state observation: a measurement that can be compared across the variation the model is expected to encounter.</p><h2><strong>Subcellular layer: organization as measured state</strong></h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EURW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EURW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png 424w, https://substackcdn.com/image/fetch/$s_!EURW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png 848w, https://substackcdn.com/image/fetch/$s_!EURW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png 1272w, https://substackcdn.com/image/fetch/$s_!EURW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EURW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png" width="577" height="433" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:433,&quot;width&quot;:577,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:327041,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://michaelkoeris.substack.com/i/208992153?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EURW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png 424w, https://substackcdn.com/image/fetch/$s_!EURW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png 848w, https://substackcdn.com/image/fetch/$s_!EURW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png 1272w, https://substackcdn.com/image/fetch/$s_!EURW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff961c937-1d25-4b5c-80cd-6ce3bf2ceec3_577x433.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The measurement question differs at the subcellular level; ergo, the problem and cost structure change. This is the first layer up, and, as per the initial thesis, this is the first layer where a parts list becomes obviously insufficient (also, again, still not completely measurable). Identity, association, and rate are still useful, but they are no longer enough. The question is, where are the things (kudos to those who remember the meme)?</p><p>The variables at this layer include organelle position and volume, membrane potential, luminal pH, local concentration, assembly state, condensate state, vesicle cargo, trafficking rate, fusion and fission, contact-site geometry, cytoskeletal alignment, local calcium dynamics, polarity, and phase behavior. Fractionation can return many of these objects to molecular assays, but the process usually destroys the intact spatial system. The factory then has to reconstruct which organization produced the measurements that were separated.</p><p>Imaging generates extraordinary raw coordinate density. Pixels are rarely the scarce resource. The expensive conversion runs from pixels to validated organization state: segmentation, compartment identity, localization, morphology, contact, motion, and trajectory. That conversion depends on labeling fidelity, live-cell compatibility, registration, batch correction, tracking, spatial resolution, temporal sampling, and control of phototoxicity or other measurement-induced changes.</p><p>Cell Painting provides a useful endpoint example. The canonical assay uses multiplexed dyes across five imaging channels to mark multiple cellular components and extracts roughly 1,500 morphological features from individual cells. The result is a dense, scalable representation of morphology and localization under chemical or genetic perturbation. Its power comes from converting images into comparable profiles; the raw image alone does not provide the biological representation. <a href="https://www.nature.com/articles/nprot.2016.105?utm_source=chatgpt.com">Nature</a></p><p>A mature endpoint-imaging assay can be inexpensive per cell while remaining demanding per interpretable organization state. Live imaging adds temporal continuity but requires instruments to remain occupied longer and introduces drift, phototoxicity, tracking errors, and trajectory attrition. Higher-resolution imaging preserves finer spatial relationships while reducing condition coverage and increasing reconstruction burden. The economics changes with the organization preserved, the duration observed, and the reliability with which those observations can be compared.</p><p>The subcellular loop is not as mature and can be described:</p><pre><code><code>choose context and perturbation
&#8594; image or otherwise preserve spatial state
&#8594; segment compartments and objects
&#8594; extract localization, assembly, and dynamic coordinates
&#8594; quality-control organization state
&#8594; update the model
&#8594; select the next context, perturbation, or observation window
</code></code></pre><p>Its unit of progress is a validated organization-state observation.</p><h2><strong>Cellular layer: identifiable decisions and responses</strong></h2><p>At the cellular level, the unit is an integrated system that receives inputs, maintains internal state, exposes an interface, and undergoes state transitions. Single-cell RNA sequencing gives a powerful projection of the internal state. Problems arise when that projection is treated as the cell itself, particularly for questions involving protein activity, interface state, kinetics, mechanics, metabolism, or response under perturbation.</p><p>A fuller internal representation should include transcript abundance and isoforms, protein abundance, post-translational modifications, chromatin accessibility, metabolic and redox state, cell-cycle position, stress state, morphology, mechanics, organelle state, and recent history. The externally visible state includes surface density, ligand occupancy, receptor activation and internalization, adhesion, secretion, uptake, conductance, force, motion, and contact. Cells around it receive an enormously reduced information flow (more on information transmission in another post). Basically, the vast majority of the cell is in a hidden state, and only select &#8220;APId&#8221; are disclosed/available.</p><p>Coordinate yield also has to be separated from operational biological information. A cell may copy, transcribe, translate, transport, and turn over enormous numbers of molecular symbols while transmitting comparatively little information through a specific decision-relevant channel. Any information-rate claim, therefore, needs a defined sender variable, receiver variable, input distribution, time window, and metric. Raw event flux, sequence throughput, operational mutual information, and thermodynamic upper bounds can differ by many orders of magnitude. Information Transfer in Cells.txtTXT The implication here is narrow but important: dense measurement can still leave the target transition weakly identified.</p><p>The useful cellular observation is usually conditional. It joins baseline state, perturbation identity, dose, duration, environment, trajectory, interface change, functional outcome, and recovery or adaptation. The experimental burden expands with the crossed condition space:</p><pre><code><code>N_conditions &#8776; N_states
             &#183; N_perturbations
             &#183; N_doses
             &#183; N_timepoints
             &#183; N_environments
             &#183; N_replicates</code></code></pre><p>This is obviously very high level but bears repeating: it&#8217;s a design ledger, NOT a prescription to enumerate every combination. Its purpose is to show why the marginal readout cost per cell can fall while the cost of learning cellular decisions remains high. Pooled methods such as Perturb-seq scale perturbation identity and high-content transcriptional readout across large numbers of cells, but dose, temporal continuity, functional outcome, and environmental context require additional experimental design. The original Perturb-seq work demonstrated the power of linking pooled CRISPR perturbations to single-cell transcriptomic states across approximately 200,000 cells. <a href="https://www.broadinstitute.org/publications/broad14056?utm_source=chatgpt.com">Broad Institute</a></p><p>Active learning belongs here as an allocation mechanism for selecting data collection goals. Under a fixed experimental budget, the system must decide which region of the condition space to measure next. That decision might balance model uncertainty, expected information gain, biological coverage, transition rarity, and experimental failure risk. A cellular factory becomes recursive when its accumulated measurements alter the allocation of the next experiment.</p><p>The cellular loop is:</p><pre><code><code>measure baseline state
&#8594; select and apply perturbation
&#8594; observe trajectory and interface
&#8594; measure functional outcome
&#8594; associate outcome with prior state and perturbation
&#8594; update the model
&#8594; choose the next region of condition space</code></code></pre><p>Its unit of progress is an identifiable perturbation-response transition.</p><h2><strong>Reduction to practice</strong></h2><p>The first three factories are parallel measurement loops rather than successive versions of one apparatus.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A3xK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A3xK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png 424w, https://substackcdn.com/image/fetch/$s_!A3xK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png 848w, https://substackcdn.com/image/fetch/$s_!A3xK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png 1272w, https://substackcdn.com/image/fetch/$s_!A3xK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A3xK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png" width="1456" height="925" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:925,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1146260,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://michaelkoeris.substack.com/i/208992153?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!A3xK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png 424w, https://substackcdn.com/image/fetch/$s_!A3xK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png 848w, https://substackcdn.com/image/fetch/$s_!A3xK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png 1272w, https://substackcdn.com/image/fetch/$s_!A3xK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92c6fcfb-5f5e-48b6-b881-7912d9651894_5795x3680.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>Figure 1. Three rudimentary data factories block schemas. The molecular loop makes the physical state legible. The subcellular loop makes organization legible. The cellular loop makes decisions and responses legible. The loops may share samples, instruments, models, automation, and quality systems, while retaining different units of progress.</p><p>The loops can share sample-handling infrastructure, robotics, instrumentation, standards, models, quality systems, and data architecture. Their outputs remain distinct. The molecular factory produces calibrated observations of physical states. The subcellular factory generates organizational states that maintain position and time. The cellular factory creates perturbation-response transitions that connect baseline, input, trajectory, interface, and outcome. This difference influences how you build and evaluate a data factory. A molecular assay that cannot resolve modification states may miss activity. A subcellular assay that ignores localization may miss the mechanism. A cellular assay that records transcripts without accounting for interfaces or outcomes may miss communication and decision-making. High throughput alone does not fix these omissions; it only scales the chosen representation.</p><p>The economics of metrology, therefore, has several coupled terms:</p><pre><code><code>measurement economics &#8776; acquisition burden
                      + calibration burden
                      + structure-preservation burden
                      + condition-space burden
                      + downstream reconstruction burden
</code></code></pre><p>This is a first-pass value-assessment device. It does not imply that every term can already be priced cleanly. It does clarify where costs move as the measurement crosses layers. Sequencing can make symbols inexpensive. LC-MS can make spectral features abundant. Imaging can make pixels abundant. Pooled screening can yield abundant perturbation readouts. The remaining factory problem is converting those outputs into identifiable observations without erasing the structure on which the task depends.</p><h2><strong>Higher layers</strong></h2><p>The same accounting becomes more difficult above the cell (but we&#8217;ll get into it!). At the multicellular layer, adjacency, gradients, matrix, mechanical coupling, and local exchange enter the state. At the organ level, measurement must connect spatially distributed cellular processes to flow, control, reserve, and failure. At the organism layer, history, environment, adaptation, and treatment trajectory become inseparable from the observation.</p><p>Those layers deserve their own treatment. The current reduction is deliberately narrower: physical state at the molecular level, organization at the subcellular level, and decision-making under perturbation at the cellular level.</p><h2><strong>Closing synthesis</strong></h2><p>This remains a Level 1 accounting. Snapshot dimensionality describes the coordinate space chosen by the measurement system; it does not determine how many independent, comparable, or task-relevant observations the experiment contains. Identifiability depends on the task, modality, calibration regime, perturbation design, and the amount of structure that survives measurement.</p><p>That changes the architecture of the first data factories. A molecular factory advances by producing validated physical-state observations. A subcellular factory advances by producing organization-state observations that preserve location and time. A cellular factory advances by identifying perturbation-response transitions across internal state, interface, trajectory, and outcome. These systems can share infrastructure, but they cannot share a single unit of progress without compressing away the reason for separating the layers.</p><p>That alters the architecture of the initial data factories. A molecular factory progresses by producing validated physical-state observations. A subcellular factory advances by generating organization-state observations that retain location and time. </p><ul><li><p>A cellular factory moves forward by identifying perturbation-response transitions across internal state, interface, trajectory, and outcome. These systems can share infrastructure, but they cannot share a single unit of progress without losing the purpose of separating the layers.</p></li><li><p>A general data platform will tend to optimize what is common across assays: sample throughput, coordinate volume, storage, and processing. The economics of metrology suggest the opposite approach. Start with the state or transition that must remain identifiable, then work backward into instruments, perturbations, standards, quality controls, and models.</p></li><li><p>A generic data platform will tend to optimize what is common across assays: sample throughput, coordinate volume, storage, and processing. The economics of metrology point in the opposite design direction. Begin with the state or transition that must remain identifiable, then engineer backward into instruments, perturbations, standards, quality controls, and models.</p></li></ul><p>Chemistry is the next instance because it exposes the problem one step earlier. The object of measurement often has to be designed, synthesized, purified, and stabilized before it can be observed.</p><p>The economics of chemical metrology, therefore, begins upstream of the instrument, with the decision about what physical object should exist.</p><p>Until next time - stay frosty!</p>]]></content:encoded></item><item><title><![CDATA[Data Is the Moat]]></title><description><![CDATA[But the Moat has to be at the National level]]></description><link>https://michaelkoeris.substack.com/p/data-is-the-moat</link><guid isPermaLink="false">https://michaelkoeris.substack.com/p/data-is-the-moat</guid><dc:creator><![CDATA[Michael Koeris]]></dc:creator><pubDate>Sun, 31 May 2026 11:23:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tBNY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364420d5-dc01-434d-a135-febb6c24dff2_4000x4000.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>A blue-chip venture firm just underwrote the biological data-infrastructure thesis. They are right about the premise, and they are reading it through the lens of drug development. Widen the lens, and you find where national capability lives.</em></p><div><hr></div><p>A biological data factory is not a volume play. It is a specification for making the right variables measurable at the right organizational layer. This is a fundamental argument about information density, not about corporate formation &#8212; and it sits squarely upstream of every business case currently being stacked on top of the stack.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://michaelkoeris.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>BVP just underwrote a neighboring conclusion from the opposite shore. In &#8220;Building biology-native data infrastructure for the AI era&#8221; (<a href="https://www.bvp.com/atlas/building-biology-native-data-infrastructure-for-the-ai-era">Hedin, Jalbut, Dai, April 13, 2026</a>), they argue that the next generation of biotech winners will be defined by three pillars: biology-native data generated at scale, agentic AI embedded in R&amp;D workflows, and closed-loop automation. It is a coherent articulation of the stack that venture capital is prepared to fund. Bonus feature: they have provided the market map to back it up - always love that.</p><p>I want to take this seriously because it is the most coherent public articulation I have seen of a thesis I have been developing within DARPA as a national capability to prevent and create strategic surprise. When a leading venture firm underwrites &#8220;data infrastructure beats model architecture,&#8221; it is a strong signal of capital-formation events. It tells founders and LPs that generating purpose-built biological data is an investable business, which changes the build-versus-buy math. It also changes the calculus for every program we think about, where the automatic assumption is/was that Uncle Sam had to seed the data layer alone.</p><p>So this piece is not a rebuttal, nor a full-throated endorsement (we don&#8217;t do that as USG anyway&#8230;); it is, however, a calibration. BVP is smart, and the thesis is good. They are also looking at it through the one lens that&#8217;s most comfortable and financially durable: drug development, which is the correct lens for a venture firm and a narrowing one for the capability as a whole. I agree with a lot of what&#8217;s being said; it matters that <em>they</em> are the ones saying it, and there are a few things that come into view when you widen the lens as we have to do here. These are the pieces I&#8217;d like to focus on because they are the most important parts of this infrastructure that will not be built by venture capital alone.</p><h2>The premise is correct, and the evidence is clean</h2><p>The core BVP claim is that the Cambrian explosion of AI biology models has already happened, and that it did not solve the problem. I would not fight a single number in their case.</p><p>By the <a href="https://epoch.ai/data/biology-ai-models-documentation">Epoch AI count</a>, more than <a href="https://epoch.ai/blog/expanding-our-analysis-of-biological-ai-models">380 biology models shipped in 2025</a>, up from fewer than ten a decade earlier. Roughly 63% of them trained on protein sequence and structure data from UniProt and the Protein Data Bank.</p><p>That last number is the whole argument compressed into one statistic.</p><p>The PDB is magnificent, and it is biased as they clearly point out: it is overwhelmingly populated by proteins that were stable enough, soluble enough, and crystallizable enough to make it through a structural pipeline designed decades ago for entirely different reasons. We trained a generation of models on a survivorship-biased sample of the proteome and then expressed surprise that they generalized poorly to the disordered, membrane-bound, and transient.</p><p>This is the part BVP nails: the bottleneck is not architecture. It is a substrate. You can stack transformers on the PDB until the heat death of the universe, and you CANNOT learn the parts of biology that the PDB never contained, which are MOST of biology.</p><p>Much value is creatable already - great. Insilico Medicine nominated a preclinical candidate for idiopathic pulmonary fibrosis after synthesizing and testing only 78 molecules, compared with the usual 1e4-6 candidates, in roughly 18 months. The fact is that worked so far as that molecule has since reported Phase IIa proof-of-concept in <em>Nature Medicine</em> (June 2025). That is the first time an AI-discovered target and an AI-designed molecule have produced a human clinical signal. It is the clearest demonstration we have that a tighter data-design loop compresses discovery.</p><p>And the market is now pricing the thesis directly. GSK committed roughly $50M to NOETIK&#8217;s virtual-cell models in January. Lilly reportedly pays Chai Discovery a mid-eight-figure annual fee for access to its biologics design platform.</p><p>Well-known other fact: In April, Anthropic acquired Coefficient Bio at around $400M (well done, both team and Dimension Cap).</p><p>So we agree on the premise. Where we part company is on what the premise is <em>for.</em></p><h2>Widening the aperture, take one: drug discovery is the application; grokking biology is the capability</h2><p>BVP&#8217;s frame is pharma R&amp;D. The stakes they cite are clinical failure rates, the cost of an approved therapy, and candidate-nomination speed. Inside that frame, every argument they make is correct. It is simply one frame.</p><p>A venture thesis is <em>supposed</em> to be narrow. It should point at the slice of a capability with the clearest near-term return and the most defensible moat. Drug discovery is that slice. It has paying customers, measurable milestones, and a willingness to license data at eight- and nine-figure levels. If I were deploying a fund, I would scope it exactly the way they did.</p><p>A national capability cannot be scoped that way. The reason to build a biological data factory is not to nominate drug candidates faster. It is to give machine intelligence the data to model biology at every scale at which biology matters: host-pathogen dynamics in a developing outbreak, tissue-level response to perturbation, the process control of a biomanufacturing run, and the ecology of an environment under stress, to name a few that traverse scale (more on scale here &#8594; </p><div class="embedded-post-wrap" data-attrs="{&quot;id&quot;:196367957,&quot;url&quot;:&quot;https://michaelkoeris.substack.com/p/on-biological-data-generation-1n&quot;,&quot;publication_id&quot;:3801181,&quot;embedding_publication_id&quot;:null,&quot;publication_name&quot;:&quot;Michael Koeris&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!tBNY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364420d5-dc01-434d-a135-febb6c24dff2_4000x4000.jpeg&quot;,&quot;title&quot;:&quot;On Biological Data Generation (1/n): More Is Different, and So Is the Data&quot;,&quot;truncated_body_text&quot;:&quot;I have been trying to get more precise about what we mean when we say &#8220;biology.&#8221;&quot;,&quot;date&quot;:&quot;2026-05-04T01:53:50.585Z&quot;,&quot;like_count&quot;:12,&quot;comment_count&quot;:4,&quot;bylines&quot;:[{&quot;id&quot;:6386039,&quot;name&quot;:&quot;Michael Koeris&quot;,&quot;handle&quot;:&quot;mkoeris&quot;,&quot;previous_name&quot;:&quot;Thinking Out Loud&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/364420d5-dc01-434d-a135-febb6c24dff2_4000x4000.jpeg&quot;,&quot;bio&quot;:&quot;Director of the Biological Technologies Office @DARPA. BOD @addgene. Serial founder Vulcan Biologics, GDMC, Sample6 (acq.), Corvium (acq.); frm. Prof @KeckGrad!&quot;,&quot;profile_set_up_at&quot;:&quot;2025-10-22T11:36:29.378Z&quot;,&quot;reader_installed_at&quot;:&quot;2025-10-22T11:36:18.054Z&quot;,&quot;publicationUsers&quot;:[{&quot;id&quot;:3875934,&quot;user_id&quot;:6386039,&quot;publication_id&quot;:3801181,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:true,&quot;publication&quot;:{&quot;id&quot;:3801181,&quot;name&quot;:&quot;Michael Koeris&quot;,&quot;subdomain&quot;:&quot;michaelkoeris&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;&quot;,&quot;logo_url&quot;:null,&quot;author_id&quot;:6386039,&quot;primary_user_id&quot;:6386039,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2025-01-19T11:10:42.264Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Koeris&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;profile&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}},{&quot;id&quot;:9130715,&quot;user_id&quot;:6386039,&quot;publication_id&quot;:8908418,&quot;role&quot;:&quot;admin&quot;,&quot;public&quot;:true,&quot;is_primary&quot;:false,&quot;publication&quot;:{&quot;id&quot;:8908418,&quot;name&quot;:&quot;Michael's Substack&quot;,&quot;subdomain&quot;:&quot;thinkingoutloudwithfriends&quot;,&quot;custom_domain&quot;:null,&quot;custom_domain_optional&quot;:false,&quot;hero_text&quot;:&quot;My personal Substack&quot;,&quot;logo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/364420d5-dc01-434d-a135-febb6c24dff2_4000x4000.jpeg&quot;,&quot;author_id&quot;:6386039,&quot;primary_user_id&quot;:null,&quot;theme_var_background_pop&quot;:&quot;#FF6719&quot;,&quot;created_at&quot;:&quot;2026-05-04T02:18:36.076Z&quot;,&quot;email_from_name&quot;:null,&quot;copyright&quot;:&quot;Michael Koeris&quot;,&quot;founding_plan_name&quot;:null,&quot;community_enabled&quot;:true,&quot;invite_only&quot;:false,&quot;payments_state&quot;:&quot;disabled&quot;,&quot;language&quot;:null,&quot;explicit&quot;:false,&quot;homepage_type&quot;:&quot;newspaper&quot;,&quot;is_personal_mode&quot;:false,&quot;logo_url_wide&quot;:null}}],&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null,&quot;status&quot;:{&quot;bestsellerTier&quot;:null,&quot;subscriberTier&quot;:1,&quot;leaderboard&quot;:null,&quot;vip&quot;:false,&quot;badge&quot;:{&quot;type&quot;:&quot;subscriber&quot;,&quot;tier&quot;:1,&quot;accent_colors&quot;:null},&quot;paidPublicationIds&quot;:[1317673,4210429,7107838,6349492],&quot;subscriber&quot;:null}}],&quot;utm_campaign&quot;:null,&quot;belowTheFold&quot;:true,&quot;type&quot;:&quot;newsletter&quot;,&quot;language&quot;:&quot;en&quot;,&quot;source&quot;:null}" data-component-name="EmbeddedPostToDOM"><a class="embedded-post" native="true" href="https://michaelkoeris.substack.com/p/on-biological-data-generation-1n?utm_source=substack&amp;utm_campaign=post_embed&amp;utm_medium=web"><div class="embedded-post-header"><img class="embedded-post-publication-logo" src="https://substackcdn.com/image/fetch/$s_!tBNY!,w_56,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364420d5-dc01-434d-a135-febb6c24dff2_4000x4000.jpeg" loading="lazy"><span class="embedded-post-publication-name">Michael Koeris</span></div><div class="embedded-post-title-wrapper"><div class="embedded-post-title">On Biological Data Generation (1/n): More Is Different, and So Is the Data</div></div><div class="embedded-post-body">I have been trying to get more precise about what we mean when we say &#8220;biology&#8230;</div><div class="embedded-post-cta-wrapper"><span class="embedded-post-cta">Read more</span></div><div class="embedded-post-meta">3 months ago &#183; 12 likes &#183; 4 comments &#183; Michael Koeris</div></a></div><p>Drug discovery is one application of a general capability. It is the application the market will overweight precisely because it pays, meaning it is also the one most likely to crowd out everything else.</p><p>This is the first place where the venture frame and the national frame genuinely diverge. Not because BVP is wrong about drugs. Because drugs are not the mission.</p><h2>Widening the aperture, take two: &#8220;more, better, multi-modal&#8221; is a stop along the way, not the destination</h2><p>BVP&#8217;s prescription is more data, better data, multi-modal data (I don&#8217;t hate that!). They enumerate the modalities: genomics, transcriptomics, proteomics, pathology, and clinical outcomes. The implicit claim is that if you assemble enough modalities at enough scale, the model will find the structure.</p><p>That phrase, &#8220;multi-modal data&#8221;, is doing a great deal of cognitive compression. A modality is a <em>projection</em> of an underlying biological state, not the state itself. Single-cell RNA-seq provides a single snapshot of a cell. It does not give you the cell. The right question is never simply &#8220;more data.&#8221; It is: more of which variables, measured at which organizational layer, at what time resolution, under what perturbation, against what notion of state?</p><p><a href="https://doi.org/10.1126/science.177.4047.393">Phil Anderson&#8217;s 1972 essay &#8220;More Is Different&#8221;</a> is the load-bearing idea here. Each organizational layer of nature introduces new variables that cannot be reduced to those of the layer below. The useful state variables of a tissue are not the concatenation of the state variables of its cells. This means you cannot assume that piling up modalities at the molecular layer will teach a model anything about the dynamics that govern the tissue layer. You can build an enormous, beautifully curated, fully multi-modal dataset that measures the wrong projections and teaches a model nothing about the layer you actually care about.</p><p>The BVP map stops one level short of &#8220;the right projection.&#8221; It enumerates instruments and modalities usefully, but does not yet ask whether those modalities are the right basis for the state they aim to predict.</p><p>That is the conceptual difference I&#8217;d like to hammer on, and why we call this concept a data factory. A factory is designed backward from the variables a model needs in order to be useful at a given layer, and you build the infrastructure to produce exactly those variables. Nobody is funding that (yet), because it is harder to define and slower to monetize than &#8220;we generate proprietary multi-modal data.&#8221;</p><h2>Widening the aperture, take three: the commons they stand on</h2><p>There is a quiet point worth surfacing in the BVP map. The foundations they rightly credit as the bedrock of computational biology (PDB, HGP, ChEMBL) were all publicly funded, pre-competitive, exquisite, and EXPENSIVE. No venture fund built them. No venture fund could have, because their returns were nowhere to be found at first, then accumulated slowly, and still remain impossible to capture.</p><p>This new layer of foundational data has the same challenging economics. Standardized perturbation-response atlases, cross-layer reference measurements, and benchmarks that let one company&#8217;s proprietary model be trusted and compared against another&#8217;s, to name just a few, are all expensive. So these high-cost, non-excludable goods are textbook examples of public goods, and that is the textbook reason the private market underinvests in them.</p><p>Venture capital will fund the proprietary slice, deep maybe, but never wide (prove me wrong), and they won&#8217;t be public (again happy to think about PPPs and so on). That is exactly what the BVP map shows, and it is exactly the right thing for venture to fund. What venture will not fund is the reference layer that makes all those proprietary slices interoperable, comparable, and trustworthy. Markets do not build commons. They build ON them. The pre-competitive data layer is not a gap in the BVP thesis by accident. It is a gap created by economic necessity, and it is the bedrock where a non-market mandate must operate.</p><h2>The strongest argument against all of this</h2><p>I owe the reader the best version of the counter-argument, including the version that wounds my own position rather than BVP&#8217;s.</p><p>The strongest objection is that the binding constraint is neither models nor discovery-stage data. It is the clinical wall. Nearly 90% of drugs that enter clinical trials fail &#8212; and that figure is a clinical-stage number, not an all-pipeline one. The failures focus on efficacy and toxicity that only manifest in humans. Insilico has reached Phase IIa, not approval. If most of the cost and most of the failure live in animal and human trials, then a faster, richer data loop compresses the cheap front third of the funnel while leaving the expensive back two-thirds untouched. On this view, both &#8220;better models&#8221; and &#8220;better data&#8221; optimize the part of the process that is neither the source of value nor the cause of failure.</p><p>This is a serious critique, and I take it seriously. My response is that it wounds the <em>drug-discovery framing</em> far more than it wounds the data-factory thesis, and that it is, in fact, an argument for the broader scope rather than against it.</p><p>If the clinical wall is the problem, the right response is to generate data that <em>predict</em> clinical failure earlier: human-relevant immunogenicity, ADMET, and perturbation responses in systems that actually resemble the patient. Those are harder projections, at higher organizational layers, exactly the ones the current data ignores. And it is an argument for all the applications of biological understanding that have no clinical trial at all (surveillance, manufacturing, environmental biology), where the clinical wall does not exist, and the data gap is just as real.</p><p>If purely computational discovery is real, an AI-native drug, from an AI-selected target, an AI-designed molecule, through an AI-designed clinical trial, likely will clear Phase III in the next few years. If by 2029 none have, we haven&#8217;t learned from all our data. It might be that no entity can generate enough private data to crack that nut.</p><p>None of that weakens the point for national-level Data Factories.</p><h2>What happens next, and what would I build?</h2><p>Three predictions, offered with a confidence modifier.</p><p>First, the frontier labs will NOT vertically integrate into data generation. This makes the commons <em>less</em> likely to be built privately, not more, because the most capable players like OAI and Anthropic will be acquiring data rather than generating it themselves. That increases the value for the public case; it does not lower it.</p><p>Second, the proprietary-data licensing market BVP describes will get VERY hot and then get (more) crowded, and the differentiator will migrate from &#8220;we have proprietary data&#8221; to &#8220;we measure the right projections of the right layer for your model (maybe it&#8217;s already there, though I don&#8217;t think so).&#8221;</p><p>The companies that internalize the Anderson point will pull away from the ones still counting modalities.</p><p>Third, the pre-competitive layer will not build itself, and if it is going to exist this decade, it will require the same kind of public, mission-driven investment that built the PDB and the Human Genome Project.</p><p>That is the part I am working on, and it is the part the venture thesis structurally cannot and should not reach.</p><p>So I will end one layer out from where the BVP thesis lands.</p><p>The data is the moat.</p><p>They are right about that. But the deepest, most valuable, and least defensible part of that moat is exactly the part the market will not fund. That part is a national capability, not a venture return.</p><p>Right thesis, through the drug-development lens. The wider scope is not a correction to their work. It is a proper division of labor.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://michaelkoeris.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[On Biological Data Generation (1/n): More Is Different, and So Is the Data]]></title><description><![CDATA[A biological data factory should not be defined by volume. It should be defined by whether it makes the right variables measurable at the right layer.]]></description><link>https://michaelkoeris.substack.com/p/on-biological-data-generation-1n</link><guid isPermaLink="false">https://michaelkoeris.substack.com/p/on-biological-data-generation-1n</guid><dc:creator><![CDATA[Michael Koeris]]></dc:creator><pubDate>Mon, 04 May 2026 01:53:50 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tBNY!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F364420d5-dc01-434d-a135-febb6c24dff2_4000x4000.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I have been trying to get more precise about what we mean when we say &#8220;biology.&#8221;</p><p>Or, more precisely, soft condensed matter self-replication systems. Same thing, but more fun.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://michaelkoeris.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>I have also been trying to get more precise about what we mean by &#8220;data factories.&#8221;</p><p>&#8220;More biological data&#8221; is true. It is also an unsatisfying, easy, intellectually lazy answer.</p><ul><li><p>More of what?</p></li><li><p>Measured at which layer?</p></li><li><p>At what time resolution?</p></li><li><p>With what perturbations?</p></li><li><p>Against what notion of state?</p></li></ul><p>The default answer is usually: whichever instruments we have available.</p><p>That is not good enough.</p><p>We undersample habitually. We undersample because of technical constraints in the soft condensed matter sciences. We quickly neck down to something human-manageable. And we often treat &#8220;state&#8221; as an afterthought.</p><p>One paper I keep coming back to is Philip Anderson&#8217;s 1972 essay, <strong>&#8220;More Is Different.&#8221; </strong>(https://fermatslibrary.com/p/146f107c for the annotated version). I do not read Anderson as making an anti-reductionist argument. The lower-level physics still matters. Molecules still obey physical law. Cells are built from molecules. Tissues are built from cells. Turtles all the way up.</p><p>But Anderson&#8217;s deeper point is that, as systems become organized, new variables become useful and sometimes necessary. Knowing the parts does not automatically give you the right description of the whole.</p><p>That feels like the right starting point for thinking about biological data factories.</p><h2><strong>The hierarchy is obvious. Measurement is not.</strong></h2><p>Biology is organized across layers:</p><ul><li><p>Molecular</p></li><li><p>Subcellular</p></li><li><p>Cellular</p></li><li><p>Multicellular</p></li><li><p>Organ</p></li><li><p>Organism</p></li></ul><p>That hierarchy sounds obvious. From a measurement perspective, it is not at all obvious. What are we measuring in each layer? How? How much? What for?</p><ul><li><p>At the molecular layer, we may care about chemical identity, binding, conformation, energy landscapes, reaction dynamics, and modifications.</p></li><li><p>At the subcellular layer, the question shifts toward organization: organelles, membranes, ribosomes, vesicles, trafficking, cytoskeleton, and localization.</p></li><li><p>At the cellular layer, we start talking about &#8220;cell state.&#8221;</p></li></ul><p>That phrase is doing a lot of compression.</p><p>A cell&#8217;s internal state includes transcripts, proteins, post-translational modifications, metabolites, chromatin, organelles, energy state, stress state, membrane composition, recent history, and a lot of things we do not routinely measure.</p><p>Single-cell RNA-seq gives us one projection of that state. It does not give us the state itself.</p><p>Maybe if we measured everything, we would have a better approximation of cell state. But would that tell us how two cells behave together?</p><p>Not yet.</p><h2><strong>N &#8805; 2 changes the data problem.</strong></h2><p>The world outside a cell does not see the cell&#8217;s full internal state.</p><p>Neighboring cells do not inspect each other&#8217;s transcriptomes, proteomes, metabolomes, chromatin states, organelle states, and histories.</p><p>They see a different information set:</p><ul><li><p>Receptors</p></li><li><p>Ligands</p></li><li><p>Glycans</p></li><li><p>Secreted factors</p></li><li><p>Metabolites</p></li><li><p>Ions</p></li><li><p>Mechanical forces</p></li><li><p>Electrical cues</p></li><li><p>Extracellular matrix remodeling</p></li><li><p>Physical contact</p></li><li><p>Other things I am probably missing</p></li></ul><p>A useful way to think about a cell is as a high-dimensional hidden state with a lower-dimensional exchange surface.</p><p>Each cell contains much more information than it can realistically export. The channel capacity is much lower than the internal state space.</p><p>And what it exports is not a neutral summary. It is a compressed, context-dependent projection of itself.</p><p>That creates an information problem:</p><ul><li><p>How much of the internal state of a cell is visible to its neighbors?</p></li><li><p>How does that visibility change?</p></li><li><p>How much depends on timing, spatial position, and prior history?</p></li><li><p>How much is lost because the interface is limited?</p></li><li><p>How much is inferred by the receiver?</p></li></ul><p>The same signal can mean proliferation, migration, activation, exhaustion, tolerance, or death depending on the receiving cell and its local context.</p><p>So, at the multicellular layer (N &#8805; 2), we are no longer just measuring the interiors of many cells.</p><p>We are measuring partially hidden systems exchanging compressed signals over space and time.</p><p>That is a different object.</p><h2><strong>The multicellular state is not the sum of cell states.</strong></h2><p>Everyone agrees that: <code>multicellular state &#8800; sum(cell states)</code></p><p>But that does not get us very far.</p><p>A better approximation might be: <code>multicellular state &#8776; cell states + interfaces + topology + time</code></p><p>That is where Anderson&#8217;s argument comes back.</p><p>If each layer has its own useful variables, then measuring the lower layer more thoroughly may still not yield the right data for the layer above.</p><p>Bulk genomic sequencing was, in some sense, the easy case of biological data extraction. DNA gave us something unusually friendly to industrialization: a stable molecule, a mostly linear alphabet, amplification, and a scalable readout.</p><p>Most of biology is not like that.</p><p>Most of biology is hidden state, partial observability, compressed exchange, spatial organization, and feedback.</p><h2><strong>Data factories should be layer-appropriate.</strong></h2><p>A biological data factory should not be defined as a machine that generates more data.</p><p>It should be defined as a machine that generates layer-appropriate data.</p><p>That means the design has to change layer by layer.</p><ul><li><p><strong>Molecular layer:</strong> capture fast physical states at massive scale.</p></li><li><p><strong>Subcellular layer:</strong> capture spatial and temporal organization, probably with some intelligent coarse-graining.</p></li><li><p><strong>Cellular layer:</strong> measure internal state plus mass transfer to and through the surface.</p></li><li><p><strong>Multicellular layer:</strong> measure exchange, topology, perturbation, and time.</p></li><li><p><strong>Organ layer:</strong> capture function, flow, architecture, innervation, perfusion, and control.</p></li><li><p><strong>Organism layer:</strong> capture longitudinal history.</p></li></ul><p>Those are not the same data problem.</p><p>They are different observability problems.</p><p>A rough mental model I am playing with is:</p><p><code>biological bandwidth &#8776; entities &#215; state variables &#215; exchange variables / time constant</code></p><p>Low in the hierarchy, biology can be fast, physical, and high-dimensional.</p><p>High in the hierarchy, biology can be slower, history-dependent, and difficult to perturb.</p><p>In the middle, where cells become tissues, the problem seems especially hard because internal state, exposed interface, spatial topology, and time all matter at once.</p><p>That middle layer may be where many AI-for-biology conversations are still underdeveloped. The model can only learn the variables that the measurement system exposes. Yes, there may be an inference possible. I do not think we are there yet.</p><p><em><strong>Change my mind.</strong></em></p><h2><strong>The engineering question</strong></h2><p>The deeper engineering question is: <strong>What would we have to build so that the right variables become measurable?</strong></p><p>A purpose-built biological data factory would turn some layer of biology into training data by selecting perturbations, measuring states, capturing exchanges, preserving spatial structure, tracking time, and using what was learned to decide what to measure next.</p><p>This is only Level 1 of the argument.</p><p>But I think Anderson gives us the right starting point.</p><p>Before asking how much biological data we need, we should ask which layer we are trying to understand.</p><p>Because more is different.</p><p><strong>Reference:</strong> Philip W. Anderson, &#8220;More Is Different: Broken Symmetry and the Nature of the Hierarchical Structure of Science,&#8221; <em>Science</em>, 1972.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://michaelkoeris.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>