BEP RESEARCH The Future of Drug Discovery ↑

Part Two · The Stack · Layer 4 of 9

The Readout Layer: Reading Biology Directly

Everything downstream in drug discovery depends on one prior question: can we actually read what a cell is doing? Before a model can predict a binding pocket or a target can be validated, biology has to be converted into data. That conversion is a physical measurement problem, and the companies that own the physics of measurement occupy the most defensible position in the stack. This is the readout layer. It turns molecules into numbers.

The readout layer splits cleanly into two worlds. Reading DNA and RNA is a solved, industrialized problem where the competition is over cost and accuracy. Reading proteins is not solved, and that gap is where most of the interesting investment turbulence lives. Understand why the second world is harder than the first and you understand the entire chapter.

How DNA Sequencing Actually Works

DNA is a chain of four bases: A, C, G, T. Sequencing is the act of determining their order. The dominant method, used by Illumina (ILMN), is called sequencing-by-synthesis. The chemistry is elegant. You take a strand of DNA, fragment it into short pieces, and anchor millions of those fragments to a glass flow cell. Then you rebuild the complementary strand one base at a time, but each incoming base carries a fluorescent tag and a chemical block that stops the reaction after a single addition. A camera photographs the flow cell, records which color lit up at each of the millions of spots, the block is cleaved off, and the cycle repeats. Read the colors across all cycles and you have read the sequence.

This approach is accurate and massively parallel, which is why Illumina's NovaSeq machines drove the cost of a human genome from roughly 100 million dollars in the mid-2000s toward the low hundreds today. The weakness is read length. Sequencing-by-synthesis works in short fragments, typically 100 to 300 bases, so the genome must be shattered and computationally reassembled like a shredded document. Repetitive or structurally complex regions are hard to reconstruct from short pieces.

A log-scale chart of cost per genome falling far faster than Moore's Law since the mid-2000s, with a flat line showing protein readout has not fallen the same way.
Reading DNA collapsed in cost faster than compute; reading proteins did not, which is where the next measurement moat sits.

That weakness created the opening for long-read sequencing, which comes in two flavors:

The competitive axes in DNA sequencing are three: read length, accuracy, and cost per genome. Illumina optimized accuracy and cost. The long-read players optimized length. And the entire field is marching toward a psychological and economic threshold, the 100-to-150-dollar genome, below which population-scale sequencing becomes routine and the addressable market expands by an order of magnitude.

Single Cell and Spatial: Reading Context, Not Just Sequence

Sequencing the genome tells you the instruction set. It does not tell you which instructions a given cell is actually running. That is the job of RNA readout, and here two refinements matter enormously.

Traditional bulk sequencing grinds up a tissue sample and reads the average gene expression across all its cells. The problem is that averaging destroys information. A tumor sample might contain cancer cells, immune cells, and healthy tissue, and the bulk average tells you nothing about which population is doing what. Single-cell sequencing solves this by partitioning a sample so that each cell's RNA is barcoded individually before sequencing. You get one expression profile per cell, which is how researchers discover rare cell types and map how a drug hits some cells and not others. 10x Genomics (TXG) built the dominant droplet-based platform for this, encapsulating each cell in an oil droplet with a uniquely barcoded bead.

The frontier beyond single cell is spatial transcriptomics: reading which genes are expressed while preserving where in the tissue each measurement came from. Location is biology. An immune cell at the edge of a tumor behaves differently from one at its core, and the moment you dissociate the tissue into a suspension you lose that geometry. Spatial methods read expression directly on an intact tissue slice, keeping the coordinate map intact. 10x competes here too, alongside Parse Biosciences (a combinatorial-barcoding approach that avoids expensive instruments, now owned by QIAGEN). QIAGEN (QGEN) itself sits upstream of all of this in sample preparation, the unglamorous but essential step of extracting clean nucleic acid before any sequencer ever sees it.

Why Proteins Are Harder Than DNA

DNA is easy to read at scale because of one trick: amplification. The polymerase chain reaction, PCR, can take a single DNA molecule and copy it billions of times, so even a vanishingly small starting sample becomes a strong, readable signal. There is no PCR for proteins. You cannot photocopy a protein directly; the workarounds all amplify a DNA tag that stands in for the protein rather than the protein itself, which is exactly what the aptamer and proximity-assay readouts described later in this chapter do. Whatever protein is in the sample is all the protein signal you will ever get.

That matters because proteins are the actual machines of biology and the direct target of most drugs. The genome is the blueprint; the proteome is the running system. But proteins are built from twenty amino acids rather than four bases, they fold into three-dimensional shapes, and they get chemically modified after they are made. There is no natural copying enzyme, no four-letter alphabet, and a dynamic range in the blood of roughly ten orders of magnitude between the most and least abundant proteins. Reading the proteome is a genuinely unsolved measurement problem, and the field has split into competing physical approaches, none yet dominant:

Reading the Field

Start with the incumbent. Illumina (ILMN) is the sequencing-by-synthesis oligopoly, and for a decade its NovaSeq installed base functioned like a toll road on genomics. It is now facing a genuine three-front attack: long-read platforms taking the high-accuracy end, low-cost challengers undercutting the middle, and a re-entering incumbent-class competitor at the top. Illumina's response has been to expand its readout surface, acquiring SomaLogic to move into proteomics and, per public reporting, absorbing PacBio short-read intellectual property. The strategic question for ILMN is whether an incumbent can defend a cost curve while simultaneously buying its way into the harder proteomics frontier.

The long-read specialists, PacBio (PACB) and Oxford Nanopore, are pure-play bets on read length and, increasingly, on real-time and point-of-care use where Nanopore's portability is unmatched. Both are smaller, single-platform names whose fortunes track adoption of long-read as the new clinical standard rather than a research luxury.

The most consequential recent development is Roche re-entering sequencing with its Axelios platform built on SBX (sequencing by expansion) nanopore chemistry, reportedly targeting a 150-dollar genome. Roche matters because it is the first challenger with incumbent-class scale, capital, and clinical channel to threaten Illumina directly. A credible Roche at 150 dollars reprices the entire oligopoly.

Below the incumbents sit the cost disruptors, all attacking cost per genome from different physics:

In single-cell and spatial, 10x Genomics (TXG) is the category-defining public name, with the same razor-and-blade instrument-plus-consumables model that made Illumina durable, though facing its own patent litigation and lower-cost entrants. Parse Biosciences inside QIAGEN (QGEN) and QIAGEN's own upstream sample-prep franchise round out the read-through: whoever prepares the sample sells a consumable regardless of which sequencer wins.

Proteomics is where the differentiated readout physics still lives inside single-platform public equities, and it is the most open field in the chapter. Olink (PEA) now sits inside Thermo Fisher and SomaLogic (aptamer) inside Illumina, so the two most mature multiplexed platforms have been absorbed by giants. What remains investable as standalone physics bets:

The Investment Shape of the Layer

Two very different regimes coexist in the readout layer, and conflating them is the common analytical error. DNA sequencing is oligopoly economics under cost-curve attack. The physics is mature, the players are known, and the fight is over dollars per genome as Roche, the private cost disruptors, and MGI compress margins beneath Illumina. That is a pricing war, and pricing wars are hard on incumbent multiples even when volumes grow.

Proteomics and spatial are the opposite: a wide-open frontier where the underlying measurement physics is not settled and where differentiated approaches still sit inside single-platform public names rather than absorbed into instrument giants. That is where a technology edge can still translate into a durable franchise, and also where the failure rate will be higher. The generalizable point for a portfolio is that the readout layer offers more competitive turbulence, and therefore more dispersion of outcomes, than the placid instrument-giant narrative suggests. The molecules that are easiest to read are already commoditizing; the molecules that matter most for drugs are the ones we still cannot read well, and that unsolved measurement problem is precisely where the asymmetric returns are hiding.

↗ Explore the Readout layer in the interactive map — every company in this layer, public and private, in one view.
↓ Download the primer (PDF, free edition) ↓ Full edition incl. the basket (PDF, password required) ↓ The research framework (CLAUDE.md — drop it in a project folder and work the way this framework works) ↗ The interactive biolab map ≡ All chapters

© BEP Holdings · Ben Pouladian. Research and commentary, not investment or medical advice.