Part Two · The Stack · Layer 4 of 9
The Readout Layer: Reading Biology Directly
Everything downstream in drug discovery depends on one prior question: can we actually read what a cell is doing? Before a model can predict a binding pocket or a target can be validated, biology has to be converted into data. That conversion is a physical measurement problem, and the companies that own the physics of measurement occupy the most defensible position in the stack. This is the readout layer. It turns molecules into numbers.
The readout layer splits cleanly into two worlds. Reading DNA and RNA is a solved, industrialized problem where the competition is over cost and accuracy. Reading proteins is not solved, and that gap is where most of the interesting investment turbulence lives. Understand why the second world is harder than the first and you understand the entire chapter.
How DNA Sequencing Actually Works
DNA is a chain of four bases: A, C, G, T. Sequencing is the act of determining their order. The dominant method, used by Illumina (ILMN), is called sequencing-by-synthesis. The chemistry is elegant. You take a strand of DNA, fragment it into short pieces, and anchor millions of those fragments to a glass flow cell. Then you rebuild the complementary strand one base at a time, but each incoming base carries a fluorescent tag and a chemical block that stops the reaction after a single addition. A camera photographs the flow cell, records which color lit up at each of the millions of spots, the block is cleaved off, and the cycle repeats. Read the colors across all cycles and you have read the sequence.
This approach is accurate and massively parallel, which is why Illumina's NovaSeq machines drove the cost of a human genome from roughly 100 million dollars in the mid-2000s toward the low hundreds today. The weakness is read length. Sequencing-by-synthesis works in short fragments, typically 100 to 300 bases, so the genome must be shattered and computationally reassembled like a shredded document. Repetitive or structurally complex regions are hard to reconstruct from short pieces.

That weakness created the opening for long-read sequencing, which comes in two flavors:
- PacBio (PACB) uses a method called HiFi. It reads a single DNA molecule in a tiny well, watching a polymerase incorporate fluorescently labeled bases in real time, and it reads the same circular molecule multiple times to average out errors. The result is long reads, often 15,000 to 20,000 bases, with accuracy that rivals short reads.
- Oxford Nanopore (ONT, London-listed) does something physically different. It threads a single DNA strand through a protein pore embedded in a membrane and measures the tiny electrical current disruption as each base passes through. Different bases block the current differently, so the sequence is read off the electrical signal. There is no synthesis, no camera, no fluorescence. This gives ultra-long reads and true portability, down to a device the size of a USB stick, at the cost of historically higher per-base error rates that have narrowed over successive chemistries.
The competitive axes in DNA sequencing are three: read length, accuracy, and cost per genome. Illumina optimized accuracy and cost. The long-read players optimized length. And the entire field is marching toward a psychological and economic threshold, the 100-to-150-dollar genome, below which population-scale sequencing becomes routine and the addressable market expands by an order of magnitude.
Single Cell and Spatial: Reading Context, Not Just Sequence
Sequencing the genome tells you the instruction set. It does not tell you which instructions a given cell is actually running. That is the job of RNA readout, and here two refinements matter enormously.
Traditional bulk sequencing grinds up a tissue sample and reads the average gene expression across all its cells. The problem is that averaging destroys information. A tumor sample might contain cancer cells, immune cells, and healthy tissue, and the bulk average tells you nothing about which population is doing what. Single-cell sequencing solves this by partitioning a sample so that each cell's RNA is barcoded individually before sequencing. You get one expression profile per cell, which is how researchers discover rare cell types and map how a drug hits some cells and not others. 10x Genomics (TXG) built the dominant droplet-based platform for this, encapsulating each cell in an oil droplet with a uniquely barcoded bead.
The frontier beyond single cell is spatial transcriptomics: reading which genes are expressed while preserving where in the tissue each measurement came from. Location is biology. An immune cell at the edge of a tumor behaves differently from one at its core, and the moment you dissociate the tissue into a suspension you lose that geometry. Spatial methods read expression directly on an intact tissue slice, keeping the coordinate map intact. 10x competes here too, alongside Parse Biosciences (a combinatorial-barcoding approach that avoids expensive instruments, now owned by QIAGEN). QIAGEN (QGEN) itself sits upstream of all of this in sample preparation, the unglamorous but essential step of extracting clean nucleic acid before any sequencer ever sees it.
Why Proteins Are Harder Than DNA
DNA is easy to read at scale because of one trick: amplification. The polymerase chain reaction, PCR, can take a single DNA molecule and copy it billions of times, so even a vanishingly small starting sample becomes a strong, readable signal. There is no PCR for proteins. You cannot photocopy a protein directly; the workarounds all amplify a DNA tag that stands in for the protein rather than the protein itself, which is exactly what the aptamer and proximity-assay readouts described later in this chapter do. Whatever protein is in the sample is all the protein signal you will ever get.
That matters because proteins are the actual machines of biology and the direct target of most drugs. The genome is the blueprint; the proteome is the running system. But proteins are built from twenty amino acids rather than four bases, they fold into three-dimensional shapes, and they get chemically modified after they are made. There is no natural copying enzyme, no four-letter alphabet, and a dynamic range in the blood of roughly ten orders of magnitude between the most and least abundant proteins. Reading the proteome is a genuinely unsolved measurement problem, and the field has split into competing physical approaches, none yet dominant:
- Aptamer-based (the technology inside SomaLogic): short synthetic DNA strands are engineered to fold into shapes that grab specific proteins. Because the binder is itself DNA, it can be counted using DNA-reading tools. High multiplexing, thousands of proteins at once.
- Proximity Extension Assay, PEA (Olink's method): each target protein is tagged by a pair of antibodies, each carrying a short DNA tag; only when both antibodies bind the same protein do the tags come close enough to be joined and amplified. Requiring two independent binding events sharply cuts false positives.
- Mass spectrometry: proteins are broken into peptide fragments and sorted by mass and charge. It identifies proteins without needing a pre-made binder for each one, but sensitivity for low-abundance proteins has been the historical constraint.
- Single-molecule protein sequencing: the aspiration to read a protein amino acid by amino acid, the way we read DNA base by base. Still early, technically brutal, and the biggest prize if it works.
Reading the Field
Start with the incumbent. Illumina (ILMN) is the sequencing-by-synthesis oligopoly, and for a decade its NovaSeq installed base functioned like a toll road on genomics. It is now facing a genuine three-front attack: long-read platforms taking the high-accuracy end, low-cost challengers undercutting the middle, and a re-entering incumbent-class competitor at the top. Illumina's response has been to expand its readout surface, acquiring SomaLogic to move into proteomics and, per public reporting, absorbing PacBio short-read intellectual property. The strategic question for ILMN is whether an incumbent can defend a cost curve while simultaneously buying its way into the harder proteomics frontier.
The long-read specialists, PacBio (PACB) and Oxford Nanopore, are pure-play bets on read length and, increasingly, on real-time and point-of-care use where Nanopore's portability is unmatched. Both are smaller, single-platform names whose fortunes track adoption of long-read as the new clinical standard rather than a research luxury.
The most consequential recent development is Roche re-entering sequencing with its Axelios platform built on SBX (sequencing by expansion) nanopore chemistry, reportedly targeting a 150-dollar genome. Roche matters because it is the first challenger with incumbent-class scale, capital, and clinical channel to threaten Illumina directly. A credible Roche at 150 dollars reprices the entire oligopoly.
Below the incumbents sit the cost disruptors, all attacking cost per genome from different physics:
- Ultima Genomics (private): a wafer-based, open-substrate architecture explicitly built to drive genome cost toward the 100-dollar floor.
- Element Biosciences (private): a benchtop instrument using avidity chemistry, which boosts signal by having multiple binding sites grab a target simultaneously, aimed at labs priced out of NovaSeq-class capital equipment.
- MGI Tech (688114, Shenzhen-listed): the sequencing arm of the BGI complex, using DNBSEQ nanoball chemistry. MGI is the price leader globally and the most disruptive force on cost, though US-listed investors should flag the geopolitical overhang, including the BIOSECURE-style scrutiny and the Department of Defense's 1260H list targeting Chinese biotech vendors.
In single-cell and spatial, 10x Genomics (TXG) is the category-defining public name, with the same razor-and-blade instrument-plus-consumables model that made Illumina durable, though facing its own patent litigation and lower-cost entrants. Parse Biosciences inside QIAGEN (QGEN) and QIAGEN's own upstream sample-prep franchise round out the read-through: whoever prepares the sample sells a consumable regardless of which sequencer wins.
Proteomics is where the differentiated readout physics still lives inside single-platform public equities, and it is the most open field in the chapter. Olink (PEA) now sits inside Thermo Fisher and SomaLogic (aptamer) inside Illumina, so the two most mature multiplexed platforms have been absorbed by giants. What remains investable as standalone physics bets:
- Seer (SEER): a nanoparticle-based front-end that tames the proteome's enormous dynamic range before mass spectrometry, effectively a sample-prep layer that makes existing mass specs see deeper.
- Nautilus Biosciences (NAUT): pursuing single-molecule protein analysis, a direct swing at the hardest version of the problem.
- Quantum-Si (QSI): building a semiconductor chip for single-molecule protein sequencing, betting that protein reading follows DNA in migrating from big instruments to cheap silicon.
- Alamar Biosciences (ALMR): its NULISA method layers a wash step onto proximity assays to strip background, chasing the ultrasensitive, low-abundance end where the most clinically interesting proteins hide.
The Investment Shape of the Layer
Two very different regimes coexist in the readout layer, and conflating them is the common analytical error. DNA sequencing is oligopoly economics under cost-curve attack. The physics is mature, the players are known, and the fight is over dollars per genome as Roche, the private cost disruptors, and MGI compress margins beneath Illumina. That is a pricing war, and pricing wars are hard on incumbent multiples even when volumes grow.
Proteomics and spatial are the opposite: a wide-open frontier where the underlying measurement physics is not settled and where differentiated approaches still sit inside single-platform public names rather than absorbed into instrument giants. That is where a technology edge can still translate into a durable franchise, and also where the failure rate will be higher. The generalizable point for a portfolio is that the readout layer offers more competitive turbulence, and therefore more dispersion of outcomes, than the placid instrument-giant narrative suggests. The molecules that are easiest to read are already commoditizing; the molecules that matter most for drugs are the ones we still cannot read well, and that unsolved measurement problem is precisely where the asymmetric returns are hiding.
↗ Explore the Readout layer in the interactive map — every company in this layer, public and private, in one view.© BEP Holdings · Ben Pouladian. Research and commentary, not investment or medical advice.