Part Two · The Stack · Layer 7 of 9
The Learn Layer — Data, Software, and the Scarcest Asset in the Lab
Every prior layer produces one thing above all else: data. The instruments measure, the sequencers read, the assays report, and every one of those events throws off a number that has to be captured, stored, made sense of, and fed back into the next design. The Learn layer is what closes the Design-Make-Test-Learn loop. It is the software and data infrastructure that turns a pile of raw instrument output into structured, reusable signal a model can actually train on. It is also, for reasons that take a moment to explain, the layer where an ordinary-sounding problem, "just store the results", hides one of the most defensible moats in the entire stack.
The reason is simple to state and hard to solve. An AI model is only as good as the data it learns from, and biological data is uniquely hostile to being learned from. Fixing that hostility is the business of this layer.
The Notebook and the Database: ELN and LIMS
Start with the two pieces of software every organized lab runs on. An Electronic Lab Notebook (ELN) is the digital replacement for the paper notebook a scientist used to scribble in: it records what experiment was run, why, by whom, with what materials, and what happened. It is the record of intent and observation. A Laboratory Information Management System (LIMS) is the operational database underneath: it tracks samples, reagents, instruments, and workflows, where every physical thing is, what state it's in, and what has been done to it. Loosely, the ELN captures the science and the LIMS captures the logistics, though modern platforms increasingly blur the line.
These sound like back-office tools, and for decades they were treated that way, necessary plumbing, bought once, rarely thought about. What changed is that the data these systems hold went from being a compliance archive to being the raw material of a model. The moment the goal became "train an AI on our accumulated experiments," the quality, structure, and connectedness of what the ELN and LIMS captured stopped being a filing question and became the whole ballgame.
The Real Problem: Instruments Don't Speak the Same Language
A single lab might run a mass spectrometer from one vendor, a sequencer from another, a plate reader from a third, and a dozen other instruments besides. Each one emits its results in its own format: a proprietary binary file, a bespoke spreadsheet layout, a vendor-specific export that means something only to that vendor's software. There are hundreds of these formats, none of them designed to talk to the others, many of them undocumented.
To an AI model, this is chaos. You cannot train on a folder of incompatible files. Before any learning can happen, all of that output has to be harmonized, pulled out of its native formats, translated into a common structure, aligned so that "concentration" means the same thing whether it came from instrument A or instrument B, and linked back to the experiment that produced it. This is unglamorous, enormously labor-intensive work, and it is the single biggest reason most pharma data sits unusable in silos. Solving it is a genuine technical and commercial franchise.
Two more ideas make the harmonized data actually useful. The first is the ontology: a controlled, agreed-upon vocabulary so that a gene, a cell type, an assay, or a unit is named consistently everywhere, rather than a dozen labs each inventing their own shorthand. The second is the set of principles known as FAIR data, Findable, Accessible, Interoperable, Reusable, a checklist for whether data can actually be located and reused later rather than being write-once and forget. Underpinning both is provenance, or lineage: the unbroken record of where a data point came from, what was done to it, and by whom. A model trained on data of unknown provenance is a model you cannot trust or defend, and in this field trust is regulatory.
Why a Generic Data Warehouse Isn't Enough
An obvious objection: the rest of the economy solved big-data storage years ago with cloud data lakes and warehouses. Why can't a lab just point one of those at its instruments and be done? The answer is that a generic data platform supplies raw storage and compute but knows nothing about biology. It does not understand that a sequence and its annotations belong together, that a plate has a spatial layout that matters, that a sample has a lineage of parents and children, or that certain modifications must be logged to satisfy a regulator. The horizontal tools give you a warehouse; they do not give you the biology-aware schema, the instrument connectors, or the domain ontologies that make the data trainable.
And then there is the regulation. Drug development runs under legal frameworks with unforgiving requirements: GxP (the family of "good practice" standards for lab, clinical, and manufacturing work) and, in the United States, 21 CFR Part 11, which governs how electronic records and electronic signatures must be handled to be legally valid. In practice this means audit trails, access controls, validated systems, and the ability to prove that a record was not altered. You cannot drop a generic consumer data tool into a regulated lab and satisfy an FDA inspector. The compliance layer is part of the product, not an afterthought, and building it is a barrier to entry all by itself.
Put the pieces together and the moat comes into focus. Whoever becomes the system of record for a lab's clean, harmonized, compliant, machine-readable experimental data controls the substrate every downstream model must train on. Switching that system of record means re-mapping every instrument, re-validating every workflow, and migrating years of provenance, a cost so high that incumbency, once established, is extremely durable. The data is stickier than the software, and the software is stickier than the hardware.
Orchestration: The Software That Actually Runs the Loop
One category in this layer does more than record; it acts. Lab orchestration software is the control system that schedules instruments, sequences the steps of a workflow, and, in the most advanced form, closes the Design-Make-Test-Learn loop automatically, deciding what to run next based on what the last run returned. This is the connective tissue between the "hands" (the robots) and the "brain" (the models). It is also where design-of-experiments logic lives: the statistical machinery for choosing which experiments to run so that each one is maximally informative. Orchestration is the software that turns a room full of instruments into a self-driving lab.
Who Holds the Data Layer
The uncomfortable truth for a public-market investor is that the best assets in this layer are private, and the market clearly knows how valuable they are, because the acquisitions are large.

- Benchling (private) is the category-defining, biology-native cloud platform, an ELN, registry, and data system built from the ground up for life science rather than retrofitted from generic tools. It is the software of record at a very large share of biotechs, reportedly valued around a $6 billion mark. Its moat is exactly the switching cost described above: once a company's science lives in Benchling, leaving is close to unthinkable. It is the clearest example of the "system of record" thesis, and it cannot be bought on any exchange.
- Dotmatics was acquired by Siemens for roughly $5.1 billion (per public reporting), an industrial software giant paying up to own a multimodal R&D data platform. That price tag is the single loudest signal in the layer: a strategic acquirer valuing lab R&D data infrastructure like a crown jewel, not a utility.
- Revvity Signals, inside Revvity (RVTY), is one of the few ways to own an ELN-and-analytics franchise on a public exchange, pairing a notebook with the Spotfire analytics layer. STARLIMS, inside Dassault Systèmes (DSY), gives the French engineering-software giant a regulated-lab LIMS position. Between them they are the readable public proxies for the informatics layer, though each is a slice of a much larger company.
- LabWare and Sapio Sciences (both private) round out the LIMS/ELN incumbents that pharma actually runs.
- On the harmonization frontier, TetraScience (private) is the flagship independent, its whole business is pulling instrument data out of its hundreds of native formats and turning it into an AI-ready "scientific data cloud." Ganymede (a "Lab-as-Code" pipeline company, acquired by Apprentice.io) and Code Ocean (reproducible computational pipelines) attack adjacent pieces of the same problem.
- Databricks (private) and Snowflake (SNOW) are the horizontal data platforms many of these biology-native tools are built on top of, the generic substrate that supplies storage and compute while the vertical players supply the biology.
- In orchestration, Artificial (private) builds whole-lab orchestration that connects wet-and-dry instruments to AI agents; Synthace (private) encodes design-of-experiments into executable protocols; Retisoft (private) does the scheduling. And NVIDIA (NVDA) reappears here too, as the model-serving layer (BioNeMo) that runs the models the data is fed into, a reminder that the same toll sits under this layer as under Design.
What It Means for a Portfolio
This is probably the most under-appreciated layer in the entire primer, and the frustration for an investor is that under-appreciation and inaccessibility travel together. The strategic prize, the harmonized, compliant, system-of-record dataset in the middle of the loop, is exactly what incumbents are paying billions to control, and yet the purest expressions of it (Benchling, TetraScience, Artificial) are private. Public exposure is real but diluted: Revvity Signals inside RVTY, STARLIMS inside Dassault, and the model-serving toll inside NVIDIA. The pattern to watch is the same one that produced the Dotmatics deal: strategics acquiring the data layer, because whoever owns the clean data at the center of the loop owns the substrate every model in every other layer has to learn from. In a field where the model is the commodity and the measurement is the moat, the organized memory of what has already been measured is the quiet compounding asset underneath both.
↗ Explore the Learn layer in the interactive map — every company in this layer, public and private, in one view.© BEP Holdings · Ben Pouladian. Research and commentary, not investment or medical advice.