We’ve made designing molecules easier and cheaper than ever, but, as I wrote about here, proving those designs work in humans is increasingly becoming a bottleneck. There are a lot of reasons for this, such as the constraints around chemical synthesis, the cost and duration of clinical trials, and how little we actually know about what happens in a lab.
Today, scientists keep notebooks and publish papers, and other scientists work from those records. But most of what makes an experiment succeed (or fail) is never actually written down.
Nowhere is that more expensive than at the handoff, when a working process has to move from the lab that developed it to a contract manufacturer, often in another country. “That is where things get screwed up,” Anna Marie Wagner told me the other week. Tech transfer is one of the most common ways manufacturing problems derail a drug launch; and manufacturing is now the single biggest cause of delays, at 64% in 2024.
Anna Marie has spent her career moving between biology and technology. She studied molecular biology, spent time as a tech investor, and returned to the sciences in 2019 unable to shake a specific discomfort: two fields she knew well, science and software, ran on entirely different economics. She’s now the co-founder of Transfyr, which builds observability infrastructure for scientific execution. The company came out of stealth last week, with The New York Times covering the launch. “The variation we see even among well-trained scientists is pretty jaw-dropping,” Anna Marie told the Times.
“In software, I can write some code, and if it’s useful, I can ship it to millions of people instantaneously, basically at no cost. But transferring a new technique to a single other lab can take months of debugging between highly trained people,” she told me. “We call that tech transfer, which is effectively the distribution cost.”
Up until now, most advancements in AI for science have focused on the parts of science that were already digitizable, such as sequences and molecular structures. And while that’s great progress, it’s not addressing the physical problems that determine whether a drug succeeds in a human and whether it can then be manufactured at scale. And it’s these problems that continue to hold back much of the industry and inhibit innovations from getting out of the lab. As Anna Marie put it, “Anything that’s getting in the way of positive scientific discoveries reaching the world is impacting everyone. It’s climate, it’s food security, it is human health, animal health. This is a global issue.”
Noise in Science
If you visit Transfyr’s lab, you’ll have the opportunity to participate in a simple exercise: a “serial dilution,” which requires diluting a solution of known concentration by a specified amount through a series of steps, in duplicate. Each participant is then scored on accuracy (how close to the target concentration they landed) and precision (how close the two samples are to each other). “We consider the ‘strike zone’ to be within 5% of the target on both accuracy and precision. The vast majority of people do not hit the strike zone. Even experienced scientists.”
The reason for this miss depends on a lot of factors. Did they pre-wet the pipette or back-pipette? Did they set their pipette to the correct volume at every step? How thoroughly did they mix the sample? Did a stray bubble get in? Did they accidentally foam the mixture? None of this is actually in the written protocol, but it all shows up in the result.
Most protocols are actually thousands of small, undocumented decisions. At scale, this invisible variation shows up in the data used to train AI models. As Anna Marie put it, “We would get data sets where the model was better at recognizing what lab did an experiment than what the experiment was designed to measure in the first place. If that data is just training on unwritten techniques we don’t even observe, is it actually teaching the model anything useful about the biology itself?”

There are surprisingly simple, undocumented ways that this lack of standardization shows up. Anna Marie gave an example of a scientist who pre-labeled all of her tubes, which, along with several other shortcuts, helped her get her protocol done two hours faster. And those two hours made a big impact on the experiment’s outcome because it turns out RNA sitting at room temperature degrades, and two hours is not a trivial amount of time. But pre-labeling tubes would never show up in a normal protocol.
“I worry that sometimes we chalk up noise in biology to ‘biology is noisy’,” Anna Marie told me. “There’s also a lot of process variability that we just accept without understanding. That’s the problem. You can accept variability, but we need to understand it so we can interpret our data under that lens.”
In other words, the problem is mistaking “we don’t know what happened” for a fact about biology, rather than a fact about the lack of tooling and observability that exists today in science.
The Minimum Viable Film Room
So then what’s actually worth measuring? There’s no shortage of variables that could theoretically matter, but Anna Marie has narrowed down the list to three that do most of the work: the operator’s actions and intent, supply chain, and then the environment. As she puts it, “Those few things comprise the minimum viable film room.”
Operator actions and intent means, first and most basically, understanding what a person (or a machine) actually did. This isn’t necessarily just visual observation, because a huge amount of what you’re working with in the lab looks identical even though it isn’t. “I’m looking at a clear liquid, and that clear liquid could be water or it could be hydrochloric acid. There’s a big difference between those two things, and you cannot tell the difference between them visually,” Anna Marie explained.
Supply chain is around the inputs, not just what a material was, but its lot number, its expiration date, who else touched it, what might have contaminated it.
The environment is about monitoring things like temperature, humidity, and CO2. Some of this might seem obvious and yet, for most of the industry today, it’s missing.
In Anna Marie’s opinion, automation around these three levers is often misunderstood. The field has sold automation on the ideas of scale and cost savings, she points out, but neither holds up especially well in practice. Utilization at automated labs across the industry has stayed frustratingly low, and scientists tend to value flexibility over throughput in ways robots can’t yet match (though I suspect this will change soon!).
However, what automation does deliver, in her view, is observability rather than scale. In her essay “On Observability,” Anna Marie makes the case using the example of self-driving cars: autonomous vehicles didn’t improve because of better maps or smarter models, but because cars were instrumented with cameras, lidar, and GPS to learn how humans actually navigated the road, years before autonomy made economic sense. In DARPA’s first driverless-car challenge in 2004, the best vehicle made it just over seven miles into a 142-mile course; a year later, after a year of watching and iterating in public, five vehicles finished the whole thing. Observability enabled autonomy.
Above is a video from a head-mounted camera recording an experiment, along with three cameras mounted above the lab bench. Source.
Lab automation works the same way: automated systems tend to log what they intended to do and what they actually did, almost as a side effect of being automated. “I would never say that autonomy is the end goal,” she told me. “It’s a tool just like any other, and that tool should be wielded by really smart, caring humans who are trying to make a difference in the world.”
Why Science Hasn’t Built Its Film Room
So why don’t we have film rooms for science already? “We have deep observability in sports, and it is so valued by the athletes. Its primary purpose is not evaluative. It is coaching and training and self-improvement,” she said. So why are we not giving elite scientists the same tools that we give elite athletes to improve their game?
Part of the answer is infrastructure. But the other part is cultural. Science still has a layer of secrecy around it. If you think you have, as she puts it, “the magic hands,” you may not be in a hurry to let everyone else have them too.
Publication incentives don’t help, either. Observability expands what’s visible, and science has not made the same peace with visible failure that sports has. “Are we incentivized in any way, shape, or form to publish comprehensive results as opposed to pretty or clean results? No, I don’t think so,” Anna Marie said.
Amber Liu made a version of this point to me a few weeks ago about AI research rather than the wet lab: workshops encouraging scientists to publish their failures exist, but adoption is slow. Amber’s observation was that the reluctance is specifically human. An AI scientist, as she put it, doesn’t carry “this burden of disclosing that they’re actually doing a lot of dumb things, that they fail a lot in the middle.”
The infrastructure that makes science observable is also the infrastructure that makes an individual scientist’s workflows and mistakes visible and public. Until the incentives around that visibility change, what’s technically solvable might not matter.
Systems Integration
Toward the end of our conversation, I asked Anna Marie what she thinks the field could look like in the next five to ten years. Her answer was more about the scientific system as a whole rather than any single breakthrough.
“There are millions of scientists who are having ideas right now all around the world. And none of them have a good way to communicate. They don’t have the language for it,” she explained.
That, to me, is the larger promise of observability in science. Beyond just cleaner data or better protocols (which matter!), it’s the possibility that knowledge from one lab can be easily transferred so that someone else, somewhere else can use it too.
“I go back to systems integration. Such a boring term, but I love it,” she joked. “Because it is the root of most advanced industries. Can we bring together components to advance products that actually solve real problems in the world, on a completely different timeline than we’re doing today?”
Systems integration is about allowing the components of the physical work of science to talk to each other. Right now, too much knowledge is siloed and unproductive, because nothing exists to carry it from where it was learned to where it is needed.
Science still hasn’t solved distribution. And until it does, better models will continue helping us imagine what to build, but they won’t tell us how to carry that knowledge into the physical world where it can actually be used.
Author’s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.


