The first time I took a Waymo, I filmed it with my mouth gaping open. Watching the steering wheel turn on its own was like something out of an Isaac Asimov novel. Looking back, I think I was in awe at the mechanics of it, what was happening to the steering wheel and blinker and brake pedal without anyone in the driver seat. I thought about that moment again recently after speaking with Wenhao Gao.
Wenhao is a postdoctoral researcher at Stanford and an incoming professor of Chemical and Biomolecular Engineering at the University of Pennsylvania. His work focuses on connecting AI molecular design to the experiments needed to test those designs.
Wenhao’s analogy for lab automation is self-driving cars. “I don’t think moving the car mechanically is the tricky part,” he told me. “Before making a decision, it needs to know the environment around it: the traffic lights, the lanes, the pedestrians, the other vehicles.”
Wenhao sees a similar challenge in chemistry. Many routine operations, including mixing, heating, and dispensing, can already be automated. “You can put A in this position and B in that position and fix them there,” he explained. “You can even hard-code those positions into the robot.”
But after the reaction, you have a mixture in a vial and you have to decide what the next step should be. “This requires first knowing what actually happened inside the vial,” Wenhao said, “this part of interpretation is the fundamental of decision making and much harder to automate than anything around the equipment.”
In our conversation, we talked about where connections are breaking in labs today, what automation can help with, and why it’s useful to limit what a model is allowed to generate.
From molecular simulation to experiments
Wenhao was initially drawn to molecular simulation and the possibility of understanding a system by modeling it. But he recognized the cost and limitations of those simulations and in 2017, he read Automatic chemical design using a data-driven continuous representation of molecules by Rafael Gómez-Bombarelli and colleagues. In this paper, molecules were represented in a continuous mathematical space, which made it possible to search for new structures by moving through that space.

“I thought, this should be the future,” Wenhao said and he shifted his focus from molecular simulation toward generative molecular design.
During this period, he came to realize the importance of experimental validation. A model can propose a molecule and predict its properties, but you still have to validate it through an experiment. This is also how you ensure you’re improving future designs.
“If you have good verification, you can close the feedback loop,” he explained.
This brought him into the wet lab where he worked on automated synthesis and saw just how much automation depends on physical resources. “In academia, you typically expect that a couple of smart ideas can fundamentally change the behavior of a system,” he said. “But lab automation is very resource intensive.”
He gave the example of an LC-MS instrument. If you have one, you have one instrument’s worth of throughput. If you want more capacity, you need more instruments. Better software can help to a degree, but you still have to account for the time and equipment each experiment needs.
He ultimately came back to reaction modeling and molecular design, but with a much clearer understanding of what models will need to connect to in order to be valuable.
Where the feedback loop breaks
Early in our conversation, I asked Wenhao to walk me through the step-by-step process of going from a desired property to a molecule in the lab.
“We start from a property profile that we want to have,” he explained. In drug discovery, that means a combination of objectives, including activity, selectivity, and safety.
From there, the process has roughly four parts:
Design. Identify or generate candidate molecules with the properties you want.
Synthesis planning. Work backwards from a candidate structure to reactions and starting materials that could produce it. This is called retrosynthetic analysis.
Execution. Run the reactions, then separate and purify the products as needed.
Readout and testing. Establish what was made and measure whether it has the properties you wanted. Use those results to decide what to try next.
“If you provide the correct information, I think we can expect current language-model agents to do a reasonably good job on some synthesis-planning steps,” he said. “That part, I’m cautiously optimistic about.”
But the harder problem is the readout.
LC-MS helps separate components of a mixture and provides information about their masses. But these outputs still require some amount of interpretation. “We’re getting indirect information, and we need to infer the molecular structure from that information,” Wenhao explained. “That requires intelligence, but it is also still an open scientific problem.”
This connects to something I’ve been exploring and wrote about recently with Anna Marie Wagner here. Recording what happened in a lab gives us information but that information rarely makes it into a written protocol. And even with a complete record of what the equipment did, we still need to understand what happened chemically.
Checking whether a reaction yielded a specific molecule is generally an easier problem versus trying to identify an unknown molecule from scratch. This is because the problem is more constrained. Recent work such as MARLIN, a July 2026 preprint, explicitly tackles structure generation without being given the correct formula.
The interpretation of an experiment determines what the system does next. An incorrect identification can send the next round of experiments in the wrong direction. This is problematic for all labs, and especially so for autonomous labs.
Designing molecules with a route to the lab
In a 2020 paper with Connor Coley, The Synthesizability of Molecules Proposed by Generative Models, Wenhao looked at how often molecular design algorithms proposed structures that a synthesis-planning program could find routes to. For some optimization tasks, the program couldn’t find a route for any of the top 100 candidates.
“The core idea was that we wanted the model to learn synthesizable chemical space only,” Wenhao explained. The research highlights how a model can perform well against a computational design objective but still produce candidates that are difficult, even impossible to test.
SynFormer, developed by Wenhao, Shitong Luo and Coley, incorporates synthesis into generation. It takes into account a set of starting materials and a defined set of reactions and then generates pathways. Though a computational route still has to be experimentally validated, the goal is that the model has already accounted for a constraint that would otherwise appear after design.

The aim is “something we can get within three, maybe up to five steps,” he said. A molecule needs to be accessible within a reasonable experimental cycle to be considered as a design candidate.
I asked what you give up by imposing that constraint. His answer was synthetically complex molecules. Some interesting molecules will fall outside the reactions and building blocks the system can use, or require much more complicated synthesis. “There is a cost,” he said. “But overall, that’s a price I would want to pay.”
You can still generate new molecules within the space, which can expand as more reactions become more reliable. You are just limiting the ways they can be assembled.
What models learn from known chemistry
“A generative model, by design, learns the distribution of its training data,” Wenhao explained. “It’s hard to expect it to generalize to something very far away from that.”
But chemistry involves creating structures that have never existed before. Reading the literature gives a model useful knowledge about what’s been done, but it doesn’t establish that the model can reason reliably about a new structure. Its predictions still need to be tested, and a lab can only run so many experiments.
The literature also makes it difficult to learn from the experiments that have already happened. As I wrote about here, scientific papers were built for human readers. “They refer to ‘molecule 2,’ and there’s a structure drawn there,” Wenhao said. Understanding the result requires connecting that drawing to the text, the reaction conditions, and the measurements. For a model to learn from that work, the context around those connections have to exist.
This means what you test is really important, and a good model should help a lab choose which candidates are worth testing.
Wenhao’s Practical Molecular Optimization benchmark looked at one part of this around how efficiently algorithms find promising candidates. It compared 25 algorithms across 23 tasks, and gave each run a budget of 10,000 computational evaluations. Under this budget constraint, most newer methods failed to outperform their predecessors.
The result highlights why evaluation budget needs to be part of how we judge a molecular design algorithm. “If you generate 10,000 designs but none of them proceed to the wet lab, a lot of compute has been spent on that generation process.” Wenhao said. But you basically get nothing.
This adds to the argument I made here: as design becomes cheaper, more of the value depends on proving what works. In chemistry, that means proposing molecules a lab can make, choosing informative experiments, and capturing results (and failures!) in a form that improves the next design.
For me, that changes how we should evaluate these systems. Given the same lab capacity, does a model help us find a molecule with the properties we want in fewer experimental cycles? And does each cycle give us new information we can use again? Those are improvements that would make cheaper molecular design translate into faster discovery.
What an autonomous lab still needs
Toward the end of our conversation, I asked Wenhao what he envisions the end state will look like. He described “something like a 3D printer for chemistry: a small unit where we can design a molecule, synthesize it, and test it.”
To get there, he thinks some operations should stay manual because they are probably easier for humans to do. “I don’t think 100% automation is needed,” he said. Automation is useful in so far as it improves throughput and makes experiments more consistent, better documented, and easier to learn from.
For the system to learn, experimental results need to feed back into molecular design. That requires a workable synthesis for a proposed molecule and a reliable readout of what the reaction produced. If the system can’t establish what happened in an experiment, it can’t reliably decide what to change or try next.
Just as with self-driving cars, automating the movements within a self-driving lab is only part of the challenge. The hardest part is being able to interpret what happened in an experiment and use that information to decide what to do next. The goal is for each experiment to help us better choose the next step, and reduce the number of steps we ultimately need to reach a useful result.
Author’s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.


