America’s Role in the Next Golden Age of Chemical Synthesis
This is something I’m still learning about. I’m early in my understanding here, so consider this an exploration of recent conversations with some experts in the field as well as what I’ve been reading and listening to.

The middle decades of the twentieth century were some of chemistry’s most productive years. The Haber-Bosch process had already begun transforming agriculture through synthetic fertilizer. DuPont invented nylon. In medicine, the period brought sulfonamides, the first synthetic antibacterials, and chlorpromazine, the drug that helped launch modern psychopharmacology. During the “golden age of natural product synthesis,” organic chemists learned to build more and more complex molecules found in nature, including alkaloids, steroids, vitamins, and antibiotics. Woodward and Corey helped turn synthesis into a planning discipline. In fact, Corey’s formalization of retrosynthetic analysis underpins many AI route-planning tools.
Today, chemistry is still driving some of the world’s biggest breakthroughs. Materials chemistry gave us lithium-ion batteries. Chemistry also shows up all over modern hardware and medicine: in semiconductor fabrication, high-performance polymers, and the materials used in sutures, catheters, and implants. It was central to the COVID-19 mRNA vaccines too: lipid nanoparticles protected the fragile mRNA and helped get it into cells.
Most approved drugs also depend on synthetic and medicinal chemistry. The GLP-1 boom is a good example: the first wave of these drugs has been peptide-based, but the major goal now is to create oral small-molecule pills. KRAS G12C inhibitors are another example; they proved that RAS, a well-known driver of tumor growth long considered “undruggable,” could be targeted.
But chemistry breakthroughs are incredibly hard-won. In drug discovery, drug development still fails roughly 90% of the time. Large pharmaceutical companies can withstand these failure rates. Most companies can’t. Meanwhile, much of the manufacturing capacity for chemistry and custom synthesis has moved overseas, especially to China and Eastern Europe, a vulnerability made more visible by the war in Ukraine. The reason looks a lot like the broader story of US industrial hollowing-out. Synthetic chemistry was relatively easy to carve out from the rest of R&D. Over time, pharma kept more of the work it saw as highest value (e.g. IP, target discovery, and commercialization), and moved more synthesis and manufacturing to places with lower labor costs, strong scientific talent, and, at least historically, looser environmental and regulatory constraints.
This has made the US chemical synthesis and manufacturing ecosystem fragile, especially for small molecules. The good news is that AI is poised to accelerate chemistry by opening new chemical space. Models are finally getting good enough to propose useful candidates, but they still lack the right data, especially failed reactions and make/test outcomes. Synthesis is how we generate that missing data, and that makes US synthesis capacity strategic.
One note before diving in, I focus mostly on small-molecule drug discovery in this piece because it is the clearest example, but the same design-make-test bottleneck exists in materials, catalysis, and energy chemistry as well.
Why Protein AI Moved First
We’ve already seen this play out in protein engineering and antibody design. The advent of de novo protein models compressed the human and experimental work of protein design into inference, opening up previously undruggable targets and novel biologics mechanisms. AlphaFold 2 made it possible to reliably predict a protein’s 3D shape. Then RFdiffusion and ProteinMPNN let you design a new protein by sketching the shape, and then finding a sequence that folds into it. The Baker lab published de novo serine hydrolase design in 2025. Researchers used AI to design antibodies from scratch. Nabla has reported de novo antibody designs against both GPCRs and peptide-MHC complexes, two hard target classes that point toward more personalized biologics.
Two things made this possible:
The experimental loop was fast: researchers can synthesize and screen thousands of protein variants in a single round, so even low initial success rates can produce useful feedback.
The data existed: decades of solved structures in the Protein Data Bank (PDB), and orders of magnitude more sequence data than structure data.
Synthetic chemistry is different. Small-molecule drugs, for example, require complex, multistep synthesis routes, and optimizing them involves weighing trade-offs between synthesis complexity and difficult-to-predict medicinal chemistry properties. In other words, a model can propose a structure, but then that structure still needs to be made and tested. Synthetic chemistry has been harder to model, harder to search, and harder to learn from. That is why it has lagged protein engineering.
Why Chemistry is Next
Structure Prediction Went Atomic.
Models now represent interactions between all molecular types: proteins, ligands, water, ions, and other molecules critical to biological function.
AlphaFold 3 extended structure prediction beyond proteins alone to proteins bound to other co-factors, including small molecules. Boltz-1 and Chai-1 brought similar capabilities into more accessible tools. At the same time, computational chemistry and materials science are getting their own general models. Machine-learned interatomic potentials, such as MACE-MP-0 and MatterSim, can simulate how atoms interact across a wider range of molecules and materials, much faster and cheaper than traditional methods.
Progress on Higher-Level Properties.
We’re advancing beyond structure to properties that are important for drug design. Things like binding affinity, synthesizability, developability, metabolism, and toxicity.
Boltz-2 added binding-affinity prediction. BoltzMol-1 pushes this further into small-molecule hit discovery by combining model-driven ranking with filters for properties like solubility, lipophilicity, and permeability. Axiom is taking a similar property-prediction approach to toxicity.
Reasoning Models Are Well-Suited to Chemistry.
Agentic models with tool use are well suited to multi-step optimization, and chemistry is fundamentally a reasoning problem. Just as protein language models benefited from treating proteins as sequences, chemistry agents may benefit from treating the synthesis task as a step-by-step reasoning problem.
In medicinal chemistry, each structural change can affect potency, selectivity, solubility, metabolism, toxicity, and synthesizability. Chemists have to reason through these trade-offs molecule by molecule, but AI could help them explore at much greater speed and scale.

ChemCrow shows how LLM agents can combine reasoning with chemistry-specific tools to improve performance and enable new capabilities to emerge. Newer benchmarks like ChemIQ show how reasoning models are becoming more capable on chemistry problems directly.
The Data Is Still Missing
Although chemistry is beginning to look modelable in a new way, we simply don’t have enough data.
The design space of synthetic chemistry is enormous. The number of drug-like chemicals reaches estimates as high as 10^60 possible molecules. Even constrained estimates are huge: GDB-17, which only includes organic small molecules of up to 17 atoms, has 166B molecules. The largest searchable libraries now approach 8.3T molecules. For comparison, ChEMBL 37, a public database of molecules and their biological activity, has only ~2.9M distinct compounds.
On the synthesis side, the problem is compounded, with synthetic route reporting of small-molecule drugs largely being contingent on drug discovery campaign success. Much of this knowledge is tacit and never publicly disclosed, with unsuccessful reactions rarely being archived, which means scraping literature won’t actually recover it. A 2022 paper on the importance of failed experiments showed how reporting bias distorts reaction models, and argued that negative results are among the most valuable missing data. The Open Reaction Database is one initiative that is trying to change this.

And the medicinal chemistry problem is even harder. Data we do actually have is heavily biased toward specific drug targets and regions of chemical space already explored. Unfortunately, decisions made during these campaigns are largely unreported in patent literature, and weighted toward molecules that survived the filters of synthesis, binding, and journal publication. Some companies are trying to close this gap by generating new data directly. Leash, for example, is generating large-scale binding data by testing millions of compounds against hundreds of proteins. But even these efforts are limited by synthesis.
If we want generative AI to discover new drugs, materials, catalysts, and industrial chemicals, we need much richer data on what we can make, what we cannot make, what failed, and why.
Synthesis Creates the Data
This data has to come from the systems that are actually trying to make new molecules. Those systems need to record what happens, and then feed those results back into the models. In other words, synthesis will unlock the learning loop.
The next generation of chemistry infrastructure therefore needs three things.
Cheaper synthesis: Novel chemistry and automation to drive down cost. Models improve with data, and the cost per molecule determines how much data we can generate.
Broader synthesis: Multistep reactions enabling exploration of novel chemistry. Models need access to chemistry beyond the existing scaffolds that make up most of today’s libraries. The goal is to make different molecules, not just more.
Faster turnaround: Quicker feedback loops for agent learning. A system that takes months to design, make, test, and learn from one molecule will not generate enough feedback for models to improve quickly.
Achieving all three requires agentic integration, meaning closed-loop learning systems. The synthesis platform itself has to be part of the learning system, which means planning routes, running reactions, capturing failures, analyzing results, and deciding alternative paths to try.
Pieces of this infrastructure are starting to emerge. Chemify, Onepot, B12, Lila, and Biosero are trying to connect route planning, reaction execution, purification, analytics, and data capture into a single make-test-learn system. Periodic Labs, CuspAI, and Radical AI are doing something similar in materials, where the hard part is connecting models to labs that can synthesize and iterate on them.

Academic research is also showing the benefits of self-driving labs with closed-loop systems. For example, Bayesian optimization outperformed expert chemists on both efficiency and consistency when benchmarked head-to-head against real experiments. A team led by Timothy Noël out of the University of Amsterdam built RoboChem, which can optimize the synthesis of ~10-20 different molecules per week. AlphaFlow, a self-driven fluidic lab, discovered a 40-parameter multi-step route that outperformed conventional sequences.
Together, these examples show that chemistry is becoming programmable. But this only matters if it is connected to the physical world. For the US, that makes synthesis capacity strategic infrastructure.
America Can’t Rent This Future
That infrastructure to connect programmable chemistry with the physical world has to be built in the US.
Synthesis capacity will become more and more strategic. The country that controls the fastest molecule-making loops will have an advantage in drug discovery, materials, agriculture, energy, and defense. We’re seeing this play out in other physical world domains, such as manufacturing and rare earths right now as well.
US pharma and biotech supply chains are deeply exposed to China. According to a 2026 CFR report, “China has both the tools and demonstrated willingness to weaponize US pharmaceutical dependence: the structural conditions enabling it run through nearly every tier of the pharmaceutical supply.” The National Security Commission on Emerging Biotechnology has warned that China is moving aggressively in biotechnology and that “the window to act is closing.” Meanwhile, China’s biopharma industry is becoming a major source of global licensing assets, with Chinese out-licensing deals having already reached 80% of 2025’s full-year total.
Washington has started responding. In December 2025, the BIOSECURE Act was signed into law. It creates a formal process for deciding which biotech suppliers pose national security risks, and for cutting them out of federal contracts and grants. In June 2026, the Pentagon added WuXi AppTec to its Section 1260H list of “Chinese military companies.” WuXi is a Chinese CRDMO whose chemistry arm took on 1,187 new small molecules in 2024, and its US revenue grew 32% last year even under the threat of restriction. The DoD is barred from working directly with those companies. Beginning in June 2027, that restriction will also cover DoD contracts that rely on their work, including US companies that use WuXi as a supplier. WuXi is suing to get off the list but there has been no ruling yet. And Section 232 tariffs went live in July 2026 at a 100% base rate on patented pharmaceuticals, their APIs, and their key starting materials from countries like China and India.
On the commercial side, companies are also starting to act. Lilly committed $27B toward their US manufacturing investment; and Novartis committed $23B to expand its US-based manufacturing and R&D footprint.
The Next Golden Age
Generative AI will create a flood of innovative ideas. The bottleneck will be turning those ideas into molecules, turning molecules into data, and turning that data back into better models.
That loop has to be built here. US dependence on foreign countries, especially China, isn’t just an inconvenience around cost or lead times anymore. It is increasingly a national security threat. The US can’t lead in AI-enabled molecular discovery while continuing to outsource the physical foundations of discovery itself.
Building AI-native synthesis capacity in the US is therefore one of the biggest opportunities of the next decade. AI and robotics can change the economics of synthesis and make domestic chemistry more attractive. For example, smaller teams can run more experiments thus leading to more data and faster feedback loops than traditional lab workflows allowed.
The private sector should build most of this infrastructure, while the government should make it easier to build. The government can do this by funding precompetitive infrastructure, creating demand through procurement, supporting domestic manufacturing credits and loan guarantees, and setting data standards so failed experiments actually become useful training data.
Rebuilding access to reagents, precursors, solvents, APIs, and specialty chemicals will take more than software and robotics, and likely more than a decade. But that is why the US has to start now. The last golden age of synthetic chemistry gave us fertilizers, polymers, antibiotics, synthetic medicines, and much of the industrial base of modern life. The next one could be even bigger.
Special thanks to Dylan Reid, Wenhao Gao, Geoffrey Smith and Paul Gamble for their help in talking through a lot of the ideas in this piece and reviewing versions of it.
Author’s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.

