<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Mixture of Experts]]></title><description><![CDATA[Conversations with the founders and researchers closing the gaps between AI and the real world – across science, industry, and the enterprise. 

Each issue is a deep dive into an idea, paper, or company.]]></description><link>https://www.mixtureofexperts.co</link><image><url>https://substackcdn.com/image/fetch/$s_!rLyP!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7350d917-bb83-466d-92a8-c396729e7a20_1280x1280.png</url><title>Mixture of Experts</title><link>https://www.mixtureofexperts.co</link></image><generator>Substack</generator><lastBuildDate>Mon, 28 Sep 2026 23:18:54 GMT</lastBuildDate><atom:link href="https://www.mixtureofexperts.co/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Annelies Gamble]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[anneliesgamble@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[anneliesgamble@substack.com]]></itunes:email><itunes:name><![CDATA[Annelies Gamble]]></itunes:name></itunes:owner><itunes:author><![CDATA[Annelies Gamble]]></itunes:author><googleplay:owner><![CDATA[anneliesgamble@substack.com]]></googleplay:owner><googleplay:email><![CDATA[anneliesgamble@substack.com]]></googleplay:email><googleplay:author><![CDATA[Annelies Gamble]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[AI Infrastructure for Manufacturing - Notes from IMTS]]></title><description><![CDATA[Two years ago I was skeptical about how big of an impact AI could really have in manufacturing.]]></description><link>https://www.mixtureofexperts.co/p/ai-infrastructure-for-manufacturing</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/ai-infrastructure-for-manufacturing</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 22 Sep 2026 15:05:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!GdkF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Two years ago I was skeptical about how big of an impact AI could really have in manufacturing. This was because back then AI meant LLMs, which meant the inputs and outputs were text-based, and text is just a wrapper around work, but it&#8217;s not the work itself in any real sense when it comes to manufacturing.</span></p><p><span>But a lot has changed in two years. A lot has changed in the last couple months!</span></p><p><span>And to say I&#8217;m no longer bearish on what AI can do in manufacturing is an understatement. I&#8217;ve never been more excited about the opportunities I&#8217;m now seeing for AI to come in and change the way we build plants and products in the manufacturing world.</span></p><p><span>Last week I moderated a panel at </span><a href="https://www.imts.com/"><span>IMTS</span></a><span> about AI infrastructure in manufacturing with </span><a href="https://www.linkedin.com/in/praveenrao/">Praveen Rao</a> (Global Head of Manufacturing at Google Cloud), <a href="https://www.linkedin.com/in/shiv-trisal/"><span>Shiv Trisal</span></a><span> (Global Lead for Industrials and Energy at Databricks), and </span><a href="https://www.linkedin.com/in/karpas/">Les Karpas</a> (Global Head of Physical AI at NVIDIA Inception)<span>. Each of them has also had to rethink something fundamental about AI in manufacturing over the last couple years.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GdkF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GdkF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!GdkF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!GdkF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!GdkF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GdkF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg" width="1456" height="1092" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1905481,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/216821811?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GdkF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg 424w, https://substackcdn.com/image/fetch/$s_!GdkF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg 848w, https://substackcdn.com/image/fetch/$s_!GdkF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!GdkF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e288729-631b-4011-aceb-06ec0839b874_4032x3024.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>It&#8217;s no longer just about the model and chasing intelligence at any cost. The models have reached a point where they are smart enough for most jobs. The constraint is now shifting toward context and the infrastructure that takes the insights generated from AI to then produce a physical outcome.</span></p><p><span>Our panel dove into these topics and here are some of my takeaways.</span></p><h2><span>From insights to action</span></h2><p><span>The panel cautioned that agent activity for the sake of generating insights is not inherently productive.</span></p><p><span>Manufacturers have become inundated with alerts. So much so that they often can&#8217;t prioritize the ones that actually matter.</span></p><p><span>What is productive, however, is closing the loop from insight to action. Thus the metric of success shifts from number of workflows or tokens to number of end-to-end business resolutions.</span></p><p><span>This isn&#8217;t a new idea but it is a lot more complicated to execute in manufacturing plants than many other industries given the nature of a factory floor and the number of legacy systems, machines and humans having to interact.</span></p><p><span>The panel advised having three foundations in place in order to move from insight to action:</span></p><ol><li><p><span>A unified data semantic across all the plant&#8217;s data.</span></p></li><li><p><span>An agentic platform with contract-based delegation to systems such as ERP and PLM.</span></p></li><li><p><span>Multimodal interfaces to allow for easier agent orchestration.</span></p></li></ol><p><span>They also suggested starting narrowly with just one problem that you can&#8217;t solve today or can only solve slowly or manually. Then work backwards from there to understand what data is needed to solve that problem. From there, you can gradually expand to more problems.</span></p><h2><span>Legacy equipment and the data problem</span></h2><p><span>Machines on a factory floor tend to be decades old, at least. They also use hundreds of different PLC vendors and protocols. And some very old machines may not even have a PLC.</span></p><p><span>These systems weren&#8217;t built for interoperability and so the data they generate often needs to be supplemented with additional context around what the measurements mean and how they relate to a process. This has to happen before AI can effectively use the data. This is one of the reasons AI adoption in manufacturing has lagged.</span></p><p><span>The panel stressed the importance of creating a unified namespace or contextualized data layer that connects all of a plant&#8217;s different data, and preserves context across handoffs.</span></p><p><span>Also, physical AI doesn&#8217;t have the internet-scale training set that LLMs had. This means the data that is available needs to be stretched further in a sense so that you can improve digital twins and simulations and then generate synthetic data to validate policies. And once deployed into production, the real-world performance needs to be fed back into the system in order to generate a continuous flywheel.</span></p><p><span>AI can&#8217;t turn bad data into good data, so this work around establishing reliable context and quality is critical before any useful actions can happen reliably. Useful and reliable being key words here!</span></p><h2><span>Own your data in open formats</span></h2><p><span>Every few weeks, there&#8217;s a new model at the top of a leaderboard. As such, the need to own your data and context, and store it in open formats is now more important than ever. This is what allows you to easily switch between the best models and not lose any of the logic you&#8217;ve been building.</span></p><p><span>Openness matters at the tooling layer too. OpenUSD is an example of a common open format for physical AI. It&#8217;s a 3D wrapper that can bring together models, time-series, language, and OT data to improve interoperability across a system. Having a common wrapper format for the 3D, time-series and OT data is what lets different systems share one representation of the plant.</span></p><h2><span>It&#8217;s a multi-AI system, not an LLM</span></h2><p><span>A lot of value on a factory floor is still driven by traditional ML. And as such, the panel stressed the importance of making sure you&#8217;re applying the right kind of AI to a problem. This also ensures you don&#8217;t fall into the trap of token maxing instead of value maxing.</span></p><p><span>One example we talked about was predictive maintenance. In this case, ML is probably best for predicting equipment failures and forecasting demand for spare parts. Whereas an LLM is best for interpreting the training manuals. And a mathematical optimization model is best for scheduling work under various constraints.</span></p><p><span>We also talked about how agents are now writing ML, which means a forecasting model can be built in under a day without needing a large data science team. This is especially impactful for small and mid-sized manufacturers who previously didn&#8217;t have the resources to build some of these tools internally.</span></p><h2><span>Edge and cloud, not edge or cloud</span></h2><p><span>A combination of cloud and edge solutions is critical on factory floors.</span></p><p><span>Edge deployments are best for anything that requires very low latency or where regulatory concerns are an issue. The cloud is best for larger workloads like analyzing large volumes of data. It&#8217;s also where the learning loop happens because of the volume of data that&#8217;s needed across both the operational and business sides of a factory.</span></p><p><span>It&#8217;s no longer a binary choice. The panel recommended having a connected architecture.</span></p><h2><span>Robotics and the case for and against humanoids</span></h2><p><span>Robotics is unquestionably the next frontier of physical AI on the factory floor. But what is less clear is what the optimal form factor will be.</span></p><p><span>The built world, including manufacturing plants, has been designed with the human body in mind. It&#8217;s thus a very natural transition to go from humans that have two arms, two legs, and ten fingers to robots that also have two arms, two legs and ten fingers. Not to mention, we&#8217;ve recorded a lot of videos of humans doing tasks. Therefore, starting with a humanoid form factor makes sense to some extent because that&#8217;s where we have the data. Then you can use cross-embodiment to transfer the policies into other bodies.</span></p><p><span>The counter argument here is that we&#8217;ve become obsessed with the body and perhaps lost sight of the basics. What matters in the physical world is the ability to move things, to take action, reliably. And the form factor to do this depends on what exactly you&#8217;re trying to move or do. The question of form factor in the abstract doesn&#8217;t make sense; you have to start with what the job to be done is.</span></p><h2><span>The adoption bottleneck</span></h2><p><span>The biggest bottleneck to AI adoption in manufacturing these days probably isn&#8217;t capability. The models are largely good enough.</span></p><p><span>What&#8217;s missing is the data and context, and the real-world applications and feedback loops.</span></p><p><span>Two things make this hard. The data problem is a legacy of machines that were never designed to interoperate. The adoption problem is that factories are measured on throughput and uptime, and production can&#8217;t absorb the disruption that standing up AI creates.</span></p><p><span>I&#8217;ve never been more excited about the opportunities ahead for physical AI in manufacturing. But to unlock its potential, both problems need to get easier. On data, this probably means starting from the problem you&#8217;re trying to solve and wiring up only the minimum set of data required, in open formats, rather than modeling the whole plant first. On adoption, this probably means setting up controlled environments for testing and very gradual scale-up only after AI has proven it can resolve an end-to-end action rather than generating yet another alert.</span></p><p><span>IMTS is an incredible conference for those building or interested in manufacturing. I used to come to it as a founder when I was building Prima. It was fun to come back this year and get to see all that has changed and continues to change for the manufacturing world!</span></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[When AI Measures What Instruments Miss]]></title><description><![CDATA[The standard role of AI in science is analyzing data an instrument has already collected. But what if AI were the instrument?]]></description><link>https://www.mixtureofexperts.co/p/when-ai-measures-what-instruments</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/when-ai-measures-what-instruments</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 15 Sep 2026 18:38:27 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7e058bb2-5474-474d-8786-d3bd07441a35_2580x1032.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OqFR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OqFR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png 424w, https://substackcdn.com/image/fetch/$s_!OqFR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png 848w, https://substackcdn.com/image/fetch/$s_!OqFR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png 1272w, https://substackcdn.com/image/fetch/$s_!OqFR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OqFR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png" width="1456" height="483" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:483,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2289493,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/215856230?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OqFR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png 424w, https://substackcdn.com/image/fetch/$s_!OqFR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png 848w, https://substackcdn.com/image/fetch/$s_!OqFR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png 1272w, https://substackcdn.com/image/fetch/$s_!OqFR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff638f98d-c2d2-439b-a615-ccf156d4323f_3664x1216.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Inside a tokamak, plasma circulates at temperatures around 100 million degrees, hotter than the core of the sun. Nothing can contain it by contact, so it&#8217;s held by magnetic fields, shaped and reshaped in real time so the plasma never reaches the wall. &#8220;It&#8217;s like you want to control a storm in a very small area, without letting this storm touch any wall.&#8221; </span><a href="https://www.linkedin.com/in/azarakhshjalalvand/"><span>Azarakhsh (Aza) Jalalvand</span></a><span>, a research scholar in Princeton&#8217;s Plasma Control Group explained to me the other week when we sat down.</span></p><p><span>Aza&#8217;s training is in computer science and AI, not fusion. His PhD was on speech analysis with neural networks, later extended to image, video, and radar data. He ended up focusing on fusion almost by accident: a professor from the applied physics department turned up at his data science lab with a pile of data from a fusion device and a request for volunteers.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>What Aza found was a field that had made remarkable progress over decades, but was still challenging due to the complexity and nonlinear behavior of plasma. &#8220;Physicists have brought us a very long way,&#8221; Aza told me. &#8220;But plasma is an extremely complex environment, and our current knowledge and models are not fast or accurate enough to predict its behavior reliably in real time, which makes monitoring and control hard. It also means we still have to rely a lot on experimentation and trial and error.&#8221;</span></p><p><span>Today, researchers have to rely heavily on sensors to estimate the plasma state and guide control decisions. Sensors provide the essential measurements of what the plasma is doing, and there&#8217;s no computational substitute for them because existing physics models can&#8217;t yet predict plasma&#8217;s behavior with sufficient speed and accuracy.</span></p><p><span>But sensors fail.</span></p><p><span>This problem is not specific to fusion. Any system with a lot of sensors (such as a satellite, a factory line, a patient monitor, etc) will eventually lose one. When that happens, everything that relied on that reading gets worse. In a well-understood system you can calculate around the loss. But when the underlying theory can&#8217;t predict the system, such as is the case in plasma, there&#8217;s no way to calculate what the dead sensor would have read. Measurement via the sensor was the only route to that number, until Aza and his colleagues built a workaround. And in doing so, they also ended up discovering measurements no instrument had ever been able to make.</span></p><p><span>In our conversation, we talked about the work that led to this breakthrough, how you trust a signal that was never recorded, and what it means when a model produces a measurement no instrument could.</span></p><h2><strong><span>Diag2Diag: its origins and applications</span></strong></h2><p><span>&#8220;If you talk to a physicist or a diagnostician and say, &#8216;Diagnostic A died. Is it possible to make its measurement by looking at diagnostic B and C?&#8217; the simple answer is no,&#8221; Aza said. &#8220;And that&#8217;s correct. They are measuring totally different things. But for me, I&#8217;m an engineer, I&#8217;m a data scientist,&#8221; he said. &#8220;All these sensors are looking at one system.&#8221;</span></p><p><span>So even though diagnostic B might not measure what diagnostic A measures, the information could still be there, distributed across the other signals in a form nobody can make sense of.</span></p><p><span>&#8220;Maybe AI can find this pattern,&#8221; Aza said. &#8220;This is where machine learning and AI shines, because these models are very good at learning and leveraging patterns that cannot be interpreted by our understanding.&#8221;</span></p><p><span>A model can learn the relationships between sensors from historical data, assuming enough of the data exists from periods when every diagnostic was running. This means that when one sensor drops out, the model can generate a synthetic version of its signal. That idea became </span><a href="https://www.nature.com/articles/s41467-025-63492-1"><span>Diag2Diag</span></a><span> (Diagnostic to Diagnostic).</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JSL5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JSL5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png 424w, https://substackcdn.com/image/fetch/$s_!JSL5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png 848w, https://substackcdn.com/image/fetch/$s_!JSL5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png 1272w, https://substackcdn.com/image/fetch/$s_!JSL5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JSL5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png" width="1322" height="730" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:730,&quot;width&quot;:1322,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:632868,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/215856230?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JSL5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png 424w, https://substackcdn.com/image/fetch/$s_!JSL5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png 848w, https://substackcdn.com/image/fetch/$s_!JSL5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png 1272w, https://substackcdn.com/image/fetch/$s_!JSL5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda8d1782-1a57-4f37-b1d3-f84e8364685a_1322x730.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Fusion reactors carry many diagnostics, each with different resolution. Diag2Diag trains on the high-resolution sensors (B and C) to reconstruct a high-resolution version of the slow one (A) &#8212; catching fast edge events the original sensor smears over. <a href="https://www.nature.com/articles/s41467-025-63492-1">Source.</a></figcaption></figure></div><p><span>The paper describes an experiment where the model is given data it hasn&#8217;t seen and asked to reconstruct a signal that was actually recorded. Its predictions are checked against what actually happened.</span></p><p><span>&#8220;We can compare and say, okay, now it&#8217;s correct, or now it&#8217;s wrong, or it&#8217;s 90% accurate,&#8221; Aza said. He added a caveat, &#8220;a model can be 90% accurate but with a lot of uncertainty. That&#8217;s also important.&#8221;</span></p><p><span>Much like any hardware sensor, the reconstructed signal is checked against known answers enough times to establish what its errors look like. After passing this validation, the model is used on live experiments.</span></p><p><span>The immediate application for this is failure mitigation.</span></p><p><span>&#8220;Whenever we have a tough system in which maintaining the hardware, maintaining the sensors, is expensive, time consuming, or even impossible,&#8221; Aza said, &#8220;synthetic measurements become interesting.&#8221;</span></p><p><span>He gave the example of a Mars rover. When a sensor dies, no one can easily fix it. But the rover has years of history from that sensor alongside all the others, and so those relationships can be learned and used to predict the missing signal.</span></p><p><span>The same idea can be applied in a lot of settings, especially where the sensors might be hard to reach such as on the seafloor, inside jet engines, in concrete, orbiting, or implanted. Like anything in AI, how well it works depends entirely on the data. &#8220;It all depends on the quality and the quantity of the data that we have collected historically,&#8221; Aza said.</span></p><h2><strong><span>AI as a new instrument</span></strong></h2><p><span>Diag2Diag was built for sensors that fail. But a sensor can be missing in another sense: it can be working perfectly and still not tell you what you need, because its measurements aren&#8217;t granular enough.</span></p><p><span>For example, Thomson Scattering is the diagnostic that measures electron temperature and density, two of the quantities physicists care most about. It fires a laser into the plasma and reads the light that scatters off free electrons. It&#8217;s also not fast enough to measure under millisecond plasma events researchers actually want to see.</span></p><p><span>So Aza&#8217;s team wanted to see if they could use Diag2Diag to reconstruct Thomson Scattering. The other much faster diagnostics became the inputs, and Thomson&#8217;s readings became the target the model learned to reproduce. Unfortunately, the model can only learn at the instants when Thomson Scattering happens to take a reading, and those instants are few and far between. This means that most of the data can&#8217;t be used.</span></p><p><span>&#8220;For example if Thomson Scattering fires at every 5 milliseconds, during training we just ignore the input data of the other diagnostics at milliseconds two, three, four, five, because we don&#8217;t have the target measurement,&#8221; Aza said. &#8220;Then we thought: what happens if after training the model we give [the model] those inputs where we don&#8217;t have the data from Thomson Scattering anyways?&#8221;</span></p><p><span>However, because there is no Thomson Scattering reading at those intermediate moments and never was, there is nothing to check against.</span></p><p><span>&#8220;That is where the data-driven validation stops,&#8221; Aza said. &#8220;We cannot do any data validation anymore. Now we have to think about physical analysis. Do these measurements physically make sense?&#8221;</span></p><p><span>This turned out to be answerable. Edge localized modes (ELMs) are eruptions at the edge of the plasma that occur when the pressure gradient builds past what the edge can hold. Aza described the approach to mitigate them to me like tapping a balloon, &#8220;the surface flattens for a fraction of a second, then springs back&#8221;. ELMs are one of the things Thomson is too slow to catch; Thomson Scattering samples on the order of milliseconds, but an ELM happens much faster than that. So a reading only happens by chance, and never enough times in a row to see the shape of the event.</span></p><p><span>Physicists had a theory that structures called magnetic islands suppressed ELMs, and they ran simulations to predict what the electron profile should look like as an island forms. The theory had no experimental confirmation because there were no instruments fast enough to validate it.</span></p><p><span>&#8220;I worked on the data and generated some outputs, but I didn&#8217;t know if they were physically meaningful,&#8221; Aza told me. So he gave them to a physicist who ran the simulations and confirmed they were a good match. In other words, the synthetic measurements, generated at moments when no hardware recorded anything, matched the simulations. &#8220;This was a very good example of an unbiased discovery,&#8221; Aza said.</span></p><p><span>AI wasn&#8217;t analyzing observations that a sensor had made. It was actually producing net new observations that no sensor had ever been able to make before.</span></p><h2><strong><span>What this changes</span></strong></h2><p><span>AI systems are already good at answering questions we know how to ask. Is this a hand? Is this signal abnormal? What is the predicted value of this measurement? But science often advances when someone notices the thing nobody was looking for.</span></p><p><span>&#8220;These AI models have been really good, sometimes almost perfect, in answering the questions that we know,&#8221; Aza said. &#8220;But in plasma fusion, the question is: what are we missing?&#8221;</span></p><p><span>Normally, answering questions we don&#8217;t know requires building new hardware to generate new data. But here it didn&#8217;t and that&#8217;s what makes this result so unusual. The evidence for the magnetic island theory was latent in data the machine had already recorded.</span></p><p><span>Potential future applications extend beyond plasma to aerospace exploration, robotic surgery, and any complex engineering or scientific systems where missing measurements can compromise safety or control.</span></p><p><span>Diag2Diag started with trying to answer what happens when a sensor is missing, broken, or too slow. But it led to discoveries much deeper. AI may not only help scientists analyze the world, it may help them observe it.</span></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Scaling Agentic Commerce]]></title><description><![CDATA[A conversation with Greg Ulrich, Mastercard's Chief AI and Data Officer]]></description><link>https://www.mixtureofexperts.co/p/scaling-agentic-commerce</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/scaling-agentic-commerce</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 08 Sep 2026 15:23:29 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vKZ5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vKZ5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vKZ5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png 424w, https://substackcdn.com/image/fetch/$s_!vKZ5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png 848w, https://substackcdn.com/image/fetch/$s_!vKZ5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png 1272w, https://substackcdn.com/image/fetch/$s_!vKZ5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vKZ5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png" width="764" height="430" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:430,&quot;width&quot;:764,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:241175,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/214745722?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vKZ5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png 424w, https://substackcdn.com/image/fetch/$s_!vKZ5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png 848w, https://substackcdn.com/image/fetch/$s_!vKZ5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png 1272w, https://substackcdn.com/image/fetch/$s_!vKZ5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f14153c-f7f3-4bc1-bcc9-b834ecd2c4fb_764x430.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>AI agents can already browse, compare, recommend, and negotiate. In fact, according to </span><a href="https://www.deloitte.com/us/en/insights/industry/retail-distribution/retail-distribution-industry-outlook.html"><span>a Deloitte study</span></a><span>, retailers are already seeing 15-20% of referral traffic coming from AI chat interfaces. And soon, they won&#8217;t just be referring traffic, but they&#8217;ll actually be the ones spending money, at scale.</span></p><p><span>But before this can happen, agent actions have to be identifiable and accountable.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>Someone has to build the infrastructure to make that possible.</span></p><p><span>Payment networks are a preview of what agentic AI in commerce will need because they&#8217;ve already solved versions of these problems once at global scale. As </span><a href="https://www.linkedin.com/in/gregu/"><span>Greg Ulrich</span></a><span>, Mastercard&#8217;s Chief AI and Data Officer, put it to me: &#8220;We have an ability to manage these systems with trust and responsibility that&#8217;s been proven over decades.&#8221;</span></p><p><span>I sat down with Greg the other week to explore what has to exist before AI agents can safely spend money, what Mastercard is doing to prepare for that, and what Greg is most bullish about for agentic commerce.</span></p><h2><strong><span>A Means to an End</span></strong></h2><p><span>Payments are much more than just the movement of money. They are identity, authorization, fraud detection, merchant trust, consumer protection, dispute resolution, compliance, and auditability. Mastercard has spent decades developing and supporting this environment.</span></p><p><span>&#8220;Data and AI are really at the foundation of a lot of what we&#8217;re doing today, and it has been for a while,&#8221; Greg told me. Mastercard sits in the ecosystem as a network provider and payment infrastructure company. &#8220;The transmission of data, the use of AI in that transaction, has been fundamental to what we&#8217;ve been doing for decades.&#8221;</span></p><p><span>Notably, AI is not a standalone strategy for Mastercard. It is embedded into how they grow payments, build services, and run their own operations. &#8220;It is still a key enabler of our strategy,&#8221; he told me. &#8220;It&#8217;s a means to an end, in my mind, as opposed to an end in and of itself.&#8221;</span></p><p><span>Greg organizes Mastercard&#8217;s AI strategy into four layers:</span></p><ol><li><p><span>The foundational layer</span></p></li><li><p><span>The value AI creates for Mastercard&#8217;s own people</span></p></li><li><p><span>The value AI creates for Mastercard&#8217;s customers</span></p></li><li><p><span>The role Mastercard plays in shaping the broader ecosystem.</span></p></li></ol><p><span>Each of these layers has the same goal, which is to make AI more capable and more trusted. Ultimately, Mastercard sees its role as helping build the infrastructure that makes AI trustworthy at scale.</span></p><p><span>Importantly, for Greg, AI capability isn&#8217;t the binding constraint. As AI becomes more capable, he believes the limit will be what people, businesses, and institutions are willing to let it do on their behalf. &#8220;How does AI work securely, responsibly and at scale for the entire ecosystem?&#8221; he said. &#8220;And how do we enable that? How do we bring in the data, the trust and the role we play to be a unique and positive force in that?&#8221;</span></p><h2><strong><span>Trusted intent to scale agentic commerce</span></strong></h2><p><span>Today, Mastercard lets consumers, merchants, acquirers, and issuers transact with trust and zero friction. Agentic commerce introduces a fifth actor.</span></p><p><span>&#8220;When you add in a new party, like the agent, how do we envision that ecosystem? How do we enable that ecosystem? How do we make sure we can identify the agents and make sure that that is a known transaction? How do we understand the intent of the consumer when you&#8217;re purchasing something so we can track it along this ecosystem to make sure that the product you have is what you asked the agent for?&#8221; Greg told me.</span></p><p><span>&#8220;In digital commerce, intent is implicit: I click buy, and the click is the intent,&#8221; he said. &#8220;Intent and action are bundled together. But in agentic commerce, they separate. Intent is an artifact that needs to be captured and verified, and potentially disputed.&#8221;</span></p><p><span>This opens up questions that didn&#8217;t exist a year ago around agent identity, delegated authority, fraud scoring for agent actions, consent sharing. Taken together, they point towards needing to extend the trust that already exists in digital commerce such that machines and agents can now operate in the system we&#8217;ve already built for people, businesses, and merchants.</span></p><p><span>Mastercard is building the standards and infrastructure to make this possible. One example Greg mentioned is </span><a href="https://www.mastercard.com/global/en/news-and-trends/stories/2026/verifiable-intent.html"><span>Verifiable Intent</span></a><span>, a standards-based trust paradigm for agentic commerce co-developed with Google. It&#8217;s a standard for safely passing information through the ecosystem so that a user&#8217;s intent can be captured, executed correctly, and traced if it isn&#8217;t.</span></p><p><span>Greg also described the importance they place on working with tech leaders, industry standards bodies, and customers in order to &#8220;build a system that&#8217;s going to work for all parties.&#8221; The ecosystem role, as he framed it, is about helping define standards, policies, and &#8220;the right balance between innovation and responsibility.&#8221;</span></p><h2><strong><span>The foundations for trust at scale</span></strong></h2><p><span>A standard is only as good as the infrastructure that enforces it. &#8220;This starts below the product surface,&#8221; Greg told me. You now need to verify delegated authority instead of just authenticating credentials.</span></p><p><span>For Mastercard, this is the foundation of trust and it comes from five connected capabilities:</span></p><ol><li><p><strong><span>Identity</span></strong><span> to know who is acting</span></p></li><li><p><strong><span>Intent</span></strong><span> to prove what was authorized</span></p></li><li><p><strong><span>Controls</span></strong><span> to define what an agent can do</span></p></li><li><p><strong><span>Trusted execution</span></strong><span> to protect and govern transactions</span></p></li><li><p><strong><span>Intelligence</span></strong><span> to assess risk and detect fraud</span></p></li></ol><p><span>None of these are products. They&#8217;re requirements, and meeting them at scale takes three things Mastercard is building.</span></p><p><strong><span>A harness.</span></strong><span> &#8220;We need standards, we need governance, we need compliance, we need observability,&#8221; Greg told me. Doing that bespoke for every AI deployment is untenable, so Mastercard is building what Greg calls an agentic factory: &#8220;Building an entire harness for all of that, which embeds all of our principles. That&#8217;s the operating system, if you will, for the agents.&#8221; As I&#8217;ve written about previously </span><a href="https://www.mixtureofexperts.co/p/from-models-to-systems"><span>here</span></a><span> and </span><a href="https://www.mixtureofexperts.co/p/the-agent-is-not-the-product"><span>here</span></a><span>, the differentiation is not so much in the model as it is in the harness around the model.</span></p><p><strong><span>Data as a governed asset.</span></strong><span> Mastercard has spent the last 18 months building what it calls its Data Commercialization Platform to &#8220;bring the data together, to democratize it, to have gold data products available to the enterprise, to have the right controls,&#8221; he said, adding: &#8220;It&#8217;s not just transaction data. It&#8217;s how we leverage all the data assets that we have with the right controls, the right governance, and the right linkages.&#8221; Those foundations become even more important in an agentic world, where agents will increasingly need access to trusted data and clear permissions.</span></p><p><strong><span>A foundation model for transactions.</span></strong><span> The most technically distinctive thing Mastercard is building is what Greg calls a </span><a href="https://www.mastercard.com/global/en/news-and-trends/stories/2026/mastercard-new-generative-ai-model.html"><span>Large Tabular Model</span></a><span>, &#8220;the equivalent of an LLM, but for transaction or tabular data. It&#8217;s a model on the data that we have that helps predict behavior and helps understand entities more effectively.&#8221; It&#8217;s built on 15 billion transactions, with a much larger version coming.</span></p><p><span>In Mastercard&#8217;s view, trust at scale is built not through a single product or model, but through the infrastructure that governs how agents operate.</span></p><h2><strong><span>The customer-facing work of trusted AI</span></strong></h2><p><span>On top of these foundations, Greg characterizes Mastercard&#8217;s customer-facing AI work into three key areas: making commerce safer, making customers smarter, and enabling more personalized experiences. Each component is connected to a piece of infrastructure agents will need.</span></p><p><strong><span>Real-time trust scoring.</span></strong><span> &#8220;We&#8217;re going to provide a score on that transaction about how likely it is to be fraudulent or legitimate,&#8221; Greg said. &#8220;It&#8217;s not about adding friction to the ecosystem. It&#8217;s about making it more seamless.&#8221; Decision Intelligence Pro, Mastercard&#8217;s real-time transaction scoring product, is one example. Greg also mentioned Mastercard&#8217;s Safety Net solution, which looks for malicious actors throughout the ecosystem, in addition to Mastercard&#8217;s </span><a href="https://investor.mastercard.com/investor-news/investor-news-details/2024/Mastercard-Finalizes-Acquisition-of-Recorded-Future/default.aspx"><span>acquisition of Recorded Future</span></a><span>, a threat intelligence company that uses AI, analytics and data to help organizations identify, prioritize and respond to cyber threats before they become attacks. Every agentic system will eventually need a real-time trust score on actions and intent, not just transactions. Fraud scoring is the mature template of this.</span></p><p><strong><span>Static to dynamic insights.</span></strong><span> &#8220;A lot of things that we used to provide in a static format are now much more dynamic,&#8221; Greg said. &#8220;We&#8217;re bringing in more data and intelligence because of what the technology allows, but also making it a lot easier to use by putting an AI interface on top that lets customers find their room for optimization, learn about their portfolio, and get better insights.&#8221; Rather than a static report, the deliverable is the interface that lets customers dynamically pull insights.</span></p><p><strong><span>Consent that travels.</span></strong><span> Greg continued: &#8220;For an agent to be able to recommend or negotiate or buy on your behalf, it needs your data. It needs to know your preferences, your purchase history and payment credentials. And then that data needs to move from you to your agent, and then from your agent to merchants and payment networks. As that data moves, your consent needs to move with it. Infrastructure is thus needed to transmit and honor those permissions across handoffs, as well as to decide who bears liability when data gets used beyond what you consented to.&#8221;</span></p><h2><strong><span>Owning the layer that differentiates</span></strong></h2><p><span>Mastercard recognizes that no single company will build the AI future alone, and instead, success will come from orchestrating a trusted ecosystem of model providers, cloud platforms, data partners and innovators, while contributing the security, governance, observability and proprietary intelligence that makes AI safe and effective at scale. Mastercard&#8217;s role is not to build every agent or model, but to help define and operate the trust layer that allows agentic commerce to function across parties, platforms and markets and allows customers to scale AI across their enterprises.</span></p><h2><strong><span>Exponential models, linear humans</span></strong></h2><p><span>At this point, the gap between what&#8217;s possible with AI and what&#8217;s actually happening within large enterprises is less about technology and more about adoption. As Greg puts it, &#8220;The quality of the models continue to go up at this exponential rate, but people&#8217;s ability to consume this and change the way they&#8217;re operating is going at a linear pace.&#8221;</span></p><p><span>For Mastercard, closing the gap has required more than just granting access to AI tools. Mastercard has helped employees develop new ways of working so that AI is now embedded into everyday workflows across the business. Engineers were among the earliest users, but Mastercard quickly realized that AI is most effective when it supports the entire product development lifecycle and broader enterprise processes.</span></p><p><span>Mastercard sees AI as a force multiplier that helps accelerate innovation and deliver better outcomes for customers. The goal isn&#8217;t just around efficiency, but instead to enable the organization to scale expertise and bring new capabilities to market more quickly.</span></p><h2><strong><span>Trust that can scale global commerce</span></strong></h2><p><span>At the end of our conversation, I asked Greg what excites him most about the next phase of AI in payments and commerce. &#8220;I think it&#8217;s around scale in a trusted way,&#8221; he said. Most of what&#8217;s happening in agentic commerce right now, even at the pace it&#8217;s taken off, &#8220;is people experimenting and seeing how it works.&#8221;</span></p><p><span>Whether cause or effect, consumers remain reluctant to transact through agents. In</span><a href="https://www.accenture.com/content/dam/accenture/final/accenture-com/document-fy26/q4/Accenture-Consumer-Pulse-2026.pdf"><span> Accenture&#8217;s 2026 survey</span></a><span> of more than 25,000 people across 16 countries, 74% said they&#8217;d let an agent handle routine tasks, but only 9% are open to letting one shop autonomously on their behalf. People are ready for agents that recommend, but they&#8217;re not ready for agents that pay.</span></p><p><span>We&#8217;ve seen this dynamic before. In a </span><a href="https://www.pewresearch.org/politics/1995/10/16/americans-going-online-explosive-growth-uncertain-destinations/"><span>1995 Pew Research report</span></a><span>, only 8% of internet users had bought anything online in the previous month. Today </span><a href="https://capitaloneshopping.com/research/online-shopping-demographics/"><span>84.3% of the US population</span></a><span> shops online at least once per year. What closed the trust gap was infrastructure: fraud scoring, zero liability, dispute rights.</span></p><p><span>Mastercard has built this kind of infrastructure for previous evolutions in commerce and is positioned to build it again. &#8220;We have a fundamental role to play in agentic commerce by building the trust infrastructure that allows agents to transact safely and at scale given where we sit in this ecosystem,&#8221; Greg told me.</span></p><p><span>He continued: &#8220;I believe agentic commerce will follow a similar arc as e-commerce: first a trust gap, then infrastructure, then adoption. The difference will be the speed of that arc. Agentic commerce will proliferate much faster in large part because the institutions that closed the last trust gap are already building for this one.&#8221; What comes next, as Greg sees it, is enterprise adoption and scale with trust, as he put it, &#8220;at the heart of everything.&#8221;</span></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Science's Observability Problem]]></title><description><![CDATA[We&#8217;ve made designing molecules easier and cheaper than ever, but, as I wrote about here, proving those designs work in humans is increasingly becoming a bottleneck.]]></description><link>https://www.mixtureofexperts.co/p/sciences-observability-problem</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/sciences-observability-problem</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 01 Sep 2026 19:08:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8LLA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nu5P!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nu5P!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png 424w, https://substackcdn.com/image/fetch/$s_!nu5P!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png 848w, https://substackcdn.com/image/fetch/$s_!nu5P!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png 1272w, https://substackcdn.com/image/fetch/$s_!nu5P!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nu5P!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png" width="1456" height="426" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:426,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nu5P!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png 424w, https://substackcdn.com/image/fetch/$s_!nu5P!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png 848w, https://substackcdn.com/image/fetch/$s_!nu5P!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png 1272w, https://substackcdn.com/image/fetch/$s_!nu5P!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab91123-44b4-4570-b36d-cce9576b1683_2048x599.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>We&#8217;ve made designing molecules easier and cheaper than ever, but, as I wrote about </span><a href="https://www.mixtureofexperts.co/p/redesigning-around-a-new-power-source"><span>here</span></a><span>, proving those designs work in humans is increasingly becoming a bottleneck. There are a lot of reasons for this, such as </span><a href="https://www.mixtureofexperts.co/p/americas-role-in-the-next-golden"><span>the constraints around chemical synthesis</span></a><span>, the cost and duration of clinical trials, and how little we actually know about what happens in a lab.</span></p><p><span>Today, scientists keep notebooks and publish papers, and other scientists work from those records. But most of what makes an experiment succeed (or fail) is never actually written down.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>Nowhere is that more expensive than at the handoff, when a working process has to move from the lab that developed it to a contract manufacturer, often in another country. &#8220;That is where things get screwed up,&#8221; </span><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Anna Marie Wagner&quot;,&quot;id&quot;:57649668,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7c4df941-17a0-4805-9759-bed440a7af0d_4762x4762.jpeg&quot;,&quot;uuid&quot;:&quot;0cb79279-cc33-4771-baa3-d6c9a3841b65&quot;}" data-component-name="MentionToDOM"></span> <span>told me the other week. Tech transfer is one of the most common ways manufacturing problems derail a drug launch; and manufacturing is now the single biggest cause of delays, at </span><a href="https://www.fiercepharma.com/manufacturing/biologics-dominate-biopharma-pipelines-more-recent-launches-tripped-manufacturing"><span>64% in 2024</span></a><span>.</span></p><p><span>Anna Marie has spent her career moving between biology and technology. She studied molecular biology, spent time as a tech investor, and returned to the sciences in 2019 unable to shake a specific discomfort: two fields she knew well, science and software, ran on entirely different economics. She&#8217;s now the co-founder of </span><a href="https://transfyr.ai/"><span>Transfyr</span></a><span>, which builds observability infrastructure for scientific execution. The company came out of stealth last week, with </span><a href="https://www.nytimes.com/2026/08/27/science/scientists-experiments-replication-ai.html"><span>The New York Times</span></a><span> covering the launch. &#8220;The variation we see even among well-trained scientists is pretty jaw-dropping,&#8221; Anna Marie told the Times.</span></p><p><span>&#8220;In software, I can write some code, and if it&#8217;s useful, I can ship it to millions of people instantaneously, basically at no cost. But transferring a new technique to a single other lab can take months of debugging between highly trained people,&#8221; she told me. &#8220;We call that tech transfer, which is effectively the distribution cost.&#8221;</span></p><p><span>Up until now, most advancements in AI for science have focused on the parts of science that were already digitizable, such as sequences and molecular structures. And while that&#8217;s great progress, it&#8217;s not addressing the physical problems that determine whether a drug succeeds in a human and whether it can then be manufactured at scale. And it&#8217;s these problems that continue to hold back much of the industry and inhibit innovations from getting out of the lab. As Anna Marie put it, &#8220;Anything that&#8217;s getting in the way of positive scientific discoveries reaching the world is impacting everyone. It&#8217;s climate, it&#8217;s food security, it is human health, animal health. This is a global issue.&#8221;</span></p><h2><strong><span>Noise in Science</span></strong></h2><p><span>If you visit Transfyr&#8217;s lab, you&#8217;ll have the opportunity to participate in a simple exercise: a &#8220;serial dilution,&#8221; which requires diluting a solution of known concentration by a specified amount through a series of steps, in duplicate.  Each participant is then scored on accuracy (how close to the target concentration they landed) and precision (how close the two samples are to each other). &#8220;We consider the &#8216;strike zone&#8217; to be within 5% of the target on both accuracy and precision. The vast majority of people do not hit the strike zone. Even experienced scientists.&#8221;</span></p><p><span>The reason for this miss depends on a lot of factors. Did they pre-wet the pipette or back-pipette? Did they set their pipette to the correct volume at every step? How thoroughly did they mix the sample? Did a stray bubble get in? Did they accidentally foam the mixture? None of this is actually in the written protocol, but it all shows up in the result.</span></p><p><span>Most protocols are actually thousands of small, undocumented decisions. At scale, this invisible variation shows up in the data used to train AI models. As Anna Marie put it, &#8220;We would get data sets where the model was better at recognizing what lab did an experiment than what the experiment was designed to measure in the first place. If that data is just training on unwritten techniques we don&#8217;t even observe, is it actually teaching the model anything useful about the biology itself?&#8221;</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8LLA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8LLA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!8LLA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!8LLA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!8LLA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8LLA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2219060,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/213744313?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8LLA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!8LLA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!8LLA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!8LLA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4cf92540-e056-450c-8a5e-4eb86649cd37_1920x1080.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Transfyr uses computer vision to identify the equipment and reagents on a bench, track the operator&#8217;s hands, and log each step as a timestamped record of what was  done. These are the unwritten parts of a protocol. <a href="https://www.nytimes.com/2026/08/27/science/scientists-experiments-replication-ai.html">Source</a>.</figcaption></figure></div><p><span>There are surprisingly simple, undocumented ways that this lack of standardization shows up. Anna Marie gave an example of a scientist who pre-labeled all of her tubes, which, along with several other shortcuts, helped her get her protocol done two hours faster. And those two hours made a big impact on the experiment&#8217;s outcome because it turns out RNA sitting at room temperature degrades, and two hours is not a trivial amount of time.  But pre-labeling tubes would never show up in a normal protocol.</span></p><p><span>&#8220;I worry that sometimes we chalk up noise in biology to &#8216;biology is noisy&#8217;,&#8221; Anna Marie told me. &#8220;There&#8217;s also a lot of process variability that we just accept without understanding. That&#8217;s the problem. You can accept variability, but we need to understand it so we can interpret our data under that lens.&#8221;</span></p><p><span>In other words, the problem is mistaking &#8220;we don&#8217;t know what happened&#8221; for a fact about biology, rather than a fact about the lack of tooling and observability that exists today in science.</span></p><h2><strong><span>The Minimum Viable Film Room</span></strong></h2><p><span>So then what&#8217;s actually worth measuring? There&#8217;s no shortage of variables that could theoretically matter, but Anna Marie has narrowed down the list to three that do most of the work: the operator&#8217;s actions and intent, supply chain, and then the environment. As she puts it, &#8220;Those few things comprise the minimum viable film room.&#8221;</span></p><p><span>Operator actions and intent means, first and most basically, understanding what a person (or a machine) actually did. This isn&#8217;t necessarily just visual observation, because a huge amount of what you&#8217;re working with in the lab looks identical even though it isn&#8217;t. &#8220;I&#8217;m looking at a clear liquid, and that clear liquid could be water or it could be hydrochloric acid. There&#8217;s a big difference between those two things, and you cannot tell the difference between them visually,&#8221; Anna Marie explained.</span></p><p><span>Supply chain is around the inputs, not just what a material was, but its lot number, its expiration date, who else touched it, what might have contaminated it.</span></p><p><span>The environment is about monitoring things like temperature, humidity, and CO2. Some of this might seem obvious and yet, for most of the industry today, it&#8217;s missing.</span></p><p><span>In Anna Marie&#8217;s opinion, automation around these three levers is often misunderstood. The field has sold automation on the ideas of scale and cost savings, she points out, but neither holds up especially well in practice. Utilization at automated labs across the industry has stayed frustratingly low, and scientists tend to value flexibility over throughput in ways robots can&#8217;t yet match (though I suspect this will change soon!).</span></p><p><span>However, what automation does deliver, in her view, is observability rather than scale. In her essay &#8220;</span><a href="https://thehardthing.substack.com/p/on-observability"><span>On Observability</span></a><span>,&#8221; Anna Marie makes the case using the example of self-driving cars: autonomous vehicles didn&#8217;t improve because of better maps or smarter models, but because cars were instrumented with cameras, lidar, and GPS to learn how humans actually navigated the road, years before autonomy made economic sense. In DARPA&#8217;s first driverless-car challenge in 2004, the best vehicle made it just over seven miles into a 142-mile course; a year later, after a year of watching and iterating in public, five vehicles finished the whole thing. Observability enabled autonomy.</span></p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;8f65fc1f-4b5a-4926-bf30-47abf948b6a6&quot;,&quot;duration&quot;:null}"></div><p><em>Above is a video from a head-mounted camera recording an experiment, along with three cameras mounted above the lab bench. <a href="https://www.nytimes.com/2026/08/27/science/scientists-experiments-replication-ai.html">Source</a>.</em></p><p><span>Lab automation works the same way: automated systems tend to log what they intended to do and what they actually did, almost as a side effect of being automated. &#8220;I would never say that autonomy is the end goal,&#8221; she told me. &#8220;It&#8217;s a tool just like any other, and that tool should be wielded by really smart, caring humans who are trying to make a difference in the world.&#8221;</span></p><h2><strong><span>Why Science Hasn&#8217;t Built Its Film Room</span></strong></h2><p><span>So why don&#8217;t we have film rooms for science already? &#8220;We have deep observability in sports, and it is so valued by the athletes. Its primary purpose is not evaluative. It is coaching and training and self-improvement,&#8221; she said. So why are we not giving elite scientists the same tools that we give elite athletes to improve their game?</span></p><p><span>Part of the answer is infrastructure. But the other part is cultural. Science still has a layer of secrecy around it. If you think you have, as she puts it, &#8220;the magic hands,&#8221; you may not be in a hurry to let everyone else have them too.</span></p><p><span>Publication incentives don&#8217;t help, either. Observability expands what&#8217;s visible, and science has not made the same peace with visible failure that sports has. &#8220;Are we incentivized in any way, shape, or form to publish comprehensive results as opposed to pretty or clean results? No, I don&#8217;t think so,&#8221; Anna Marie said.</span></p><p><a href="https://x.com/JIACHENLIU8"><span>Amber Liu</span></a><span> made </span><a href="https://www.mixtureofexperts.co/p/philosophical-transactions-of-the"><span>a version of this point to me a few weeks ago</span></a><span> about AI research rather than the wet lab: workshops encouraging scientists to publish their failures exist, but adoption is slow. Amber&#8217;s observation was that the reluctance is specifically human. An AI scientist, as she put it, doesn&#8217;t carry &#8220;this burden of disclosing that they&#8217;re actually doing a lot of dumb things, that they fail a lot in the middle.&#8221;</span></p><p><span>The infrastructure that makes science observable is also the infrastructure that makes an individual scientist&#8217;s workflows and mistakes visible and public. Until the incentives around that visibility change, what&#8217;s technically solvable might not matter.</span></p><h2><strong><span>Systems Integration</span></strong></h2><p><span>Toward the end of our conversation, I asked Anna Marie what she thinks the field could look like in the next five to ten years. Her answer was more about the scientific system as a whole rather than any single breakthrough.</span></p><p><span>&#8220;There are millions of scientists who are having ideas right now all around the world. And none of them have a good way to communicate. They don&#8217;t have the language for it,&#8221; she explained.</span></p><p><span>That, to me, is the larger promise of observability in science. Beyond just cleaner data or better protocols (which matter!), it&#8217;s the possibility that knowledge from one lab can be easily transferred so that someone else, somewhere else can use it too.</span></p><p><span>&#8220;I go back to systems integration. Such a boring term, but I love it,&#8221; she joked. &#8220;Because it is the root of most advanced industries. Can we bring together components to advance products that actually solve real problems in the world, on a completely different timeline than we&#8217;re doing today?&#8221;</span></p><p><span>Systems integration is about allowing the components of the physical work of science to talk to each other. Right now, too much knowledge is siloed and unproductive, because nothing exists to carry it from where it was learned to where it is needed.</span></p><p><span>Science still hasn&#8217;t solved distribution. And until it does, better models will continue helping us imagine what to build, but they won&#8217;t tell us how to carry that knowledge into the physical world where it can actually be used.</span></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Routing Intelligence: Prateek Jain on Long-Horizon Agents]]></title><description><![CDATA[An interview with Prateek Jain, Distinguished Scientist at Google DeepMind]]></description><link>https://www.mixtureofexperts.co/p/routing-intelligence-prateek-jain</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/routing-intelligence-prateek-jain</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 25 Aug 2026 17:58:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!P04d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!P04d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!P04d!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg 424w, https://substackcdn.com/image/fetch/$s_!P04d!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg 848w, https://substackcdn.com/image/fetch/$s_!P04d!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!P04d!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!P04d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg" width="764" height="430" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:430,&quot;width&quot;:764,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:110931,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/212736163?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!P04d!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg 424w, https://substackcdn.com/image/fetch/$s_!P04d!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg 848w, https://substackcdn.com/image/fetch/$s_!P04d!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!P04d!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05b672a-822f-4edb-af26-733ab5b32581_764x430.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>There are two high-level ways to make an AI agent work on something hard for a long time: give it enough context to hold more of the problem in mind, or build scaffolding around it so it doesn&#8217;t have to. Scaffolding can include retrieval, memory, planning, tools, tests, checkpoints, and sub-agents.</span></p><p><span>I recently sat down with </span><a href="https://www.prateekjain.org/"><span>Prateek Jain</span></a><span> to talk through what the right combination of these two approaches is. Prateek is a Distinguished Scientist at Google DeepMind, where he works on new architectures for frontier models with a focus on efficiency and elasticity, and co-leads Gemini&#8217;s long-term innovation area that focuses on risky and ambitious research projects.</span></p><p><span>Recently Prateek has been focusing a lot on long-horizon agents, systems that stay on a task longer than most models are trained to do right now. Prateek explains the goal in terms of human effort: &#8220;Let&#8217;s say there is a task that requires ten experts to be at a problem for five days, or ten days, or maybe a month. We want to see if models can go and solve those tasks, which means that our models should know how to plan for a month. They should know how to delegate. They should know how to share information across all these tasks and subtasks and sub-agents.&#8221;</span></p><p><span>But before this can happen, a few things need to be solved first:</span></p><ol><li><p><span>Cost: Attention cost grows quadratically with context length, so doubling what a model holds in mind roughly quadruples what it costs to serve.</span></p></li><li><p><span>Handoff: When a main agent delegates to a sub-agent, the sub-agent typically only sees what was written down for it. This means the reasoning that produced the subtask is thrown out.</span></p></li><li><p><span>Rationing: Anyone running agents at scale has to manage compute costs. But it&#8217;s still unclear how an agent should manage the budget it gives each sub-agent it spins up.</span></p></li></ol><p><span>There are tradeoffs across each of these. More context gives the model more room to work, but it also makes the system more expensive to run. A text-only handoff is cheaper, but the sub-agent only gets the summary, instead of all the work that led up to it. A tight compute budget for the sub-agent might result in the sub-agent failing and thus the work needing to be redone.</span></p><p><span>In our conversation we explored these tradeoffs, and what has to change architecturally in order for long-horizon agents to work at scale.</span></p><h2><strong><span>What is a long-horizon agent, anyway?</span></strong></h2><p><span>&#8220;Long-horizon&#8221; is a term that gets thrown around a lot but what actually makes something &#8220;long-horizon&#8221;? I asked Prateek and he said he defines it along two axes: task difficulty and training distributions.</span></p><p><span>On difficulty: a long-horizon task is one that would take multiple human experts multiple days (think ten experts working for a week or a month).</span></p><p><span>On training distribution: standard RL post-training have multiple rollouts where the model has practiced full attempts at a task with trajectories of somewhat limited length. A long-horizon task runs past anything in that range, so the model is improvising instead of drawing on something it has done before. As Prateek put it: &#8220;The model was trained for a certain rollout length range, and now you are asking the model to do maybe 10x more work,&#8221; at which point the task is &#8220;out of distribution&#8221; with respect to training data.</span></p><h2><strong><span>Long context is necessary, but not sufficient</span></strong></h2><p><span>When I asked Prateek whether he thinks we should give models ever-longer context windows, or build agentic scaffolding around shorter-context models, he said he sees them as a complement. As tasks get harder, &#8220;the amount of context we will need to solve the problem is going to be extremely high. So having agentic scaffolding, as well as ideas like retrieval or maybe separate memory, will be critical.&#8221;</span></p><p><span>But squeezing everything through a short context window breaks down, because hard tasks are underspecified: &#8220;For very hard tasks you might not already know what all you need to put in the context.&#8221;</span></p><p><span>He gave an example: &#8220;If you are trying to ask Gemini about a fairly underspecified problem &#8216;oh, my wife&#8217;s birthday is tomorrow, can you please plan the whole day?&#8217; Gemini would need to understand how your relationship is, what are the things you like to do for fun. It might have some rough idea: okay, let&#8217;s find pictures where both of these people are together, or emails around itinerary planning.  But it might not be able to pinpoint the exact things apriori.&#8221;</span></p><p><span>Pure retrieval only works if you know what to look for, which can be hard when the task is underspecified. Meanwhile, long context is the opposite because instead of deciding upfront what to look for, you bring in a broad field of information and then let the model decide what&#8217;s relevant and how the pieces connect.</span></p><p><span>Long context simplifies orchestration. With a short context window, you need more retrieval, more sub-agent calls, and more tool calls. This is all extra orchestration burden, which, as Prateek put it, &#8220;is not ideal. You want to give as much context to the model as possible.&#8221;</span></p><p><span>The goal instead is &#8220;a really clean, simple solution which is also fairly general and can solve very challenging long-horizon tasks&#8221;: a capable long-context model that needs less machinery around it, not more.</span></p><h2><strong><span>The efficiency wall, elastic models, and routing</span></strong></h2><p><span>Every token of context has a price. &#8220;The serving cost, especially with super long context, is going to grow quadratically, which is challenging,&#8221; Prateek acknowledged. Inference economics is an architectural constraint for agents doing long-running work at scale.</span></p><p><span>He&#8217;s optimistic, however, that active research in frontier labs and academia &#8220;might be able to bring down the cost significantly.&#8221;</span></p><p><span>Some of Prateek&#8217;s work is focused here, on model efficiency, but from the angle of how much model each token uses. &#8220;Models are sort of monolithic. For every task and every phase, you are basically passing the data through the same number of layers and doing the same amount of work.&#8221; This means every step, no matter how hard or easy it is, costs the same.</span></p><p><a href="https://arxiv.org/abs/2310.07707"><span>MatFormer</span></a><span>, one of Prateek&#8217;s best-known projects, asks whether a model can instead be elastic, dialing its power up or down continuously. In practice, Prateek explained, that could mean adjusting the model &#8220;based on the complexity of the task or maybe based on how loaded your servers are or what is the cost currently of your tokens.&#8221;</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MZtK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MZtK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png 424w, https://substackcdn.com/image/fetch/$s_!MZtK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png 848w, https://substackcdn.com/image/fetch/$s_!MZtK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png 1272w, https://substackcdn.com/image/fetch/$s_!MZtK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MZtK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png" width="1456" height="663" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:663,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MZtK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png 424w, https://substackcdn.com/image/fetch/$s_!MZtK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png 848w, https://substackcdn.com/image/fetch/$s_!MZtK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png 1272w, https://substackcdn.com/image/fetch/$s_!MZtK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe118cdc0-6e2a-4b0b-b93d-919e7ee29ba4_1980x902.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">MatFormer makes the feed-forward part of a transformer &#8220;nested,&#8221; so one trained model can be run at several different sizes. At inference time, the system can use a smaller version for cheaper, simpler tasks, or a larger version when it needs more capability. <a href="https://arxiv.org/pdf/2310.07707">Source</a>.</figcaption></figure></div><p><span>&#8220;If agents can say, hey, this is a simple task, I can delegate it to a small model within this large family of models versus, oh, this is a slightly harder task, let me go one notch higher. That can enable high quality solutions at relatively low cost,&#8221; Prateek said.</span></p><p><span>This decision-making requires routing, meaning a policy has to decide which model to call, whether to spawn a sub-agent, what budget it gets, and when to keep retrying versus stop.</span></p><p><span>&#8220;As agents scale to really large numbers of tasks per second,&#8221; Prateek said, &#8220;we will need to start rationing how much compute you are giving to these agents.&#8221; The risk, however, is that you pick a sub-agent too small for the task and it fails, forcing the main agent to evaluate and redo the work. In this case, you end up spending more than if you had just had a more generous upfront routing budget to begin with.</span></p><p><span>Getting routing right means jointly assessing &#8220;the complexity of the task and the capability of the sub-agent.&#8221; And budgets cascade, since the sub-agent then decides for itself how much thinking its allotment buys.</span></p><p><span>This kind of elastic routing hasn&#8217;t fully arrived. &#8220;There is still more research to be done&#8221; on the quality-versus-cost tradeoff. But the direction feels inevitable to him: &#8220;As agents percolate to pretty much everything we do on a day-to-day basis, then having to do this rationing of compute and of capabilities becomes unavoidable.&#8221;</span></p><h2><strong><span>The handoff problem</span></strong></h2><p><span>Delegating to a smaller model requires planning a precise sub-task for the smaller model along with enough context/hints to the smaller model so that it can solve it. In today&#8217;s main-agent/sub-agent pattern, it&#8217;s mostly whatever the main model serialized into text (e.g. a task description rather than the internal state that produced it). The main model does a bunch of work, writes up a subtask, and hands it off. &#8220;The sub-agent&#8217;s view is only the subtask that has been given. It doesn&#8217;t know what all work the primary model would have done,&#8221; Prateek explained. Everything the main model figured out in its KV cache is essentially thrown out at the handoff.</span></p><p><a href="https://arxiv.org/abs/2402.08644"><span>Tandem Transformers</span></a><span>, another of Prateek&#8217;s projects, is focused on this tension. The work pairs a small model with a large one, trained together so the small model consumes the large model&#8217;s representations rather than a text description of the task. &#8220;Whatever work the main model has been doing in its full form, in the KV-cache format itself, can be transferred to the sub-agent,&#8221; Prateek said. &#8220;Which hopefully can provide it with a much richer context so that it can solve the task in a much better fashion.&#8221;</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LGRN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123306b9-2329-46e7-8808-30e822509786_1614x712.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LGRN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123306b9-2329-46e7-8808-30e822509786_1614x712.png 424w, https://substackcdn.com/image/fetch/$s_!LGRN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123306b9-2329-46e7-8808-30e822509786_1614x712.png 848w, https://substackcdn.com/image/fetch/$s_!LGRN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123306b9-2329-46e7-8808-30e822509786_1614x712.png 1272w, https://substackcdn.com/image/fetch/$s_!LGRN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123306b9-2329-46e7-8808-30e822509786_1614x712.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LGRN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123306b9-2329-46e7-8808-30e822509786_1614x712.png" width="1456" height="642" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/123306b9-2329-46e7-8808-30e822509786_1614x712.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:642,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LGRN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123306b9-2329-46e7-8808-30e822509786_1614x712.png 424w, https://substackcdn.com/image/fetch/$s_!LGRN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123306b9-2329-46e7-8808-30e822509786_1614x712.png 848w, https://substackcdn.com/image/fetch/$s_!LGRN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123306b9-2329-46e7-8808-30e822509786_1614x712.png 1272w, https://substackcdn.com/image/fetch/$s_!LGRN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F123306b9-2329-46e7-8808-30e822509786_1614x712.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">In a Tandem Transformer, the large model (left) reads the text in chunks and the small model (right) generates tokens one at a time. The blue arrows are the handoff: the small model attends to the large model's internal representations of everything that came before, so it inherits the bigger model's understanding instead of working from a summary. <a href="https://arxiv.org/pdf/2402.08644">Source</a>.</figcaption></figure></div><p><span>In short, the sub-agent gets the main agent&#8217;s working memory, making the delegation much more flexible. The main model no longer has to fully specify the subtask+hints/context in text before passing it along.</span></p><h2><strong><span>Team players and strategic thinkers</span></strong></h2><p><span>Prateek believes some of the most important domains for long-horizon agents are agentic coding, particularly for machine learning and LLMs, and STEM research, especially in biotech and material science research. Success in these areas requires judgment, &#8220;this kind of work requires the development of taste and personality itself in the model, along with domain expertise and general intelligence,&#8221; noted Prateek.</span></p><p><span>Five years out, Prateek expects the main job of these systems to be organizational. &#8220;What I see is the base models becoming not only very intelligent and powerful, but also having a lot of capabilities around the ability to plan, delegate, coordinate. The model can be put in a variety of harnesses. You can attach memory to them, you can attach knowledge bases to them, and make them work.&#8221;</span></p><p><span>And underneath all of that, the model has to know its own limits. &#8220;Does it know that this is beyond its capability so it can bring in other ideas? Is it able to recognize that maybe it should spin up a new subagent with focus on specific tools?&#8221; Prateek sees these as the key questions. Every routing decision depends on that: sizing a sub-agent, setting a budget, deciding whether to escalate. A model that can&#8217;t tell hard from easy, can&#8217;t ask for help from humans or call potentially expensive tools, and can&#8217;t ration effectively.</span></p><p><span>Today, the models are great ICs, especially when supplied with clear tasks and strategies to solve the task. What they&#8217;re still learning is how to be, as Prateek put it, &#8220;really good team players, strategic thinkers, and leaders.&#8221;</span></p><p><em>Thanks to Divy Thakkar for making this connection!</em></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged. Prateek is speaking in his personal capacity. The views expressed here are his own and do not represent those of his company.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[America’s Role in the Next Golden Age of Chemical Synthesis]]></title><description><![CDATA[This is something I&#8217;m still learning about.]]></description><link>https://www.mixtureofexperts.co/p/americas-role-in-the-next-golden</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/americas-role-in-the-next-golden</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 18 Aug 2026 17:56:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q_YT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><span>This is something I&#8217;m still learning about. I&#8217;m early in my understanding here, so consider this an exploration of recent conversations with some experts in the field as well as what I&#8217;ve been reading and listening to.</span></em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Q_YT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Q_YT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif 424w, https://substackcdn.com/image/fetch/$s_!Q_YT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif 848w, https://substackcdn.com/image/fetch/$s_!Q_YT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif 1272w, https://substackcdn.com/image/fetch/$s_!Q_YT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Q_YT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif" width="800" height="450" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:450,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:997098,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/211743412?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Q_YT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif 424w, https://substackcdn.com/image/fetch/$s_!Q_YT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif 848w, https://substackcdn.com/image/fetch/$s_!Q_YT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif 1272w, https://substackcdn.com/image/fetch/$s_!Q_YT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe0357f4e-635d-4909-88d7-1d3066263e45_800x450.gif 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Himastatin, a complex natural-product antibiotic synthesized by MIT chemists. AI can propose new molecules, but synthesis is what turns molecular ideas into matter, measurements, and data. <a href="https://news.mit.edu/2022/himastatin-synthesis-chemical-0224">Source</a>.</figcaption></figure></div><p><span>The middle decades of the twentieth century were some of chemistry&#8217;s most productive years. The </span><a href="https://www.nature.com/articles/501006a"><span>Haber-Bosch process</span></a><span> had already begun transforming agriculture through synthetic fertilizer. </span><a href="https://www.sciencehistory.org/stories/magazine/nylon-a-revolution-in-textiles/"><span>DuPont invented nylon</span></a><span>. In medicine, the period brought </span><a href="https://www.sciencehistory.org/education/scientific-biographies/gerhard-domagk/"><span>sulfonamides</span></a><span>, the first synthetic antibacterials, and</span><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC2655089/"><span> chlorpromazine</span></a><span>, the drug that helped launch modern psychopharmacology. During the &#8220;</span><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC6010076/"><span>golden age of natural product synthesis</span></a><span>,&#8221; organic chemists learned to build more and more complex molecules found in nature, including alkaloids, steroids, vitamins, and antibiotics.</span><a href="https://en.wikipedia.org/wiki/Robert_Burns_Woodward"><span> Woodward</span></a><span> and </span><a href="https://en.wikipedia.org/wiki/Elias_James_Corey"><span>Corey</span></a><span> helped turn synthesis into a planning discipline. In fact, Corey&#8217;s formalization of retrosynthetic analysis underpins many AI route-planning tools.</span></p><p><span>Today, chemistry is still driving some of the world&#8217;s biggest breakthroughs. Materials chemistry gave us lithium-ion batteries. Chemistry also shows up all over modern hardware and medicine: in semiconductor fabrication, high-performance polymers, and the materials used in sutures, catheters, and implants. It was central to the COVID-19 mRNA vaccines too: lipid nanoparticles protected the fragile mRNA and helped get it into cells.</span></p><p><a href="https://link.springer.com/article/10.1007/s10822-023-00529-x"><span>Most approved drugs</span></a><span> also depend on synthetic and medicinal chemistry. The GLP-1 boom is a good example: the first wave of these drugs has been peptide-based, but the major goal now is to create oral small-molecule pills. KRAS G12C inhibitors are another example; they proved that RAS, a well-known driver of tumor growth long considered &#8220;</span><a href="https://www.nature.com/articles/nrd4389"><span>undruggable</span></a><span>,&#8221; could be targeted.</span></p><p><span>But chemistry breakthroughs are incredibly hard-won. In drug discovery, </span><a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC9293739/"><span>drug development still fails roughly 90% of the time</span></a><span>. Large pharmaceutical companies can withstand these failure rates. Most companies can&#8217;t. Meanwhile, much of the manufacturing capacity for chemistry and custom synthesis has moved overseas, especially to </span><a href="https://www.cfr.org/reports/the-pharma-choke-point"><span>China</span></a><span> and Eastern Europe, </span><a href="https://www.science.org/doi/10.1126/science.abq7841"><span>a vulnerability made more visible</span></a><span> by the war in Ukraine. The reason looks a lot like the broader story of US industrial hollowing-out. Synthetic chemistry was relatively easy to carve out from the rest of R&amp;D. Over time, pharma kept more of the work it saw as highest value (e.g. IP, target discovery, and commercialization), and moved more synthesis and manufacturing to places with lower labor costs, strong scientific talent, and, at least historically, looser environmental and regulatory constraints.</span></p><p><span>This has made the US chemical synthesis and manufacturing ecosystem fragile, especially for small molecules. The good news is that AI is poised to accelerate chemistry by opening new chemical space. Models are finally getting good enough to propose useful candidates, but they still lack the right data, especially failed reactions and make/test outcomes. Synthesis is how we generate that missing data, and that makes US synthesis capacity strategic.</span></p><p><span>One note before diving in, I focus mostly on small-molecule drug discovery in this piece because it is the clearest example, but the same design-make-test bottleneck exists in materials, catalysis, and energy chemistry as well.</span></p><h2><strong><span>Why Protein AI Moved First</span></strong></h2><p><span>We&#8217;ve already seen this play out in protein engineering and antibody design. The advent of de novo protein models compressed the human and experimental work of protein design into inference, opening up previously undruggable targets and novel biologics mechanisms. </span><a href="https://www.nature.com/articles/s41586-021-03819-2"><span>AlphaFold 2</span></a><span> made it possible to reliably predict a protein&#8217;s 3D shape. Then </span><a href="https://www.nature.com/articles/s41586-023-06415-8"><span>RFdiffusion</span></a><span> and </span><a href="https://www.science.org/doi/10.1126/science.add2187"><span>ProteinMPNN</span></a><span> let you design a new protein by sketching the shape, and then finding a sequence that folds into it. The Baker lab published </span><a href="https://www.bakerlab.org/wp-content/uploads/2025/02/science.adu2454.pdf"><span>de novo serine hydrolase design</span></a><span> in 2025. Researchers used AI to design </span><a href="https://www.nature.com/articles/d41586-024-00846-7"><span>antibodies from scratch</span></a><span>. Nabla has reported de novo antibody designs against both </span><a href="https://doi.org/10.1101/2025.05.28.656709"><span>GPCRs</span></a><span> and </span><a href="https://nabla-public.s3.us-east-1.amazonaws.com/2026_Nabla_JAM2_multispecific_pMHC.pdf"><span>peptide-MHC complexes</span></a><span>, two hard target classes that point toward more personalized biologics.</span></p><p><span>Two things made this possible:</span></p><ol><li><p><span>The experimental loop was fast: researchers can synthesize and screen thousands of protein variants in a single round, so even low initial success rates can produce useful feedback.</span></p></li><li><p><span>The data existed: decades of solved structures in the </span><a href="https://www.rcsb.org/"><span>Protein Data Bank</span></a><span> (PDB), and orders of magnitude more sequence data than structure data.</span></p></li></ol><p><span>Synthetic chemistry is different. Small-molecule drugs, for example, require complex, multistep synthesis routes, and optimizing them involves weighing trade-offs between synthesis complexity and difficult-to-predict medicinal chemistry properties. In other words, a model can propose a structure, but then that structure still needs to be made and tested. Synthetic chemistry has been harder to model, harder to search, and harder to learn from. That is why it has lagged protein engineering.</span></p><h2><strong><span>Why Chemistry is Next</span></strong></h2><p><strong><span>Structure Prediction Went Atomic.</span></strong></p><p><span>Models now represent interactions between all molecular types: proteins, ligands, water, ions, and other molecules critical to biological function.</span></p><p><a href="https://www.nature.com/articles/s41586-024-07487-w"><span>AlphaFold 3</span></a><span> extended structure prediction beyond proteins alone to proteins bound to other co-factors, including small molecules. </span><a href="https://boltz.bio/boltz1"><span>Boltz-1</span></a><span> and </span><a href="https://www.biorxiv.org/content/10.1101/2024.10.10.615955v2"><span>Chai-1</span></a><span> brought similar capabilities into more accessible tools. At the same time, computational chemistry and materials science are getting their own general models. </span><a href="https://www.nature.com/articles/s41570-025-00793-5"><span>Machine-learned interatomic potentials</span></a><span>, such as </span><a href="https://doi.org/10.1063/5.0297006"><span>MACE-MP-0</span></a><span> and </span><a href="https://arxiv.org/abs/2405.04967"><span>MatterSim</span></a><span>, can simulate how atoms interact across a wider range of molecules and materials, much faster and cheaper than traditional methods.</span></p><p><strong><span>Progress on Higher-Level Properties.</span></strong></p><p><span>We&#8217;re advancing beyond structure to properties that are important for drug design. Things like binding affinity, synthesizability, developability, metabolism, and toxicity.</span></p><p><a href="https://jclinic.mit.edu/boltz-2-towards-accurate-and-efficient-binding-affinity-prediction/"><span>Boltz-2</span></a><span> added binding-affinity prediction. </span><a href="https://www.biorxiv.org/content/10.64898/2026.07.04.736485v1"><span>BoltzMol-1</span></a><span> pushes this further into small-molecule hit discovery by combining model-driven ranking with filters for properties like solubility, lipophilicity, and permeability. </span><a href="https://axiombio.ai/"><span>Axiom</span></a><span> is taking a similar property-prediction approach to toxicity.</span></p><p><strong><span>Reasoning Models Are Well-Suited to Chemistry.</span></strong></p><p><span>Agentic models with tool use are well suited to multi-step optimization, and chemistry is fundamentally a reasoning problem. Just as protein language models benefited from treating proteins as sequences, chemistry agents may benefit from treating the synthesis task as a step-by-step reasoning problem.</span></p><p><span>In medicinal chemistry, each structural change can affect potency, selectivity, solubility, metabolism, toxicity, and synthesizability. Chemists have to reason through these trade-offs molecule by molecule, but AI could help them explore at much greater speed and scale.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CDRq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CDRq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png 424w, https://substackcdn.com/image/fetch/$s_!CDRq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png 848w, https://substackcdn.com/image/fetch/$s_!CDRq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png 1272w, https://substackcdn.com/image/fetch/$s_!CDRq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CDRq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png" width="1308" height="950" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:950,&quot;width&quot;:1308,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CDRq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png 424w, https://substackcdn.com/image/fetch/$s_!CDRq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png 848w, https://substackcdn.com/image/fetch/$s_!CDRq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png 1272w, https://substackcdn.com/image/fetch/$s_!CDRq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcd879ce4-d428-446e-bb2b-43f49667c3bd_1308x950.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">ChemCrow connects an LLM to chemistry-specific tools and robotic synthesis. The system can reason through a task, select tools such as retrosynthesis or procedure prediction, and execute experiments through a connected lab platform. <a href="https://www.nature.com/articles/s42256-024-00832-8">Source</a>.</figcaption></figure></div><p><a href="https://www.nature.com/articles/s42256-024-00832-8"><span>ChemCrow</span></a><span> shows how LLM agents can combine reasoning with chemistry-specific tools to improve performance and enable new capabilities to emerge. Newer benchmarks like </span><a href="https://pubmed.ncbi.nlm.nih.gov/41411158/"><span>ChemIQ</span></a><span> show how reasoning models are becoming more capable on chemistry problems directly.</span></p><h2><strong><span>The Data Is Still Missing</span></strong></h2><p><span>Although chemistry is beginning to look modelable in a new way, we simply don&#8217;t have enough data.</span></p><p><span>The design space of synthetic chemistry is </span><a href="https://www.science.org/content/blog-post/hideous-numbers-compounds"><span>enormous</span></a><span>. The number of drug-like chemicals reaches estimates as high as </span><a href="https://pubs.rsc.org/en/content/articlehtml/2010/md/c0md00020e"><span>10^60 possible molecules</span></a><span>. Even constrained estimates are huge: </span><a href="https://pubs.acs.org/doi/10.1021/ci300415d"><span>GDB-17</span></a><span>, which only includes organic small molecules of up to 17 atoms, has 166B molecules.  The </span><a href="https://www.biosolveit.de/chemical-spaces/#explore"><span>largest searchable libraries</span></a><span> now approach 8.3T molecules. For comparison, </span><a href="https://www.ebi.ac.uk/chembl/beta/"><span>ChEMBL 37</span></a><span>, a public database of molecules and their biological activity, has only ~2.9M distinct compounds.</span></p><p><span>On the synthesis side, the problem is compounded, with synthetic route reporting of small-molecule drugs largely being contingent on drug discovery campaign success. Much of this knowledge is tacit and never publicly disclosed, with unsuccessful reactions rarely being archived, which means scraping literature won&#8217;t actually recover it. A 2022 paper on </span><a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/anie.202204647"><span>the importance of failed experiments</span></a><span> showed how reporting bias distorts reaction models, and argued that negative results are among the most valuable missing data. The </span><a href="https://open-reaction-database.org/about"><span>Open Reaction Database</span></a><span> is one initiative that is trying to change this.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MVcf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MVcf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg 424w, https://substackcdn.com/image/fetch/$s_!MVcf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg 848w, https://substackcdn.com/image/fetch/$s_!MVcf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!MVcf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MVcf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg" width="724" height="396" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:396,&quot;width&quot;:724,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MVcf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg 424w, https://substackcdn.com/image/fetch/$s_!MVcf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg 848w, https://substackcdn.com/image/fetch/$s_!MVcf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!MVcf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30776596-6a46-4cf0-aaca-9f8acdce7256_724x396.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Reaction-yield distributions in the Open Reaction Database. The orange bars show USPTO-mined patent data, which is heavily skewed toward high-yield reactions. In other words, patents mostly report reactions that worked. The blue bars show curated non-USPTO data, much of it from high-throughput experiments, which includes many failed or low-yield reactions. <a href="https://www.linkedin.com/posts/open-reaction-database_openreactiondatabase-opendata-machinelearning-activity-7482877615381360640-2iRf?utm_source=social_share_send&amp;utm_medium=member_desktop_web&amp;rcm=ACoAAASzfa4BsHI9ivDAufqzSA_uupgQfMmrV-I">Source</a>.</figcaption></figure></div><p><span>And the medicinal chemistry problem is even harder. Data we do actually have is heavily biased toward specific drug targets and regions of chemical space already explored. Unfortunately, decisions made during these campaigns are largely unreported in patent literature, and weighted toward molecules that survived the filters of synthesis, binding, and journal publication. Some companies are trying to close this gap by generating new data directly. </span><a href="https://www.leash.bio/"><span>Leash</span></a><span>, for example, is generating large-scale binding data by testing millions of compounds against hundreds of proteins. But even these efforts are limited by synthesis.</span></p><p><span>If we want generative AI to discover new drugs, materials, catalysts, and industrial chemicals, we need much richer data on what we can make, what we cannot make, what failed, and why.</span></p><h2><strong><span>Synthesis Creates the Data</span></strong></h2><p><span>This data has to come from the systems that are actually trying to make new molecules. Those systems need to record what happens, and then feed those results back into the models. In other words, synthesis will unlock the learning loop.</span></p><p><span>The next generation of chemistry infrastructure therefore needs three things.</span></p><ol><li><p><strong><span>Cheaper synthesis</span></strong><span>: Novel chemistry and automation to drive down cost. Models improve with data, and the cost per molecule determines how much data we can generate.</span></p></li><li><p><strong><span>Broader synthesis</span></strong><span>: Multistep reactions enabling exploration of novel chemistry. Models need access to chemistry beyond the existing scaffolds that make up most of today&#8217;s libraries. The goal is to make different molecules, not just more.</span></p></li><li><p><strong><span>Faster turnaround</span></strong><span>: Quicker feedback loops for agent learning. A system that takes months to design, make, test, and learn from one molecule will not generate enough feedback for models to improve quickly.</span></p></li></ol><p><span>Achieving all three requires agentic integration, meaning closed-loop learning systems. The synthesis platform itself has to be part of the learning system, which means planning routes, running reactions, capturing failures, analyzing results, and deciding alternative paths to try.</span></p><p><span>Pieces of this infrastructure are starting to emerge. </span><a href="https://www.chemify.io/technology"><span>Chemify</span></a><span>, </span><a href="https://www.onepot.ai/announcement"><span>Onepot</span></a><span>, </span><a href="https://b12-labs.com/"><span>B12</span></a><span>, </span><a href="https://www.lila.ai/"><span>Lila</span></a><span>, and </span><a href="https://biosero.com/"><span>Biosero</span></a><span> are trying to connect route planning, reaction execution, purification, analytics, and data capture into a single make-test-learn system. </span><a href="https://periodic.com/"><span>Periodic Labs</span></a><span>, </span><a href="https://cusp.ai/"><span>CuspAI</span></a><span>, and </span><a href="https://www.radical-ai.com/"><span>Radical AI</span></a><span> are doing something similar in materials, where the hard part is connecting models to labs that can synthesize and iterate on them.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!PSgn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!PSgn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png 424w, https://substackcdn.com/image/fetch/$s_!PSgn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png 848w, https://substackcdn.com/image/fetch/$s_!PSgn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png 1272w, https://substackcdn.com/image/fetch/$s_!PSgn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!PSgn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png" width="1216" height="652" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:652,&quot;width&quot;:1216,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!PSgn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png 424w, https://substackcdn.com/image/fetch/$s_!PSgn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png 848w, https://substackcdn.com/image/fetch/$s_!PSgn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png 1272w, https://substackcdn.com/image/fetch/$s_!PSgn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4ee7d7c9-f094-4abf-9f7e-0922b8dde872_1216x652.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>The design-make-test-learn loop in chemistry. In AlphaFlow, an automated lab system runs experiments, returns data to a learning agent, and uses that feedback to choose the next experiment. </span><a href="https://www.nature.com/articles/s41467-023-37139-y">Source</a>.</figcaption></figure></div><p><span>Academic research is also showing the benefits of self-driving labs with closed-loop systems. For example, Bayesian optimization </span><a href="https://www.nature.com/articles/s41586-021-03213-y"><span>outperformed expert chemists</span></a><span> on both efficiency and consistency when benchmarked head-to-head against real experiments. A team led by Timothy No&#235;l out of the University of Amsterdam built </span><a href="https://cen.acs.org/physical-chemistry/computational-chemistry/Robot-assistant-optimizes-photochemistry/102/i3"><span>RoboChem</span></a><span>, which can optimize the synthesis of ~10-20 different molecules per week. </span><a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10015005/"><span>AlphaFlow</span></a><span>, a self-driven fluidic lab, discovered a 40-parameter multi-step route that outperformed conventional sequences.</span></p><p><span>Together, these examples show that chemistry is becoming programmable. But this only matters if it is connected to the physical world. For the US, that makes synthesis capacity strategic infrastructure.</span></p><h2><strong><span>America Can&#8217;t Rent This Future</span></strong></h2><p><span>That infrastructure to connect programmable chemistry with the physical world has to be built in the US.</span></p><p><span>Synthesis capacity will become more and more strategic. The country that controls the fastest molecule-making loops will have an advantage in drug discovery, materials, agriculture, energy, and defense. We&#8217;re seeing this play out in other physical world domains, such as </span><a href="https://www.axios.com/2026/05/27/us-manufacturing-china-imports"><span>manufacturing</span></a><span> and </span><a href="https://www.csis.org/analysis/rare-earth-export-restrictions-one-year-later"><span>rare earths</span></a><span> right now as well.</span></p><p><span>US pharma and biotech supply chains are deeply exposed to China. According to a </span><a href="https://www.cfr.org/reports/the-pharma-choke-point"><span>2026 CFR report</span></a><span>, &#8220;China has both the tools and demonstrated willingness to weaponize US pharmaceutical dependence: the structural conditions enabling it run through nearly every tier of the pharmaceutical supply.&#8221; The National Security Commission on Emerging Biotechnology </span><a href="https://www.biotech.senate.gov/wp-content/uploads/2026/06/US-v-China-Action-on-NSCEB-Recommendations.pdf"><span>has warned</span></a><span> that China is moving aggressively in biotechnology and that &#8220;the window to act is closing.&#8221; Meanwhile, China&#8217;s biopharma industry is becoming a major source of global licensing assets, with Chinese out-licensing deals having </span><a href="https://www.reuters.com/legal/litigation/china-innovative-drug-out-licensing-deal-value-reaches-new-high-first-half-2026-2026-07-13/"><span>already reached 80% of 2025&#8217;s full-year total</span></a><span>.</span></p><p><span>Washington has started responding. In December 2025, </span><a href="https://www.lw.com/en/insights/biosecure-act-becomes-law-limiting-grants-with-biotechnology-companies-of-concern"><span>the BIOSECURE Act </span></a><span>was signed into law. It creates a formal process for deciding which biotech suppliers pose national security risks, and for cutting them out of federal contracts and grants. In June 2026, the Pentagon added WuXi AppTec to its Section 1260H list of &#8220;Chinese military companies.&#8221; WuXi is a Chinese CRDMO whose chemistry arm took on 1,187 new small molecules in 2024, and its US revenue grew 32% last year even under the threat of restriction. The DoD is barred from working directly with those companies. Beginning in June 2027, that restriction will also cover DoD contracts that rely on their work, including US companies that use WuXi as a supplier. WuXi is suing to get off the list but there has been no ruling yet. And </span><a href="https://www.bhfs.com/insight/trump-administration-announces-section-232-tariffs-on-pharmaceuticals/"><span>Section 232 tariffs</span></a><span> went live in July 2026 at a 100% base rate on patented pharmaceuticals, their APIs, and their key starting materials from countries like China and India.</span></p><p><span>On the commercial side, companies are also starting to act. Lilly committed </span><a href="https://investor.lilly.com/news-releases/news-release-details/lilly-plans-more-double-us-manufacturing-investment-2020"><span>$27B</span></a><span> toward their US manufacturing investment; and Novartis committed </span><a href="https://www.novartis.com/us-en/news/media-releases/novartis-plans-expand-its-us-based-manufacturing-and-rd-footprint-total-investment-23b-over-next-5-years"><span>$23B</span></a><span> to expand its US-based manufacturing and R&amp;D footprint.</span></p><h2><strong><span>The Next Golden Age</span></strong></h2><p><span>Generative AI will create a flood of innovative ideas. The bottleneck will be turning those ideas into molecules, turning molecules into data, and turning that data back into better models.</span></p><p><span>That loop has to be built here. US dependence on foreign countries, especially China, isn&#8217;t just an inconvenience around cost or lead times anymore. It is increasingly a national security threat. The US can&#8217;t lead in AI-enabled molecular discovery while continuing to outsource the physical foundations of discovery itself.</span></p><p><span>Building AI-native synthesis capacity in the US is therefore one of the biggest opportunities of the next decade. AI and robotics can change the economics of synthesis and make domestic chemistry more attractive. For example, smaller teams can run more experiments thus leading to more data and faster feedback loops than traditional lab workflows allowed.</span></p><p><span>The private sector should build most of this infrastructure, while the government should make it easier to build. The government can do this by funding precompetitive infrastructure, creating demand through procurement, supporting domestic manufacturing credits and loan guarantees, and setting data standards so failed experiments actually become useful training data.</span></p><p><span>Rebuilding access to reagents, precursors, solvents, APIs, and specialty chemicals will take more than software and robotics, and likely more than a decade. But that is why the US has to start now. The last golden age of synthetic chemistry gave us fertilizers, polymers, antibiotics, synthetic medicines, and much of the industrial base of modern life. The next one could be even bigger.</span></p><p><em><span>Special thanks to Dylan Reid, Wenhao Gao, </span>Geoffrey Smith<span> and Paul Gamble for their help in talking through a lot of the ideas in this piece and reviewing versions of it.</span></em></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Frictions That Make AI Forecasting Hard]]></title><description><![CDATA[To better forecast the impact of AI on any industry, you must understand its overall impact on an integrated ecosystem.]]></description><link>https://www.mixtureofexperts.co/p/the-frictions-that-make-ai-forecasting</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/the-frictions-that-make-ai-forecasting</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 11 Aug 2026 19:24:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Sa7e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!L2Jk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!L2Jk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png 424w, https://substackcdn.com/image/fetch/$s_!L2Jk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png 848w, https://substackcdn.com/image/fetch/$s_!L2Jk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png 1272w, https://substackcdn.com/image/fetch/$s_!L2Jk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!L2Jk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png" width="5495" height="1548" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1548,&quot;width&quot;:5495,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4767130,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/210797287?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ab7b379-e8c3-4b51-949a-6bcb07c3dc4f_5556x1548.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!L2Jk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png 424w, https://substackcdn.com/image/fetch/$s_!L2Jk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png 848w, https://substackcdn.com/image/fetch/$s_!L2Jk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png 1272w, https://substackcdn.com/image/fetch/$s_!L2Jk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e87717-9b53-4b2f-bd40-4c7812e51e98_5495x1548.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>To better forecast the impact of AI on any industry, you must understand its overall impact on an integrated ecosystem. It may streamline a piece of the workflow but how will this play out as the other steps digest this acceleration? Where will the bottlenecks be created and can they be cleared? If not, the overall throughput will not increase.</span></p><p><span>AI forecasting is about predicting what AI systems will be capable of, as well as how those capabilities will impact industries, institutions, and everyday life. That second part is much harder.</span></p><p><span>&#8220;It&#8217;s not something you can get from first principles,&#8221; </span><a href="https://substack.com/@abio"><span>Abi Olvera</span></a><span> told me when we spoke last week about AI forecasts. I first met Abi after I reached out because of </span><a href="https://secondthoughts.ai/p/why-arent-bioweapons-common"><span>one of her recent articles</span></a><span> in which she talks about AI forecasting from the lens of biosecurity.</span></p><p><span>Abi spent more than seven years as a State Department diplomat working across crisis preparedness, national security, China, international coordination, and emerging technology risks. She served in Dakar and Cairo, where she focused on Egypt&#8217;s $12 billion IMF program. She also worked on cybersecurity and law enforcement issues on the China Desk. Across her career, she&#8217;s had to think about poverty, national security, crisis preparedness, energy, biosecurity, cybersecurity, and the institutions that have to respond when technology changes faster than policy can. Today, Abi writes the </span><a href="https://abio.substack.com/"><span>Positive Sum</span></a><span> Substack and is launching an organization focused on state capacity and bottlenecks in the economy. She is also a Special Advisor at Golden Gate Institute, and an affiliate at the </span><a href="https://www.iaps.ai/"><span>Institute for AI Policy and Strategy</span></a><span>.</span></p><p><span>In our conversation, Abi and I discussed where she sees the most risk around forecasting gaps, the drivers of those gaps, and what can be done to help fix them.</span></p><h2><strong><span>Forecasting Gaps</span></strong></h2><p><span>Forecasting gaps happen when people assume having access to information is the same as  having the ability to act on it. AI makes a domain easier to understand, but doesn&#8217;t necessarily make the work easier to execute.</span></p><p><span>Take radiology for example. In 2016, </span><a href="https://www.dotmed.com/news/story/39033"><span>Geoffrey Hinton told a Toronto machine learning conference</span></a><span> that &#8220;people should stop training radiologists,&#8221; because it was &#8220;completely obvious&#8221; deep learning would outperform them within five years. And while AI has become very valuable in radiology (the </span><a href="https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices"><span>FDA has cleared hundreds of AI-enabled medical imaging tools</span></a><span>), the labor-market forecast was very wrong. Radiology did not disappear; in fact demand </span><em><span>grew</span></em><span>. At </span><a href="https://www.beckershospitalreview.com/radiology/mayo-clinic-radiology-leads-in-ai-use/"><span>the Mayo Clinic</span></a><span>, one of the more aggressive AI adopters, radiology headcount reportedly grew 55% since 2016 while the department built a 40-person AI team and used more than 250 AI models.</span></p><p><span>&#8220;There&#8217;s been amazing progress, but these AI tools for the most part look for one thing,&#8221; said Dr. Charles E. Kahn Jr., a professor of radiology at the University of Pennsylvania&#8217;s Perelman School of Medicine and editor of the journal </span><a href="https://pubs.rsna.org/journal/ai"><span>Radiology: Artificial Intelligence</span></a><span>. As </span><a href="https://www.nytimes.com/2025/05/14/technology/ai-jobs-radiologists-mayo-clinic.html"><span>The New York Times</span></a><span> wrote, &#8220;Radiologists do far more than study images. They advise other doctors and surgeons, talk to patients, write reports and analyze medical records. After identifying a suspect cluster of tissue in an organ, they interpret what it might mean for an individual patient with a particular medical history, tapping years of experience.&#8221;</span></p><p><span>AI improved the task of interpreting medical images, but radiologists still integrate clinical context, communicate with physicians, guide treatment decisions, perform procedures, and carry accountability. The capability arrived, but the forecast failed because interpreting images was only the tip of the iceberg of the &#8220;radiology job&#8221; - and as AI became more proficient in this task at the top of the workflow, more radiologists were needed to keep up with the increase in related work.</span></p><p><span>Abi believes forecasts focus too much on what technology makes possible and not enough on the realities of domains and social incentives governing whether, where, and why people will actually use it. We need to better understand both sides: where does AI change the workflow, and how does the world respond.</span></p><h2><strong><span>Biases and Blind Spots</span></strong></h2><p><span>Part of what makes AI forecasting so hard is that the people closest to the technology are often farthest from the domains where the forecasts will play out. As Abi put it, &#8220;Because AI forecasting has tended to be from the more technical, Bay Area-based community, it systematically falls down more in any domain that has to do with social behavior, adoption, or any large real-world or physical world component.&#8221; In San Francisco, for example, people tend to start from capability.  Software engineers often use AI for tasks where the model can complete much of the workflow end to end. In other domains, however, the workflow extends beyond text into physical systems or interactions with outside actors, which makes AI&#8217;s impact both harder to see and harder to forecast.</span></p><p><span>&#8220;People in DC are less likely to be using Claude Cowork or Codex,&#8221; she said. &#8220;If people use it [in DC], a lot of times they might be using the chatbot version, which is great, but that doesn&#8217;t really unleash the parts of AI that are crazy surprising.&#8221;</span></p><p><span>On the other hand, people in San Francisco often underestimate frictions around AI adoption because they are focused on the vision more so than the nitty gritty work that translates vision to reality in the real world.</span></p><p><span>In truth, neither perspective is complete. San Francisco may be closer to the frontier of technological capabilities, but further from the systems where those capabilities are deployed. DC (or other centers of non-technical power), meanwhile, is closer to those systems where the rubber meets the road, but further from the true possibility of AI technology. Unfortunately, public discourse rarely closes this gap. The people most motivated to speak usually have a stake in the outcome (e.g. a company trying to sell a product or an institution trying to defend its role). Advocacy creates the resources and incentives to shape the conversation, which is valuable, but can also mean the loudest arguments are not always the most representative view of practitioners. And because the content people consume is usually the content that already fits their worldview, public discourse increasingly serves as echo chambers rather than bridges.</span></p><h2><strong><span>Cross-Framework Research</span></strong></h2><p><span>This is why AI forecasting needs more contact with practitioners. They are the ones with the most grounded view of where AI can change the work vs where it can&#8217;t, and which parts of the job outsiders are likely to overlook from their AI ivory towers.</span></p><p><span>Abi believes we need more research that brings together both sides &#8211; the technical and the practical. She calls this cross-framework research.</span></p><p><span>One of her favorite examples of this kind of research is </span><a href="https://arxiv.org/abs/2602.16703"><span>a study</span></a><span> that tested whether AI actually helped novices perform molecular biology tasks in a lab. The study compared people with internet access to people with internet plus frontier AI models, then measured whether they could complete hands-on wet-lab tasks over eight weeks. The result was mixed: AI seemed to help on some intermediate steps, but it didn&#8217;t significantly increase the number of people who could complete the full lab workflow from start to finish. The study tried to evaluate the delta between simply having better instructions versus being able to execute a difficult physical workflow.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Sa7e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Sa7e!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png 424w, https://substackcdn.com/image/fetch/$s_!Sa7e!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png 848w, https://substackcdn.com/image/fetch/$s_!Sa7e!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png 1272w, https://substackcdn.com/image/fetch/$s_!Sa7e!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Sa7e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png" width="1456" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1900069,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/210797287?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Sa7e!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png 424w, https://substackcdn.com/image/fetch/$s_!Sa7e!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png 848w, https://substackcdn.com/image/fetch/$s_!Sa7e!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png 1272w, https://substackcdn.com/image/fetch/$s_!Sa7e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5dc4c1fc-d786-4e58-8807-a68b16118341_4584x1984.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Study design from an eight-week wet-lab experiment comparing novices with internet access to novices with internet plus frontier LLMs. The setup tests a key forecasting question: when AI improves access to instructions and troubleshooting, does it actually change the bottleneck in a hands-on workflow? <a href="https://arxiv.org/pdf/2602.16703">Source</a>.</figcaption></figure></div><p><span>For many industries, a key question with AI is whether it can replace enough tacit knowledge, troubleshooting ability, and hands-on competence to change what a novice can do.</span></p><h2><strong><span>What the World Lets AI Do</span></strong></h2><p><span>This has implications for policy too. If governments don&#8217;t know yet exactly how a technology will evolve, rushing to write rigid and detailed rules isn&#8217;t helpful. Abi believes more focus needs to go toward building the capacity to respond quickly as evidence gets clearer.</span></p><p><span>Abi has argued that governments should focus on </span><a href="https://x.com/Abi0lvera/status/2047416816641704338"><span>building ecosystem capacity</span></a><span>: better reporting, better evaluations, and more technical expertise inside institutions. This doesn&#8217;t require policymakers to perfectly forecast every future use case, but rather gives them the ability to notice what&#8217;s happening, test assumptions, and act effectively as they get better information and clearer directions emerge.</span></p><p><span>AI will transform science, security, industry, and government. But the path from capability to transformation runs through the parts of the world that are hardest to model from the outside. As AI accelerates parts of a workflow, more attention needs to shift to whatever is still slow, physical, tacit, or institutionally constrained. As Abi put it: &#8220;Bottlenecks become more important when everything else gets automated.&#8221;</span></p><p><span>The question, then, is not just what AI can do. It&#8217;s whether that capability changes the bottleneck. That&#8217;s the difference between what AI can technically do and what it can actually do.</span></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p>]]></content:encoded></item><item><title><![CDATA[Measuring the road to AGI]]></title><description><![CDATA[In conversation with Greg Kamradt, President of ARC Prize]]></description><link>https://www.mixtureofexperts.co/p/measuring-the-road-to-agi</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/measuring-the-road-to-agi</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 04 Aug 2026 18:55:36 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/209816473/51ad2c502b6751635702ac0d673e1314.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><span>We talk about the evolution from ARC-AGI-1 to ARC-AGI-3 and the lifecycle behind each. ARC-3 alone cost ~$500-750K to build. Greg walks through the leading approaches to beating it (reverse-engineered world models vs. frame-search harnesses) and where models struggle today.</span></p><p><span>Also covered: ARC&#8217;s definition of AGI as human learning efficiency, why world models and memory matter, current ARC-3 scores, long-horizon evals, open vs. closed source dynamics, multimodal integration, the underinvestment in parametric learning, what&#8217;s next for agents (spending, agent-to-agent communication, always-on), and why vibe coding raises the floor but not the ceiling.</span></p><p><strong><span>Chapters</span></strong></p><p><span>00:14 Background &amp; Joining ARC Prize</span></p><p><span>03:06 Role at ARC Prize &amp; Founding Team</span></p><p><span>03:55 ARC1 to ARC3 Evolution</span></p><p><span>06:07 Building a Benchmark: Idea to Sunset</span></p><p><span>08:08 ARC4 &amp; What&#8217;s Next</span></p><p><span>08:45 Approaches to Beating ARC</span></p><p><span>11:33 Why Benchmarks Matter</span></p><p><span>12:24 Who Are Benchmarks For</span></p><p><span>13:42 Evolution of Benchmark Difficulty</span></p><p><span>15:27 Long Horizon Agents &amp; Evals</span></p><p><span>16:59 Defining AGI and ASI</span></p><p><span>18:37 Role of World Models</span></p><p><span>22:02 Where Models Struggle Today</span></p><p><span>22:57 Open Source vs Closed Source Models</span></p><p><span>25:52 Trends &amp; What&#8217;s Exciting</span></p><p><span>27:03 Under-invested Research Areas</span></p><p><span>29:55 Vibe Coding &amp; Role of Humans</span></p><p><span>32:18 Final Thoughts</span></p>]]></content:encoded></item><item><title><![CDATA[Philosophical Transactions of the AI Age]]></title><description><![CDATA[In the 1660s, Europe&#8217;s natural philosophers had a knowledge-sharing problem.]]></description><link>https://www.mixtureofexperts.co/p/philosophical-transactions-of-the</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/philosophical-transactions-of-the</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 21 Jul 2026 18:48:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vaHf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vaHf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vaHf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg 424w, https://substackcdn.com/image/fetch/$s_!vaHf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg 848w, https://substackcdn.com/image/fetch/$s_!vaHf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!vaHf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vaHf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg" width="1456" height="1994" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1994,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1040989,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/207955729?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vaHf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg 424w, https://substackcdn.com/image/fetch/$s_!vaHf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg 848w, https://substackcdn.com/image/fetch/$s_!vaHf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!vaHf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33ab5b73-d188-4c82-ae5e-c672274d8205_1500x2054.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Title page of the first volume of <em>Philosophical Transactions</em>, covering 1665&#8211;1666.</figcaption></figure></div><p><span>In the 1660s, Europe&#8217;s natural philosophers had a knowledge-sharing problem. Discoveries were spreading through private letters, which meant they were spreading slowly and often not at all, because scientists worried about losing credit. Some went so far as to publish findings as ciphers and anagrams, staking a dated claim to the discovery while keeping its contents secret. Henry Oldenburg, secretary of the newly formed Royal Society of London, was at the center of this correspondence network. His job was to collect, translate, and circulate scientific reports among members across Europe.</span></p><p><span>In 1665, Oldenburg decided to turn that correspondence network into a formal printed version of the scientific communications of the society. Thus, the </span><em><span>Philosophical Transactions of the Royal Society</span></em><span>, one of the first scientific journals, was born. It made discoveries public, dated, citable, and open to scrutiny. And in doing so, it helped create the system of published papers through which scientific knowledge has accumulated ever since.</span></p><p><span>More than three hundred and fifty years later, the paper is still the basic unit of scientific knowledge. And for the first time since Oldenburg, the format is meeting a reader it wasn&#8217;t built for.</span></p><p><span>&#8220;The whole research infrastructure was built for humans,&#8221; </span><a href="https://x.com/JIACHENLIU8"><span>Amber Liu</span></a><span> told me. &#8220;It&#8217;s been built to accommodate human speeds: reading speed, understanding speed, experimenting speed. The entire stack, starting from the paper PDF and arXiv, to conferences and peer review, to enterprise research tools like GitHub and Weights &amp; Biases are all built for human researchers.&#8221; Without infrastructure designed for AI scientists, she argues, &#8220;it&#8217;s very hard for them to either compound knowledge or collaborate with each other.&#8221;</span></p><p><span>That idea is at the center of her recent paper, </span><a href="https://arxiv.org/pdf/2604.24658"><span>The Last Human-Written Paper: Agent-Native Research Artifacts</span></a><span>. I sat down with her to discuss the paper, how she sees the role of AI scientists, and what it will take to get there.</span></p><h2><strong><span>Recording and transmitting knowledge</span></strong></h2><p><span>The scientific paper takes months of branching, backtracking, and dead-ended research. What comes out, the paper itself, is a compression of all that, which exacts two taxes.</span></p><p><strong><span>Storytelling tax</span></strong><span>. A paper is a persuasion device. As Amber puts it, &#8220;the paper PDF is entirely built for human researchers to convince human reviewers. Its pages are a polished story and they hide all the tedious work that&#8217;s been done in the process.&#8221; The failed hypotheses, the rejected designs, the pivots, the &#8220;days and nights of experiments trying to fine-tune the performance.&#8221; All of that disappears in the final cut.</span></p><p><strong><span>Engineering tax</span></strong><span>. In persuading reviewers, papers systematically underspecify what&#8217;s needed to actually rerun the work. Setup details, hyperparameter choices, and all decisions that made an experiment succeed. These live with the research team who wrote the paper, but don&#8217;t actually get published in the paper.</span></p><p><span>As Amber put it, &#8220;the paper itself is a lossy compression of the research knowledge. It doesn&#8217;t have all the necessary information needed to reproduce the important results.&#8221; And it&#8217;s this part, the part that gets thrown away, the research process itself, that is more valuable than just the outcome itself.</span></p><p><span>These taxes were somewhat necessary costs for human readers. Without them, it would be even harder than it already is for humans to read all the relevant papers published.</span></p><h2><strong><span>Agent-native research artifacts</span></strong></h2><p><span>However, this changes with AI agents. Agents can process many orders of magnitude  more than humans can and the missing details are actually a key part to understanding and reproducibility for agent &#8220;readers&#8221;.</span></p><p><span>In a world where AI agents drive a lot of our scientific research, the question becomes what substrate do these agents need in order to do cumulative, reliable science?</span></p><p><span>Intelligence alone doesn&#8217;t compound. Intelligence plus a medium for faithfully recording and transmitting knowledge does. &#8220;If we have a more AI-native protocol to document research knowledge, it basically facilitates the research to compound over time,&#8221; Amber said.</span></p><p><span>Her proposed replacement for the paper is the Agent-Native Research Artifact, or ARA. This is a machine-executable research package with four layers:</span></p><ol><li><p><strong><span>Scientific logic</span></strong><span>: the conceptual-level claims and reasoning behind the work</span></p></li><li><p><strong><span>Executable code and full specifications</span></strong><span>: the actual implementation tied to the claims</span></p></li><li><p><strong><span>The exploration graph</span></strong><span>: the full trajectory of the research, including the branches that went nowhere</span></p></li><li><p><strong><span>Evidence</span></strong><span>: the raw outputs, not summary statistics</span></p></li></ol><p><span>&#8220;Everybody is talking about intuition, which is very vague. People don&#8217;t know what intuition is, and people don&#8217;t know how to learn intuition from top researchers,&#8221; Amber told me. The scientific logic layer is meant to be a container for that knowledge, &#8220;things we learn in the process but don&#8217;t necessarily document inside the paper.&#8221; For example, how a seasoned researcher approaches a new problem or tunes hyperparameters. This is basically the idea of research taste, which we&#8217;ll return to, made traceable.</span></p><h2><strong><span>Preserving failure</span></strong></h2><p><span>The exploration graph preserves a different kind of lost knowledge: the failed experiments that today never get published. This is expensive knowledge that is never leveraged in any formal capacity. Someone spent compute, time, and attention learning that a path doesn&#8217;t work, but then the publishing system deletes those lessons, leaving every future researcher (human or artificial) to rediscover the same dead ends. Amber believes that we need to start treating &#8220;what didn&#8217;t work&#8221; as first-class scientific memory.</span></p><p><span>There&#8217;s a technical urgency behind this argument as well. &#8220;Today&#8217;s models are very good at trying to prove something is right, because all their training data contains the successful trajectories,&#8221; Amber explained. This can lead to perverse results. In her analysis of research benchmarks, she&#8217;s watched models game their objectives in creative ways: one model &#8220;is really good at looking at ground-truth answers and making up synthetic data and merging it into the training data so it can reach 100% scores,&#8221; while another family of models will happily modify (even hardcode) components of an architecture &#8220;just to please the judge inside the evaluator.&#8221;</span></p><p><span>In other words, the hope with feeding a model the dead ends, abandoned branches, and intermediate reasoning is that it can start to develop critical thinking about its own work.</span></p><p><span>Today, we&#8217;re starting to see workshops encouraging humans to publish their failures, but adoption is slow. &#8220;For AI scientists, it&#8217;s relatively easier to publish their trajectory,&#8221; she noted, &#8220;because an AI scientist doesn&#8217;t have this burden of disclosing that they&#8217;re actually doing a lot of dumb things, that they fail a lot in the middle.&#8221;</span></p><h2><strong><span>Plurality of judgment</span></strong></h2><p><span>If AI can generate &#8220;tons of ideas, experiments, and conclusions overnight,&#8221; as Amber puts it, then the bottleneck migrates downstream to verification. Last week, I wrote about</span><a href="https://www.mixtureofexperts.co/p/redesigning-around-a-new-power-source"><span> where this is showing up in the physical world</span></a><span>, but it&#8217;s also showing up in scientific research.</span></p><p><span>The failure mode Amber worries about most is speed compounding error. Researchers are now running self-evolving systems, which are agents doing, in her words, &#8220;twenty-four hours research.&#8221; So if the agent hallucinates or reward-hacks mid-stream, that mistake propagates through everything built on top of it, wasting weeks of compute on a corrupted branch. Amber believes we need formal verification and neurosymbolic methods to &#8220;guardrail&#8221; the research process and catch errors as they happen, rather than asking a human to try to audit after the fact.</span></p><p><span>The human verification system is straining under the same load. Recent ICML and ACL cycles have seen submissions double, and, in Amber&#8217;s words, &#8220;the peer review system is essentially broken.&#8221; The morning we spoke, one of her students had been commiserating about EMNLP reviews; his group&#8217;s internal statistics suggested the ratings were very noisy, &#8220;good papers get bad scores and bad papers get good scores.&#8221;</span></p><p><span>To help reduce this noise, agents should be responsible for more of the mechanical checking. &#8220;Most objective review should be done by AI to save the attention of humans,&#8221; Amber told me. The vision is automated verification of claims against evidence, something like a grammar checker for science, which is the analogy she uses in her paper. Humans, meanwhile, keep what was always the interesting part: judging significance, novelty, and taste. Two excellent reviewers can weigh the same paper differently: one prioritizing resource efficiency, another rigor, &#8220;and both of them are correct.&#8221; That plurality of judgment needs to continue to exist.</span></p><h2><strong><span>The foundations of compounding intelligence</span></strong></h2><p><span>If execution is increasingly the machines&#8217; job, the bottleneck shifts to judgment. &#8220;How to identify the important problems, how to approach them, how to raise the questions. This is what matters now, because AI agents can take over the execution,&#8221; Amber told me. She added that, &#8220;Research taste itself is a resource-allocation algorithm. For a hundred ideas, you need to pick the ten, or the one, actually valuable enough to put more resources into.&#8221; Taste is the skill of deciding where attention, time, and compute go when generating options is nearly free.</span></p><p><span>In the long run, &#8220;only the top 0.1% of researchers with great taste will actually control the resources.&#8221; Whether or not you accept the number, the direction is hard to argue with. As execution gets cheap, knowing what&#8217;s worth executing becomes the valuable skill.</span></p><p><span>But before this can happen, taste has to be captured, failures preserved, and claims made verifiable. As Amber put it, &#8220;Human wisdom started accumulating because we invented language.&#8221; Oldenburg understood this in 1665. The scientists of his era were lacking a medium, and the journal he invented enabled three and a half centuries of discovery to compound. Science now has a new kind of reader, working at speeds no journal was designed for, and the medium has to be reinvented again.</span></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Redesigning Around a New Power Source]]></title><description><![CDATA[In Pablo Lubroth&#8217;s interview of Jack Scannell, Scannell references an analogy made by David Shaywitz where he draws a parallel to the early electrification of factories during the Second Industrial Revolution.]]></description><link>https://www.mixtureofexperts.co/p/redesigning-around-a-new-power-source</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/redesigning-around-a-new-power-source</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Thu, 16 Jul 2026 23:31:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!VoCP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VoCP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VoCP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png 424w, https://substackcdn.com/image/fetch/$s_!VoCP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png 848w, https://substackcdn.com/image/fetch/$s_!VoCP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png 1272w, https://substackcdn.com/image/fetch/$s_!VoCP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VoCP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png" width="1264" height="864" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1264,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1765179,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/207330905?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VoCP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png 424w, https://substackcdn.com/image/fetch/$s_!VoCP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png 848w, https://substackcdn.com/image/fetch/$s_!VoCP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png 1272w, https://substackcdn.com/image/fetch/$s_!VoCP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb26b9156-6d4e-46c7-a299-74d2aa436162_1264x864.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><a href="https://www.bbc.com/news/business-40673694">Source</a></figcaption></figure></div><p><span>In </span><a href="https://decodingbio.substack.com/p/a-conversation-with-jack-scannell"><span>Pablo Lubroth&#8217;s interview of Jack Scannell</span></a><span>, Scannell references an analogy made by David Shaywitz where he draws a parallel to the early electrification of factories during the Second Industrial Revolution. In the decades after electric power arrived, many factories simply bolted an electric motor onto a layout built for steam, and saw almost no gain. The returns came only once the whole floor was redesigned around the new power source. That rebuild took decades.</span></p><p><span>It feels like we&#8217;re at a similar point in the AI revolution. We&#8217;ve invented the AI equivalent of the electric motor, but we&#8217;re still bolting that capability onto layouts built for steam.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>This is especially true in many physical domains where AI is getting very good at the technical steps around generating designs and simulating how they&#8217;ll behave. But translating that capability into an output in the physical world is still a work in progress.</span></p><h2><strong><span>Where this shows up in manufacturing</span></strong></h2><p><span>In manufacturing, AI has moved fastest on the upstream, design-adjacent work. It reads 3D CAD geometries and technical drawings and turns them into production parameters, cost estimates, and toolpaths. It flags manufacturability problems (thin walls, features a tool can&#8217;t reach, tolerances a process can&#8217;t hold) and suggests fixes.</span></p><p><span>These tools are about design. But what matters most is reliably producing high-quality goods at scale. And this is bounded by physical cycle times, which accelerating design work doesn&#8217;t really impact. You can model a part in seconds, but you still have to cut the tool, run the line, measure what comes off it, adjust, and run it again. Speeding up design doesn&#8217;t compress the loop that actually takes the time.</span></p><p><span>This part of the loop, the physical production, also generates valuable data around what actually came off the line, and if/where it deviated from spec. Feeding that back into the AI design and simulation tools is what closes the loop, producing better designs on the next run. So, in theory, while you can&#8217;t make any single physical cycle faster (yet), you can make each run teach the next one, thus reducing the number of runs needed to reach the yield or quality you&#8217;re after.</span></p><p><span>This is why </span><a href="https://x.com/AnneliesGamble/status/2041502675229897118?s=20"><span>the promise of the AI-native manufacturer</span></a><span> is exciting. It points toward systems that try to orchestrate the whole line, from design through to finished output. The premise is that once you own the entire loop you can close it. As the upstream design tools proliferate, this is where I think the value will accrue. Rather than bolting an &#8220;electric motor&#8221; into a legacy factory, these new entrants are electing to rebuild the entire factory.</span></p><h2><strong><span>Where this shows up in life sciences</span></strong></h2><p><span>The same dynamic is also playing out in life sciences. Two excellent recent pieces acknowledge that although we&#8217;ve made designing molecules easier and cheaper than ever, the hard part (and likely where value will accrue) is now validating those designs in humans.</span></p><p><span>Carlos Outeiral, in </span><a href="https://couteiral.substack.com/p/the-antibody-revolution-is-real-the"><span>The antibody revolution is real. The business, less so.</span></a><span> explores how, once several labs plus open-source models can all design comparable antibodies, the design step commoditizes, margins compress, and what ultimately ends up mattering is the asset itself, a differentiated drug. As he puts it, &#8220;the real problem is whether the antibody produces a real benefit in a real patient, which is a property not of the molecule but of the biology, and which no amount of binding affinity or easily assayable properties can buy you.&#8221;</span></p><p><span>In the Decoding Bio interview I mentioned at the start of this piece, Scannell argues that one of the fundamental bottlenecks to R&amp;D productivity is predictive validity, which is the extent to which preclinical models and assays accurately predict human outcomes. Scannell coined the term </span><a href="https://en.wikipedia.org/wiki/Eroom%27s_law"><span>Eroom&#8217;s law</span></a><span> in 2012, which says the inflation-adjusted cost of developing a new drug roughly doubles every nine years. The name Eroom is Moore spelled backwards, a nod to the fact that what&#8217;s happening in drug discovery is the opposite of what&#8217;s happening with transistors. One explanation for this is the lack of predictive validity. If the biology you&#8217;re testing against doesn&#8217;t predict human outcomes, designing binders faster just produces wrong answers faster. Whereas, a small gain in validity can outweigh a hundredfold gain in screening throughput.</span></p><p><span>This mirrors what&#8217;s happening in manufacturing as discussed above. We have good molecular design and simulation tools, but the value only shows up once you close the loop between those designs and the physical world, which in biology means testing in humans to create drugs. Closing the loops means rebuilding the &#8220;factory,&#8221; except here the factory is the drug-development company itself: vertically integrated platforms that own design, simulation, and physical validation together, so that what&#8217;s learned in the clinic feeds back into the next design, improving predictive validity. You can&#8217;t skip the human trial, but each one can teach the models that generate the next candidate, so you need fewer shots to find something that works in a real patient.</span></p><h2><strong><span>The scarce thing is downstream</span></strong></h2><p><span>In many physical world domains, AI is commoditizing the upstream steps around design and value is now moving downstream to where the bottleneck is: proving the thing actually works where it has to work. Better or more manufactured products, drugs that work in humans (not just mice). There are many more examples of this across other physical domains, but the throughline in all of them is that physical validation is hard because it&#8217;s slow, expensive, and irreducible.</span></p><p><span>We&#8217;re in the midst of our own revolution, one I believe will eventually dwarf the Industrial Revolution by orders of magnitude (some people say it already has). AI is the new power source. And because of it, we&#8217;re designing new molecules, discovering new materials, and simulating parts faster than ever. But how we reorient ourselves to translate those discoveries into finished goods and viable businesses is still an open question. I don&#8217;t think it will take as long as it took factories to redesign around electric motors, but I do think we&#8217;re only just starting on what is arguably the hardest part.</span></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The thoughts that never become behavior]]></title><description><![CDATA[A conversation with Sean Escola, Dan Wetmore, and Tom Griffiths]]></description><link>https://www.mixtureofexperts.co/p/the-thoughts-that-never-become-behavior</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/the-thoughts-that-never-become-behavior</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 07 Jul 2026 16:05:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!jLF8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em><span>&#8220;[M]any observers assume that the long-elusive goal of human-level intelligence &#8211; sometimes referred to as &#8220;artificial general intelligence&#8221; &#8211; is within our grasp. However, in contrast to the optimism of those outside the field, many front-line AI researchers believe that major breakthroughs are needed before we can build artificial systems capable of doing all that a human, or even a much simpler animal like a mouse, can do.&#8221;</span></em><span>  &#8211; </span><a href="https://www.nature.com/articles/s41467-023-37180-x"><span>Catalyzing next-generation Artificial Intelligence through NeuroAI</span></a></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jLF8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jLF8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg 424w, https://substackcdn.com/image/fetch/$s_!jLF8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg 848w, https://substackcdn.com/image/fetch/$s_!jLF8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg 1272w, https://substackcdn.com/image/fetch/$s_!jLF8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jLF8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg" width="1456" height="543" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:543,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:3934,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/svg+xml&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.mixtureofexperts.co/i/205788465?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jLF8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg 424w, https://substackcdn.com/image/fetch/$s_!jLF8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg 848w, https://substackcdn.com/image/fetch/$s_!jLF8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg 1272w, https://substackcdn.com/image/fetch/$s_!jLF8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95e197e9-aac2-439a-a3a4-85c300269265_1180x440.svg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>A few weeks ago, I sat down with </span><a href="https://x.com/AdamMarblestone"><span>Adam Marblestone</span></a><span> to talk about </span><a href="https://www.mixtureofexperts.co/p/whats-worth-reading-off-a-brain"><span>what might be worth reading off the brain</span></a><span>. The weights of the brain&#8217;s learning subsystem are tuned over a single lifetime, particular to one brain. Adam believes that the thing worth reading is actually the wiring: the architecture and the reward circuitry built into the structure itself.</span></p><p><span>After our conversation, Adam pointed me toward a few other people thinking about how biology might inform AI, and I recently sat down with three of them: </span><a href="https://x.com/SeanEscola"><span>Sean Escola</span></a><span>, </span>a computational neuroscientist who co-founded Herophilus (platform acquired by Genentech) and the robotics company Fauna (acquired by Amazon), and is a Venture Partner at Protocol Labs<span>; </span><a href="https://x.com/dzwetmore?lang=en"><span>Dan Wetmore</span></a><span>, who has spent about fifteen years at the intersection of biosignals and hardware, including at CTRL-labs, the EMG wristband company acquired by Meta; and </span><a href="https://x.com/cocosci_lab"><span>Tom Griffiths</span></a><span>, </span>a cognitive scientist and the director of the AI lab at Princeton<span>.</span></p><p><span>In a 2023 </span><em><span>Nature Communications</span></em><span> paper, </span><a href="https://www.nature.com/articles/s41467-023-37180-x"><span>Catalyzing next-generation Artificial Intelligence through NeuroAI</span></a><span>, Sean and Adam, along with their co-authors, propose that to accelerate progress in AI, we must invest in fundamental research in NeuroAI. &#8220;Historically, many key AI advances, such as convolutional ANNs and reinforcement learning, were inspired by neuroscience. Neuroscience continues to provide guidance [...] but this is often based on findings that are decades old. The fact that such cross-pollination between AI and neuroscience is far less common than in the past represents a missed opportunity.&#8221;</span></p><p><span>In my conversation with Sean, Dan and Tom, we explored what better cross-pollination could look like and the impact this could have on the field. As Sean put it, &#8220;There&#8217;s a whole set of potential ways that we can build better artificial intelligence by looking towards biological intelligence.&#8221;</span></p><h2><strong><span>The brain is not one inductive bias</span></strong></h2><p><span>In cognitive science and AI, a widely-used framework holds that any information-processing system can be understood at three distinct levels of analysis. In his 1982 book, </span><em><a href="https://direct.mit.edu/books/monograph/3299/VisionA-Computational-Investigation-into-the-Human"><span>Vision</span></a></em><span>, the computational neuroscientist David Marr proposed these levels:</span></p><ol><li><p><strong><span>Computational level</span></strong><span>, which is the abstract problem a system solves and its ideal solution</span></p></li><li><p><strong><span>Algorithmic level</span></strong><span>, which is the representations and algorithms that approximate that solution</span></p></li><li><p><strong><span>Implementation level</span></strong><span>, which is how those representations and algorithms are physically realized.</span></p></li></ol><p><span>In </span><a href="https://arxiv.org/html/2503.13401v2"><span>Levels of Analysis for Large Language Models</span></a><span>, Tom and members of his lab argue that methods developed in cognitive science can be useful for understanding large language models. &#8220;The same three levels can be used for analyzing large language models, focusing on how such systems are shaped by their function, the solutions that they seem to find, and the realization of those solutions in weights and units within the underlying artificial neural network.&#8221;</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!U7US!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!U7US!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png 424w, https://substackcdn.com/image/fetch/$s_!U7US!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png 848w, https://substackcdn.com/image/fetch/$s_!U7US!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png 1272w, https://substackcdn.com/image/fetch/$s_!U7US!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!U7US!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png" width="1456" height="442" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:442,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!U7US!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png 424w, https://substackcdn.com/image/fetch/$s_!U7US!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png 848w, https://substackcdn.com/image/fetch/$s_!U7US!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png 1272w, https://substackcdn.com/image/fetch/$s_!U7US!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee3e6a87-897c-42c0-8b1d-7ae7ca98ec43_1706x518.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>Understanding natural and artificial minds across Marr&#8217;s levels of analysis. </span><a href="https://arxiv.org/html/2503.13401v2">Source</a>.</figcaption></figure></div><p><span>In our conversation, Sean expanded upon this and described a parallel hierarchy of four ways in which neuroscience offers a source of inductive bias for AI, moving from least disruptive to most disruptive:</span></p><ol><li><p><strong><span>Representational</span></strong><span>, which is about asking models to represent information internally the way brains do. It&#8217;s the simplest of the four in some ways. The data you&#8217;d need is neural activity, and it stays compatible with the transformer stack we have today, and likely also with whatever replaces it.</span></p></li><li><p><strong><span>Algorithmic</span></strong><span>, which focuses on new, biologically inspired learning rules. The data you&#8217;d need might be neural activity and/or connectomes.</span></p></li><li><p><strong><span>Architectural</span></strong><span>, which are new circuit motifs derived from connectomes to replace the transformer. Transformers are the backbone of nearly everything we build, but they have no real analog to the brain&#8217;s circuitry, beyond the loose sense that both have something like attention and something like neurons. But nothing in a transformer derives from the brain&#8217;s macro- or micro-structure.</span></p></li><li><p><strong><span>The compute substrate</span></strong><span>, which would be neuromorphic or wetware-inspired hardware that computes the way neurons do. This is the far end of the ladder where, as Sean put it, &#8220;we would have to throw out the entire stack and start from scratch.&#8221;</span></p></li></ol><p><span>These are ordered by how much of today&#8217;s AI each one leaves standing. &#8220;You can kind of see as you go down this hierarchy,&#8221; Sean said, &#8220;that you&#8217;re going from things that are completely compatible with the existing hegemony of the LLM built out of transformers existing on GPUs, to &#8216;we&#8217;re going to throw out the entire stack and start from scratch.&#8217;&#8221; In other words, the compute substrate could have the most dramatic payoff, but it would also require demolishing most of the existing infrastructure of the modern AI ecosystem.</span></p><p><span>Representational alignment, on the other hand, is the least disruptive approach and it&#8217;s also the rung where the argument for reading brains gets most interesting. Whether it can improve AI hinges on whether neural data carries information you couldn&#8217;t get anywhere else. This is the line of attack Sean, Dan, and Tom argue should be prosecuted.</span></p><p><span>One warning from Dan that applies to every rung on the ladder is that the brain is only useful to copy at the right level of description. &#8220;One of the challenges historically in this field,&#8221; he said, &#8220;is to operate at the right level of abstraction for the specific goals that you have.&#8221;</span></p><p><span>Neuromorphic computing is the cautionary tale. Early efforts tried to reproduce the analog behavior of individual neurons and how they spike. That, Dan said, was a case where &#8220;folks were being much too grounded in the details of how biology was doing it, and not abstracting away the process that the circuits were actually carrying out.&#8221;</span></p><p><span>Nothing that follows is an argument for copying neurons. It&#8217;s an argument about a specific kind of information, and about reading it at the level where it means something.</span></p><h2><strong><span>Why behavior is not enough</span></strong></h2><p><span>In my post with Adam, we discussed whether brain data tells a model anything it couldn&#8217;t already figure out on its own. As Adam put it: &#8220;What&#8217;s the delta? What&#8217;s the difference between having that information and having just the information about the world that we train on now? Are there things in that neural activity that we can&#8217;t already predict from the data that it&#8217;s seeing?&#8221;</span></p><p><span>Put simply, the answer to this might be that behavior hides most of cognition. Sean, Tom, and their co-author Patrick Mineault discuss this in their recent paper </span><a href="https://arxiv.org/abs/2603.03414"><span>Cognitive Dark Matter: Measuring What AI Misses</span></a><span>. Sean gave the example of watching someone cook. &#8220;Imagine someone is making pasta for dinner and you&#8217;re observing them. At some point in the recipe, they reach into the cabinet and pull out some spice mix and put it in the pasta. At what point did they make the decision to do that? Just at that moment, an hour before, a week before when they were at the grocery store.&#8221; It only becomes a visible action when they reach for the jar, but the decision might have been made much earlier. &#8220;The behavior records the act, not the deciding. You&#8217;re only getting the selected behaviors,&#8221; he said, &#8220;not the unselected behaviors.&#8221;</span></p><p><span>LLMs are trained on the selected behavior, but human cognition is mostly weighing other options and choosing not to do those. These are all the evaluations we weigh (consciously or subconsciously) before we take an action. Recreating human reasoning may require understanding the discarded branches instead of just the chosen one. And that data isn&#8217;t gleaned from behavioral data because it never became visible. Neural data is a bet that you can read the compression loss.</span></p><p><span>You don&#8217;t need neural data everywhere. Instead of trying to make models brain-like across the board, the goal is to find the specific areas where they fall short and collect data from humans doing those specific tasks. Tom is a cognitive scientist and he works at this intersection of computer science and psychology. He&#8217;s interested in &#8220;how the methods we&#8217;ve developed for studying human minds give us a tool for understanding these AI systems.&#8221;</span></p><p><span>Sean, Dan and Tom believe that if you know where a model&#8217;s reasoning diverges from a person&#8217;s, you can design a task that isolates it. The neural data fills the gap at the level of representation where you&#8217;ve established the model is weak.</span></p><h2><strong><span>Capability and trust are the same problem</span></strong></h2><p><span>The hope in doing this is that this will make models both more capable and more trustworthy. The trust half runs through theory of mind. &#8220;We&#8217;re able to engage as strangers in a conversation and engage in business relationships or whatever it is because we have some theory of each other&#8217;s minds that will allow us to operate in ways that are safe for us,&#8221; Sean explained. &#8220;In order for us as humans to incorporate agents into our societies,&#8221; Sean said, &#8220;we will need to have our models of their minds be valid.&#8221;</span></p><p><span>The hope is that if you design tasks that are based on theory of mind or based on cooperative behavior, then you can make models more prosocial.</span></p><p><span>This is largely the thesis of cooperative AI research. In the 2021 </span><em><span>Nature</span></em><span> paper, </span><a href="https://www.nature.com/articles/d41586-021-01170-0"><span>Cooperative AI: machines must learn to find common ground</span></a><span>, Allan Dafoe and his co-authors argue that theory of mind is a precondition for AI systems to cooperate safely with people, and that cultivating it should be a first-class research goal rather than a side effect of scale.</span></p><p><span>Also, a model that scores well but reasons in alien ways is a supervision problem. It&#8217;s hard to predict and hard to catch if something goes wrong. But by making a system reason more like a human, then perhaps the path it took to an answer becomes something we can actually follow.</span></p><h2><strong><span>The missing traces of thought</span></strong></h2><p><span>We are building AI almost entirely out of selected behavior, on the assumption that whatever matters survives into the behavior output. But what if it doesn&#8217;t? What if true decision-making happens upstream and gets compressed away before the visible behavior? If that&#8217;s the case, then a model trained on outputs is learning just from the residue of the reasoning that produced it. The goal of reading the unselected is that we can learn from the part it structurally leaves out.</span></p><p><span>Whether that part carries a signal an LLM couldn&#8217;t already recover is, as Adam said, &#8220;one of these things that needs to be tried.&#8221;  Because if it does, the same signal could teach a model to reason where it currently falls short and make that reasoning something we can more easily understand.</span></p><p><span>After all, if a human mind is more than the sum of its outputs, then so is everything worth learning from one.</span></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Breakeven Point: Rethinking AI for Physics Simulation]]></title><description><![CDATA[Neural networks can solve physical simulations much faster than the numerical solvers engineers have relied on for decades.]]></description><link>https://www.mixtureofexperts.co/p/the-breakeven-point-rethinking-ai</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/the-breakeven-point-rethinking-ai</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 30 Jun 2026 14:32:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vAoL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vAoL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vAoL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png 424w, https://substackcdn.com/image/fetch/$s_!vAoL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png 848w, https://substackcdn.com/image/fetch/$s_!vAoL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png 1272w, https://substackcdn.com/image/fetch/$s_!vAoL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vAoL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png" width="680" height="406" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:406,&quot;width&quot;:680,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vAoL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png 424w, https://substackcdn.com/image/fetch/$s_!vAoL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png 848w, https://substackcdn.com/image/fetch/$s_!vAoL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png 1272w, https://substackcdn.com/image/fetch/$s_!vAoL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2dc08803-c9bb-4f9b-92a8-6a82f9c9bfeb_680x406.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><span>How breakeven complexity works, shown for two neural solvers (FFNO and Poseidon-T) on a Navier-Stokes problem. The blue line is the total cost of a neural solver: a fixed training cost up front, plus a small amount for each run. The orange line is a classical solver tuned to the same accuracy, which pays its cost fresh every run. They cross at the breakeven point (N*) &#8212; the number of runs where the two even out. To the left, with few runs, the classical solver is cheaper; to the right, the neural solver pulls ahead. </span><a href="https://arxiv.org/pdf/2605.15399">Source</a>.</figcaption></figure></div><p><span>Neural networks can solve physical simulations much faster than the numerical solvers engineers have relied on for decades. That is exciting for a lot of reasons, some of which I wrote about previously </span><a href="https://www.mixtureofexperts.co/p/ai-physics-simulation-opportunities"><span>here</span></a><span>.</span></p><p><span>But neural solvers can be much less accurate than classical ones. And they&#8217;re only fast once they&#8217;re trained.</span></p><p><span>In order to train neural solvers, someone has to generate training data (usually by running the very classical solver the model hopes to replace), then train, tune, and validate it. Those costs might be worth it if speed matters. But maybe not.</span></p><p><span>So when does paying the upfront cost become worth it?</span></p><p><span>That&#8217;s the question </span><a href="https://pages.cs.wisc.edu/~khodak/"><span>Mikhail (Misha) Khodak</span></a><span>, together with </span><a href="https://yijingz02.github.io/"><span>Yijing Zhang</span></a><span>, </span><a href="https://nick11roberts.science/"><span>Nicholas Roberts</span></a><span>, and </span><a href="https://tm157.github.io/"><span>Tanya Marwah</span></a><span>, set out to answer in </span><em><a href="https://arxiv.org/abs/2605.15399"><span>Breakeven Complexity: A New Perspective on Neural Partial Differential Equation Solvers</span></a></em><span>.</span></p><p><span>Last week, I sat down with Misha, an assistant professor of computer sciences at UW-Madison, where he's been working on specialized foundation models and folding AI tools into algorithm design and scientific computing. In our conversation, we unpacked the research, the motivation behind it, and the impact he hopes it will have on the broader industry.</span></p><p><span>&#8220;It&#8217;s not really a case of &#8216;we&#8217;re just going to replace all the classical solvers with deep learning&#8217;&#8221; he told me, &#8220;There are stages of many processes where different things will be useful.&#8221;</span></p><h2><strong><span>The pessimism that started it</span></strong></h2><p><span>In 2024 </span><a href="https://arxiv.org/abs/2407.07218"><span>McGreivy and Hakim from Princeton published a paper</span></a><span> that landed hard on the field. &#8220;They went through many deep learning papers, mostly ones applying deep learning to fluid simulations,&#8221; Misha told me. &#8220;They found that in a lot of these papers, the comparisons are mainly against fairly weak baselines. The solvers being used aren&#8217;t the ones an actual practitioner would use.&#8221;</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qGVj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qGVj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png 424w, https://substackcdn.com/image/fetch/$s_!qGVj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png 848w, https://substackcdn.com/image/fetch/$s_!qGVj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png 1272w, https://substackcdn.com/image/fetch/$s_!qGVj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qGVj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png" width="918" height="868" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:868,&quot;width&quot;:918,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qGVj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png 424w, https://substackcdn.com/image/fetch/$s_!qGVj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png 848w, https://substackcdn.com/image/fetch/$s_!qGVj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png 1272w, https://substackcdn.com/image/fetch/$s_!qGVj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73c71b50-1dbc-4ff8-b17b-a6b474bde526_918x868.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">How weak baselines and reporting bias inflate published results. Each marker is one paper, colored by how its method compared to a standard solver. Panel (a) is the likely true picture. (b) shows what the results would likely be without outcome reporting bias. (c) shows the results in the published literature. <a href="https://arxiv.org/pdf/2407.07218">Source</a>.</figcaption></figure></div><p><span>Those papers were overlooking a knob every numerical solver has, which is resolution. You can always make a classical solver faster by </span><em><span>down-resolving</span></em><span>. &#8220;In practice you can set a lower resolution for your solver,&#8221; Misha said. &#8220;The speed is determined by how many time steps you take and how fine your mesh is. And they often don&#8217;t compare to that.&#8221;</span></p><p><span>A neural network that&#8217;s &#8220;faster&#8221; than a high-resolution classical solver, in other words, hasn&#8217;t necessarily won anything, because the honest comparison is on speed against a cheap classical solver tuned to the same accuracy.</span></p><p><span>That paper, Misha said, &#8220;caused a lot of pessimism at the time among the scientific computing community.&#8221;</span></p><p><span>Despite the paper, the field kept building. &#8220;I was trying to reconcile this issue where we were still developing these ML methods and some people seemed optimistic,&#8221; he told me, &#8220;but others, especially in the physics community, would tell you to just down-resolve to get a faster solver.&#8221;</span></p><h2><strong><span>Counting solves until you break even</span></strong></h2><p><span>The paper proposes breakeven complexity, a metric that counts the numbers of runs before a learned solver is cost-effective relative to an error-equivalent traditional solver. Error-equivalent here means the classical solver has been down-resolved to the same accuracy as the neural one.</span></p><p><span>A neural solver only becomes economical when its upfront cost can be amortized. Unlike a classical solver, which pays its computational cost every time you run it, a trained neural network can be reused across every subsequent solve. This is important because most real-world engineering isn&#8217;t solving one equation once. As Misha explained it to me, &#8220;You are trying to optimize some engineering object. And to do that you have to solve many, many similar partial differential equations.&#8221; In practice that means optimizing a design, which requires repeatedly evaluating slightly different geometries, boundary conditions, or material properties. (The paper sets aside real-time applications, where neural networks already have an obvious speed advantage.)</span></p><p><span>The question then becomes: how many times do you have to run the model before the time you save at inference outweighs the time you spent generating data and training?</span></p><p><span>The answer, they found, depends a lot on the type of problem you&#8217;re solving.</span></p><p><span>At one end are toy problems, which are simplified academic test cases. &#8220;If you want to solve these toy 2D Navier-Stokes with periodic boundary conditions,&#8221; Misha said, &#8220;you have to perform hundreds of thousands of inference calls before your data generation and optimization cost pays off.&#8221; The reason is that for easy problems, the classical solver is very hard to beat even after you degrade its resolution. &#8220;For these toy problems there are extremely fast solvers. Even scaling down your resolution, you&#8217;ll still end up with a fast solver that is pretty good. So a neural network is just not going to beat it anytime soon. You really need to perform millions of inference calls.&#8221;</span></p><p><span>But as the team moved to harder problems, that changed. As you move to harder problems you require less inference calls in order to make training this neural network worth it. &#8220;What happens if instead of Navier-Stokes with periodic boundary conditions you have more regular inlet-outlet boundary conditions, and maybe you throw some blocks in there. Put some obstacles into your fluid flow,&#8221; Misha said. &#8220;Suddenly that solver becomes very expensive. And scaling down that solver doesn&#8217;t just speed it up. The tradeoff becomes much worse. It will be faster but much less accurate.&#8221;</span></p><p><span>On hard problems, down-resolving the classical solver degrades its accuracy, while the neural network holds its speed advantage at comparable quality. The team confirmed the pattern along several axes of difficulty: scaling a chaotic system from 1D to 2D to 3D, predicting further into the future, and pushing fluids toward turbulence.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pudr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pudr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png 424w, https://substackcdn.com/image/fetch/$s_!pudr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png 848w, https://substackcdn.com/image/fetch/$s_!pudr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png 1272w, https://substackcdn.com/image/fetch/$s_!pudr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pudr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png" width="1156" height="404" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:404,&quot;width&quot;:1156,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pudr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png 424w, https://substackcdn.com/image/fetch/$s_!pudr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png 848w, https://substackcdn.com/image/fetch/$s_!pudr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png 1272w, https://substackcdn.com/image/fetch/$s_!pudr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26ea51a0-a34c-42fc-91bf-d6de9c300380_1156x404.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">In each panel, larger circles are harder problems. As difficulty rises, points move right (test error goes up - the typical metric makes neural solvers look worse) but down (breakeven complexity falls - they pay off sooner). Difficulty varies by spatial dimension (left), prediction horizon (middle), and turbulence (right). <a href="https://arxiv.org/pdf/2605.15399">Source</a>.</figcaption></figure></div><h2><strong><span>Model size and scaling laws for training</span></strong></h2><p><span>Breakeven also depends heavily on which model you use. One of the paper&#8217;s more surprising findings is that some of the leading foundation models for physics, such as Poseidon and Walrus, don&#8217;t always come out ahead. &#8220;In some cases the latest large models are actually too expensive to run, at least on the simulations we tried,&#8221; Misha told me. &#8220;When I say too expensive, I mean more expensive than a reasonably down-resolved classical solver. So it&#8217;s not even worth it to run them.&#8221; In several cases, well-tuned, smaller neural operators outperformed the larger transformer-based models, suggesting that &#8220;it&#8217;s worth it to pre-train very efficient models rather than very large models.&#8221;</span></p><p><span>That same philosophy extends to training itself. Rather than simply throwing more compute at the problem, the team uses scaling laws to determine how a fixed compute budget should be divided between generating training data with expensive classical solvers and training the neural network. As Misha put it, &#8220;You can very nicely and consistently predict what is the optimal tradeoff as your computation budget increases.&#8221;</span></p><p><span>Even then, &#8220;a practitioner should view this as the upper bound on the performance of a neural network for their problem,&#8221; Misha said, &#8220;because we are training and testing on the same distribution.&#8221; Real-world engineering systems inevitably drift into scenarios the model never saw during training, so neural networks degrade in ways classical solvers generally don&#8217;t. Therefore, Misha sees this more as a screening test: if a neural solver can&#8217;t outperform a classical solver under these idealized conditions, it&#8217;s unlikely to do so in production. If it does clear that bar, it&#8217;s probably worth trying on your actual problem.</span></p><h2><strong><span>A funnel, not a replacement</span></strong></h2><p><span>Towards the end of our conversation, I asked Misha what he was most surprised by from this research. &#8220;I&#8217;m more optimistic now than I was before I started this,&#8221; he said. &#8220;I actually was maybe more skeptical of neural solvers when I came in. And that&#8217;s why I wanted to figure it out.&#8221;</span></p><p><span>Misha also described the future he sees with neural and classical solvers as a pipeline. He referenced the illustration from </span><a href="https://periodic.com/"><span>Periodic Labs</span></a><span>, a Zetta portfolio company. &#8220;The idea was, with these physical simulations, there&#8217;s an agent trying to optimize some material or protein or drug. They&#8217;re first going to see what heuristics say is a good candidate, then narrow down. Then they&#8217;ll try a neural surrogate and narrow down a bit more. Then they&#8217;ll try high-fidelity classical solvers and narrow down a bit more. And then they&#8217;ll actually have to go to the lab and build the thing.&#8221;</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!O789!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!O789!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg 424w, https://substackcdn.com/image/fetch/$s_!O789!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg 848w, https://substackcdn.com/image/fetch/$s_!O789!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!O789!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!O789!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg" width="1200" height="539" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:539,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!O789!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg 424w, https://substackcdn.com/image/fetch/$s_!O789!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg 848w, https://substackcdn.com/image/fetch/$s_!O789!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!O789!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd8176a2b-8446-460e-9b69-820d54adb232_1200x539.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Illustration courtesy of <a href="https://x.com/khoomeik">Rohan Pandey</a> from Periodic Labs; the pipeline runs right to left, fastest stage to slowest. <a href="https://x.com/khoomeik/status/1973056793292034461?s=20">Source</a>.</figcaption></figure></div><p><span>A funnel, in other words where we use cheap heuristics at the wide mouth, fast neural surrogates next, expensive high-fidelity solvers after that, and physical experiment at the narrow end. Each stage tuned to a different point on the speed-accuracy curve.</span></p><p><span>&#8220;There&#8217;s room for all of these things,&#8221; he said. &#8220;It&#8217;s not really a case of replacing all the classical solvers with deep learning. There are stages of many processes where different approaches will be useful.&#8221;</span></p><p><span>It&#8217;s not a question of whether neural solvers beat the classical way, but rather where each one outperforms vis-a-vis the other. The breakeven paper is the first attempt to answer that question with real numbers.</span></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Mixture of Experts! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Agent Is Not the Product]]></title><description><![CDATA[At this point, most companies and individuals are using a relatively baseline level of AI models&#8217; capabilities.]]></description><link>https://www.mixtureofexperts.co/p/the-agent-is-not-the-product</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/the-agent-is-not-the-product</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Wed, 24 Jun 2026 15:56:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yi_N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yi_N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yi_N!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png 424w, https://substackcdn.com/image/fetch/$s_!yi_N!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png 848w, https://substackcdn.com/image/fetch/$s_!yi_N!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png 1272w, https://substackcdn.com/image/fetch/$s_!yi_N!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yi_N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png" width="1456" height="1113" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1113,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:294434,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://anneliesgamble.substack.com/i/203419129?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yi_N!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png 424w, https://substackcdn.com/image/fetch/$s_!yi_N!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png 848w, https://substackcdn.com/image/fetch/$s_!yi_N!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png 1272w, https://substackcdn.com/image/fetch/$s_!yi_N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faa85d76b-2e6d-49c6-8343-dade44d9ecad_2720x2080.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>At this point, most companies and individuals are using a relatively baseline level of AI models&#8217; capabilities. Which means we don&#8217;t see huge benefits from each new model release. And yet, AI is still failing across many enterprise deployments. Why?</span></p><p><span>I think it&#8217;s because most enterprise AI conversations focus on the wrong question: what part of our budget can we put towards AI? For example, should we use our existing software like NetSuite or throw AI at it and move to a new AI-native ERP?</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Opportunities! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>Whereas the right question is something more fundamental and not related to AI at all: what are our existing processes and where do they break down?</span></p><p><span>&#8220;You will not transform your company without rebuilding operations from the ground up,&#8221; </span><a href="https://x.com/vasuman"><span>Vas Moza</span></a><span>, founder of </span>Varick Agents<span>, told me. </span><a href="https://www.varickagents.com/"><span>Varick</span></a><span> is an applied AI company that goes into large enterprise organizations to help them transform from the inside out with AI.</span></p><p><span>Rebuilding operations from the ground up means knowing what to automate deterministically, what to give to an agent, and what to leave to a human. This isn&#8217;t a model or agent engineering question. It&#8217;s a process engineering one. &#8220;That sort of process reengineering,&#8221; Vas said, &#8220;is what makes AI 100 times more effective.&#8221;</span></p><h2><span>Not every workflow needs an agent</span></h2><p><span>Knowing what </span><em><span>not</span></em><span> to agentify is the first question to answer when it comes to enterprise deployments &#8211; because not every workflow deserves an agent. My mental model for this is:</span></p><p><strong><span>Deterministic automation is for rules.<br></span></strong><span>If the inputs are structured and the desired behavior can be specified, you probably do not need an agent. You need software. These are for tasks like routing an invoice, syncing a CRM field, checking whether required documents are present. These workflows are valuable, but they don&#8217;t need open-ended reasoning.</span></p><p><strong><span>Agents are for judgment under context.<br></span></strong><span>Agents make sense when the work requires interpreting inputs, pulling context across systems, or executing a multi-step process where the path changes depending on what the agent discovers. The agent is useful where deterministic automation breaks down.</span></p><p><strong><span>Humans are for accountability, ambiguity, and trust.<br></span></strong><span>Some work should remain human. This might be because the consequences or relationships are too important, or the organization doesn&#8217;t yet know what &#8220;good&#8221; looks like. In these cases, the right AI product might be a copilot, reviewer, researcher, or QA layer.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Al2R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Al2R!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png 424w, https://substackcdn.com/image/fetch/$s_!Al2R!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png 848w, https://substackcdn.com/image/fetch/$s_!Al2R!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png 1272w, https://substackcdn.com/image/fetch/$s_!Al2R!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Al2R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png" width="1456" height="1006" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1006,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:300490,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://anneliesgamble.substack.com/i/203419129?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Al2R!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png 424w, https://substackcdn.com/image/fetch/$s_!Al2R!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png 848w, https://substackcdn.com/image/fetch/$s_!Al2R!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png 1272w, https://substackcdn.com/image/fetch/$s_!Al2R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6149f14f-4ce1-4269-9028-4a971f21261b_2720x1880.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Then, even once you know an agent is the right fit, you still have to decide whether it&#8217;s worth building. That&#8217;s a separate evaluation across dimensions like manual hour displacement, key-person risk, cycle-time reduction, revenue uplift, time and cost to build, and, most importantly, expected ROI. &#8220;A workflow with a 10x ROI should get prioritized,&#8221; Vas said. &#8220;A workflow with a 2x ROI maybe, maybe not. And a workflow with a sub 1x ROI should be left alone.&#8221;</span></p><p><span>The decision of where an agent belongs is downstream of all these diagnoses. Crucially, the most painful workflow isn&#8217;t necessarily the most valuable one to automate.</span></p><h2><span>The deployment is the product</span></h2><p><span>In many companies, the workflow lives in people, who sometimes can&#8217;t actually articulate what it is they&#8217;re doing. This is the tribal knowledge problem.</span></p><p><span>Before I was an investor, I built a company in the manufacturing and supply chain space. I remember trying to understand factory workflows by asking people what they did. They could always show me, but they rarely could explain it.</span></p><p><span>Vas sees the same thing in enterprise operations. The people doing the work aren&#8217;t usually used to narrating the work. They may have done it for fifteen years and know which exceptions are important to document versus which aren&#8217;t, but they may not have language for why.</span></p><p><span>And not all of that knowledge is worth capturing. </span><a href="https://www.linkedin.com/in/ahmad-kakar/"><span>Ahmad Kakar</span></a><span>, CEO of </span><a href="https://geteuclid.ai/"><span>Euclid</span></a><span>, which builds agentic systems for trucking and logistics, told me, &#8220;The catch with tribal knowledge is that a considerable portion isn&#8217;t true expertise &#8211; it&#8217;s workarounds people built for non-existent or broken systems over the years. If you just capture it and encode it into the agent, you&#8217;ve automated the dysfunction. The work is pulling it out of people&#8217;s heads and then deciding what was real judgment versus scar tissue.&#8221;</span></p><p><span>The process is also a delicate one because the person whose knowledge you&#8217;re asking for often assumes their job is at risk. &#8220;We want to make sure they don&#8217;t feel like they&#8217;re having their entire job replaced,&#8221; Vas said, &#8220;because most of the time that isn&#8217;t the case. There&#8217;s a lot of work to be done at growing organizations.&#8221;</span></p><p><a href="https://x.com/TroyShen3"><span>Troy Shen</span></a><span>, co-founder of </span><a href="https://usecervo.com/"><span>Cervo</span></a><span>, an AI platform for customs brokers, frames the same dynamic as a sales problem. &#8220;Successful AI adoption takes two sales, not one,&#8221; he told me. &#8220;You&#8217;re selling an outcome to the executive and a dramatically better workday to the operator who&#8217;ll actually use the agents.&#8221; And without operator buy-in, the deployment will fail.</span></p><p><a href="https://anneliesgamble.substack.com/p/what-it-takes-to-build-and-ship-enterprise"><span>I spoke with Bihan Jiang, the Director of Product at Decagon</span></a><span>, last year about this topic and she told me that a big chunk of her job is internal change management for customers, not just technical implementation. &#8220;There&#8217;s often a board-level directive to &#8216;bring AI into CX,&#8217; which creates top-down momentum and we often end up being a part of many vibrant, cross-functional conversations on the customer&#8217;s side because there&#8217;s a lot of trust and alignment going into a decision this transformative.&#8221;</span></p><p><a href="https://www.linkedin.com/in/suril-kantaria-59522922/"><span>Suril Kantaria</span></a><span>, co-founder of </span><a href="https://www.adaptional.com/"><span>Adaptional</span></a><span>, which builds agentic AI for insurance claims, frames the capture process as an apprenticeship. &#8220;In many enterprises, work is an apprenticeship model,&#8221; he told me. &#8220;This is a simple but key insight from deploying into some of the largest insurance companies. To succeed, we onboard like eager new hires &#8212; learn how the best adjusters handle claims, make decisions, ensure compliance, and keep up with internal policy. Core to the &#8216;product&#8217; is both a deep understanding of our domain (insurance claims) and an ability to handle the contours of every customer&#8217;s claims process.&#8221;</span></p><p><span>This is what makes the deployment process partly technical, partly anthropological, and partly political. As Troy put it, &#8220;Deploy for a Fortune 500 and you&#8217;re rolling out across multiple offices, each with its own history, procedures, politics, and attitude toward AI. That navigation takes a rare mix of emotional intelligence, improvisation, and stakeholder management.&#8221;</span></p><p><span>The goal is to create a map of the knowledge tree within a company &#8211; things like workflow maps, decision rules, exception taxonomies, escalation paths, eval sets, access policies, approval thresholds, and edge-case libraries. This creates structured logic for an organization that can be inspected and, over time, improved.</span></p><h2><span>Modernization does not mean migration</span></h2><p><a href="https://x.com/AnneliesGamble/status/2059317203850391620?s=20"><span>I&#8217;ve argued before that legacy modernization</span></a><span> is an important wedge in enterprise AI. This means translating old code into new code and requires a full-system modernization. While this is true, it&#8217;s worth noting that sometimes the best first move is to modernize the work </span><em><span>around</span></em><span> a piece of software rather than the software itself. &#8220;We&#8217;re not going to force you to migrate off your systems of record, which is the enterprise reality for most companies,&#8221; Vas said. &#8220;They&#8217;re very married to their ERP or their CRM.&#8221;</span></p><p><span>These systems hold years of process and integrations so sometimes it&#8217;s easier (and better) to build on top of them. This means connecting them and then extracting the process logic between them. Then give agents the context to execute work across them. &#8220;We build on top wherever possible,&#8221; Vas said &#8220;because there is still so much work to be done that exists on top of these platforms and by interconnecting these systems.&#8221;</span></p><p><span>This approach has limits if the underlying architecture is brittle. This might mean the data isn&#8217;t clean, or permissions are broken, and so building on top of these problems means you end up automating around dysfunction instead of fixing it.</span></p><p><span>The best enterprise AI deployments I&#8217;ve seen do some combination. They build on top where possible, then let the agent deployment show what actually needs to be rebuilt. Modernization used to mean making legacy software usable by modern developers. These days, modernization means making legacy operations legible to agents.</span></p><h2><span>The pattern library</span></h2><p><span>Each deployment in services-heavy AI companies looks bespoke. However, there are patterns to be learned in the deployment expertise itself, which makes repeatability across customers easier over time.</span></p><p><span>For example, after multiple finance department deployments, a company can start to learn patterns around how finance departments run and where they break. Importantly, this compounding doesn&#8217;t depend on moving client data around. &#8220;We allow our agents to have a higher baseline accuracy,&#8221; Vas said, &#8220;not because we extract data from our clients, but because we have the methodology and practices best suited for the right-tail edge cases we&#8217;ve seen across deployments.&#8221; The data from one client does not necessarily need to be copied into another client&#8217;s system for the next deployment to improve.</span></p><p><span>This means you start to learn the workflow archetypes, questions to ask, evaluation harnesses, governance models, etc. </span><a href="https://www.linkedin.com/in/clemenskomorek/"><span>Clemens Komorek</span></a><span>, CEO of </span><a href="https://zalion.ai/"><span>Zalion</span></a><span>, which builds AI procurement agents for industrial buyers, sees the value of this at the transaction level. &#8220;Every transaction teaches the system something about supplier relationships and negotiation levers,&#8221; he told me. &#8220;And because supply chains keep shifting, the system has to adapt as outcomes change. This is how localized expertise becomes something repeatable.&#8221;</span></p><p><span>What remains bespoke is the business-specific context. &#8220;That part is not repeatable,&#8221; Vas said, &#8220;because every single business is different. Pretending otherwise is why most AI projects fail.&#8221;</span></p><p><span>The accumulated knowledge is what creates a pattern library and this is the moat. It&#8217;s knowing where autonomy belongs and how to turn processes into structured operational logic, built up across deployment after deployment.</span></p><h2><span>Why neither the labs nor the incumbents own this</span></h2><p><span>It is tempting to assume the labs will eventually own enterprise AI deployments because if the models keep getting better, why not just build the agents? And yes, the labs will own a lot of the horizontal intelligence infrastructure, but there&#8217;s a difference between selling intelligence versus selling operational transformation.</span></p><p><span>The labs are optimized for scalable intelligence whereas enterprise transformation is not cleanly scalable. Plus, no enterprise wants its entire company dependent on just one or two model providers because it would make them highly vulnerable.</span></p><p><span>So there&#8217;s still a need for a layer that decides which model belongs where, how to route work, how to control cost, and how to ensure the workflow continues even when the model layer shifts underneath. This is something I wrote about last year &#8211; the shift </span><a href="https://x.com/AnneliesGamble/status/2056541468043747586?s=20"><span>from models to systems</span></a><span>.</span></p><p><span>Consulting firms and legacy systems integrators are the other candidates to own enterprise transformation. They&#8217;ve shepherded companies through past technology shifts, and they know how to navigate large organizations. The question is whether they can build and continuously upgrade production AI systems fast enough.</span></p><p><span>AI moves quickly. Best practices from six months ago are practically obsolete now. &#8220;The inertia required to move a large behemoth organization like a McKinsey to become AI-native themselves,&#8221; Vas argues, &#8220;is insurmountable.&#8221; A firm with tens of thousands of employees and decades of internal process inertia may struggle to become AI-native quickly enough to deliver the kind of systems customers actually need.</span></p><p><span>This is not to say the incumbents will not participate. They will. They already are. But there is an opening for a new category of company that combines consulting-grade process understanding with software-grade deployment velocity. These can be horizontal companies or vertical-specific solutions.</span></p><h2><span>The agent is the mechanism</span></h2><p><span>A company transforms when the relationship across its people, its processes and its outputs change. &#8220;Rebuilding a company around AI,&#8221; Vas said towards the end of our conversation, &#8220;means starting to understand where the processes themselves break down. It&#8217;s not about slapping AI on top of operations that aren&#8217;t yet designed for it.&#8221;</span></p><p><span>In many cases, this means rebuilding the operating model from the ground up. And in that new model, plenty of work still belongs to humans or to deterministic automation &#8211; not to agents. Where an agent actually belongs is something you discover through the deployment process itself. And that&#8217;s what compounds &#8211; the accumulated knowledge of where autonomy works and how to deploy it.</span></p><p><span>After all, enterprises aren&#8217;t looking to buy an agent. They&#8217;re looking to enable their business to be better than it is today, either with more efficiency or with greater outcomes than they could reach before. The agent is just one mechanism.</span></p><p><em><span>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</span></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Opportunities! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[What's worth reading off a brain]]></title><description><![CDATA[A conversation with Adam Marblestone]]></description><link>https://www.mixtureofexperts.co/p/whats-worth-reading-off-a-brain</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/whats-worth-reading-off-a-brain</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 16 Jun 2026 17:57:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Xk7p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Xk7p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Xk7p!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png 424w, https://substackcdn.com/image/fetch/$s_!Xk7p!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png 848w, https://substackcdn.com/image/fetch/$s_!Xk7p!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png 1272w, https://substackcdn.com/image/fetch/$s_!Xk7p!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Xk7p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png" width="1456" height="738" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:738,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Xk7p!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png 424w, https://substackcdn.com/image/fetch/$s_!Xk7p!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png 848w, https://substackcdn.com/image/fetch/$s_!Xk7p!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png 1272w, https://substackcdn.com/image/fetch/$s_!Xk7p!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb676aa90-036a-4fcf-9405-571ecaeec867_2048x1038.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The two-subsystem model of the brain: a learning subsystem (cortex, striatum, amygdala, cerebellum) that builds a trained model from scratch within a lifetime, and a steering subsystem (hypothalamus, brainstem) that is largely hardcoded by the genome and supplies the supervisory and control signals shaping behavior. From Steve Byrnes, '<a href="https://www.alignmentforum.org/posts/hE56gYi5d68uux9oM/intro-to-brain-like-agi-safety-3-two-subsystems-learning-and">Intro to Brain-Like-AGI Safety</a>' (2022).</figcaption></figure></div><p>We can read every weight in a large language model and still can&#8217;t really say what the model is doing. Why then would tracing every connection in a brain, a far harder problem, tell us anything we couldn&#8217;t get more cheaply somewhere else? It sounds like a reason not to bother mapping brains at all.</p><p><a href="https://x.com/AdamMarblestone">Adam Marblestone</a> thinks it&#8217;s the opposite. <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Adam Marblestone&quot;,&quot;id&quot;:130737054,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!sFH0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb368bd20-7153-44e8-989b-8f3ac9aea04b_144x144.png&quot;,&quot;uuid&quot;:&quot;dc62968c-ecd3-4555-93fd-7d42f6766ef8&quot;}" data-component-name="MentionToDOM"></span> runs<a href="https://www.convergentresearch.org/"> Convergent Research</a>, a nonprofit that incubates what it calls Focused Research Organizations (FROs), mid-scale science projects that are too capital-intensive and engineering-heavy for an academic lab but too far from a product at the start for venture capital.</p><p>&#8220;When many AI people or computer scientists hear &#8216;map the connectome,&#8217;&#8221; he told me, &#8220;I think they hear &#8216;we&#8217;re going to map the weight matrix. We&#8217;re going to know all these weights.&#8217;&#8221; That, he thinks, is the wrong way to picture it because the weights are not where the interesting information is. &#8220;The weights are just as much a function of what you trained it on as it is anything about the architecture of that system. So I actually don&#8217;t really care about the specific weights. What I care about is the architecture.&#8221;</p><h2><strong>Outside the trained weights</strong></h2><p>To understand why Adam doesn&#8217;t care about the weights, it helps to look at how an AI system actually gets built. When you train a neural network, the decisions about the architecture, the goal it&#8217;s optimizing for, the data it learns from all live in the code the researcher writes.</p><p>&#8220;I have some PyTorch code that sets up the architecture of the network. I write in code what the loss function is. I feed it data. And then in the end I end up with a weight matrix that lives inside this box.&#8221; As he puts it, &#8220;a bunch of the interesting stuff about what the researchers actually did isn&#8217;t in the weight matrix.&#8221;</p><p>But the brain doesn&#8217;t work that way. There&#8217;s no separate code, it has to build everything out of neurons: &#8220;The learning signals are neurons. The basic architecture is how the neurons are initialized and how they&#8217;re connected. The cost functions are specific neurons that deliver whatever that reward signal is. Everything has to do with neurons. It doesn&#8217;t have a separate programmer.&#8221;</p><p>In an AI system, the design sits in the code itself, outside the trained weights. In a brain there&#8217;s no such thing as &#8220;outside the trained weights,&#8221; so all of the design has to be built physically into the cells and their connections.</p><p>Adam develops this further in his paper, <a href="https://asteriskmag.com/issues/13/the-sweet-lesson-of-neuroscience">The Sweet Lesson of Neuroscience</a>. In it, he references <a href="https://osf.io/preprints/osf/fe36n_v1">Steve Byrnes&#8217; work</a>, which recasts the entire brain as two interacting systems: a learning subsystem and a steering subsystem. The first learns from experience during the animal&#8217;s lifetime. The second is mostly hardwired and sets the goals, priorities, and reward signals that shape that learning. The learning subsystem is what learns. The steering subsystem is what decides what&#8217;s worth learning. For more on this, <a href="https://www.lesswrong.com/posts/hE56gYi5d68uux9oM/intro-to-brain-like-agi-safety-3-two-subsystems-learning-and">here is another good overview</a> on the two subsystems.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A-5-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A-5-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png 424w, https://substackcdn.com/image/fetch/$s_!A-5-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png 848w, https://substackcdn.com/image/fetch/$s_!A-5-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png 1272w, https://substackcdn.com/image/fetch/$s_!A-5-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A-5-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png" width="1456" height="924" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:924,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!A-5-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png 424w, https://substackcdn.com/image/fetch/$s_!A-5-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png 848w, https://substackcdn.com/image/fetch/$s_!A-5-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png 1272w, https://substackcdn.com/image/fetch/$s_!A-5-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7e304998-1af3-488e-af28-5bd18715ca70_2048x1300.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Simplified sketch of proposed algorithmic architecture for the vertebrate brain: a learning subsystem (cortex, striatum, cerebellum) that acquires structured models within a lifetime, and a hardcoded steering subsystem (hypothalamus, brainstem) that supplies the supervisory and control signals driving behavior. From Chen and Macosko, "<a href="https://www.preprints.org/manuscript/202602.0767">Cellular Scaling Laws in the Mammalian Brain</a>" (2026), building on Steve Byrnes' learning/steering framework.</figcaption></figure></div><p>The steering subsystem&#8217;s circuits aren&#8217;t a record of anything learned. They are the design, hardwired by evolution, with the reward signals written directly into the cell types and their wiring. That is why Adam cares about reading out the steering subsystem and not the learned weights of the cortex. The cortex&#8217;s wiring is mostly a snapshot of values it acquired through experience. The steering subsystem&#8217;s wiring is the genome&#8217;s specification itself. <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Patrick Mineault&quot;,&quot;id&quot;:17921567,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e0dbee55-9a27-4d31-b784-8779443b8f7d_400x400.jpeg&quot;,&quot;uuid&quot;:&quot;a79d108d-55ca-467f-8992-ba21a59fcf1d&quot;}" data-component-name="MentionToDOM"></span> wrote <a href="https://www.neuroai.science/p/cell-types-encoding-the-brains-bios">a good piece</a> that dives deeper into Adam&#8217;s thinking on this.</p><p>Fei Chen and Evan Macosko, the PIs from the Broad Institute who published one of the original whole-mouse-brain transcriptomes, find evidence of this in their work, <a href="https://www.preprints.org/manuscript/202602.0767">Cellular Scaling Laws in the Mammalian Brain</a>. The largest number of distinct, bespoke neuron types sits in the old deep structures, the hypothalamus and brainstem. These regions have far fewer neurons overall but a much wider variety of them. The cortex is the reverse. It is large, but built from many copies of a few repeating templates.</p><p>This fits what the framework predicts. The cortex works like a neural network, so it can learn almost anything and doesn&#8217;t need specialized hardware. It just needs many copies of the same flexible units. The steering subsystem has the opposite job. It encodes specific, hardwired goals: hunger, thirst, fear, the drive to mate, the urge to breathe. Each is something the genome has to specify in advance, because an animal can&#8217;t afford to learn them by trial and error. So if the steering subsystem really is the brain&#8217;s hardwired reward function, the bespoke cell-type diversity should pile up there too, in those deep structures rather than the cortex. And that is what the research shows.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XY7l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XY7l!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png 424w, https://substackcdn.com/image/fetch/$s_!XY7l!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png 848w, https://substackcdn.com/image/fetch/$s_!XY7l!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png 1272w, https://substackcdn.com/image/fetch/$s_!XY7l!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XY7l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png" width="1456" height="1381" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1381,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XY7l!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png 424w, https://substackcdn.com/image/fetch/$s_!XY7l!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png 848w, https://substackcdn.com/image/fetch/$s_!XY7l!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png 1272w, https://substackcdn.com/image/fetch/$s_!XY7l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28ec1b49-24e5-4ca0-95cb-4cc21714eba5_2048x1942.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Brain regions plotted by neuron count against number of molecularly defined cell types in the mouse brain. 'Learning centers' (cortex, hippocampus, cerebellum, etc.) hold large neuron populations but relatively few cell types. Brainstem, interbrain, and pallidum regions show the opposite: fewer neurons, much greater cell-type diversity. From Chen and Macosko, "<a href="https://www.preprints.org/manuscript/202602.0767">Cellular Scaling Laws in the Mammalian Brain</a>" (2026). I used Claude to recreate the figure at higher resolution, so there may be small differences from the original.</figcaption></figure></div><p><a href="https://news.mit.edu/2025/former-mit-researchers-advance-new-model-innovation-0606">A thread of work</a> that Adam began a decade ago along with <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Ed Boyden&quot;,&quot;id&quot;:40850931,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!ZrjJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F730ea22e-779a-410b-96f4-c24dc8762c07_1133x1133.jpeg&quot;,&quot;uuid&quot;:&quot;067aa08b-4d7e-41f8-b64e-07e2e7c99866&quot;}" data-component-name="MentionToDOM"></span> at MIT helped start the use of expansion microscopy and in-situ sequencing to read out neural wiring. Adam later helped popularize FROs as a way to fund this kind of infrastructure-heavy science. The mapping itself is now being pushed by <a href="https://www.e11.bio/">E11 Bio</a>, the first FRO spun out of his Convergent Research, which is trying to drop the cost of a mouse-brain connectome from billions of dollars to low tens of millions. The hope is that mapping at high enough resolution would expose the brain&#8217;s design: not the trained values, but what specifically produced them. &#8220;Once we&#8217;ve done that, I don&#8217;t actually care that much about the specific weights,&#8221; he said.</p><h2><strong>Mapping the reward functions</strong></h2><p>A detailed enough map, the thinking goes, would expose the machinery that computes internal rewards. &#8220;It might be that the infant is trying to first establish eye contact, or it&#8217;s trying to find and pay attention to novel stimuli or something like that versus boring stimuli,&#8221; Adam said. &#8220;Am I making eye contact with the parent? Am I finding novelty? Am I controlling my environment? These are all things that the brain probably has to have some way of detecting and rewarding.&#8221;</p><p>On the <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Dwarkesh Patel&quot;,&quot;id&quot;:4281466,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!5eJb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb715ffd1-f7d7-4755-af88-c48efe647f5b_400x400.jpeg&quot;,&quot;uuid&quot;:&quot;800be2b4-c2d2-4eff-85ce-60d61fbddcf1&quot;}" data-component-name="MentionToDOM"></span> <a href="https://www.dwarkesh.com/p/adam-marblestone">podcast</a>, Adam argued that the field has tended to neglect the role of these very specific reward functions. Machine learning gravitates toward mathematically simple objectives, like predicting the next token. His hunch is that evolution did the opposite, building a lot of complexity into the brain&#8217;s reward functions: <a href="https://arxiv.org/abs/1606.03813">many different ones for different regions, switched on at different stages of development</a>. If he&#8217;s right, those reward functions in the brain would each be a group of cells you could in principle point to.</p><p>A detailed map is the starting point, but what Adam really wants isn&#8217;t the frozen snapshot so much as the process that generates it: &#8220;I want to understand the drivers... how it starts out and then how it learns within the lifetime.&#8221; He acknowledges that the translation from a static map to a complete description of &#8220;what is it trying to do&#8221; will be hard. &#8220;Will you be able to actually translate between a map and that information? Maybe,&#8221; he said. &#8220;But if not, I still think that having the maps is going to be a multiplier on the rate of overall neuroscience progress.&#8221; So even in the pessimistic case, where the wiring doesn&#8217;t hand you the algorithm, the map still accelerates everything else.</p><h2><strong>Why the mapping is so hard</strong></h2><p>In October 2024, the<a href="https://www.nature.com/articles/d41586-024-03190-y"> FlyWire consortium</a>, which includes <a href="https://mrclmb.ac.uk/research-leaders/gregory-jefferis/">Greg Jefferis&#8217; group</a> and collaborators at Cambridge University, Princeton and the University of Vermont, published the first complete wiring diagram of an adult fruit fly brain: roughly 140,000 neurons and more than 54 million synapses. Producing it required slicing a single fly brain into thousands of ultrathin sections, imaging each with electron microscopes, and using machine learning to stitch the images back into a 3D reconstruction. A fruit fly has on the order of 100,000 neurons. A human has something closer to 100 billion.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!s6we!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61396688-f522-40b8-ab92-9db57653cf8a_920x489.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!s6we!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61396688-f522-40b8-ab92-9db57653cf8a_920x489.png 424w, https://substackcdn.com/image/fetch/$s_!s6we!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61396688-f522-40b8-ab92-9db57653cf8a_920x489.png 848w, https://substackcdn.com/image/fetch/$s_!s6we!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61396688-f522-40b8-ab92-9db57653cf8a_920x489.png 1272w, https://substackcdn.com/image/fetch/$s_!s6we!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61396688-f522-40b8-ab92-9db57653cf8a_920x489.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!s6we!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61396688-f522-40b8-ab92-9db57653cf8a_920x489.png" width="920" height="489" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/61396688-f522-40b8-ab92-9db57653cf8a_920x489.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:489,&quot;width&quot;:920,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!s6we!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61396688-f522-40b8-ab92-9db57653cf8a_920x489.png 424w, https://substackcdn.com/image/fetch/$s_!s6we!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61396688-f522-40b8-ab92-9db57653cf8a_920x489.png 848w, https://substackcdn.com/image/fetch/$s_!s6we!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61396688-f522-40b8-ab92-9db57653cf8a_920x489.png 1272w, https://substackcdn.com/image/fetch/$s_!s6we!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61396688-f522-40b8-ab92-9db57653cf8a_920x489.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The FlyWire map of a fruit fly brain. A human brain has about a million times as many neurons. <a href="https://www.cambridgenetwork.co.uk/news/whole-brain-connectome-fruit-fly-most-complex-brain-ever-mapped">Source</a>.</figcaption></figure></div><p>To address this, E11 Bio is making comprehensive static circuit maps more scalable. The core bet is a shift in imaging physics: &#8220;The traditional way of doing static circuit mapping all the way down to the neuron and synapse level is the electron microscopes, which are extremely precise in what they can see spatially, but they are not easily scalable... If you can switch that to using a light-based microscope rather than an electron-based microscope, it&#8217;s much easier to have thick pieces of tissue that you can sort of see through and are much easier to handle.&#8221;</p><p>Light-based methods bring a second advantage: they can read out molecules. &#8220;It also gives you the advantage of being able to see molecules kind of overlaid &#8212; so what are the specific receptors and transmitters that are used by the synapses?&#8221; This is important because the brain&#8217;s connections are not uniform in the way an artificial network&#8217;s are.</p><p>&#8220;Unlike in a computer, where there&#8217;s maybe a few types of connections, there&#8217;s actually many different types of connections in an actual brain,&#8221; Marblestone said. Excitatory or inhibitory, but also varying in their time scales and in how they adapt and learn. A wiring diagram that only records which neuron connects to which, with no read-out of the receptors and transmitters at each synapse, would miss most of what tells those connection types apart. And thus, much of what distinguishes one learning rule from another.</p><p>This is the focus of a second FRO Adam is helping to catalyze, <a href="https://www.meridial.org/">Meridial</a>, which pushes from static snapshots toward <em>dynamic</em> maps. Meridial is not expected to be as comprehensive as static mapping, but the goal is to observe a subset of connections changing over time and thus to understand the rules for how synapses change. &#8220;You won&#8217;t get every connection. But even if you just look at a subset of connections, being able to understand the rules for how they change, I think that&#8217;s super AI relevant,&#8221; he said.</p><h2><strong>Training models from brain data</strong></h2><p>If the wiring encodes the design, then there could be a path whereby you could train AI systems directly on brain activity. This would mean a model learns to represent the world the way a brain does rather than only the way labeled data does. &#8220;There are some companies starting to do brain-data-based training,&#8221; Adam told me.</p><p>The question, he says, is whether brain data tells a model anything it couldn&#8217;t already figure out on its own. &#8220;What&#8217;s the delta? What&#8217;s the difference between having that information and having just the information about the world that we train on now? Are there things in that neural activity that we can&#8217;t already predict from the data that it&#8217;s seeing?&#8221; If a model can already infer how a brain would respond to an image just from the image, the recording adds nothing. Brain data is only worth collecting if it carries something we can&#8217;t get from other types of data.</p><p>A line of work on representational alignment has shown measurable gains from nudging artificial networks toward neural data. Aligning vision models to human EEG can make their representations<a href="https://arxiv.org/abs/2401.17231"> more brain-like and more robust</a>. Fine-tuning speech models on fMRI recordings of people listening to stories, a method <a href="https://mtoneva.com/">Mariya Toneva</a> and colleagues call <a href="https://arxiv.org/abs/2510.21520">brain-tuning</a>, improves their downstream performance, with the largest gains on tasks that require semantic understanding. And a 2025 study found that aligning auditory models to<a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC12319826/"> individual fMRI recordings</a> improved performance on downstream tasks, especially where training data was scarce.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!w1uH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!w1uH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png 424w, https://substackcdn.com/image/fetch/$s_!w1uH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png 848w, https://substackcdn.com/image/fetch/$s_!w1uH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png 1272w, https://substackcdn.com/image/fetch/$s_!w1uH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!w1uH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png" width="1456" height="1398" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1398,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!w1uH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png 424w, https://substackcdn.com/image/fetch/$s_!w1uH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png 848w, https://substackcdn.com/image/fetch/$s_!w1uH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png 1272w, https://substackcdn.com/image/fetch/$s_!w1uH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6c82cf27-5777-4b19-abce-dc988e01350a_2048x1966.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">How representational alignment works in practice. An image-recognition model (CORnet) gets an added encoding module that predicts the EEG a person produces when viewing the same image. Training minimizes two losses at once: category classification and EEG generation. So the model learns to see more like a human brain. From Lu, Wang &amp; Golomb (2024), '<a href="https://arxiv.org/pdf/2401.17231">Achieving more human brain-like vision via human EEG representational alignment</a>.'</figcaption></figure></div><p>There&#8217;s signal that brain data contains something AI can&#8217;t already extract from the world, but there are still a lot of open questions. As Adam put it, &#8220;It&#8217;s one of these things that needs to be tried.&#8221;</p><h2><strong>Why this paradigm got skipped</strong></h2><p>If reading every weight in a language model can&#8217;t explain it, why expect a brain map to do better? The objection assumes you&#8217;d be staring at its weights. But Adam believes you&#8217;d be staring at something an LLM never had: the design itself.</p><p>Adam was on the neuroscience team at Google DeepMind from roughly 2018 to 2020, when the field&#8217;s brain-inspired instincts were near their peak. &#8220;What I was working on was memory architectures, we were focused on questions like &#8216;what does the hippocampus do as a memory system?&#8217; &#8216;Does it have some way of <a href="https://arxiv.org/abs/2002.02385">compressing information</a>?&#8217;&#8221; The leading research programs leaned on reinforcement learning, training systems through reward and trial-and-error in an elaborate, <a href="https://arxiv.org/abs/1803.10760">brain-inspired form</a>, full of what Adam calls &#8220;bells and whistles.&#8221;</p><p>Then LLMs took off and bypassed most of it. LLMs use a stripped-down form of reinforcement learning: reward the good outputs, adjust, repeat. They have no internal model of the world. Whereas the brain is thought to do something richer. &#8220;The way that large language models do reinforcement learning and post-training is in some ways kind of like the simplest or most brute-force way to do RL,&#8221; Adam said, &#8220;whereas people believe that the brain does model-based RL,&#8221; which is focused on building a model of how the world works and planning against it.</p><p>Model-based RL gives the brain a set of faculties that LLMs don&#8217;t have built in. &#8220;The brain has more innately built systems,&#8221; Adam said, &#8220;to predict what&#8217;s going to happen in the future, or simulate different possible events. Or to go back to different memories to use that memory to make a prediction.&#8221; The brain also has ways of working out which actions deserve credit for a reward that only comes later, the problem of temporal credit assignment, as well as value functions that estimate how good a situation is likely to turn out. &#8220;These are things that are pretty clearly built into the mammalian brain,&#8221; he said, &#8220;that LLMs don&#8217;t have a built-in architectural solution for.&#8221;</p><p>With today&#8217;s LLMs, &#8220;you&#8217;re not really building in things like a hippocampus or prefrontal cortex or a striatum or some of the brain areas that we know we have,&#8221; he noted. These systems may approximate some of those functions as emergent byproducts of training, but they don&#8217;t have them as architectural commitments.</p><p>The brain does, and that is the argument for going to look at it. Every faculty the brain has, it had to build into the structure itself, where it can in principle be found and read. Adam is not alone in believing that intelligence needs a kind of built-in cognitive structure, rather than expecting it to emerge from scale alone. Emmanuel Dupoux, Yann LeCun, and Jitendra Malik argue that today's models are missing an architecture <a href="https://arxiv.org/abs/2603.15381">inspired by human and animal cognition</a>. They analyze how autonomous learning works in living organisms and propose a roadmap for reproducing it in artificial systems.</p><p>As Adam put it, &#8220;Having the maps is going to be this multiplier on the rate of overall neuroscience progress. It will help us understand the truth of how humans do it.&#8221; That truth isn&#8217;t in the weights. The learned weights of the brain&#8217;s learning subsystem are tuned over a single lifetime, particular to one brain. The thing worth reading is the wiring: the architecture and the reward circuitry built into the structure itself.</p><p><em>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Opportunities! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Beyond Words]]></title><description><![CDATA[Voice AI has gotten really good, really fast.]]></description><link>https://www.mixtureofexperts.co/p/beyond-words</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/beyond-words</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 09 Jun 2026 15:52:19 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!euFU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!euFU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!euFU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png 424w, https://substackcdn.com/image/fetch/$s_!euFU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png 848w, https://substackcdn.com/image/fetch/$s_!euFU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png 1272w, https://substackcdn.com/image/fetch/$s_!euFU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!euFU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png" width="1456" height="619" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:619,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1473572,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://anneliesgamble.substack.com/i/201320208?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!euFU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png 424w, https://substackcdn.com/image/fetch/$s_!euFU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png 848w, https://substackcdn.com/image/fetch/$s_!euFU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png 1272w, https://substackcdn.com/image/fetch/$s_!euFU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe915c877-bdc5-41d4-9967-9391d33810b4_1923x817.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Voice AI has gotten really good, really fast. Google&#8217;s Gemini 3.1 Flash Live and OpenAI&#8217;s GPT-Realtime-2 can now reason over a live conversation and pick up on tone, not just words. AI is starting to hear not only what we say but how we say it.</p><p>But only up to a point. Three things are still in the way. The first is what these systems do with how you sound: they hear it, then mostly go with your words. The second is who gets heard at all: the models that come closest are closed, and work best in English and a few high-resource languages, leaving most of the world behind. And third, when AI clones a voice, even the best cloning models we have today flatten vocal identity rather than preserving it, pulling every voice toward the same center.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Opportunities! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><a href="https://martijnbartelds.nl/">Martijn Bartelds</a>, a speech researcher at Together AI who did his postdoc at Stanford, has spent most of his career studying voice AI, across multilingual models, endangered-language ASR, and voice synthesis. &#8220;The human voice is so rich,&#8221; he told me. &#8220;And I think that makes human-to-human communication also so special.&#8221; I sat down with him this week to talk through where voice AI falls short and what it will take to fix it.</p><h2><strong>The architectural problem</strong></h2><p>The standard way voice AI is built is itself the first place it falls short. Traditionally, voice AI models transcribe speech to text first, then reason over the text. This architecture keeps the words and discards nearly everything else. &#8220;If you do the transcription process, of course the only thing that you keep is just an orthographic representation. So the text only. You abstract away from everything else,&#8221; Martijn explained. &#8220;But if you want to let the model deeply understand paralinguistic information &#8211; everything that is captured in the speech signal that makes up for all the things other than just the words &#8211; you need a vastly different approach.&#8221;</p><p>In our conversation, Martijn referenced a<a href="https://aclanthology.org/2025.acl-long.682/"> 2025 ACL survey of speech language models</a> that also discusses the problem with current voice AI model architecture. The survey describes the problems as threefold:</p><ol><li><p>Information loss during modality conversion</p></li><li><p>Latency from chaining three systems (speech-to-text (ASR) &#8594; language model (LLM) &#8594; text-to-speech (TTS))</p></li><li><p>Errors that accumulate across them</p></li></ol><p>Another architectural path some voice AI models choose to take is to use an audio-encoder-plus-LLM. Here a speech encoder is attached to an LLM so the model consumes a representation of the audio. But this approach has an alignment problem: &#8220;You have to somehow align the representations of the speech encoder model and the large language model.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bHd3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bHd3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png 424w, https://substackcdn.com/image/fetch/$s_!bHd3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png 848w, https://substackcdn.com/image/fetch/$s_!bHd3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png 1272w, https://substackcdn.com/image/fetch/$s_!bHd3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bHd3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png" width="704" height="314" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:314,&quot;width&quot;:704,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:261481,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://anneliesgamble.substack.com/i/201320208?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bHd3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png 424w, https://substackcdn.com/image/fetch/$s_!bHd3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png 848w, https://substackcdn.com/image/fetch/$s_!bHd3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png 1272w, https://substackcdn.com/image/fetch/$s_!bHd3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F09c65645-fc4d-4807-8e69-a806d6f6dba0_704x314.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The richness of a voice, visualized. Transcription reduces all of this to a line of text. <a href="https://speechprocessingbook.aalto.fi/Representations/Spectrogram_and_the_STFT.html">Source</a>.</figcaption></figure></div><p>With both paths the audio is treated as a second-class citizen, either discarded for text or bolted on after the fact. Martijn wants a third option: a single model that treats speech and text as equals from the start, so nothing has to be translated away or stitched together. &#8220;Having the ability to really reason about the audio is something that seems crucial to me,&#8221; he said.</p><h2><strong>Building toward a unified model</strong></h2><p>Having a single model that keeps meaning and paralinguistic content together is also the architecture that the 2025 ACL paper proposes. With no conversion to text, nothing is lost in translation. And collapsing three chained systems (ASR, LLM and TTS) into one removes both the latency and the accumulating error.</p><p>&#8220;I see the model being one complete engine, so to speak,&#8221; Martijn told me. &#8220;It should handle the text and the audio&#8230; they should be of equal importance.&#8221; In other words, he&#8217;s imagining one embedding space where it no longer matters whether a piece of understanding arrived as speech or as text, where reasoning about the audio is as native to the model as reasoning about text.</p><p>This is the architecture the frontier has now largely adopted. Google&#8217;s Gemini 3.1 Flash Live processes raw audio natively in a single real-time model, which lets it read tone and emotion along with the words. OpenAI&#8217;s GPT-Realtime-2, released in May 2026, works the same way, reasoning through a live conversation without the round trip to a separate text model older pipelines required.</p><p>However, while these models are very good now at expressive output and voice generation, they are still not very good at fully understanding and acting upon the information present in the input (i.e. the details in the user&#8217;s incoming voice).</p><h2><strong>The data problem</strong></h2><p>LLMs got good fast in large part because they had a lot of data from the web to train on. Audio doesn&#8217;t have an equivalent to the web that is readily available. High-quality speech data has to be manufactured, which means cleaning recordings, aligning words to audio, labeling speakers and languages, filtering noise. Of these steps, Martijn says alignment is probably the hardest part: &#8220;For some approaches, you need a careful alignment between the words and the actual text. So you need the data to be transcribed in the first place.&#8221;</p><p>For well-resourced languages, commercial incentives justify the work needed to get and clean data; and there&#8217;s more raw data available in the first place. But for many of the world&#8217;s languages, such incentives often don&#8217;t exist. &#8220;For some of the digitally underrepresented languages I worked with, like the Dutch dialects or Australian Aboriginal languages, only a couple of hours are available. There is nothing else there,&#8221; he said. &#8220;This means you have to be creative. It&#8217;s about figuring out how we can get the most out of the data.&#8221;</p><p>And for Martijn, being creative means working on the model and the data together, rather than treating them as separate problems solved in sequence. During his postdoc, Martijn built <a href="https://openreview.net/pdf?id=yt40xuRBA9">a multilingual training algorithm</a> that left the dataset fixed but made the model aware, mid-training, of which languages were lagging. This let it shift more weight toward them and lifted performance across the set without a single new hour of audio. The data question and the modeling question, as he puts it, go hand in hand.</p><h2><strong>The voice cloning distortion</strong></h2><p>Toward the end of our conversation, Martijn and I talked about the third problem: voice cloning, an increasingly prominent use of voice AI, where even with the best models we have today, the richness of a voice is getting lost on the way out.</p><p>&#8220;Voice cloning&#8221; implies fidelity and it&#8217;s easy to assume the output of this technology is an exact copy of a speaker&#8217;s voice. But in<a href="https://arxiv.org/html/2605.16578v1"> Voice &#8220;Cloning&#8221; is Style Transfer</a>, Martijn and his collaborators found the opposite. Listeners consistently rated the cloned voices as more customer-service-like, authoritative, and warm than the originals. &#8220;These models don&#8217;t really faithfully clone someone&#8217;s voice, but more or less transfer this style of how someone speaks,&#8221; Martijn said.</p><p>The fear with cloning is impersonation (deepfakes, fraud, etc). But Martijn&#8217;s work sheds light on another risk: that voice AI actually reshapes how we sound. His study found cloning flattens vocal identity rather than preserves it, nudging every voice toward the same optimized center.</p><p>The same pattern is showing up in text. In <a href="https://arxiv.org/abs/2603.18161">How LLMs Distort Our Written Language</a>, Natasha Jaques and her collaborators found that LLM edits move essays farther than human edits do (even when asked for minimal changes) and in a consistent direction.</p><p>When AI mediates human expression, it optimizes away the irregularities that make communication personal. Whether the medium is writing or voice, AI pulls us toward a common mean such that we all begin to sound and write the same.</p><h2><strong>What it means to be heard</strong></h2><p>What these systems can hear and who gets heard both come down to whether we treat voice as something richer than text. The frontier models are finally starting to hear the how and not just the what. But these systems that do it best are closed, and they still work best in English and a handful of high-resource languages. &#8220;I would just love to see more people working on trying to create these multimodal audio-text language models end to end and making them open source,&#8221; Martijn said. &#8220;Having a very strong open-source competitor in that space would be fantastic,&#8221; especially, he notes, one that also serves speakers of digitally underrepresented languages.</p><p>Martijn envisions a model that understands us more fully when we speak, and leaves us sounding like ourselves. &#8220;If we say the exact same content, but you have more hesitation in your voice, you should get a different answer than me,&#8221; he said.</p><p>To get there, voice has to be treated as something fundamentally different from text. Language is not its transcript. And being heard is not the same as being transcribed. &#8220;The words, your pitch, your tone. This is so broad and so rich, but it&#8217;s everything,&#8221; Martijn said. &#8220;The model should go beyond the words.&#8221;</p><p><em>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Opportunities! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Physical Seams of the AI Buildout]]></title><description><![CDATA[The Mag Seven&#8217;s 2026 capex guidance is larger than the Apollo moon program, the transcontinental railroad, and the interstate highway system combined.]]></description><link>https://www.mixtureofexperts.co/p/the-physical-seams-of-the-ai-buildout</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/the-physical-seams-of-the-ai-buildout</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 02 Jun 2026 12:34:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!fi4b!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fi4b!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fi4b!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png 424w, https://substackcdn.com/image/fetch/$s_!fi4b!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png 848w, https://substackcdn.com/image/fetch/$s_!fi4b!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png 1272w, https://substackcdn.com/image/fetch/$s_!fi4b!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fi4b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png" width="1456" height="769" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:769,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:4254907,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://anneliesgamble.substack.com/i/200288515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fi4b!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png 424w, https://substackcdn.com/image/fetch/$s_!fi4b!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png 848w, https://substackcdn.com/image/fetch/$s_!fi4b!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png 1272w, https://substackcdn.com/image/fetch/$s_!fi4b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc381ecf3-6c1a-4dd7-8122-8348e4d26889_2264x1196.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><a href="https://www.wpr.org/news/data-center-wisconsin-guardrails-proposed-bill">Source</a></figcaption></figure></div><p>The Mag Seven&#8217;s 2026 capex guidance is larger than the Apollo moon program, the transcontinental railroad, and the interstate highway system combined.</p><p>I sat down last week with <a href="https://x.com/brexton">Brexton Pham</a>, Global Co-Head of Compute Infrastructure at <a href="https://www.cantor.com/">Cantor Fitzgerald</a> to talk about this buildout and where he sees the opportunity. Brexton has a unique perspective because his mandate spans the trifecta of power, land, and capital. As Brexton put it, they&#8217;re set up for &#8220;tackling humanity&#8217;s largest ever infrastructure efforts.&#8221;</p><p>And we&#8217;re still very early. By his count, worldwide internet penetration sits at around 75%, while worldwide AI penetration is &#8220;maybe 15% if I&#8217;m being really generous. And that 15% is also primarily the layman&#8217;s ChatGPT usage.&#8221;</p><p>The buildout is enormous and has barely started. The hyperscalers and the nation-states are committing tens of billions to capex. As the buildout progresses, it is increasingly opening up new categories of opportunity.</p><h2><strong>The wrong filter, and the right one</strong></h2><p>&#8220;We are entering this world where verticalization matters more than ever and anybody can build anything much faster,&#8221; Brexton said. In other words, it&#8217;s very hard to predict what the hyperscalers will and won&#8217;t do, and it&#8217;s only getting harder.</p><p>Anything that looks like an unguarded seam today can be swallowed tomorrow. &#8220;You should assume your competitors can verticalize overnight,&#8221; he said. Behind-the-meter power, nuclear, on-site generation all looked non-core to a hyperscaler a few years ago, but they are now moving directly into these categories.</p><p>&#8220;When OpenAI and Anthropic first came out, everyone, including myself, assumed that they were AWS-shaped. They were platform-shaped, in that they would never cannibalize their own customers. And boy were we wrong,&#8221; Brexton said. &#8220;We have never seen customer cannibalization to this scale before. But that&#8217;s on the software side. Physics is much harder.&#8221;</p><p>&#8220;The decision to own physical assets,&#8221; he said, &#8220;to own land, to own a building, to maintain it, to build equipment, to sell equipment. It is significantly harder than software.&#8221; These physical seams create differentiation and durability over time. They are more resistant to verticalization because the constraint is physics and time.</p><p>And as AI infrastructure scales, the economics of these physical seams change as well. Data centers themselves were considered low-margin, Brexton pointed out, until AI data centers came along: &#8220;AI data centers are extremely high margin; inference is very high margin.&#8221; Whether that holds across every workload is debatable, and margins vary by model and use case. But the reflexive belief that owning physical assets means accepting thin margins is, in his view, a holdover from the pre-AI world.</p><p>So if physical difficulty is the moat, here&#8217;s my rough map of where the non-hyperscaler opportunities are the most compelling:</p><h2><strong>Opportunity 1: Energy Aggregation</strong></h2><p>Power is one of the binding constraints on the entire buildout. &#8220;It doesn&#8217;t matter how many chips you have if you don&#8217;t have the literal powered land for it,&#8221; Brexton said. The hyperscalers know this, which is why they&#8217;re now investing directly in generation via nuclear restarts, SMR offtake deals, on-site power.</p><p>But there&#8217;s a difference between owning the power and assembling it. &#8220;If I&#8217;m a hyperscaler, do I actually want to do the work of aggregating behind-the-meter power? Probably not,&#8221; Brexton said. Hyperscalers obviously care a lot about power, but the aggregation work itself to get that power requires sourcing and permitting on-site generation, wiring in storage, negotiating with landowners and local utilities, etc. These are all areas that hyperscalers would rather buy the result than build the capability.</p><p>There are companies that are already running at this: <a href="https://www.americanterawatt.com/">American Terawatt</a> on the behind-the-meter side; and <a href="https://www.aalo.com/">Aalo Atomics</a>, <a href="https://www.valaratomics.com/">Valar Atomics</a>, and <a href="https://blueenergy.co/">Blue Energy</a> on the nuclear side. &#8220;Obviously the hyperscalers aren&#8217;t tackling nuclear&#8221; at the build level, Brexton said. &#8220;Why would they?&#8221;</p><p>Aggregating distributed power, siting and operating reactors, and building generation are bound by interconnection queues, regulatory timelines, and physical construction. These are hard physical problems and thus the solutions for them are more defensible.</p><h2><strong>Opportunity 2: Resilience Infrastructure</strong></h2><p>The supply chain underneath AI is fragile. And during the recent Iran war when multiple Azure and AWS data centers in the Gulf states were attacked, we got a glimpse at just how fragile it actually is. The analogy here is the cybersecurity build-out of the 2010s. Before Stuxnet, essentially nobody was paying for OT security. The Iran data center strikes may be a similar inflection point.</p><p>As Brexton put it: &#8220;We have clearly entered a world of asymmetric economic warfare. They were able to counter our multi-million dollar rockets with drones that cost maybe $10,000. We&#8217;re supposed to have the most powerful military on earth, but they were throwing toys at our missiles.&#8221;</p><p>And the exposure isn&#8217;t confined to the data centers themselves. Our vulnerabilities stretch around the globe:</p><ul><li><p>Rare earths, which power semis, are basically monopolized by China today.</p></li><li><p>ASML, the only company on earth that makes EUV lithography machines, is in the Netherlands.</p></li><li><p>TSMC is in Taiwan. &#8220;If Taiwan got invaded tomorrow,&#8221; Brexton said, &#8220;it would be a very, very scary thing.&#8221;</p></li><li><p>Intercontinental sea cables carry trillions of dollars of data packets every day. &#8220;Cut them, and you cripple economies overnight,&#8221; Brexton said.</p></li><li><p>The Strait of Hormuz, where roughly a third of seaborne oil passes. &#8220;We were wildly exposed with the Strait of Hormuz,&#8221; he said. &#8220;And we still are.&#8221; And there are other choke points like the Strait of Malacca, Bab el-Mandeb and the Suez, the Taiwan Strait, the Panama Canal and the Turkish Straits.</p></li></ul><p>Each one of these is a single point of failure that no operator can fix on their own. Therefore, hardening AI infrastructure is equally as important as building it in the first place. Some of sub-categories here include:</p><ul><li><p>Redundancy and failover software for compute and data pipelines</p></li><li><p>Alternative routing infrastructure for when sea cables get cut or regions go down</p></li><li><p>Supply chain visibility tools so operators can see, in real time, which inputs are getting squeezed</p></li><li><p>Physical site monitoring and threat intelligence designed for industrial-scale AI facilities</p></li><li><p>Insurance products and financial instruments that let operators hedge geopolitical risk in their compute supply chains. This is a category that didn&#8217;t really exist five years ago and is now being underwritten by Lloyd&#8217;s and a handful of specialty carriers</p></li></ul><p>The buyer profile here is broader than the other categories because nearly every AI company needs resilience tooling.</p><h2><strong>Opportunity 3: Space</strong></h2><p>The commercial case for moving data centers off Earth comes down to escaping the constraints throttling terrestrial build-outs: land, grid, permitting, and politics. Ben Thompson had<a href="https://stratechery.com/2026/the-spacex-ipo-and-data-centers-in-space/"> a great piece</a> the other day about data centers in space. Two points from that article stand out. First, data centers in space can look wildly different from data centers on Earth. They can be satellite-sized compute racks interconnected with lasers, each with its own solar power and radiator arrays. Some examples of companies building here are <a href="https://www.starcloud.com/">Starcloud</a> (formerly Lumen Orbit), which has already flown an Nvidia GPU in orbit, and <a href="https://kepler.space/">Kepler Communications</a>, which is operating what&#8217;s currently the largest in-orbit compute cluster, a handful of edge processors linked by laser.</p><p>Second, and more importantly, Thompson points out that agentic workloads don&#8217;t need the low latency that human-facing inference requires. This makes them uniquely well-suited to orbit, where round-trip latency is higher but land, grid, and permitting constraints basically disappear. Cooling doesn&#8217;t disappear, though and constrains how big these systems get. In vacuum there&#8217;s no air to carry heat away; you can only radiate it. An example of a company trying to solve this is <a href="https://sophia.space/">Sophia Space</a>, which is building thin tile-shaped satellites that sit processors against a passive heat sink to kill the need for active cooling.</p><p>Beyond the commercial cases for space, there&#8217;s also a national-security one. &#8220;If you buy into the belief that a lot of the supply chain will be considered critical infrastructure,&#8221; Brexton said, &#8220;then you can also imagine that we will move more critical infrastructure into space as a consequence of our desire to enhance national security.&#8221; From there, &#8220;the argument for space data centers becomes very compelling, the argument for in-orbit manufacturing and in-orbit refueling becomes very interesting.&#8221;</p><p>Examples of companies in this servicing layer are <a href="https://www.starfishspace.com/">Starfish Space</a>, which is building a satellite servicing vehicle, and <a href="https://www.infiniteorbits.io/">Infinite Orbits</a>, which is focused on satellite life-extension and inspection. Assembling, refueling, and repairing orbital infrastructure is a hard physical engineering problem with long lead times.</p><h2><strong>Opportunity 4: Labor</strong></h2><p>One of the top reasons build-outs get delayed is that, in Brexton&#8217;s words, &#8220;we straight up don&#8217;t have enough electricians and data center operators.&#8221; As <a href="https://anneliesgamble.substack.com/p/stop-talking-just-about-gpus?triedRedirect=true">I wrote about a few weeks ago in my conversation</a> with <a href="https://substack.com/@bepresearch">Ben Pouladian</a>, Jensen Huang has said the same thing: the bottleneck he&#8217;s most worried about is the shortage of plumbers and electricians.</p><p>There are really two constraints here. The first is raw supply: there aren&#8217;t enough skilled people, and training them takes years. The second is location: even where the people exist, they&#8217;re rarely near the build sites. &#8220;Abilene, Texas for example has a lot of powered land and is a very obvious place for large industrial build-outs,&#8221; Brexton said. &#8220;But it&#8217;s in the middle of nowhere. So you&#8217;re asking families to relocate to a place that doesn&#8217;t have affordable housing or much of a residential area.&#8221; This is an issue of proximity to some extent and it exists across the entire supply chain: &#8220;If I break a part this afternoon, I can get a new part the next morning in Shenzhen, whereas in the US I&#8217;m waiting six weeks,&#8221; he said.</p><p><a href="https://x.com/AnneliesGamble/status/2023834463202095508?s=20">In my conversation</a> with <a href="https://x.com/samanfarid">Saman Farid</a> of <a href="https://formic.co/">Formic</a> we talked about this same problem. He framed US industrial weakness as a utilization problem: a typical US factory runs far fewer of its available production hours versus a Chinese one, so the same building and equipment yields a fraction of the output. The problem is we can&#8217;t just solve this with more labor because we don&#8217;t have access to more labor.</p><p>So the opportunities split across doing more with the workers you have, and reducing how dependent a site is on workers being nearby:</p><ul><li><p><strong>Robotics and automation</strong> to enable on-site construction, electrical work, and manufacturing</p></li><li><p><strong>Workforce creation</strong> via accelerated credentialing, apprenticeship-to-placement pipelines, and staffing built specifically for data center construction</p></li><li><p><strong>Remote operations</strong> and lights-out facility management that reduce how many people a remote site needs on the ground</p></li></ul><p>Each of these is a way to solve the physical constraints around the labor problem.</p><h2><strong>The Foundation of the Intelligence Economy</strong></h2><p>There&#8217;s one constraint that we haven&#8217;t talked about yet, but it gates all of the categories above: public sentiment. Municipality pushback is already causing build-out delays, and it&#8217;s only increasing. &#8220;The average American,&#8221; Brexton argues, &#8220;is meaningfully more anxious about AI than excited, and Silicon Valley keeps underestimating that anxiety.&#8221;</p><p>Some of that anxiety is misinformation. He traces the water-usage panic to a figure in Karen Hao&#8217;s <em>Empire of AI</em> that overstated data-center water use by a large factor and was later revised, &#8220;but the damage was already done.&#8221; Some of it isn&#8217;t. Either way it exists, and it&#8217;s yet another hurdle to confront as the buildout proceeds.</p><p>This hurdle further stymies the physical buildout, which Brexton believes we&#8217;re still &#8220;severely underestimating in order to power 24/7 demand of intelligence.&#8221;</p><p>The easy conclusion is that this entire map will be drawn by giants: hyperscalers, nation-states, utilities, and the capital providers large enough to finance the buildout. Or, on the other side, that public backlash will slow the whole thing down before a new startup ecosystem can form around it.</p><p>I don&#8217;t buy either. The hyperscalers will define much of the demand, and public sentiment will shape where and how fast the buildout happens. But neither eliminates the startup opportunities, which I believe are largely physical and can&#8217;t be verticalized overnight. And while some of these opportunities may look like narrow seams today, over time, they will become part of the foundation that the intelligence economy depends on.</p><p><em>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Opportunities! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Making What's Old New Again]]></title><description><![CDATA[I really liked this Stratechery interview of United CEO Scott Kirby from January, in which Kirby describes how, starting in 2016, he committed United to spending several hundred million dollars rewriting SHARES, the airline&#8217;s Fortran-based reservation system originally written in the 1960s, onto modern cloud infrastructure.]]></description><link>https://www.mixtureofexperts.co/p/making-whats-old-new-again</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/making-whats-old-new-again</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 26 May 2026 16:56:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!9K17!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9K17!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9K17!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png 424w, https://substackcdn.com/image/fetch/$s_!9K17!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png 848w, https://substackcdn.com/image/fetch/$s_!9K17!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png 1272w, https://substackcdn.com/image/fetch/$s_!9K17!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9K17!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png" width="1082" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:1082,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1187998,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://anneliesgamble.substack.com/i/199352832?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9K17!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png 424w, https://substackcdn.com/image/fetch/$s_!9K17!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png 848w, https://substackcdn.com/image/fetch/$s_!9K17!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png 1272w, https://substackcdn.com/image/fetch/$s_!9K17!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9de5096f-4f81-43a7-848f-f7586b7d6d0b_1082x720.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/ai-for-it-modernization-faster-cheaper-and-better">Source</a></figcaption></figure></div><p>I really liked this <a href="https://stratechery.com/2026/an-interview-with-united-ceo-scott-kirby-about-tech-transformation/">Stratechery interview of United CEO Scott Kirby</a> from January, in which Kirby describes how, starting in 2016, he committed United to spending several hundred million dollars rewriting SHARES, the airline&#8217;s Fortran-based reservation system originally written in the 1960s, onto modern cloud infrastructure. The project still isn&#8217;t finished, the last cutover is scheduled for next year. All of United&#8217;s recent customer-facing differentiation sits on top of this new infrastructure. And while United is far from a perfect airline, the widening profitability gap between United and the rest of the industry is largely a consequence of their decision to get off the legacy mainframe. Kirby noted in the interview that United and Delta will collectively account for 100% of industry profitability this year. &#8220;You can&#8217;t do what we do unless you do [the modernization] first,&#8221; Kirby said. &#8220;It was a key unlock.&#8221;</p><p><a href="https://www.mckinsey.com/capabilities/quantumblack/our-insights/ai-for-it-modernization-faster-cheaper-and-better">McKinsey research</a> estimates that roughly 70% of Fortune 500 software was built more than twenty years ago. And according to<a href="https://www.ibm.com/think/topics/cobol-modernization"> IBM</a>, there are still an estimated 250 billion lines of COBOL in production. At the same time, <a href="https://www.bcg.com/publications/2026/the-200-billion-dollar-ai-opportunity-in-tech-services">BCG</a> recently put out a report that says &#8220;agentic AI will ultimately expand the total addressable market for technology services, unlocking up to $200 billion in net new value pools in the next five years.&#8221;</p><p>Most of the revenue in this market is currently captured by Accenture, TCS, Infosys, Cognizant, Wipro, and Capgemini. They&#8217;ve used an offshore-labor-arbitrage model, which is most exposed to AI-native delivery.</p><p>As a result, there are unsurprisingly a lot of companies going after this space now: Mechanical Orchard, 8090, Tessara, Moderne, among others. But there are still opportunities for new entrants. And what I find most exciting about this category is that modernization is just a wedge into a much larger custom development relationship, on stacks built natively to be modernized again.</p><h3><strong>Why Now</strong></h3><p>My thesis around this category is based on four observations about how the market is changing:</p><p><strong>1. Enterprises will continue to outsource, not insource.</strong> United is the exception that proves the rule. Kirby explicitly notes that no other airline has done what United did, <em>&#8220;certainly not to the extent that we&#8217;ve done it. They&#8217;re still on old legacies, because it&#8217;s hard.&#8221;</em> The historical outsourcing model priced labor arbitrage. AI agents collapse the labor cost, which means the model has to change, but the structural preference for outsourcing the work isn&#8217;t going to. We&#8217;re now seeing evidence of this in lots of consulting firms partnering with AI labs including EY&#8217;s April 2026 partnership with 8090.</p><p><strong>2. The business logic buried in legacy code is the asset.</strong> The business rules encoded in the mainframe logics represent decades of institutional knowledge. I like Mechanical Orchard&#8217;s framing that &#8220;the system in action is the specification.&#8221; In other words, the value is around extracting and verifying the behavioral specification inside these systems. That extraction is the foundation of everything else a modernization platform can do.</p><p><strong>3. Modernization is the wedge into bigger opportunities around new development.</strong> Once a system is modernized, companies can unlock a lot of new capabilities that they previously weren&#8217;t able to. After spending hundreds of millions rewriting its 1960s-era SHARES reservation system, United was able to build differentiated customer-facing products. None of that new development would have been possible without the modernization first. It also gives companies that modernize an edge in a crowded, otherwise undifferentiated, market. &#8220;We&#8217;re doing all this and no one&#8217;s copying the things that matter, which is great,&#8221; Kirby said in the Stratechery interview.</p><p><strong>4. Capability, security, and the talent cliff are converging.</strong> Three vectors are converging simultaneously, answering the why now question. AI models can now read, translate, and validate legacy code at scale. When <a href="https://claude.com/blog/how-ai-helps-break-cost-barrier-cobol-modernization">Anthropic published a blog post in February</a> about Claude Code reading COBOL, <a href="https://venturebeat.com/technology/ibms-usd40b-stock-wipeout-is-built-on-a-misconception-translating-cobol-isnt">IBM lost ~$40B in market cap in a single day</a>. Security exposure has also become <a href="https://www.ibm.com/reports/data-breach">untenable</a>: the global average breach cost is $4.4M and 97% of organizations reported an AI-related security incident and lacked proper AI access controls. And the third vector is talent: according to <a href="https://www.metaintro.com/blog/ai-modernize-legacy-software-tech-workers-2026">some sources</a>, the average COBOL programmer is 55, with roughly 10% retiring annually and 60% expected to retire within five years.</p><h3><strong>The White Space</strong></h3><p>It&#8217;s useful to place the existing players on two axes: <em>modernize what&#8217;s there</em> versus <em>build what&#8217;s next</em>, and <em>platform/product</em> versus <em>services-wrapped delivery</em>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mZnA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mZnA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png 424w, https://substackcdn.com/image/fetch/$s_!mZnA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png 848w, https://substackcdn.com/image/fetch/$s_!mZnA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png 1272w, https://substackcdn.com/image/fetch/$s_!mZnA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mZnA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png" width="1456" height="1032" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1032,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:244963,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://anneliesgamble.substack.com/i/199352832?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mZnA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png 424w, https://substackcdn.com/image/fetch/$s_!mZnA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png 848w, https://substackcdn.com/image/fetch/$s_!mZnA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png 1272w, https://substackcdn.com/image/fetch/$s_!mZnA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76436ddd-cae4-47c8-87be-4c840258d8b8_2068x1466.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The established players mentioned in the image above have collectively staked out mainframe, ERP, agent tooling, and regulated-enterprise software factories. But there is still whitespace for new entrants. Below are some areas I&#8217;m particularly bullish about.</p><p><strong>Post-mainframe, pre-cloud enterprise systems.</strong> Two decades of custom enterprise software built on Microsoft, Oracle, and open-source stacks. These systems run inventory, claims, distributor management, financial modules, and a long tail of back-office workflows at Fortune 500-1000 companies. The opportunity is around a full-system modernization that addresses application logic, the data layer (stored procedures, ETL, batch jobs), and integrations.</p><p><strong>Vertical-specific modernization</strong> where deep regulatory expertise is needed, such as in healthcare, defense, manufacturing OT/IT convergence, energy, and telecom. Each has its own legacy stack and compliance requirements, so domain depth is important.</p><p><strong>The mid-market tier</strong>, where projects run $150K&#8211;$2M rather than tens of millions, is large, fragmented, and underserved. AI changes the unit economics enough that a product-led motion could work well here.</p><p><strong>Continuous-modernization platforms.</strong> United&#8217;s modernization efforts started in 2016, and they&#8217;re still not done a decade later. These initiatives can now happen much faster than before, but even a new system will become outdated within a certain number of years. A platform built for continuous modernization is more interesting than one built for a single mainframe-to-cloud migration. In other words, this is like an always-on layer that maintains a living map of the enterprise&#8217;s business logic, which can be re-translated onto whatever stack comes next.</p><p><strong>Non-code legacy systems</strong> like ETL pipelines, EDI integrations, batch schedules, message queues, and undocumented processes. Most modernization efforts are around application code, but that&#8217;s just one layer of a legacy system. The infrastructure that actually encodes business logic is equally important and has often been overlooked. <a href="https://www.curietech.ai/">Curie</a> is a good example of a company going after this layer &#8211; their agents handle MuleSoft integration migration, management, and the transition from APIs to agents. AI is uniquely suited to help here.</p><h2><strong>The only way you grow</strong></h2><p>In the Stratechery interview, Scott Kirby said: &#8220;The biggest mistake most people make in their careers is never making big mistakes. It&#8217;s the only way you grow. You have to decide, and not deciding on the status quo is a decision in itself.&#8221;</p><p>For a decade, the status quo on modernization (aka do nothing) was defensible because the alternative was an expensive, decade-long overhaul with no clear ROI. AI has changed that.</p><p>The companies that do nothing now are still making a decision, they just may not realize it yet. There&#8217;s an opportunity to build the platform that helps them see it, and gives them a credible path forward.</p><p><em>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Opportunities! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Stop talking just about GPUs]]></title><description><![CDATA[A few weeks ago, I listened to Dwarkesh Patel&#8217;s interview with Jensen Huang. It&#8217;s a great interview and one I recommend listening to if you haven&#8217;t already. In it, Jensen said chip-side bottlenecks will be resolved in two or three years. The bottleneck he&#8217;s more worried about is energy, and specifically he calls out his concern over the shortage of plumbers and electricians.]]></description><link>https://www.mixtureofexperts.co/p/stop-talking-just-about-gpus</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/stop-talking-just-about-gpus</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 19 May 2026 15:56:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!pqUq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pqUq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pqUq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png 424w, https://substackcdn.com/image/fetch/$s_!pqUq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png 848w, https://substackcdn.com/image/fetch/$s_!pqUq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png 1272w, https://substackcdn.com/image/fetch/$s_!pqUq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pqUq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png" width="1456" height="1030" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1030,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1595720,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://anneliesgamble.substack.com/i/198428758?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pqUq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png 424w, https://substackcdn.com/image/fetch/$s_!pqUq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png 848w, https://substackcdn.com/image/fetch/$s_!pqUq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png 1272w, https://substackcdn.com/image/fetch/$s_!pqUq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fffbd9926-0afa-4ce2-9df4-24fb98fe4976_1491x1055.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><a href="https://bepresearch.com/">Source</a></figcaption></figure></div><p>A few weeks ago, I listened to <a href="https://www.dwarkesh.com/p/jensen-huang">Dwarkesh Patel&#8217;s interview with Jensen Huang</a>. It&#8217;s a great interview and one I recommend listening to if you haven&#8217;t already. In it, Jensen said chip-side bottlenecks will be resolved in two or three years. The bottleneck he&#8217;s more worried about is energy, and specifically he calls out his concern over the shortage of plumbers and electricians.</p><p>This was on my mind when I sat down the other week with <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Ben Pouladian&quot;,&quot;id&quot;:11157401,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/596c2fbb-f0ce-42ac-a8b3-735bca99c9ad_3024x3024.jpeg&quot;,&quot;uuid&quot;:&quot;d02ab7c3-3dbe-47d0-93d1-88a19d1beea2&quot;}" data-component-name="MentionToDOM"></span>, founder of <a href="https://bepresearch.com/">BEP Research</a>, an independent research shop covering GPUs, memory, optical interconnects, and data center power as one converging system. His work has become required reading for hedge funds, asset managers, and the engineers building the stack, and has been cited everywhere from the <a href="https://www.wsj.com/tech/ai/ai-is-using-so-much-energy-that-computing-firepower-is-running-out-156e5c85?eafs_enabled=false">WSJ</a> to sell-side research desks.</p><p>Most of the AI infrastructure conversation right now is about GPUs. Who has them and who can get them. But GPUs are just one ingredient in a much more complex supply chain, and I don&#8217;t think they&#8217;re the most urgent constraint.</p><p>&#8220;The biggest constraints are energy or electricity, finding powered land,&#8221; Ben told me. &#8220;And then once you find that powered land, finding the people and the money to help build that data center.&#8221;</p><p>This is a physical problem, and it&#8217;s slower and more operationally complex than buying chips. But if AI is going to scale anything close to what current capex commitments imply, there&#8217;s a massive opportunity in building the physical stack underneath it and the software and hardware layers that coordinate it.</p><h2><strong>The bottleneck is a chain, not a point</strong></h2><p><a href="https://www.pjm.com/">PJM</a>, the grid operator covering 67 million people across the mid-Atlantic, just came out with <a href="https://www.pjm.com/-/media/DotCom/library/reports-notices/special-reports/2026/20260506-powering-reliability-through-market-design.pdf">a report about rising demand and constrained supply</a> where they say &#8220;we are facing a possible decade-long structural reality where demand growth will continually threaten to outpace supply additions.&#8221;</p><p>The GPU shortage is an important part of the story, but it isn&#8217;t the full story. AI infrastructure comes online as a sequence, not a single event. So the idea that the constraint is a single choke point oversimplifies how the stack actually gets built.</p><p>In fact, one of the last things that goes into the data center is the compute hardware. &#8220;First, you need to actually build the thing and make sure there&#8217;s power and it works,&#8221; Ben said.</p><p>And even once the GPUs arrive, the bottleneck just moves one layer deeper to memory architecture. Things like High Bandwidth Memory and the KV cache that holds an inference in working memory gate how much intelligence you can extract per watt. A GPU starved of memory bandwidth draws full power but delivers a fraction of the output. So the constraint is also around getting the right memory onto the silicon once it&#8217;s in the data center.</p><p>Working backwards, this means you need software orchestration, racks, chips, cooling and electrical, construction, community acceptance, permitting, grid interconnection, powered land, and electricity.</p><p>Each of these layers has a different production timeline. As Ben put it: &#8220;Chips are scarce this quarter. Power is scarce this decade.&#8221; Interconnection queues are years long, and grid capacity is already under pressure from electrification, EV charging, and manufacturing reshoring. On top of that, permitting is slow and depends heavily on the speed of local politics.</p><p>And then there is the issue of labor. Tradespeople take decades to train. &#8220;It&#8217;s purely physical, human, blue-collar labor,&#8221; Ben said. &#8220;You can&#8217;t spin it up like an AWS instance.&#8221; This is what Jensen meant when he said plumbers and electricians are <em>the</em> most challenging bottleneck right now.</p><h2><strong>Manufacturing intelligence</strong></h2><p>Ben kept using the phrases &#8220;<a href="https://www.nvidia.com/en-us/glossary/ai-factory/">AI factory</a>&#8220; and &#8220;<a href="https://developer.nvidia.com/blog/scaling-token-factory-revenue-and-ai-efficiency-by-maximizing-performance-per-watt/">token factory</a>,&#8221; borrowing a framing that Jensen Huang has used many times when describing the next generation of data centers.</p><p>Traditional data centers hosted software; they stored data and ran enterprise workloads. AI data centers are production plants; they turn energy and data into tokens. &#8220;The modern factory is not making metal,&#8221; Ben said. &#8220;It&#8217;s making intelligence.&#8221;</p><p>And although the output is tokens instead of parts, the production-system questions are the same as any factory: throughput, yield, energy efficiency, utilization, predictive maintenance. The units of output are metrics like tokens per watt or tokens per dollar of capex. These are financial metrics. Every watt of power and every dollar of capex now has a token-denominated yield attached to it. As inference workloads grow faster than training workloads, these questions become more acute.</p><p>The need to efficiently convert watts into intelligence becomes even more urgent.</p><p>This drives what kind of opportunities need to be built next. Traditional factories spawned entirely new categories of software and hardware. AI factories will need the equivalents, but for tokens-as-output. Almost none of that exists yet.</p><h2><strong>The opportunity is at the seams</strong></h2><p>Production systems are complex, and coordination between the layers is mostly manual or opaque. There&#8217;s a massive labor shortage on top of that. Wherever coordination is fragmented, slow, or expensive, there&#8217;s room for new companies. Software, hardware, and everything in between.</p><p>Starting at the bottom of the stack is the need to <strong>find viable powered land</strong>. And once you find it, the procurement process is often painful. <a href="https://www.tapestryenergy.com/en">Tapestry</a>, which spun out of Alphabet&#8217;s X moonshot factory, is essentially building Google Maps for the grid. It&#8217;s a knowledge graph that helps developers and utilities operate at a much higher speed and resolution than they can today.</p><p>As an aside, there are bets trying to escape these constraints entirely by moving data centers to space. Google&#8217;s Project Suncatcher and the recent SpaceX talks are the most visible. They sidestep some of the problems (land, grid interconnection) but not all of them. Most coverage focuses on launch costs, which would need to fall by an order of magnitude before any of this is viable at scale. But the harder constraint is thermal. In vacuum there&#8217;s no air to carry heat away (you can only radiate it) and that physics is what really gates the architecture.</p><p>Once you have powered land, <strong>the factory itself needs to be built and operated</strong>. This is the layer Jensen was pointing at when he talked about plumbers and electricians. Scheduling tradespeople, sequencing trades on site, managing lead times are all still very manual processes. As mentioned earlier and as I&#8217;ve written about previously, such as in my <a href="https://x.com/AnneliesGamble/status/2023834463202095508?s=20">conversation with Saman Farid</a>, the founder of <a href="https://formic.co/">Formic</a>, we have a massive labor shortage. Training tradespeople takes decades. We don&#8217;t have decades. There&#8217;s an opportunity to leverage robotics to do a lot of the manual work that humans have historically done. Companies like <a href="https://watneyrobotics.com/">Watney Robotics</a> are examples of types of companies I&#8217;m very excited about here.</p><p><a href="https://www.crusoe.ai/">Crusoe</a> is an example of what a <strong>vertically integrated AI factory</strong> company looks like. They source their own energy, build their own modular data centers, manufacture them in their own facility, and run a cloud layer on top. Every layer of the stack (power, building envelope, cooling, hardware, software) is something they&#8217;re either building or coordinating.</p><p>Once the factory is running, <strong>routing power</strong> matters because every watt matters. Power routed to cooling is power not routed to compute. As power becomes the binding constraint, there&#8217;s an opportunity to optimize the thermal and electrical envelope in real time.</p><p>Above the silicon, <strong>the orchestration layer </strong>is just as early. GPUs sit idle for a meaningful share of their lives. <a href="https://amppublic.com/">AMP</a>, an Alphabet-affiliated public benefit corporation, is pooling compute across independent AI labs to smooth utilization across the field. So when one lab is in a training run and another is in deployment mode, the aggregate demand curve is much smoother than any individual workload.</p><p>And above the orchestration layer, sits <strong>the financial layer.</strong> Compute is becoming the most important commodity of the decade. Like oil, we need market infrastructure to enable a liquid market for buyers and sellers of compute to transact. This means we need tooling for GPU pricing, hedging, and financing. As I wrote about <a href="https://x.com/AnneliesGamble/status/2046618812548800647?s=20">here</a>, compute will eventually become a more liquid market where capacity is procured on demand rather than primarily through long-dated bilateral contracts. <a href="https://ornn.com/">Ornn</a> is one company I&#8217;m excited about that is building in this space.</p><p>&#8220;The investment surface around power is deep,&#8221; Ben said, &#8220;and the layers on top of it are almost entirely greenfield.&#8221;</p><h2><strong>Why this matters</strong></h2><p>This buildout isn&#8217;t going to stop next year. &#8220;We spent 15 years building regular data centers with CPUs to run regular websites,&#8221; Ben said. &#8220;This is not the same thing.&#8221;</p><p>The stack is more physical, more tightly coupled across layers, and built around producing something rather than hosting it. Some of the most important companies of the last generation were born during the cloud buildout. The next generation is getting built now.</p><p>The opportunities are in the seams between power and compute, between construction and capital, between watts and tokens.</p><p><em>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Opportunities! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI for the Real World: A conversation with Yann LeCun]]></title><description><![CDATA[Are today&#8217;s language models the path towards machine intelligence, or are they just a commercially viable local maximum?]]></description><link>https://www.mixtureofexperts.co/p/ai-for-the-real-world-a-conversation</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/ai-for-the-real-world-a-conversation</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 12 May 2026 15:24:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hbio!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Hbio!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Hbio!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png 424w, https://substackcdn.com/image/fetch/$s_!Hbio!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png 848w, https://substackcdn.com/image/fetch/$s_!Hbio!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png 1272w, https://substackcdn.com/image/fetch/$s_!Hbio!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Hbio!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png" width="1536" height="614" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/de1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:614,&quot;width&quot;:1536,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1576331,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://anneliesgamble.substack.com/i/197365430?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9e3865e0-0f1e-40af-90e6-5764c5d8e28b_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Hbio!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png 424w, https://substackcdn.com/image/fetch/$s_!Hbio!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png 848w, https://substackcdn.com/image/fetch/$s_!Hbio!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png 1272w, https://substackcdn.com/image/fetch/$s_!Hbio!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde1d3429-6fa9-4fa7-b59b-c3833491d8d1_1536x614.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Are today&#8217;s language models the path towards machine intelligence, or are they just a commercially viable local maximum?</p><p>Yann LeCun is one of the clearest and most consistent voices arguing for the latter. In his view, LLMs are not intelligent, however useful they may be. Systems trained to predict sequences of discrete tokens don&#8217;t have an understanding of the world, which is a fundamental building block of intelligence.</p><p>I sat down with Yann a couple weeks ago to explore this idea and his vision for the future.</p><p>&#8220;There&#8217;s one question of whether the models we have today are useful? Is there a market for them? Yes.&#8221; But on the bigger question, &#8220;Will these models take us to human-level intelligence or something similar to it? Absolutely no.&#8221;</p><p>Yann recently founded <a href="https://amilabs.xyz/">AMI Labs</a>, a Zetta portfolio company, to build what he thinks the alternative will look like: world models that can understand the physical world and predict the consequences of actions.</p><h2><strong>Why language isn&#8217;t intelligence</strong></h2><p>&#8220;Much of human knowledge and thought has nothing to do with language,&#8221; Yann said. And yet we credit anything that speaks fluently with understanding. &#8220;We&#8217;re biased towards attributing intelligence to things that can express themselves through language.&#8221;</p><p>He walked me through a calculation he&#8217;s done before. A four-year-old has been awake for roughly 16,000 hours. The optic nerve carries about one byte per second per fiber, with roughly a million fibers per eye. If you multiply it out, you get something on the order of 10^14 bytes of visual data reaching the brain in the first four years of life, roughly the same order of magnitude as the entire text corpus used to pretrain a modern LLM.</p><p>&#8220;It would take any of us something like 400,000 years to read through that,&#8221; he said. In other words, a small child has already absorbed, through vision alone, about as much raw information as the largest language models see in training. &#8220;We&#8217;re never going to get to human-level AI by just training on text. It&#8217;s just not going to happen.&#8221;</p><p>What LLMs do have is an ability to accumulate and retrieve declarative knowledge. This means they look smarter over time without developing deeper models of reality. They simply become more familiar with the kinds of questions people ask.</p><p>&#8220;If you want a system to act intelligently,&#8221; he said, &#8220;it has to be able to predict the consequences of its actions. And LLMs are completely incapable of doing this.&#8221;</p><p>Yann believes in language models for two specific domains: coding and math. &#8220;The reason why it works so well in these two domains is because these are domains where the mere manipulation of symbols is actually kind of the substrate of reasoning.&#8221; But these are narrow cases. &#8220;For everyday things that require a little bit of common sense reasoning and certainly planning, they&#8217;re just never going to get there.&#8221;</p><h2><strong>What the alternative looks like</strong></h2><p>The alternative is what Yann has been working toward for over 15 years. It&#8217;s a system that learns how the world evolves, and can predict what the consequence of a sequence of actions is going to be.</p><p>&#8220;This is the only way to build an agentic system that is reliable,&#8221; he said. &#8220;I do not understand how people can even think of building agentic systems that do not have this ability of predicting the consequences of their actions before they do them.&#8221;</p><p>The hard part is learning such a model from real-world data. Next-token prediction works because symbols are discrete and compressible. The physical world is not. &#8220;I&#8217;ve been working on this for over 15 years, and essentially failing the first 10 years, because I was using generative architectures trying to predict what&#8217;s going to happen in the video at the pixel level. This kind of data is just not predictable.&#8221;</p><p>He gave the example of a pen balanced on your hand. If you let go, you can predict that it will fall. But you can&#8217;t predict the exact direction it will fall, or the precise configuration of every pixel in the next frame. If you train a system to predict all of those details, you&#8217;re forcing it to model noise and contingency as though they were the essence of intelligence. &#8220;When you try to train a system to predict every detail in a situation, you kind of kill it because you try to train it to do something that&#8217;s impossible.&#8221;</p><p>His proposed alternative is Joint Embedding Predictive Architecture (JEPA). Rather than predicting every pixel, the system learns an abstract representation of the world and makes its predictions there. &#8220;All the details about the input that are not predictable, all the noise, all the complexities of it are basically going to be eliminated from the representation so that the prediction can be reliable.&#8221; You learn the latent state that matters for planning, even if you can&#8217;t regenerate a photorealistic frame from it.</p><p>Once you have an abstract world model, reasoning becomes search through that model. That&#8217;s what LLMs can&#8217;t do, because they don&#8217;t have a model to search through. &#8220;The idea that reasoning is a kind of search is really fundamental,&#8221; he said. &#8220;LLMs don&#8217;t do this. They don&#8217;t have any ability to really search for an answer. They just produce an answer, a token.&#8221; Chain-of-thought, in his view, is a workaround: &#8220;a very, very inefficient way of coercing autoregressive prediction systems to basically approach reasoning.&#8221; Real reasoning, he argues, is internal simulation. This means manipulating mental models, running counterfactuals, planning hierarchically the way a human plans a trip to Paris (aka not at the level of muscle commands, but refining subgoals from the top down).</p><p>This is why he prefers the term <a href="https://arxiv.org/abs/2602.23643">Superhuman Adaptable Intelligence</a> to AGI. &#8220;The true property of intelligence is to solve new problems you&#8217;ve not been trained to solve.&#8221;</p><h2><strong>AMI Labs and World Models</strong></h2><p>That thesis is now Yann&#8217;s company: AMI Labs, Advanced Machine Intelligence, (pronounced &#8220;ah-mee,&#8221; just like the French word for friend).</p><p>AMI is building AI for the real world. &#8220;A lot of industry is just running things, right? Like physical things. And this is where current AI technology falls short,&#8221; he told me. The company&#8217;s stated focus is industrial process control, automation, wearable devices, robotics, and healthcare.</p><p>A huge portion of the economy depends on running physical systems (factories, supply chains, power grids, biological systems, transportation networks). These are environments where text is often the interface <em>around</em> the work, but not the work itself. &#8220;AMI is building generic foundation models that can be applied to any situation where you need an intelligent system to run something physical,&#8221; Yann said.</p><p>The physical-economy layer of AI will be built on a different stack from what most companies are using today. Rather than predicting the next token, this is about predicting the next state.</p><p>There are a number of other companies also trying to build versions of world models. The approaches differ on what the model tries to predict: pixels and geometry versus abstract state.</p><p>Fei-Fei Li&#8217;s<a href="https://www.worldlabs.ai/"> World Labs</a> is building, according to their website, &#8220;world models that can perceive, generate, reason, and interact with the 3D world.&#8221; Their first product,<a href="https://www.worldlabs.ai/labs"> Marble</a>, turns text, images, or video into 3D environments that designers can open in different creative tools. Google DeepMind&#8217;s<a href="https://deepmind.google/discover/blog/genie-3-a-new-frontier-for-world-models/"> Genie 3</a> takes a different approach to a similar problem, generating interactive worlds in real time that users can navigate frame by frame.</p><p>1X and Generalist are building video-pretrained world models specifically for humanoid robotics.<a href="https://www.1x.tech/"> 1X</a>&#8216;s model learns from internet video first, then from footage shot from a human&#8217;s point of view, and uses a second model to turn its predictions of &#8220;what should happen next&#8221; into robot movements.<a href="https://generalistai.com/"> Generalist</a> combines ideas from world models and VLAs, training on roughly 500,000 hours of real-world physical interaction data collected from wearables worn by humans doing everyday tasks.</p><p>NVIDIA&#8217;s<a href="https://github.com/nvidia-cosmos"> Cosmos</a> is building a platform to &#8220;help developers build customized world models for their Physical AI setups.&#8221; Meanwhile, Tesla is building a single AI model that can drive cars and control humanoid robots, treating both as different bodies running the same underlying intelligence.</p><p>What distinguishes AMI is the architectural bet around JEPA-style abstract representation rather than pixel-level generation. Pixel-perfect prediction is computationally expensive and, as Yann argued for years before the field caught up, trying to predict the unpredictable actively degrades the model&#8217;s grip on what matters. Abstract representation preserves the causally relevant structure while removing the noise. If it works, it&#8217;s both a better model of physics and a cheaper one to deploy.</p><h2><strong>Why this matters</strong></h2><p>For robotics specifically, the implications are significant. The dominant approach today, vision-language-action models that map observations directly to motor commands, runs into two well-understood ceilings.</p><p>The first is data. Teleoperated robot data is the highest-quality source but doesn&#8217;t parallelize. It&#8217;s bounded by the number of robots you own and the hours skilled operators can work. Researchers have developed workarounds: hand-held grippers like UMI that let humans collect demos without a robot, wearable rigs that record everyday activity, cross-embodiment datasets that pool data across robot types, and simulation pipelines. But there is an embodiment gap for each that has to be bridged. Meanwhile, the largest available corpus by far, human video on the internet, is hard to exploit directly because the actions aren&#8217;t labeled. Recent work on inverse dynamics and latent action models is starting to unlock it, which is part of why world models have gained momentum.</p><p>The second is embodiment lock-in. Observation-to-action mapping tends to couple learned knowledge to a specific robot body. Transfer across embodiments is possible but imperfect. A policy trained on one arm typically needs significant adaptation to work on another. Knowledge ends up captured at the level of &#8220;how this robot should move in this specific setting&#8221; rather than &#8220;what should happen in the world.&#8221;</p><p>World models attack both problems at once. If you learn an abstract representation of how the world evolves (how objects fall, how contact propagates, how liquids behave), you&#8217;ve learned something that is true regardless of which body is acting in it. That knowledge can be absorbed from video without action labels, because the goal isn&#8217;t to predict the next motor command but to predict the next state. A model that understands physics can then be adapted to whatever embodiment is available, with calibration rather than retraining.</p><p>The opportunity extends well beyond robotics. &#8220;There are tons and tons of applications of this type,&#8221; Yann told me. &#8220;You want to control anything in the real world: manufacturing plant, turbojet engine, chemical process. A human cell. You want to plan a sequence of treatment for a patient to, I don&#8217;t know, control blood sugar. If you have a good predictive model of at least some aspect of the state of the patient, you might be able to do this kind of planning on a personalized basis.&#8221;</p><h2><strong>A system that thinks</strong></h2><p>It&#8217;s easy, in a moment like this one, to mistake the shape of the market for the shape of the problem. LLMs are producing extraordinary value, and they will keep doing so in cases where symbolic manipulation is the actual work.</p><p>But most of the economy doesn&#8217;t run on words and symbols. It runs on physical systems, environments where text serves as a wrapper, but isn&#8217;t the work itself. The systems capable of operating in those environments will need something current models don&#8217;t have: a base-level understanding of the world, the ability to predict the consequences of actions, and the capacity to adapt to problems they weren&#8217;t trained on.</p><p>Intelligence is much more than language. Future AI systems will still use language, but language will no longer be their only substrate.</p><p>As Yann put it, &#8220;language will serve as an interface to a system that thinks.&#8221;</p><p><em>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Opportunities! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Who Gets to Solve the Physical World?]]></title><description><![CDATA[When the iPhone app store launched, it wasn&#8217;t clear at the time how or even why hundreds of thousands of apps would get built.]]></description><link>https://www.mixtureofexperts.co/p/who-gets-to-solve-the-physical-world</link><guid isPermaLink="false">https://www.mixtureofexperts.co/p/who-gets-to-solve-the-physical-world</guid><dc:creator><![CDATA[Annelies Gamble]]></dc:creator><pubDate>Tue, 05 May 2026 16:09:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0pQr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0pQr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0pQr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png 424w, https://substackcdn.com/image/fetch/$s_!0pQr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png 848w, https://substackcdn.com/image/fetch/$s_!0pQr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png 1272w, https://substackcdn.com/image/fetch/$s_!0pQr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0pQr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png" width="1144" height="704" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:704,&quot;width&quot;:1144,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:912057,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://anneliesgamble.substack.com/i/196559463?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0pQr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png 424w, https://substackcdn.com/image/fetch/$s_!0pQr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png 848w, https://substackcdn.com/image/fetch/$s_!0pQr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png 1272w, https://substackcdn.com/image/fetch/$s_!0pQr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4e2bbe3-3020-4fc2-ae08-a54b0112addd_1144x704.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Flyby Robotics</figcaption></figure></div><p>When the iPhone app store launched, it wasn&#8217;t clear at the time how or even why hundreds of thousands of apps would get built. After all, phones were for calls. But it turned out that the platform created an economy in which anyone who had a problem could build an app to solve it. These apps addressed markets that were often too small to justify a company, but still many people found value using them.</p><p>That same dynamic is starting to happen in physical AI.</p><p>DJI alone supposedly has more than 100,000 developers building applications on its drones, mostly in China and on problems that no large company would pick up. The DJI ecosystem exists despite DJI not really making it easy. The hardware is largely closed. Onboard AI compute is minimal. Developer experience is an afterthought. The fact that 100,000 developers built on it anyway tells you what the latent demand looks like.</p><p>If that&#8217;s what is possible on closed hardware with poor tooling, imagine what happens when the platform is actually built for it. The ceiling on who gets to solve physical-world problems moves.</p><p>I recently sat down with Jason Lu, the founder of <a href="https://www.flybyrobotics.com/">Flyby Robotics</a>, to explore why this shift is happening and what opportunities this now unlocks.</p><h2><strong>The old market structure was a tax on distance</strong></h2><p>Physical-world problems have always been solved by whoever could afford to solve them, which mostly meant whoever could justify the cost of solving them at scale.</p><p>At the top, megacorps like DJI, Skydio, and Anduril picked off the problems large enough to move a multi-billion-dollar needle. Below them, venture-backed startups went after problems big enough to justify raising their next round of funding. Below them, smaller specialized companies went after the problems too narrow for the tier above. And below all of them sat people who actually lived the problems day-to-day, but got to solve essentially nothing.</p><p>Jason says this is because &#8220;the juice wasn&#8217;t worth the squeeze&#8221; for the people with the resources to solve them. Every layer of distance between the problems and the person with the authority to solve them acts as a filter. As Jason put it, &#8220;If you&#8217;re the 1,000th employee at Skydio, it&#8217;s unlikely you&#8217;re going to really care about these &#8216;small&#8217; problems, like offshore oil rig corrosion or crop yield boosting 20% with multispectral imagery.&#8221;</p><p>Problems that aren&#8217;t big enough or profitable enough don&#8217;t get solved. This doesn&#8217;t mean those problems aren&#8217;t worth solving, it&#8217;s just that the unit economics of solving them has required a scale that most physical-world problems can&#8217;t reach.</p><p>This asymmetry has defined digital versus physical for a long time. Software ate the world. At first, it was mostly the parts of the world where one solution could be sold to millions of customers. But eventually with the app store and now vibe coding, anyone can build a piece of software to enable whatever problem they have, no matter how small.</p><p>This hasn&#8217;t yet been feasible for the physical world, but that is changing</p><h2><strong>Three things became true at the same time</strong></h2><p>What&#8217;s changing is that three independent curves are crossing at roughly the same time.</p><p>The first is compute on the edge. Until recently, you couldn&#8217;t put a lot of processing power on a small flying robot. &#8220;Imagine you&#8217;re trying to solve this for coding, but there are no computers. There are no servers. There&#8217;s no mouse,&#8221; Jason said. The drones Flyby Robotics are bringing to market carry 150 to 300 trillion operations per second of onboard compute, and in some cases even more. RAM is going from 2 gigabytes to 16 to, eventually, 128 or more.</p><p>Drawing on the iPhone analogy again, Jason said &#8220;the iPhone could only have happened because of advancements in RAM technology, in microprocessors, and in touchscreen technology. Those applications built on top of the iPhone were possible because these technical unlocks happened. That same transition is happening in the physical AI space.&#8221;</p><p>The second is abstraction over hardware. Controlling a gimbal or routing camera data into a model or deploying that model efficiently on a GPU &#8211; all of that used to require a specialist. There was no equivalent of an OS for physical AI, but that&#8217;s finally starting to change. Flyby for example is building the layer that lets a developer move a camera or a drone without ever touching the underlying SDK.</p><p>The third curve is coding agents. Claude Code, Codex, Cursor, and the broader category of AI-assisted development have dropped the skill floor for software work by an order of magnitude. Now everyone is empowered to describe what they want to build in natural language and AI will write the code for it.</p><p>Any one of these shifts on their own wouldn&#8217;t be enough. But the combination of capable hardware, accessible abstractions, and AI coding makes this possible for the first time.</p><h2><strong>The Opportunity</strong></h2><p>This changes what gets built.</p><p>The opportunity I&#8217;m most excited about is the platform that makes the long tail possible. This is both the hardware, and the developer layer that sits on top of it.</p><p>Right now you mostly can&#8217;t buy a drone or a robot with enough onboard compute to run a model that&#8217;s also open enough to add your own software. Solving this is a hardware problem that requires having the right chips, sensors, batteries, airframes, manufacturing lines, supply chains, and firmware.</p><p>Sitting on top of that hardware is the platform layer, which includes the abstractions, the developer experience, the model deployment tooling, the integration with coding agents. Without this platform, only specialists can use the hardware, and the ecosystem stays small. But with this platform, thousands of hyper-specific applications can get built by people closest to the problems.</p><p>This is the vision Flyby is pursuing in aerial robotics, and I suspect we&#8217;ll see others take on adjacent pieces of the work in different physical domains such as ground robots, underwater systems, manipulation, sensing. The shape of the opportunity is the same in each case: build the full stack that lets the people closest to a problem actually solve it, and capture value across the entire ecosystem of solutions that emerges on top of you.</p><p>The interesting thing about the App Store, in retrospect, wasn&#8217;t any single app. It was that the platform created the conditions for problems to get solved by the people who had them. Physical AI is approaching its version of that moment. The substrate is starting to come together and the distance between a problem and the person who can solve it is collapsing.</p><p><em>Author&#8217;s note: An LLM was used for light copy editing only (spelling, grammar, and clarity). Content, meaning, tone, and structure remain unchanged.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.mixtureofexperts.co/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading AI Opportunities! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>