We talk about the evolution from ARC-AGI-1 to ARC-AGI-3 and the lifecycle behind each. ARC-3 alone cost ~$500-750K to build. Greg walks through the leading approaches to beating it (reverse-engineered world models vs. frame-search harnesses) and where models struggle today.
Also covered: ARC’s definition of AGI as human learning efficiency, why world models and memory matter, current ARC-3 scores, long-horizon evals, open vs. closed source dynamics, multimodal integration, the underinvestment in parametric learning, what’s next for agents (spending, agent-to-agent communication, always-on), and why vibe coding raises the floor but not the ceiling.
Chapters
00:14 Background & Joining ARC Prize
03:06 Role at ARC Prize & Founding Team
03:55 ARC1 to ARC3 Evolution
06:07 Building a Benchmark: Idea to Sunset
08:08 ARC4 & What’s Next
08:45 Approaches to Beating ARC
11:33 Why Benchmarks Matter
12:24 Who Are Benchmarks For
13:42 Evolution of Benchmark Difficulty
15:27 Long Horizon Agents & Evals
16:59 Defining AGI and ASI
18:37 Role of World Models
22:02 Where Models Struggle Today
22:57 Open Source vs Closed Source Models
25:52 Trends & What’s Exciting
27:03 Under-invested Research Areas
29:55 Vibe Coding & Role of Humans
32:18 Final Thoughts


