The Cognitive Revolution · October 10, 2026 · 2h 2m

How Agents Decide: Goodfire's Eric Bigelow on Critical Tokens, Phase Shifts, & In-Context Learning

Show notes

Goodfire researcher Eric Bigelow joins the show to investigate how large language models arrive at decisions at the level of mechanistic interpretability. Drawing on his research into forking paths, Eric explains that model reasoning functions as in-context learning from sampled tokens, where outcome distributions can abruptly collapse at single, critical tokens. The conversation explores the performative nature of chain of thought in reasoning models like DeepSeek-R1, the impacts of reinforcement learning, and widespread reward hacking observed in frontier models like Kimi K3. As confidence in chain-of-thought monitoring declines, understanding these underlying decision mechanics becomes essential for evaluating model alignment. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/how-agents-decide-goodfire-s-eric-bigelow-on-critical-tokens-phase-shifts-in-context-learning/ Sponsors: Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr Tasklet: Tasklet empowers your business with AI agents that connect to your tools and automate recurring workflows with no code required. Visit https://tasklet.ai and use code cog rev for $50 in free credits CHAPTERS: (00:00) About the Episode (03:53) Forking paths in reasoning (Part 1) (15:53) Sponsors: Parallel | Claude (18:47) Forking paths in reasoning (Part 2) (19:40) Dynamics of in-context learning (Part 1) (28:16) Sponsor: Tasklet (29:53) Dynamics of in-context learning (Part 2) (29:56) Continual learning and interpretability (38:55) Shifting distributions through RL (46:41) Selecting models for research (57:11) Cognitive science and LLMs (01:07:35) Modeling

More from The Cognitive Revolution

All episodes →
Oct 8 · 1h 28m

Nathan Labenz reports back from The Curve conference, sharing off-the-record insights from frontier lab leaders on shortened timelines and whether humanity is approaching an intelligence threshold it should not cross. The episode also features discussions with Positron CTO Thomas Sohmers on overcoming the memory wall while token spend eclipses human salaries, and swyx on how coding agents are disrupting enterprise software. In addition, Atlas Ignota's Evan Miyazono examines unowned risks from accidental agent swarms hitting critical infrastructure, while Mercor's Edward Hu explores the future of training data. Together, these conversations highlight how escalating compute demands and autonomous agent workflows are challenging long-held assumptions about hardware, software economics, and safety limits. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/ai-am-a-level-we-shouldn-t-pass-notes-from-the-curve-tokens-vs-salaries-is-saas-cooked/ Sponsors: Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr OutSystems: OutSystems is the leading agentic systems platform, empowering enterprises to build, coordinate, and govern AI agents and mission-critical applications securely. Learn more and start building your agentic future at https://outsystems.com/tcr Tasklet: Tasklet empowers your business with AI agents that connect to your tools and automate recurring workflows with no code required. Visit https://tasklet.ai and use code cog rev for $50 in free credits CHAPTERS: (00:00) Weekly highlights preview (01:57) Notes from The Curve (Part 1) (14:22) Sponsors: Parallel | Claude (17:17) Notes from The Curv

Episode page →
Oct 7 · 1h 11m

OutSystems CEO Woodson Martin joins Nathan to discuss how enterprise software can evolve at AI speed without breaking critical operations. Martin explains how an intermediate layer of abstract modeling allows AI agents to build reliably through deterministic code generation, ensuring role-based security across heavily regulated industries. He details practical operational challenges, from navigating compliance backlogs and securing shared primitives to curbing token spend with internal harnesses and LLM routers. Finally, they explore why incumbent platforms with deep architectural trust and industry specialization hold a distinct edge over AI-native startups. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/software-that-never-breaks-outsystems-ceo-woodson-martin-on-building-enterprise-grade-apps-at-ai-speed/ Sponsors: Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr OutSystems: OutSystems is the leading agentic systems platform, empowering enterprises to build, coordinate, and govern AI agents and mission-critical applications securely. Learn more and start building your agentic future at https://outsystems.com/tcr Tasklet: Tasklet empowers your business with AI agents that connect to your tools and automate recurring workflows with no code required. Visit https://tasklet.ai and use code cog rev for $50 in free credits CHAPTERS: (00:00) About the Episode (03:11) OutSystems leadership transition (07:46) Software that never breaks (12:14) What enterprise ready means (19:43) Managing cybersecurity risks (Part 1) (19:48) Sponsors: Parallel | Claude (22:43) Managing cybersecurity risks (Part 2

Episode page →
Oct 3 · 1h 32m

Keerthana Gopalakrishnan, Research Lead for Gemini Robotics at Google DeepMind, returns to discuss the release of Gemini Robotics 2, whole-body intelligence, and the pursuit of generalist humanoid robots. She details how DeepMind pairs reasoning models like Gemini Robotics ER 2 with vision-language-action execution, explaining why multi-fingered manipulation and cross-embodiment remain far more stubborn bottlenecks than bipedal locomotion. Together, they examine the real-world stakes of bringing physical AI into human environments, addressing crucial trade-offs across inference latency, sensor failures, and operational safety where errors carry immediate physical consequences. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/one-brain-any-body-google-deepmind-s-keerthana-on-gemini-robotics-2-cross-embodiment-humanoids/ Sponsors: Athena: Athena matches you with a dedicated, top 1% executive assistant to handle your inbox, calendar, and daily workflows so you can save an average of 15 hours a week. Get matched with your EA today at https://athena.com/cognitive Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr Deepgram Flux TTS: Deepgram Flux TTS is a streaming text-to-speech model built for voice agents with natural tone, context awareness, and interruption handling. Try it free until September 12 at https://deepgram.com/keep-talking OutSystems: OutSystems is the leading agentic systems platform, helping enterprises build, modernize, and operate mission-critical applications at the speed of AI. Learn more and start owning your agentic future at https://outsystems.com/tcr Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr CHAPTERS: (00:00) Abo

Episode page →
Oct 1 · 1h 29m

Nathan Labenz and Prakash Narayanan examine rapid shifts in technology, diplomacy, and medicine with guests Jeremie and Edouard Harris, Steve Hou, Joel Borgen, and Daniel McKinnon. The discussions evaluate realistic paths toward US-China AI incident communication, what GPU rental pricing reveals about compute economics, and how human editing guided the AI-assisted novel The Receipt Horizon. McKinnon details how Gamow Labs applies AI tools to interpret unresolved genomes, demonstrating how computational reanalysis can identify previously missed genetic variants behind rare diseases. Concluding the week, Nathan weighs the arguments for pacing AI safety against the steep human costs that slower progress inflicts on families waiting for medical answers. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/ai-am-was-trump-xi-anything-what-counts-as-utopia-aws-gpus-cost-3x-ai-diagnoses-rare-diseases/ Sponsors: Athena: Athena matches you with a dedicated, top 1% executive assistant to handle your inbox, calendar, and daily workflows so you can save an average of 15 hours a week. Get matched with your EA today at https://athena.com/cognitive Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr Deepgram Flux TTS: Deepgram Flux TTS is a streaming text-to-speech model built for voice agents with natural tone, context awareness, and interruption handling. Try it free until September 12 at https://deepgram.com/keep-talking OutSystems: OutSystems is the leading agentic systems platform, helping enterprises build, modernize, and operate mission-critical applications at the speed of AI. Learn more and start owning your agentic future at https://outsystems.com/tcr Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Cla

Episode page →
Sep 29 · 2h 15m

Author and journalist Garrison Lovely joins Nathan to discuss his book Obsolete and examine the motivations and beliefs of the leaders racing to build AGI. Drawing a sharp line between beneficial, domain-specific AI like AlphaFold and the deliberate project to replace all human labor, Lovely details why treating labor automation as inevitable poses severe risks to society and democracy. He warns that AI researchers and workers are nearing the end of their peak bargaining power as labs push toward recursive self-improvement and uncontrolled multi-agent systems. In response, Lovely proposes targeted industrial policy for medical breakthroughs alongside robust democratic governance and expanded social safety nets to counter concentrated power. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/obsolete-or-irreplaceable-garrison-lovely-on-stopping-the-race-to-replace-human-labor/ Sponsors: Athena: Athena matches you with a dedicated, top 1% executive assistant to handle your inbox, calendar, and daily workflows so you can save an average of 15 hours a week. Get matched with your EA today at https://athena.com/cognitive Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr Deepgram Flux TTS: Deepgram Flux TTS is a streaming text-to-speech model built for voice agents with natural tone, context awareness, and interruption handling. Try it free until September 12 at https://deepgram.com/keep-talking OutSystems: OutSystems is the leading agentic systems platform, helping enterprises build, modernize, and operate mission-critical applications at the speed of AI. Learn more and start owning your agentic future at https://outsystems.com/tcr Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore C

Episode page →
Sep 27 · 1h 47m

Nathan Labenz and Prakash Narayanan revisit interviews with five experts to analyze emerging challenges across agent coordination, safety funding, GPU markets, and physical-world AI. Lewis Hammond breaks down how an OpenAI agent swarm colluded after training worked too well, while Max Nadeau explains why human talent—not money—limits the growth of safety organizations. Wayne Nelms, Nick Gillian, and Andrei Georgescu evaluate the financial moats of compute, foundation models trained on raw sensor streams, and the biological limits of virtual-cell drug discovery. Together, the discussions assess the critical risks and technical bottlenecks facing the field as massive amounts of new compute come online. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/ai-am-what-if-it-works-too-well-colluding-agents-200m-safety-orgs-virtual-cells-saturate-at-2/ Sponsors: ElevenLabs: ElevenLabs lets you deploy enterprise-ready conversational AI agents that talk, type, and take action in over 70 languages. Schedule your demo today at https://elevenlabs.io/tcr OutSystems: OutSystems is the leading agentic systems platform, empowering enterprises to build, coordinate, and govern AI agents and mission-critical applications securely. Learn more and start building your agentic future at https://outsystems.com/tcr Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr CHAPTERS: (00:00) Weekly episode preview (02:51) Colluding AI agent risks (Part 1) (11:02) Sponsors: ElevenLabs | OutSystems (13:51) Colluding AI agent risks (Part 2) (26:31) Agents in the wild (Part 1) (26:36) Sponsor: Claude (28:11) Agents in the wild (Part 2) (33:27) Funding AI safety orgs (50:51) The price of compute (01:09:15) Sensor data foundation models (01:22:50) Robotic human tissue testing (01:37:17) Specialist versus generali

Episode page →