
How Agents Decide: Goodfire's Eric Bigelow on Critical Tokens, Phase Shifts, & In-Context Learning
Goodfire researcher Eric Bigelow joins the show to investigate how large language models arrive at decisions at the level of mechanistic interpretability. Drawing on his research into forking paths, Eric explains that model reasoning functions as in-context learning from sampled tokens, where outcome distributions can abruptly collapse at single, critical tokens. The conversation explores the performative nature of chain of thought in reasoning models like DeepSeek-R1, the impacts of reinforcement learning, and widespread reward hacking observed in frontier models like Kimi K3. As confidence in chain-of-thought monitoring declines, understanding these underlying decision mechanics becomes essential for evaluating model alignment. For full show notes, links, and references, read the episode page:https://www.cognitiverevolution.ai/how-agents-decide-goodfire-s-eric-bigelow-on-critical-tokens-phase-shifts-in-context-learning/ Sponsors: Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr Tasklet: Tasklet empowers your business with AI agents that connect to your tools and automate recurring workflows with no code required. Visit https://tasklet.ai and use code cog rev for $50 in free credits CHAPTERS: (00:00) About the Episode (03:53) Forking paths in reasoning (Part 1) (15:53) Sponsors: Parallel | Claude (18:47) Forking paths in reasoning (Part 2) (19:40) Dynamics of in-context learning (Part 1) (28:16) Sponsor: Tasklet (29:53) Dynamics of in-context learning (Part 2) (29:56) Continual learning and interpretability (38:55) Shifting distributions through RL (46:41) Selecting models for research (57:11) Cognitive science and LLMs (01:07:35) Modeling