
Loading

Last call for regular tickets for AI Engineer NYC ! As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps - for new tickets only, no refunds! See you in 2 weeks! While we tend to cover industry on the pod, every so often we celebrate a clearly emerging superstar PhD. In 2024 we featured Shunyu Yao , who went on to build Operator at OpenAI and is now Chief AI Scientist of Tencent . In 2025 we featured Jack Morris , who went on to cofound Engram at $600m and is now a leading voice on continual learning . This year we are proud to feature the work of Alex Zhang of MIT. From GPU kernels and KernelBench to Recursive Language Models , Mismanaged Geniuses , and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems. RLMs took over the timeline early this year: and an RLM based harness was the first to ~solve ARC-AGI-3 before OpenAI’s Astra: and is even today, influencing new research that has more extreme implications than RLMs: We go deep on GPU Mode and AI-written kernels, research taste and why academics should take bets industry labs won’t, GEV and alternatives to the standard autoregressive language model, and the idea of harnesses as compositional generalizers. Alex explains RLMs, context offloading, programmatic subagent calling, Prime Agent, persistent subagents, and why the “language model” of the future may actually be an invisible swarm of agents underneath a simple interface. We also discuss OpenAI’s massive agent experiments, Kimi swarms, open-ended research at Sakana AI, speculative programmatic tool calling, capability overhang, Neuralese, and where Alex thinks the next big research opportunities may lie. We discuss: * Why AI-generated GPU kernels still leave substantial room for human expertise * How one expert insight can potentially replace enormous amounts of brute-force token search * Why PhD students
Episode page →It’s hard to believe that Periodic was only launched last September : One year later, it is considered one of the pre-eminent AI scientist labs, with dizzying talent density and astonishing progress in the autonomous lab buildout: Most people are familiar with the standard credentials of Liam and Dogus , but we found an incredible “talent slope” while learning more about Periodic, where each successive employee seems more impressive than the last: From building AI systems that reason over noisy physical experiments to creating laboratories where every instrument can become intelligent, Periodic Labs is betting that the next frontier of AI won’t come from simply training on more internet data , it will come from letting models experiment with the real world. In this episode, Periodic Labs’ Liam Fedus and Ekin Dogus Cubuk join swyx and Brandon to explain why scientific discovery is fundamentally different from math and coding, and what it takes to build AI scientists that can actually discover new materials. We go deep on Periodic’s vision for “synthesis superintelligence” : reinforcement learning grounded in physical experiments, AI-powered materials characterization, simulations and density functional theory, high-throughput labs, and systems that learn from the entire process of doing science rather than only its published results. Liam and Dogus also explain why frontier models will still need experiments , why failed experiments may be some of the most valuable training data, what it means to give every piece of lab equipment “140 IQ,” and how autonomous experimentation could compress decades of scientific trial-and-error into months. We discuss: * Why intelligence alone isn’t enough for scientific discovery * How reinforcement learning changes when the environment is the physical world * Why science requires reasoning under uncertainty, noise, and missing information * Prediction, synthesis, and characterization in the materials discovery loop * Why physics and
Three months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks: We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a big platform. However, we were at Anthropic for the Computer Use launch , there for Claude Cowork with the first big podcast on it, organized the first Computer Use track at AIE presenting the state of the art, and were close to the OpenAI-Sky Software acquisition that now powers the complete domination of computer use that Codex enjoys today. This is why we’re excited to bring you today’s first guest, Ari Weinstein , cofounder of Sky and now leading all the amazing CUA progress that casuals might miss: Ari explains why Computer Use is now “180 degrees different” from where it was months ago, how agents are learning to debug and recover from failures, why combining screenshots with accessibility data, the DOM, Playwright, and generated code changes the speed equation, and why the next frontier is making agents literally superhuman at using software. OpenAI clones Jev In the second half, Nikunj Handa from OpenAI’s API team breaks down the new developer stack: async tool calling, mid-turn steering, WebSockets, UltraFast inference, the Decisions API, prompt caching, pre-warming, compaction, and the Agents API . Given that we were the first Jev podcast , we particularly focus on the unusually fast sprint on the Decisions API: And why it is just a Luna wrapper for now but the team is motivated and egoless enough to clone what they consider to be good patterns. We discuss: * Why OpenAI thinks Computer Use has changed dramatically in just the last few months * Dots and what changes when every agent gets its own Linux computer * Why Computer Use can now complete some tasks faster than the average human * The path from human-level to “literally superhuman” computer use * Why moder
Episode page →We are excited to have Anthropic share their latest AI x Finance work at AI Engineer New York , coming up in 2 weeks! In case you’ve been under a rock, here’s a non-exhaustive list of what Anthropic has been shipping since closing the largest fundraise of all time in May at $47B ARR: * June: Launched Claude Tag and Sonnet 5 and Fable 5 * July: Opus 5 , /checkup . crossed $65B ARR * Last month: Fable/Mythos 5.1 , and EFS (upcoming pod) * IPO target $2T , end 2026 ARR estimated $100B * Cowork/chat merged before did * Claude Mods * Dario endorses the same Pacing the Frontier message cosigned by all labs * Last week: Opus 5.5 , Plugins portal , Cloud Sessions/ Claude Projects * Today: Sonnet 5.5 ! Today’s episode should catch you up, with Thariq Shihipar , the explainer-king of Anthropic, who we last caught up on Fable launch day with The Field Guide to Fable : The Future of Mutable Software Pay special attention to Claude Mods (especially the cheatsheet ): In general this is also the inverse of the other viral tweet from Thariq: Cloud Brain, Local Hands And give a try to Claude Projects : The “hands” terminology is not just an analogy for the local/cloud paradigm that is being built up at frontier coding agent companies like Cognition , but is ALSO particularly relevant to the safety systems discussions that we’ll be discussing with Anthropic in an upcoming episode as they prepare to pace to frontier with responsible AI deployment. For those who want Thariq’s writing tips we teased at the start of the pod, watch the full video here: From the rapid rise of Claude Code to a future where agents can rewrite their own harnesses, collaborate across teams, and operate across cloud and local environments, the way we build software is changing extraordinarily fast. In this episode, Anthropic’s Thariq Shihipar joins swyx and Vibhu to unpack how power users are actually working with Claude Code today , why prompting remains a high-skill discipline, and where Anthropic thinks the
Episode page →From the earliest days of open-weight models to becoming the neutral routing layer for more than 10 million developers , OpenRouter is one of the clearest bets that the future of AI will be multi-model. In this episode, OpenRouter co-founder & CEO Alex Atallah, with AMP’s Anjney Midha returning with swyx to unpack how OpenRouter emerged from the first wave of Llama , Alpaca, Mistral, and Midjourney, why model diversity mattered before it was consensus, and how a company dismissed as “just a wrapper” became critical infrastructure for the AI ecosystem. We go deep on the product and distribution lessons behind OpenRoute r: why model labs can spend billions training a checkpoint and still struggle to get it into developers’ hands, how Mistral helped prove the value of a competitive inference marketplace, why OpenRouter chose focus over expanding into fine-tuning, memory, and other adjacent products, and how its rankings became a real-time map of how AI usage was changing. Alex also explains OpenRouter’s early experiments with model fusion , why they deleted the first version and brought it back years later, and how the platform grew to more than 10 trillion tokens per day. Finally, Anjney explains why Stripe and OpenRouter fit together , why token fraud may become one of the defining security problems of the AI economy , and why the next wave of fraud won’t just come from humans but from autonomous agents attacking increasingly valuable token flows . We discuss: * Why OpenRouter bet early that no single AI model would win everything * Alpaca, Llama, and open models becoming impossible to ignore * Why Discord’s early AI deployments exposed the limitations of closed models * Why model labs can spend billions on training and still fail at distribution * How OpenRouter became a neutral distribution layer for model developers * Why VCs dismissed OpenRouter as “just a marketplace” or “just a wrapper” * The Mistral price war and the first real proof of an inference marketplac
Episode page →Earlier this month, world model company Runway introduced GWM Worlds 2 , a research preview that “turns high-fidelity video and audio generation into real-time interactive simulation .” Runway calls this an “autoregressive diffusion” model; with autoregressive describing how it generates over time. One new feature in particular caught our eye: WorldPrompt , a proposed input format for specifying a generated world and the actions within it. It allows you to fix some aspects of a simulated environment — including the first frame — and then create a series of timestamped events . The events, or actions, can even be prompted in real-time. To understand the implications of WorldPrompt, we spoke to Kamil Sindi , Runway’s CTO, and Robin Kahlow , its Principal Research Scientist for generative video and multimodal AI. We also have exclusive comments from Anastasis Germanidis , co-founder & co-CEO of Runway, courtesy of a podcast swyx and Vibhu did with him. Who’s building real-time interactive world models? First, some context about world models that can generate interactive video and audio in real-time . Runway is reportedly valued at $5.3 billion , based on its most recent fund raise of $315 million in February . Its first release, GWM Worlds, was launched last December. Alongside Runway, there are several other notable projects in this domain: Google DeepMind’s Genie 3 (which also generates at 720p and 24 fps), Odyssey-2 Pro , and World Labs’ RTFM (Real-Time Frame Model). We’ve summarized their differences in the following table: Given the complexity and massive latency demands of real-time video and audio generation (which we’ll get into below), all of the projects listed above have limitations . For instance, Google notes that Genie 3 “can currently support a few minutes of continuous interaction, rather than extended hours.” But as our interviews with Runway show, real progress is being made. The central idea of WorldPrompt WorldPrompt, a new feature in GWM Worlds 2,
Episode page →The OpenAI → Hugging Face attack has people asking “what else do we need to worry about?” and Anthropic’s filters flag two things: cyber-security and biology. The natural question is: what about bio-security, then? Clem Delangue argues that cyber-warfare defensive capabilities need to be open and to keep pace with frontier models’ attack capabilities Radical Numerics co-founder Eric Nguyen sat down with us and explained why the same models that increase biological capability can also keep defense from falling behind. Building a virus from scratch While he was at Stanford, Eric couldn’t get traction on Genomic Language Models (GLMs) for a long time. Biologists didn’t believe it would work, didn’t think they could verify the output, and didn’t see important applications beyond what they could already do. He kept pushing, eventually helping lead the development of Evo and contributing to Evo 2 at Arc Institute. Those models were later used by a separate Arc/Stanford team to generate entire bacteriophage genomes that were synthesized into functional viruses ! Long context unlocks biological intelligence Early ChatGPT spit out poems and email, and early DNA language models like Evo and Evo-2 could build a genome from scratch. DNA is different, however, from natural language in that it has a very small alphabet (4 characters ACTG) and that its sequences are very long: * 60K for an average human gene * long being up to 2.3M * the whole human genome around 3B. Innovation in long-context models made this possible about 3 years ago (footnote: striped hyena), long before the frontier labs were building 1M+ context models. Now Eric and other AI x Bio luminaries have founded Radical Numerics to build and scale GLMs to tack a wide range of biological problems, extending well beyond generating DNA. Thinking in DNA Their GLMs already do pretty well with RNA and protein because there are clear markers in the DNA sequence for genes (RNA sequences the perform many functions) and speci
Episode page →