OpenAI DevDay 2026 live blog
In this newsletter:
OpenAI DevDay 2026 live blog
We’re going to need default hard budget caps on pretty much everything
Plus 1 link, 3 quotations, 2 notes, 2 releases, 1 research report, 1 tool, 1 museum, and 1 comment
Thanks for reading Simon Willison’s Newsletter! Subscribe for free to receive new posts and support my work.
Sponsor message: Voice agents sound off when TTS has no idea what was just said. Flux TTS carries context across turns and picks the right tone. As low as 80ms to first response. Works with the voice agent tools you already use. Hear the demo.
OpenAI DevDay 2026 live blog - 2026-09-29
I spent Tuesday at OpenAI DevDay, in Fort Mason, San Francisco. Same as last year I live blogged the keynote and some other notes during the day.
OpenAI gave me a free ticket and a seat in the “creator” area for the keynote.
I won’t include the full turn-my-turn liveblog notes in this newsletter, but here are the most interesting announcements from the keynote:
dots is OpenAI’s new “personal AI agent” product, and the signature feature of the entire event - all of DevDay’s branding was based around dots. They’re planning to provide “specialist dots” for legal, finance, etc as well.
ChatGPT Space is a new collaboration platform within ChatGPT, a little bit like Google Drive crossed with Google Docs.
GPT-6.1 Sol described as GPT-6 Astra but cheaper and faster - 1/5th of the price of Astra.
Ultrafast mode: 8x faster - up to 300 tokens/second. Available in the API, ChatGPT, and Codex. 6x the price of standard.
Pro 500 plan: A new subscription plan for $500/month: This gives you access to Ultrafast, and 25x the usage of Plus.
Decisions API (preview): Their repsonse to Jev - this builds on top of Luna to provide a fast response from a predefined set of options.
Codex Cloud: An updated version of their cloud version of Codex.
Codex Security Cloud: With “Daybreak Blue access” - scheduled scans, automatic de-duplication of detected issues.
Sign in with ChatGPT: lets people sign into your app and use the tokens they are paying for already. I’ve wanted this one for years!
We’re going to need default hard budget caps on pretty much everything - 2026-10-03
Here’s a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps. I’m talking about the feature of pay-by-usage services and APIs that lets you say “after $X/month, cut this thing off and return errors”. These need to be hard limits. Soft caps, “after $X/month, send me a warning email”, will not cut it.
Coding agents, and personal agents (coding agents wrapped in a less threatening UI), greatly reduce the friction of spinning up code that can do useful things. Sometimes those things cost money - calls to paid APIs, or hosted web applications, or systems that can bill for additional storage and compute.
Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage.
An argument against this is that businesses don’t want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.
I think hard budget caps need to be the default. If someone wants to live dangerously they should be able to do that, but it needs to be on an opt-in basis. Have a nice, clear checkbox somewhere prominent:
Remove the budget cap. My application will not be shut down if I exceed the configured budget limit, and I will be responsible for subsequent charges.
The service I most want to see this from is AWS. I’ve heard plenty of stories from people who refuse to use AWS for personal projects out of (justified) fear that a runaway service might bankrupt them. I’ve also heard stories from people who didn’t anticipate this and ended up seriously burned.
... and it turns out AWS finally launched spending limits a few weeks ago! From their announcement New AWS experience helps builders get started and ship faster on 16th September:
When you’re ready to upgrade to a paid plan, you can set a monthly spend limit for your project based on your usage patterns so that you stay within your budget. If a project’s usage reaches its spend limit, your project is paused for that month.
See also Create a spend limit in AWS Settings, though that page warns that “We’re currently releasing our new experience to a limited number of customers.” Here’s hoping that hits general availability for existing accounts soon.
Google Cloud launched a similar feature in July, called Spend Caps, which lets you “set a monthly financial cap on specific services within a project”. Looks like this is becoming a trend!
In an ideal world, our agents could help with this. It would be great if agents started biasing towards recommending providers with hard budget caps, and warning new and inexperienced builders against deploying applications using uncapped services that might get them into trouble.
Quote 2026-09-28
To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem. [...]
So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know what to do when something goes wrong? Do I have the right incident response? The right comms and messaging? Do I have the right people ready to go when capabilities jump?
@joedaroo, Agent Security at OpenAI, identity confirmed by The Information’s Rocket Drew
Link 2026-09-28 Claude Sonnet 5.5:
New Sonnet model from Anthropic. They say it “runs 30%+ faster, and costs up to 30% less for most work” - it’s priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well.
Here are some pelicans riding bicycles. Sonnet 5.5 suffered from the same bug as Opus 5.5: the “max” thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG.
Here’s the pelican it gave me for thinking effort “xhigh”, at a cost of 5.74 cents and taking 41 seconds:
Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks.
The most interesting thing about Sonnet 5.5 is that it’s now the model used for the free tier on claude.ai. OpenAI’s ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering.
I ran this prompt against that free tier:
build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL
And got back this page, which is a solid effort.
Anthropic’s announcement reiterates that Haiku 5.5 will be available “in the coming weeks”. I really hope that one is price-competitive with GPT-6 Luna!
Release: llm-anthropic 0.30
In addition to Claude Sonnet 5.5, this release adds the ability to run llm anthropic refresh to refresh the list of Anthropic models directly from their API - which means I don’t need to push a new release just to add support for a newly released model.
I also added an llm anthropic count command which can use their free token counting API to return a count of tokens that will be used by any prompt, before you send that prompt.
Tool: Photo Scrubber — local face blur & metadata removal
I took a photograph of some protesters, then thought about how I don’t like sharing photographs of strangers with identifiable faces. I had GPT-6 Astra build this experimental tool that would identify faces and automatically blur them out.
It uses Google’s MediaPipe C++ library, compiled to WebAssembly via @mediapipe/tasks-vision, plus the BlazeFaceface detection model.
comment: GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
I’m a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv...
Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
They’re not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr...
Quote 2026-09-29
We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.
Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities
Note 2026-09-30
I visited the Museum of the City of New York today and got to see He Built This City: Joe Macken’s Model, the 50 x27 feet model of the city built over a 21 year period from balsa wood and cardboard.
It exceeded my already high expectations. The exhibition closes on 12th October so you should absolutely make a priority to see it if you get the chance.
Quote 2026-10-01
[...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs.
Matthew Green, Is sandboxing sufficient to contain rogue agents?
Release: pwasm 0.2a0
pwasm is one of my folly projects - an entirely vibe-coded pure Python WebAssembly engine that I built in January during my first bout of AI mania.
I hadn’t touched it since January, so I decided to let Claude Opus 5.5 loose on it and see if it could make any significant improvements:
Evaluate current state of pwasm - then consider what it would take to get the MicroPython and micro JavaScript experiments from the research repo working under it - and what it would take to speed it up
42 commits later (with minimal follow-up prompting) it now handles almost all of the WASM specification and the wheel from PyPI bundles working WASM builds of MicroPython, QuickJS and Micro QuickJS.
I wouldn’t trust this thing at all - hence the alpha version tag - but it’s interesting seeing how today’s models can improve on the work of models from 10 months ago.
Museum: Rex’s Dino Store
Located just before the turnstiles in the Grand Army Plaza subway station at the north end of Brooklyn’s Prospect Park is this former newsstand which is now operated by a dinosaur.
The density of dinosaur puns is exceptional.
Note 2026-10-03
I just sent the September edition of my sponsors-only monthly newsletter. If you are a sponsor (or start a sponsorship now) you can access it here.
This month:
More Fable class models
A pricing war
3D graphics, Blender, and pixel art
LLMs come for mathematics
So many more accidental cyberattacks
The vulnapocalypse comes for Datasette
What I’m using right now
My software releases this month
2026 in LLMs (so far)
Here’s a copy of the August newsletter as a preview of what you’ll get. Pay $10/month to stay a month ahead of the free copy!
Research: Qwen3.8 27B addition in words
Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could “compute the sum but return the answer in words” across increasingly large numbers. Here’s the chart he shared of those results:
I’m confident GPT-4o didn’t cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment.
I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf. Here’s the result for a run of 30 attempts per combination with reasoning disabled:
Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%:
It got the right answer in 167 out of 169 attempts, and since these were one-shot I’m confident a second run would produce different results here.
Here’s a version of the report that includes the reasoning traces from some of those larger calculations, which include text like this:
Wait, let me redo this more carefully.
4,299,366,105,622
6,088,794,067,970
Let me align them:
4 2 9 9 3 6 6 1 0 5 6 2 2
6 0 8 8 7 9 4 0 6 7 9 7 0
Adding from right to left:
Position 1 (units): 2 + 0 = 2
Position 2 (tens): 2 + 7 = 9
Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1If you find this newsletter useful, please consider sponsoring me via GitHub. $10/month and higher sponsors get a monthly newsletter with my summary of the most important trends of the past 30 days - here are previews from Juneand July and August.
Thanks for reading Simon Willison’s Newsletter! Subscribe for free to receive new posts and support my work.



