I gave Opus 5.5 one prompt and six hours to visualize Invisible Cities
Piotr Migdał of Quesma gave GPT-6 Astra in Codex and Claude Opus 5.5 in Claude Code the same prompt: build a three.js visualization of every city in Calvino's Invisible Cities, with up to six hours to work. Opus ran six subagents in parallel for 1 hour 25 minutes. Astra finished in 53 minutes at medium effort and used roughly $10 in API tokens. Both results are live demos with published source code. The test copies an interactive optics demo Ryan Sael posted on X. In that demo, Opus 5.5 built a camera focus explainer in one shot, taking 1 hour 26 minutes and $25.66 in tokens. Migdał found that Astra picked up what amounts to the Claude visualization house style, with beige backgrounds and numbered sections. For founders, the gap is a useful benchmark: similar finished work cost about seven times more on one model than the other. Unattended, long-running agent sessions are becoming a real way to produce software, but how fast they finish, what they cost and how good the design is still vary a lot by model. Teams need to measure all three on their own tasks.