Anthropic ships Claude Haiku 5.5 at $0.10 per million input tokens, a 90% cut from Haiku 4.5 on short prompts
The model went live on October 7, 2026, and is available on the Claude Platform, Amazon Web Services, Google Cloud, and Microsoft Azure. Anthropic puts the average saving versus Haiku 4.5 at about 75%, because about 90% of Haiku 4.5 requests fell into the lower tier. Cache reads cost $0.01 per million tokens on shorter prompts. The context window grows from 200K to 1M tokens, and max output doubles from 64K to 128K. The 72.4% OSWorld result counts partial credit. On the same test, the rate for completing every checkpoint is 37.1%. The new price is in line with rival OpenAI's smallest model, GPT-6 Luna. Anthropic also cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens, which it says should make most agentic tasks about 20% cheaper. One catch: a new tokenizer eats into some of those savings by consuming more tokens per task. For founders, routing, extraction and subagent calls just got about ten times cheaper per token. Before migrating, check where real workloads land relative to the 100,000-token line, and how many tokens the new tokenizer actually uses on your own prompts.