Claude Haiku 5.5 and GPT-6 Luna: the small-model price war hits $0.10 per million tokens
Small models have become nearly free to run. Anthropic launched Claude Haiku 5.5 on October 7, at the same price as OpenAI's GPT-6 Luna.
Pricing
| Per million tokens | Haiku 5.5 (prompt < 100,000) | Haiku 5.5 (prompt > 100,000) | GPT-6 Luna | |---|---|---|---| | Input | $0.10 | $0.50 | $0.10 | | Output | $0.50 | $2.50 | $0.50 |
Haiku 5.5 is available on Anthropic's platform, AWS, Google Cloud and Azure. Anthropic estimates workloads cost about 75% less than on Haiku 4.5, accounting for request sizes and token consumption, and says roughly 90% of Haiku 4.5 requests fall in the short tier.
Two caveats
- The sticker price isn't the real cost. Neowin's headline: Haiku 5.5 has aggressive pricing but consumes far more tokens. As with Sonnet 5.5, the right unit is cost per completed task, not price per token.
- The 100,000-token threshold matters. Beyond it, the rate is five times higher. An agent that accumulates context (history, documents, tool results) can slip into the expensive tier unnoticed.
Why it matters for agents
Small models do most of the work in an agent architecture: classification, routing, extraction, summarization, checks. At this price you can run those steps continuously and reserve large models for decisions that need them. It's also an argument for a single gateway (such as LiteLLM): with two vendors at the same price, being able to switch without rewriting workflows becomes a real lever.
What about self-hosting?
A local model that's free to run is still attractive for privacy, not for price. At $0.10 per million tokens, a VPS or dedicated GPU only pays off at very high volumes. Redo the math with your own numbers before migrating.
Sources: StreetInsider, Neowin, VentureBeat, Technology.org, iClarified.