# claude-haiku-5-5 — API, Pricing & Context Window | Vivgrid

> Source: https://vivgrid.com/docs/models/claude-haiku-5-5

Run Anthropic's claude-haiku-5-5 on Vivgrid: the latest and fastest Haiku model for high-volume, latency-sensitive work, with a 1M-token context window, 128K max output, and geo-distributed acceleration.

`claude-haiku-5-5` is the latest model in Anthropic's Haiku tier and the fastest model in the current Claude lineup, built for high-volume, latency-sensitive work such as classification, routing, extraction, and subagent tasks. Succeeding `claude-haiku-4-5`, it moves to a **1M-token context window** and **128K-token max output** at a lower per-token price, and runs with adaptive thinking on by default.

On Vivgrid, `claude-haiku-5-5` is served through the native **Messages API** (`/messages`) with geo-distributed acceleration across **AMER, EMEA, and APAC**. It uses the same unified Vivgrid API key as the rest of the catalog, so moving an agent from `claude-haiku-4-5` is a model-string change. It uses Anthropic's newer tokenizer, so the same text counts as roughly 30% more tokens than on `claude-haiku-4-5`.

<ModelSpec slug="claude-haiku-5-5" />

Prompt caching is billed at **$0.125 per 1M cache-write tokens** on the 5-minute TTL and **$0.20 per 1M** on the 1-hour TTL, with cached reads at $0.01.

Prompts over **100K tokens** (uncached input, cache reads, and cache writes combined) are billed at the long-context rate: $0.50 input / $0.05 cached read / $0.625 5-minute cache write / $1.00 1-hour cache write / $2.50 output per 1M tokens.

## Ideal use cases

- High-volume classification, extraction, and routing
- Subagents and fast steps inside larger agent pipelines
- Real-time assistants where latency matters most
- Teams on `claude-haiku-4-5` moving to a 1M-token context at a lower price

## Related models

- [claude-haiku-4-5](/docs/models/claude-haiku-4-5) — the prior Haiku release
- [claude-sonnet-5-5](/docs/models/claude-sonnet-5-5) — the latest Sonnet model, a step up in capability
- [gpt-6-luna](/docs/models/gpt-6-luna) — the fast, low-cost gpt-6 model
