# gpt-6-luna — API, Pricing & Context Window | Vivgrid

> Source: https://vivgrid.com/docs/models/gpt-6-luna

gpt-6-luna on Vivgrid: OpenAI's fast, low-cost gpt-6 model with a 1.05M-token context window for high-volume agent workloads.

`gpt-6-luna` is the fast, low-cost model in OpenAI's gpt-6 family. It keeps the generation's **1.05M-token context window** and **128K-token max output** at $0.10 per 1M input tokens, making it the cheapest way on Vivgrid to run long-context agents at scale.

On Vivgrid, `gpt-6-luna` is served through a single OpenAI-compatible endpoint with **geo-distributed acceleration across AMER, EMEA, and APAC**, reachable with the same unified key as the rest of the catalog. It is available on both **Chat Completions** (`/chat/completions`) and the **Responses API** (`/responses`).

<ModelSpec slug="gpt-6-luna" />

Prompt caching is billed at **$0.125 per 1M cache-write tokens**, with cached reads at $0.01.

Requests beyond **272K input tokens** are billed at the long-context rate: $0.20 input / $0.02 cached input / $0.25 cache write / $0.75 output per 1M tokens.

## Ideal use cases

- High-volume agent and content pipelines where cost dominates
- Long-context retrieval, summarization, and classification
- Fast interactive assistants and chat products
- Fallback or draft model alongside gpt-6-sol / astra

## Related models

- [gpt-6-sol](/docs/models/gpt-6-sol) — the balanced gpt-6 model
- [gpt-6-astra](/docs/models/gpt-6-astra) — the gpt-6 generation flagship
- [gpt-5.6-luna](/docs/models/gpt-5.6-luna) — the prior-generation low-cost model
