# gpt-6.1-sol — API, Pricing & Context Window | Vivgrid

> Source: https://vivgrid.com/docs/models/gpt-6.1-sol

Run OpenAI's gpt-6.1-sol on Vivgrid: the latest balanced gpt-6 model with a 1.05M-token context window, 128K max output, geo-distributed acceleration, and a unified API.

`gpt-6.1-sol` is the latest balanced model in OpenAI's gpt-6 family, succeeding `gpt-6-sol`. It keeps the generation's **1.05M-token context window** and **128K-token max output** at the same per-token price as `gpt-6-sol`, making it the new default for production coding agents that need more than a small model.

On Vivgrid, `gpt-6.1-sol` is served through a single OpenAI-compatible endpoint with **geo-distributed acceleration across AMER, EMEA, and APAC**, so requests are routed to the nearest compute region for lower latency. Moving an agent from `gpt-6-sol` is a model-string change in the Console, with no application code to touch.

It is available on both OpenAI surfaces — **Chat Completions** (`/chat/completions`) and the **Responses API** (`/responses`) — so coding CLIs that expect the Responses wire format, such as [Codex](/docs/tutorials/codex) and [Pi](/docs/tutorials/pi), work against the same key.

<ModelSpec slug="gpt-6.1-sol" />

Prompt caching is billed at **$2.50 per 1M cache-write tokens**, with cached reads at $0.20.

Requests beyond **272K input tokens** are billed at the long-context rate: $4.00 input / $0.40 cached input / $5.00 cache write / $15.00 output per 1M tokens.

## Ideal use cases

- Production coding agents balancing quality, speed, and cost
- Coding CLIs on the Responses API (Codex, Pi, OpenCode)
- Whole-repository analysis within a 1.05M-token context
- Teams on `gpt-6-sol` upgrading at the same price

## Related models

- [gpt-6-sol](/docs/models/gpt-6-sol) — the prior balanced gpt-6 release
- [gpt-6-astra](/docs/models/gpt-6-astra) — the gpt-6 generation flagship
- [gpt-6-luna](/docs/models/gpt-6-luna) — the fast, low-cost gpt-6 model
