# glm-5.3-flash — API, Pricing & Context Window | Vivgrid

> Source: https://vivgrid.com/docs/models/glm-5.3-flash

glm-5.3-flash on Vivgrid: Zhipu AI's fast, low-cost multimodal model with image, video, and audio input, a 1M-token context window, and 384K output.

`glm-5.3-flash` is Zhipu AI's lightweight sibling to `glm-5.3`. It trades the flagship's coding-benchmark tuning for speed and price, and adds image, video, and audio input — so the same request can reason over text alongside media. It keeps the line's **1M-token context window** and **384K-token max output**.

On Vivgrid, `glm-5.3-flash` is reachable through the unified, OpenAI-compatible endpoint, so it integrates cleanly alongside models from OpenAI, Anthropic, Google, and DeepSeek under one API key.

## Specifications

| Field | Value |
| --- | --- |
| Provider | Zhipu AI |
| Model ID | `glm-5.3-flash` |
| Best for | Coding |
| Available on | Chat Completions, Vibe coding |
| Context window | 1,000,000 tokens |
| Max output | 384,000 tokens |
| Modalities | Text, Image, Video, Audio |
| Tool / function calling | Yes |
| Knowledge cutoff | 2026-01 |
| Acceleration | Global (centralized) |

## Pricing

Pricing in USD per 1M tokens, matching the provider's rates.

| Input | Cached input | Output |
| --- | --- | --- |
| $0.15 | $0.04 | $0.50 |

### What one request costs

The rates above applied to a 20K-token prompt with a 2K-token reply — a typical single agent turn, small enough that every request is billed at the base rate.

| Request | Cost |
| --- | --- |
| Cold prompt, nothing cached | $0.00400 |
| Warm prompt, 90% of the input served from cache | $0.00202 |


## Use glm-5.3-flash in your coding agent

`glm-5.3-flash` works with the same Vivgrid key and `https://api.vivgrid.com/v1` endpoint as every other model in the catalog. Grab your API key from the [Vivgrid Console](https://console.vivgrid.com), then drop one of the configurations below into your agent of choice.

### OpenCode

[OpenCode](https://opencode.ai) has Vivgrid built in, so you can use `glm-5.3-flash` without writing any config: open `/models`, search for `vivgrid`, paste your API key and pick the model. The [OpenCode tutorial](/docs/tutorials/opencode) walks through it step by step.

Prefer to configure it by hand? Install OpenCode:

```bash
npm i -g opencode-ai
```

Declare `glm-5.3-flash` under the Vivgrid provider in your global OpenCode config, `~/.config/opencode/opencode.json`. The `@ai-sdk/openai-compatible` provider talks to Vivgrid over **Chat Completions** (`/v1/chat/completions`):

```json title="~/.config/opencode/opencode.json" lines icon="json" highlight={3,6,9,10}
{
  "$schema": "https://opencode.ai/config.json",
  "model": "vivgrid/glm-5.3-flash",
  "provider": {
    "vivgrid": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Vivgrid",
      "options": {
        "baseURL": "https://api.vivgrid.com/v1",
        "apiKey": "{env:VIVGRID_API_KEY}"
      },
      "models": {
        "glm-5.3-flash": {
          "name": "glm-5.3-flash",
          "tool_call": true,
          "attachment": true,
          "modalities": { "input": ["text", "image", "video", "audio"], "output": ["text"] },
          "limit": { "context": 1000000, "output": 384000 }
        }
      }
    }
  }
}
```

Export your key and launch OpenCode:

```bash
export VIVGRID_API_KEY="viv-xxxxxxxxxxxxx"
opencode
```

Use `/models` to switch between `glm-5.3-flash` and other Vivgrid models mid-session. See the full [OpenCode tutorial](/docs/tutorials/opencode).

### Pi

Install [Pi](https://pi.dev):

```bash
npm install -g --ignore-scripts @earendil-works/pi-coding-agent
```

Add `glm-5.3-flash` to the Vivgrid provider in `~/.pi/agent/models.json`. Pi talks to Vivgrid over **Chat Completions** (`/v1/chat/completions`):

```json title="~/.pi/agent/models.json" lines icon="json" highlight={4,5,9}
{
  "providers": {
    "vivgrid": {
      "baseUrl": "https://api.vivgrid.com/v1",
      "apiKey": "viv-xxxxxxxxxxxxx",
      "api": "openai-completions",
      "models": [
        {
          "id": "glm-5.3-flash",
          "input": ["text", "image"],
          "contextWindow": 1000000,
          "maxTokens": 384000
        }
      ]
    }
  }
}
```

Make it the model Pi starts with in `~/.pi/agent/settings.json`:

```json title="~/.pi/agent/settings.json" lines icon="json" highlight={2,3}
{
  "defaultProvider": "vivgrid",
  "defaultModel": "glm-5.3-flash"
}
```

Then verify and start Pi:

```bash
pi --list-models vivgrid
pi
```

Or run a one-shot prompt without entering the TUI:

```bash
pi -p "explain the auth flow in this repo" --model vivgrid/glm-5.3-flash
```

See the full [Pi tutorial](/docs/tutorials/pi).

## API surfaces

The Vivgrid Console catalogs `glm-5.3-flash` on 2 surfaces, all behind the same API key. The endpoint shown is the one the Console previews for that surface.

| Surface | Primary endpoint | Use it for |
| --- | --- | --- |
| Chat Completions | `/chat/completions` | Agent projects, where Vivgrid injects the model, system prompt and tools server-side, and Model API calls that name the model per request. |
| Vibe coding | `/chat/completions` | Coding CLIs and IDE agents that drive the model themselves — point the tool at Vivgrid and keep your own loop. |

## Quick start

Get an API key from the Vivgrid Console (https://console.vivgrid.com), then call `glm-5.3-flash` directly.

```bash
curl https://api.vivgrid.com/v1/chat/completions \
  -H "Authorization: Bearer $VIVGRID_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [
      { "role": "user", "content": "Say hello in English, Chinese and Spanish." }
    ],
    "stream": true
  }'
```

## Choosing within the GLM 5.3 family

Prices are USD per 1M tokens at the base rate.

| Model | Context | Max output | Input | Output |
| --- | --- | --- | --- | --- |
| [`glm-5.3`](/docs/models/glm-5.3) | 1,000,000 | 128,000 | $1.20 | $4.20 |
| `glm-5.3-flash` (this page) | 1,000,000 | 384,000 | $0.15 | $0.50 |


## Ideal use cases

- High-volume, cost-sensitive agent traffic that still needs tool calling
- Multimodal steps that mix text with image, video, or audio input
- Whole-repository reasoning within a 1M-token context at low cost
- Draft or first-pass generation ahead of a flagship model's final pass

## Related models

- [gpt-6-astra](/docs/models/gpt-6-astra) — OpenAI's frontier gpt-6 coding model
- [glm-5.3](/docs/models/glm-5.3) — the flagship, text-only coding model this is derived from
- [deepseek-v4-flash-vision-exp](/docs/models/deepseek-v4-flash-vision-exp) — alternative low-cost multimodal model
