glm-5.3-flash — API, Pricing & Context Window | Vivgrid
glm-5.3-flash on Vivgrid: Zhipu AI's fast, low-cost multimodal model with image, video, and audio input, a 1M-token context window, and 384K output.
glm-5.3-flash is Zhipu AI's lightweight sibling to glm-5.3. It trades the flagship's coding-benchmark tuning for speed and price, and adds image, video, and audio input — so the same request can reason over text alongside media. It keeps the line's 1M-token context window and 384K-token max output.
On Vivgrid, glm-5.3-flash is reachable through the unified, OpenAI-compatible endpoint, so it integrates cleanly alongside models from OpenAI, Anthropic, Google, and DeepSeek under one API key.
Specifications
| Provider | Zhipu AI |
| Model ID | glm-5.3-flash |
| Best for | Coding |
| Available on | Chat Completions, Vibe coding |
| Context window | 1,000,000 tokens |
| Max output | 384,000 tokens |
| Modalities | Text, Image, Video, Audio |
| Tool / function calling | Yes |
| Knowledge cutoff | 2026-01 |
| Acceleration | 🌐 Global (Centralized) |
Pricing
Pricing in USD per 1M tokens, matching the provider's rates.
| Input | Cached input | Output |
|---|---|---|
| $0.15 | $0.04 | $0.50 |
What one request costs
The rates above applied to a 20K-token prompt with a 2K-token reply — a typical single agent turn, small enough that every request is billed at the base rate.
| Request | Cost |
|---|---|
| Cold prompt, nothing cached | $0.00400 |
| Warm prompt, 90% of the input served from cache | $0.00202 |
Use glm-5.3-flash in your coding agent
glm-5.3-flash works with the same Vivgrid key and https://api.vivgrid.com/v1 endpoint as every other model in the catalog. Grab your API key from the Vivgrid Console, then drop one of the configurations below into your agent of choice.
OpenCode
OpenCode has Vivgrid built in, so you can use glm-5.3-flash without writing any config: open /models, search for vivgrid, paste your API key and pick the model. The OpenCode tutorial walks through it step by step.
Prefer to configure it by hand? Install OpenCode:
npm i -g opencode-aiDeclare glm-5.3-flash under the Vivgrid provider in your global OpenCode config, ~/.config/opencode/opencode.json. The @ai-sdk/openai-compatible provider talks to Vivgrid over Chat Completions (/v1/chat/completions):
{
"$schema": "https://opencode.ai/config.json",
"model": "vivgrid/glm-5.3-flash",
"provider": {
"vivgrid": {
"npm": "@ai-sdk/openai-compatible",
"name": "Vivgrid",
"options": {
"baseURL": "https://api.vivgrid.com/v1",
"apiKey": "{env:VIVGRID_API_KEY}"
},
"models": {
"glm-5.3-flash": {
"name": "glm-5.3-flash",
"tool_call": true,
"attachment": true,
"modalities": { "input": ["text", "image", "video", "audio"], "output": ["text"] },
"limit": { "context": 1000000, "output": 384000 }
}
}
}
}
}Export your key and launch OpenCode:
export VIVGRID_API_KEY="viv-xxxxxxxxxxxxx"
opencodeUse /models to switch between glm-5.3-flash and other Vivgrid models mid-session. See the full OpenCode tutorial.
Pi
Install Pi:
npm install -g --ignore-scripts @earendil-works/pi-coding-agentAdd glm-5.3-flash to the Vivgrid provider in ~/.pi/agent/models.json. Pi talks to Vivgrid over Chat Completions (/v1/chat/completions):
{
"providers": {
"vivgrid": {
"baseUrl": "https://api.vivgrid.com/v1",
"apiKey": "viv-xxxxxxxxxxxxx",
"api": "openai-completions",
"models": [
{
"id": "glm-5.3-flash",
"input": ["text", "image"],
"contextWindow": 1000000,
"maxTokens": 384000
}
]
}
}
}Make it the model Pi starts with in ~/.pi/agent/settings.json:
{
"defaultProvider": "vivgrid",
"defaultModel": "glm-5.3-flash"
}Then verify and start Pi:
pi --list-models vivgrid
piOr run a one-shot prompt without entering the TUI:
pi -p "explain the auth flow in this repo" --model vivgrid/glm-5.3-flashSee the full Pi tutorial.
API surfaces
The Vivgrid Console catalogs glm-5.3-flash on 2 surfaces, all behind the same API key. The endpoint shown is the one the Console previews for that surface.
| Surface | Primary endpoint | Use it for |
|---|---|---|
| Chat Completions | /chat/completions | Agent projects, where Vivgrid injects the model, system prompt and tools server-side, and Model API calls that name the model per request. |
| Vibe coding | /chat/completions | Coding CLIs and IDE agents that drive the model themselves — point the tool at Vivgrid and keep your own loop. |
Quick start
Get an API key from the Vivgrid Console, then call glm-5.3-flash directly.
curl https://api.vivgrid.com/v1/chat/completions \
-H "Authorization: Bearer $VIVGRID_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3-flash",
"messages": [
{ "role": "user", "content": "Say hello in English, Chinese and Spanish." }
],
"stream": true
}'Choosing within the GLM 5.3 family
Prices are USD per 1M tokens at the base rate. Rows marked this page are the model described here.
| Model | Context | Max output | Input | Output |
|---|---|---|---|---|
glm-5.3 | 1,000,000 | 128,000 | $1.20 | $4.20 |
glm-5.3-flash — this page | 1,000,000 | 384,000 | $0.15 | $0.50 |
Ideal use cases
- High-volume, cost-sensitive agent traffic that still needs tool calling
- Multimodal steps that mix text with image, video, or audio input
- Whole-repository reasoning within a 1M-token context at low cost
- Draft or first-pass generation ahead of a flagship model's final pass
Related models
- gpt-6-astra — OpenAI's frontier gpt-6 coding model
- glm-5.3 — the flagship, text-only coding model this is derived from
- deepseek-v4-flash-vision-exp — alternative low-cost multimodal model