Models

glm-5.3-flash — API, Pricing & Context Window | Vivgrid

glm-5.3-flash on Vivgrid: Zhipu AI's fast, low-cost multimodal model with image, video, and audio input, a 1M-token context window, and 384K output.

glm-5.3-flash is Zhipu AI's lightweight sibling to glm-5.3. It trades the flagship's coding-benchmark tuning for speed and price, and adds image, video, and audio input — so the same request can reason over text alongside media. It keeps the line's 1M-token context window and 384K-token max output.

On Vivgrid, glm-5.3-flash is reachable through the unified, OpenAI-compatible endpoint, so it integrates cleanly alongside models from OpenAI, Anthropic, Google, and DeepSeek under one API key.

Specifications

ProviderZhipu AI
Model IDglm-5.3-flash
Best forCoding
Available onChat Completions, Vibe coding
Context window1,000,000 tokens
Max output384,000 tokens
ModalitiesText, Image, Video, Audio
Tool / function callingYes
Knowledge cutoff2026-01
Acceleration🌐 Global (Centralized)

Pricing

Pricing in USD per 1M tokens, matching the provider's rates.

InputCached inputOutput
$0.15$0.04$0.50

What one request costs

The rates above applied to a 20K-token prompt with a 2K-token reply — a typical single agent turn, small enough that every request is billed at the base rate.

RequestCost
Cold prompt, nothing cached$0.00400
Warm prompt, 90% of the input served from cache$0.00202

Use glm-5.3-flash in your coding agent

glm-5.3-flash works with the same Vivgrid key and https://api.vivgrid.com/v1 endpoint as every other model in the catalog. Grab your API key from the Vivgrid Console, then drop one of the configurations below into your agent of choice.

OpenCode

OpenCode has Vivgrid built in, so you can use glm-5.3-flash without writing any config: open /models, search for vivgrid, paste your API key and pick the model. The OpenCode tutorial walks through it step by step.

Prefer to configure it by hand? Install OpenCode:

npm i -g opencode-ai

Declare glm-5.3-flash under the Vivgrid provider in your global OpenCode config, ~/.config/opencode/opencode.json. The @ai-sdk/openai-compatible provider talks to Vivgrid over Chat Completions (/v1/chat/completions):

~/.config/opencode/opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "model": "vivgrid/glm-5.3-flash",
  "provider": {
    "vivgrid": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Vivgrid",
      "options": {
        "baseURL": "https://api.vivgrid.com/v1",
        "apiKey": "{env:VIVGRID_API_KEY}"
      },
      "models": {
        "glm-5.3-flash": {
          "name": "glm-5.3-flash",
          "tool_call": true,
          "attachment": true,
          "modalities": { "input": ["text", "image", "video", "audio"], "output": ["text"] },
          "limit": { "context": 1000000, "output": 384000 }
        }
      }
    }
  }
}

Export your key and launch OpenCode:

export VIVGRID_API_KEY="viv-xxxxxxxxxxxxx"
opencode

Use /models to switch between glm-5.3-flash and other Vivgrid models mid-session. See the full OpenCode tutorial.

Pi

Install Pi:

npm install -g --ignore-scripts @earendil-works/pi-coding-agent

Add glm-5.3-flash to the Vivgrid provider in ~/.pi/agent/models.json. Pi talks to Vivgrid over Chat Completions (/v1/chat/completions):

~/.pi/agent/models.json
{
  "providers": {
    "vivgrid": {
      "baseUrl": "https://api.vivgrid.com/v1",
      "apiKey": "viv-xxxxxxxxxxxxx",
      "api": "openai-completions",
      "models": [
        {
          "id": "glm-5.3-flash",
          "input": ["text", "image"],
          "contextWindow": 1000000,
          "maxTokens": 384000
        }
      ]
    }
  }
}

Make it the model Pi starts with in ~/.pi/agent/settings.json:

~/.pi/agent/settings.json
{
  "defaultProvider": "vivgrid",
  "defaultModel": "glm-5.3-flash"
}

Then verify and start Pi:

pi --list-models vivgrid
pi

Or run a one-shot prompt without entering the TUI:

pi -p "explain the auth flow in this repo" --model vivgrid/glm-5.3-flash

See the full Pi tutorial.

API surfaces

The Vivgrid Console catalogs glm-5.3-flash on 2 surfaces, all behind the same API key. The endpoint shown is the one the Console previews for that surface.

SurfacePrimary endpointUse it for
Chat Completions/chat/completionsAgent projects, where Vivgrid injects the model, system prompt and tools server-side, and Model API calls that name the model per request.
Vibe coding/chat/completionsCoding CLIs and IDE agents that drive the model themselves — point the tool at Vivgrid and keep your own loop.

Quick start

Get an API key from the Vivgrid Console, then call glm-5.3-flash directly.

curl https://api.vivgrid.com/v1/chat/completions \
  -H "Authorization: Bearer $VIVGRID_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [
      { "role": "user", "content": "Say hello in English, Chinese and Spanish." }
    ],
    "stream": true
  }'

Choosing within the GLM 5.3 family

Prices are USD per 1M tokens at the base rate. Rows marked this page are the model described here.

ModelContextMax outputInputOutput
glm-5.31,000,000128,000$1.20$4.20
glm-5.3-flash — this page1,000,000384,000$0.15$0.50

Ideal use cases

  • High-volume, cost-sensitive agent traffic that still needs tool calling
  • Multimodal steps that mix text with image, video, or audio input
  • Whole-repository reasoning within a 1M-token context at low cost
  • Draft or first-pass generation ahead of a flagship model's final pass

On this page