GLM 5.3 Flash

glm-5.3-flash · Z.AI

GLM-5.3-Flash is a high-efficiency multimodal model from Z.AI. It supports a context window of roughly 1 million tokens, along with text, image, and video inputs, and includes tool-calling capabilities. It is primarily designed for coding agents, complex reasoning, and long-horizon software engineering tasks. Built on the existing GLM technology stack, the model has been further post-trained and optimized to deliver strong performance while placing greater emphasis on inference efficiency, responsiveness, and cost.The model is offered at a limited-time 50% discount; users are welcome to try it.

API Pricing

Input$0.113$0.056 / 1M tokens
Output$0.394$0.197 / 1M tokens
Cache read$0.028$0.014 / 1M tokens

50% off

Specifications

Context1.05M tokens
Max output131K tokens
Modalitiestext, image, video
CapabilitiesThinking, Streaming, Tool calling, Prompt caching

Frequently asked questions

What is GLM 5.3 Flash?

GLM-5.3-Flash is a high-efficiency multimodal model from Z.AI. It supports a context window of roughly 1 million tokens, along with text, image, and video inputs, and includes tool-calling capabilities. It is primarily designed for coding agents, complex reasoning, and long-horizon software engineering tasks. Built on the existing GLM technology stack, the model has been further post-trained and optimized to deliver strong performance while placing greater emphasis on inference efficiency, responsiveness, and cost.The model is offered at a limited-time 50% discount; users are welcome to try it.

What is the context length of GLM 5.3 Flash?

GLM 5.3 Flash has a 1,048,576 token context window. It supports up to 131,072 output tokens.

How much does GLM 5.3 Flash cost?

On AIHubMix, GLM 5.3 Flash costs $0.056 per million input tokens and $0.197 per million output tokens. Cached input reads are billed at $0.014 per million tokens. These are current promotional rates — list price is $0.113 input and $0.394 output per million tokens.

What modalities does GLM 5.3 Flash support?

GLM 5.3 Flash accepts text, image and video input.

What capabilities does GLM 5.3 Flash support?

GLM 5.3 Flash supports thinking, tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.

How do I call GLM 5.3 Flash via API?

GLM 5.3 Flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to glm-5.3-flash — no other code changes needed.

Who created GLM 5.3 Flash?

GLM 5.3 Flash is developed by Z.AI. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

When was GLM 5.3 Flash released?

GLM 5.3 Flash was released on August 26, 2026 by Z.AI.

More models from Z.AI

See all Z.AI models →

GLM 5.3

by Z.AI

GLM-5.3 is Z.AI’s coding and agentic reasoning model, built for complex software…

$1.127 $1.014/1M in · $3.944 $3.549/1M out
10% off
1,000,000 tokens context

Coding GLM 5.3

by Z.AI

GLM-5.3 is Z.ai’s reasoning model for coding and agentic workflows, designed for complex…

$0.06/1M in · $0.22/1M out

Coding GLM 5.3 Flash (free)

by Z.AI

coding-glm-5.3-flash-free is the open and free version of coding-glm-5.3-flash. To ensure…

Coding GLM 5.3 Flash

by Z.AI

Coding GLM 5.3 Flash is a dedicated version of GLM 5.3 Flash built for AI coding and…

$0.028/1M in · $0.099/1M out
1,000,000 tokens context

Coding GLM 5.3 (free)

by Z.AI

coding-glm-5.3-free is the open and free version of coding-glm-5.3. To ensure stable…

Ox Alpha

by Z.AI

This model actually points to glm-5.3-flash; if you need to use it in production, you can…

1,048,576 tokens context

Use GLM 5.3 Flash via the AIHubMix unified API — one interface for every major LLM.