GLM-5.3-FlashX

GLM-5.3-FlashX 智谱多模态理解模型,1M上下文,推理速度达 200 tokens/s,提供更快、更流畅的模型体验。

Model ID: glm-5.3-flashx · Type: chat · Provider: Zhipu

Endpoints: /v1/chat/completions · /v1/messages · /v1/responses

Pricing

Input (per 1M tokens)$0.1665 USD
Output (per 1M tokens)$0.5625 USD
Cache read (per 1M tokens)$0.03375 USD
from openai import OpenAI

client = OpenAI(api_key="sk-...", base_url="https://api.router.ai/v1")
resp = client.chat.completions.create(
    model="glm-5.3-flashx",
    messages=[{"role": "user", "content": "Hello"}],
)
print(resp.choices[0].message.content)

FAQ

How much does GLM-5.3-FlashX cost on openbat.ai?

GLM-5.3-FlashX (`glm-5.3-flashx`) is billed per usage at $0.1665/1M in · $0.5625/1M out, in USD. Current pricing is always listed at https://openbat.ai/models/glm-5.3-flashx.

How do I call GLM-5.3-FlashX through openbat.ai?

Send a request to https://api.router.ai/v1/v1/chat/completions with the header `Authorization: Bearer <your API key>` and `"model": "glm-5.3-flashx"`. The API is OpenAI-compatible, so any OpenAI SDK works by changing base_url to https://api.router.ai/v1 — no other code change.

Which endpoints does GLM-5.3-FlashX support?

GLM-5.3-FlashX can be called on: /v1/chat/completions; /v1/messages; /v1/responses.

What is GLM-5.3-FlashX's context window?

GLM-5.3-FlashX accepts up to 1,048,576 input tokens and can return up to 131,072 output tokens. Requests exceeding the input limit are rejected before reaching the model.

What can GLM-5.3-FlashX do?

GLM-5.3-FlashX supports: vision, function_calling, prompt_caching.

Who makes GLM-5.3-FlashX?

GLM-5.3-FlashX is a chat model from Zhipu, available through the openbat.ai gateway with the same API key as every other model.

Call it through the openbat.ai OpenAI-compatible endpoint (API base: https://api.router.ai/v1). AI agents can discover and call every model on this gateway through MCP (https://mcp.router.ai/mcp) with no manual integration.

API reference · All models