Z.ai · Model identity
GLM 5.3 FlashX
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
1pricing route
What published route metadata supports
CachingReasoningStructured outputTool useVideo inputVision
Input modalities
ImageTextVideo
Supported API parameters
include_reasoningmax_tokensreasoningreasoning_effortresponse_formattemperaturetool_choicetoolstop_ktop_p
Only source-backed fields are filled
Release dateNot independently verified
Knowledge cutoffNot independently verified
Open weightsNot independently verified
OpenRouter catalog date2026-09-18
LLMPrice does not infer a release date, knowledge cutoff, or open-weight status from a model name, Hugging Face link, or pricing route.
Could this workload cost less?Compare GLM 5.3 FlashX with cheaper models using the same workload.
Find cheaper alternatives →| Endpoint / route | Model ID | Input | Cached | Output | Context | Source |
|---|---|---|---|---|---|---|
| OpenRouterStandard | z-ai/glm-5.3-flashx | $0.37 | $0.075 | $1.25 | 1.05M | OpenRouter API ↗ |
0 verified price changes since Aug 26, 2026
1 route availability change recorded. No earlier prices are inferred.
- Sep 19, 2026OpenRouter · Standard
New pricing route added