Back to models
NewChat Models

Z.ai

GLM-5.3-FlashX

Fastest

Faster-serving GLM-5.3-Flash tier with native multimodal inputs, 1M context and separately priced API access.

Fast servingMultimodal

At a glance

Know the model before you prompt.

Z.ai API specifications
Context window
1M
Maximum output
128K
Inputs → output
Text + Images + Video + Files → Text
Knowledge cutoff
Not verified

API model ID: glm-5.3-flashx

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Faster serving for the native multimodal GLM-5.3-Flash model
  • Text, images, video and files as documented inputs
  • 1M context and 128K maximum output
  • Always-on reasoning, function calling and streaming in supported integrations

Before you choose

  • The provider reports 200 tokens/second for FlashX; that is not a guarantee of end-to-end speed, time to first token or performance inside EZ Ai Assist.
  • The shared guide says FlashX is not yet included in the GLM Coding Plan. Direct API and subscription access must be checked separately.
  • Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
  • Visual and file inputs support understanding with text output, not native image, video or audio generation. File format, size and ingestion requirements depend on the endpoint and app integration.
  • Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.

Choose the reasoning effort

lowhighmax · default

Thinking is always enabled and cannot be disabled. Set reasoning_effort to low, high or max; max is the default. A disabled thinking.type request is not supported. Lower effort changes the reasoning budget, not the model into a non-thinking endpoint.

  • Use the distinct glm-5.3-flashx API ID. The cached-input rate is $0.075 per million tokens, not $0.07 or the standard Flash rate.
  • Text parameters follow GLM-5.3; reasoning stays enabled. Do not interpret faster serving as permission to send a disabled-thinking request.
  • Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
  • For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Benchmark a visual review queue

Measure speed and accuracy on the same work.

Design a comparison of Flash and FlashX for the supplied screenshot-review tasks. Keep prompts, images and acceptance criteria identical. Define how to measure first-response latency, total duration, token usage, missed defects and rework. Return a blank results table and decision rule without inventing benchmark measurements.

Workflow 02

Triage incoming document packs

Use fast analysis without losing source traceability.

For each supplied document pack, identify its topic, missing required evidence and the next reviewer action. Include document and page references, mark unreadable content, and avoid filling gaps from general knowledge. Return a prioritized queue with a short reason per item; do not approve or submit anything.

Workflow 03

Compare design revisions quickly

Focus on observable differences.

Compare these before-and-after product screenshots using the supplied requirements. List only visible changes, regressions and ambiguous areas, identifying the image region for each. Keep stylistic preferences separate from requirement failures. Finish with the smallest interactive test set needed to validate the revision.

Developer reference

Z.ai API pricing

These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.37
Cached input$0.075
Output$1.25
  • Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
  • Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
  • The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.

Common questions

A few things worth knowing.

Is 200 tokens per second guaranteed?

No. It is a provider-reported throughput figure. Prompt size, reasoning, input processing, service load and app integration all affect experienced latency.

Are its context and reasoning different from Flash?

The shared guide gives both IDs a 1M context and 128K maximum output, with always-on thinking and GLM-5.3 text parameters. FlashX is the faster serving tier.

Why is its cached-input price shown to three decimals?

The official rate is $0.075 per million cached-input tokens. Rounding it to two decimals would obscure the actual listed rate.

Is FlashX included in the GLM Coding Plan?

The reviewed guide says it is not yet included, unlike standard Flash. This is a Z.ai product distinction and does not establish access through EZ Ai Assist.

Does the context window guarantee complete recall?

No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.

Are these prices the cost of my EZ Ai Assist plan?

No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.

Are all of these capabilities available in the app?

Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider