Back to models
Chat Models

Z.ai

GLM-4.5-X

Premium high-speed GLM-4.5 text tier with 128K context, hybrid thinking and separate API rates.

Fast serving128K context

At a glance

Know the model before you prompt.

Z.ai API specifications
Context window
128K
Maximum output
96K
Inputs → output
Text → Text
Knowledge cutoff
Not verified

API model ID: glm-4.5-x

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • High-speed tier named separately in the GLM-4.5 guide
  • 128K context in the official model overview
  • 96K maximum output documented for the GLM-4.5 text series
  • Hybrid thinking and function-oriented text workflows

Before you choose

  • Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
  • This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
  • Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.

Extended thinking

enabled · defaultdisabled

The GLM-4.5 family uses hybrid thinking by default: with thinking.type enabled, the model decides whether reasoning is needed. Set disabled for the non-thinking path. The enabled setting is not a guarantee of a long reasoning trace on every request.

  • Use glm-4.5-x, not glm-4.5. The X pricing row is substantially higher and is not an automatic acceleration option on the base ID.
  • The family guide describes ultra-fast response but does not establish a request-level speed guarantee. Benchmark your actual workload.
  • Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
  • For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Measure a premium serving tier

Decide whether speed earns its cost.

Plan a comparison of base GLM-4.5 and its X tier for these time-sensitive text tasks. Hold prompts and scoring criteria constant, record total duration, token usage and rework, and define a cost-per-successful-task calculation. Provide a results template rather than invented measurements.

Workflow 02

Prepare a rapid incident brief

Stay factual when time is limited.

Use the incident timeline, logs and status notes below to write a concise briefing. Separate confirmed impact, active hypotheses and next diagnostic steps, citing each source. Flag conflicting timestamps and unknown scope. Do not claim recovery or assign blame without evidence.

Workflow 03

Review a deadline-sensitive change

Keep urgency from expanding the patch.

Review this urgent change against its stated acceptance criteria. List concrete failure modes and the smallest tests needed before release. Identify any proposed changes outside scope and explain what can safely wait. Do not approve deployment or imply that unrun checks have passed.

Developer reference

Z.ai API pricing

These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$2.20
Cached input$0.45
Output$8.90
  • Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
  • Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
  • The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.

Common questions

A few things worth knowing.

What is the X tier for?

The provider presents it as a high-performance, ultra-fast-response option in the GLM-4.5 family. It has a separate API ID and premium rates.

Does the X label guarantee better answers?

No. Faster serving is not a quality guarantee. Compare correctness and rework as well as latency on the tasks you intend to run.

How are its limits verified?

The overview lists 128K context for GLM-4.5-X. The API reference documents 96K maximum output for the GLM-4.5 text series, which includes the X ID.

Is X a vision model?

No. It is a text-family tier. GLM-4.5V is the separately documented vision model.

Does the context window guarantee complete recall?

No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.

Are these prices the cost of my EZ Ai Assist plan?

No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.

Are all of these capabilities available in the app?

Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider