Back to models
Chat Models

Z.ai

GLM-4.5-AirX

Faster-serving lightweight GLM-4.5 tier with 128K context and higher API rates than standard Air.

LightweightFast serving

At a glance

Know the model before you prompt.

Z.ai API specifications
Context window
128K
Maximum output
96K
Inputs → output
Text → Text
Knowledge cutoff
Not verified

API model ID: glm-4.5-airx

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Lightweight high-speed tier in the GLM-4.5 family
  • 128K context explicitly listed for AirX
  • 96K maximum output for the documented text series
  • Hybrid thinking for coding and structured text workflows

Before you choose

  • Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
  • This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
  • Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.

Extended thinking

enabled · defaultdisabled

The GLM-4.5 family uses hybrid thinking by default: with thinking.type enabled, the model decides whether reasoning is needed. Set disabled for the non-thinking path. The enabled setting is not a guarantee of a long reasoning trace on every request.

  • AirX has a distinct identifier and pricing row. Do not round or replace its $0.22 cached-input rate with Air's $0.03 rate.
  • The guide's speed positioning is not a guarantee under every input length or concurrent workload. Keep latency and quality measurements separate.
  • Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
  • For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Optimize an interactive text queue

Use measured latency and quality thresholds.

Given this interactive text workload and service targets, propose an evaluation of Air and AirX. Define representative prompts, first-response and completion timing, error categories and a human-review threshold. Include token and retry costs without inventing measurements or assuming a fixed speedup.

Workflow 02

Draft a concise operator handoff

Keep a quick response tied to evidence.

Turn these operations notes into a short handoff with current state, confirmed issues, next actions and owners explicitly named in the source. Include evidence references and mark missing owners or deadlines. Do not turn tentative ideas into assigned commitments.

Workflow 03

Validate structured summaries

Catch omissions before automated use.

Compare these generated summaries with their supplied source texts and target schema. Identify missing facts, unsupported values and schema violations separately. Return corrected fields only when the source establishes them, and flag ambiguous records for review instead of guessing.

Developer reference

Z.ai API pricing

These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$1.10
Cached input$0.22
Output$4.50
  • Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
  • Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
  • The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.

Common questions

A few things worth knowing.

How does AirX differ from Air?

It is the faster-serving lightweight tier described in the family guide, with a separate API identifier and higher listed prices.

Does it keep Air's context size?

The overview explicitly lists 128K for both. The API documents a 96K maximum output for the GLM-4.5 text series.

Can I use Air's price for an AirX estimate?

No. AirX has its own input, cached-input and output rows. Include the actual selected endpoint and retry usage in cost estimates.

Is faster serving always the better choice?

Not necessarily. Evaluate end-to-end task duration, correctness and total cost. A faster draft that requires more review may not improve the workflow.

Does the context window guarantee complete recall?

No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.

Are these prices the cost of my EZ Ai Assist plan?

No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.

Are all of these capabilities available in the app?

Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider