Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Lightweight GLM-4.5 text-family model
128K context and the documented 96K text-series output limit
Hybrid thinking and a non-thinking option
Coding, summarization and tool-oriented text workflows
Before you choose
Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.
Extended thinking
enabled · defaultdisabled
The GLM-4.5 family uses hybrid thinking by default: with thinking.type enabled, the model decides whether reasoning is needed. Set disabled for the non-thinking path. The enabled setting is not a guarantee of a long reasoning trace on every request.
The provider overview separately lists Air's 128K context, and the API documents the GLM-4.5 text-series output ceiling.
Air and AirX have different API IDs and rates. AirX is the faster-serving option, not a free speed setting on Air.
Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Classify a support queue
Use a clear schema and uncertainty label.
Classify each supplied support message using only the categories and definitions below. Return message ID, category, supporting phrase and a needs-review flag. Use unknown when no category fits; do not invent customer history or take account actions. Include a short list of ambiguous cases for a human.
Workflow 02
Summarize meeting decisions
Keep decisions separate from discussion.
Turn this meeting transcript into decisions, assigned actions and open questions. Cite the speaker and passage for each decision, preserve explicit deadlines, and mark unassigned work as unassigned. Do not convert suggestions into commitments or fill missing names from context.
Workflow 03
Draft a constrained customer reply
Produce a useful draft without unsupported promises.
Draft a concise reply to this customer using the supplied policy and case facts. Answer only what those sources establish, acknowledge missing information and suggest the next permitted step. Do not promise refunds, timelines or product capabilities unless explicitly supported. Return the draft plus a short fact-check list.
Developer reference
Z.ai API pricing
These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.
Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.
Common questions
A few things worth knowing.
Why choose Air instead of base GLM-4.5?
Air is positioned as the lightweight value model and has lower listed token rates. Whether it is suitable depends on accuracy, rework and latency in your own tasks.
Does Air keep 128K context?
Yes, the provider overview lists 128K for Air. The API documents a 96K maximum output for the GLM-4.5 text series.
Is AirX simply an automatic mode?
No. AirX is a separate API identifier and pricing row. Selecting Air does not automatically grant AirX serving.
Can low-cost classification be trusted without review?
No. Test against labeled examples, preserve an unknown category and review ambiguous or consequential outcomes before acting.
Does the context window guarantee complete recall?
No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.
Are these prices the cost of my EZ Ai Assist plan?
No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.
Are all of these capabilities available in the app?
Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.