Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
200K text context and 128K maximum output
Long-horizon planning and iterative engineering workflows
Thinking enabled by default with a documented off switch
Function calling and streaming for supported integrations
Before you choose
Z.ai reports sustained execution for up to eight hours in its evaluations. That is a provider claim about an agent setup, not a guaranteed model response duration or EZ Ai Assist background feature.
Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.
Extended thinking
enabled · defaultdisabled
Use thinking.type enabled or disabled. Thinking is enabled by default; the exact decision to reason depends on the model and request. Evaluate both modes using the same acceptance criteria. Do not assume GLM-5.3's low/high/max effort controls apply to this model.
Use glm-5.1 rather than a newer model's ID when comparing results. Keep prompt, supplied evidence and evaluation conditions fixed.
The central API describes reasoning_effort for GLM-5.2 and newer. This guide does not assign those levels to GLM-5.1.
Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Plan an optimization experiment
Choose measurable improvements and stopping rules.
Review these performance traces, code excerpts and workload constraints. Identify the leading bottleneck hypotheses, then design one small experiment for each with success criteria and rollback steps. Cite the evidence, distinguish predicted improvements from measured ones, and define when to stop optimizing.
Workflow 02
Continue a stalled debugging task
Preserve useful history without repeating guesses.
Use this debugging log and current code to summarize confirmed facts, ruled-out hypotheses and remaining uncertainties. Propose the next three discriminating tests in order. For each, explain how different outcomes change the diagnosis. Do not repeat completed tests unless new evidence justifies it.
Workflow 03
Write a resumable task handoff
Make a long workflow safe to continue.
Convert this engineering transcript into a concise handoff: objective, constraints, files reviewed, changes actually made, test results and unresolved work. Cite the relevant evidence and separate planned actions from completed actions. End with the next safe step and any approvals needed before continuing.
Developer reference
Z.ai API pricing
These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.
Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.
Common questions
A few things worth knowing.
Is eight-hour autonomous work guaranteed?
No. The official guide describes provider evaluations of long-horizon tasks. Actual execution requires an agent runtime and suitable tools, and still needs supervision and verification.
What limits does GLM-5.1 document?
It lists 200K context and 128K maximum output. Do not inherit the 1M context of GLM-5.2 or GLM-5.3.
Does it use GLM-5.3 effort settings?
The central API documents reasoning_effort for GLM-5.2 and newer. This guide only presents the verified GLM-5.1 thinking switch.
Is it a direct vision model?
The model guide specifies text input and text output. Analyze extracted text or choose a separately documented multimodal model for images and video.
Does the context window guarantee complete recall?
No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.
Are these prices the cost of my EZ Ai Assist plan?
No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.
Are all of these capabilities available in the app?
Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.