Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
200K text context, correcting the older 128K catalog shorthand
Published maximum output of 128K tokens
Optional thinking for reasoning and coding tasks
Function calling, structured output and streaming through the API
Before you choose
Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.
Extended thinking
enabled · defaultdisabled
Use thinking.type enabled or disabled. Thinking is enabled by default; the exact decision to reason depends on the model and request. Evaluate both modes using the same acceptance criteria. Do not assume GLM-5.3's low/high/max effort controls apply to this model.
The official model overview and dedicated guide both list 200K context. The 128K figure is its maximum output, not its context window.
Use the glm-5 API ID. Turbo, vision Turbo and GLM-5.3 are distinct entries with separate documentation.
Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Review an API design
Find contract problems before implementation.
Review these API routes, payload examples and requirements. Identify incompatible fields, ambiguous error semantics and missing edge cases with references to the supplied definitions. Suggest the smallest contract changes and a test matrix. Separate blocking defects from optional improvements and do not assume undocumented client behavior.
Workflow 02
Synthesize a technical briefing
Ground recommendations in supplied material.
Using only the reports below, write a technical briefing with the decision, supporting evidence, risks and unresolved questions. Cite a report and section for every material claim. Where sources conflict, show the disagreement instead of silently choosing one. Keep recommendations separate from established facts.
Workflow 03
Design a regression suite
Translate failures into focused checks.
From this bug history and relevant code, propose a regression suite grouped by behavior. For each test explain the input, expected result and past failure it would catch. Include error paths and boundary values, and identify any fixtures still needed. Do not claim the tests pass before they have been implemented and run.
Developer reference
Z.ai API pricing
These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.
Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.
Common questions
A few things worth knowing.
Is GLM-5's context 128K or 200K?
The official guide lists 200K context. Its maximum output is 128K; those are separate limits.
Does base GLM-5 accept images?
Its dedicated guide specifies text input and output. Vision support belongs to separately documented models such as GLM-5V-Turbo or GLM-5.3-Flash.
Can thinking be turned off?
The thinking documentation allows disabled thinking for GLM-5. Do not apply GLM-5.3's forced-thinking restriction to it.
Is this the Turbo endpoint?
No. glm-5 and glm-5-turbo are distinct identifiers. Their positioning and current pricing verification differ.
Does the context window guarantee complete recall?
No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.
Are these prices the cost of my EZ Ai Assist plan?
No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.
Are all of these capabilities available in the app?
Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.