Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
128K text context and 96K maximum output
Hybrid thinking enabled by default
Coding, instruction following and multi-step text analysis
Function calling for integrated agent workflows
Before you choose
Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.
Extended thinking
enabled · defaultdisabled
The GLM-4.5 family uses hybrid thinking by default: with thinking.type enabled, the model decides whether reasoning is needed. Set disabled for the non-thinking path. The enabled setting is not a guarantee of a long reasoning trace on every request.
The API documents 96K maximum output for the GLM-4.5 text series. GLM-4.5V is a separate vision model with different limits.
Use thinking.type enabled or disabled. Hybrid thinking may choose not to produce a reasoning trace for a simple request.
Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Review a business rule implementation
Test behavior against explicit rules.
Compare this implementation with the supplied business rules. Identify concrete mismatches and missing edge cases, citing the relevant rule and function. Suggest minimal changes and a table of tests including invalid input. Keep unresolved policy questions separate from code defects.
Workflow 02
Create a grounded operating procedure
Make each step traceable to approved material.
Using only the policy and process documents below, draft an operating procedure with prerequisites, steps, exceptions and escalation points. Cite the source section for each requirement. Flag conflicting instructions and missing approvals rather than silently resolving them.
Workflow 03
Check a function-call response
Separate valid structure from correct intent.
Review these user requests, tool schemas and proposed function calls. Check schema validity, intent alignment and permission boundaries separately. For each issue explain the exact field and supporting evidence, then propose a corrected payload only when the intent is unambiguous. Do not execute any call.
Developer reference
Z.ai API pricing
These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.
Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.
Common questions
A few things worth knowing.
What does hybrid thinking mean?
With thinking enabled, the model can decide whether reasoning is needed. Disable thinking for the non-thinking path; neither setting guarantees factual correctness.
Does the text model accept photos?
No. GLM-4.5 is documented as text input. GLM-4.5V has its own visual-input guide and smaller limits.
Is Air the same model identifier?
No. GLM-4.5-Air is the lightweight tier, with its own API ID and lower listed rates. Compare task quality rather than treating the names as aliases.
Does the output budget replace the context limit?
No. The 96K output ceiling and 128K context describe different constraints. Keep requests within the endpoint's total budget and inspect actual usage.
Does the context window guarantee complete recall?
No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.
Are these prices the cost of my EZ Ai Assist plan?
No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.
Are all of these capabilities available in the app?
Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.