Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Video, image, text and file understanding
128K context with a 32K output ceiling in the API reference
Native function calling in the GLM-4.6V family
Thinking-mode switching and streaming responses
Before you choose
Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
Visual and file inputs support understanding with text output, not native image, video or audio generation. File format, size and ingestion requirements depend on the endpoint and app integration.
Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.
Extended thinking
enabled · defaultdisabled
Use thinking.type enabled or disabled. Thinking is enabled by default; the exact decision to reason depends on the model and request. Evaluate both modes using the same acceptance criteria. Do not assume GLM-5.3's low/high/max effort controls apply to this model.
The model page emphasizes context; the central vision API provides the 32K maximum output for the GLM-4.6V series.
A model-generated function call needs an executor. Validate extracted visual data before allowing it to drive consequential actions.
Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Extract chart evidence
Keep uncertain visual readings visible.
Read the supplied chart images and accompanying notes. Extract the labeled values with units and chart references, marking any unreadable or estimated value explicitly. Compare the trends without inventing missing data. End with the exact source values a person should confirm before using the analysis.
Workflow 02
Review a document packet
Connect conclusions across pages.
Analyze these supplied document pages for the stated review question. Build an evidence table containing each finding, page reference, conflicting statement and missing information. Keep quoted facts separate from interpretation and flag pages whose resolution prevents a reliable reading.
Workflow 03
Validate a visual tool handoff
Check perception before proposing an action.
Given these screenshots, extracted fields and proposed tool-call payloads, check whether each payload is supported by visible evidence. Identify mismatched amounts, names and identifiers. Return corrected proposals only where the source is clear, and mark all uncertain or consequential actions for human review without executing them.
Developer reference
Z.ai API pricing
These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.
Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.
Common questions
A few things worth knowing.
Where does the 32K output limit come from?
The central Chat Completion API documents 32K maximum output for the GLM-4.6V series. The model overview separately lists 128K context.
Does it support native function calling?
Yes, the family guide and API document it. The surrounding application must still execute and validate any proposed call.
Can it process an entire hour-long video reliably?
The provider gives long-video examples, but those are not guarantees for every encoding, request or task. Follow input limits and evaluate missed details on representative material.
Does it generate images?
No. This guide describes multimodal inputs with text output. Image generation is a separate model capability.
Does the context window guarantee complete recall?
No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.
Are these prices the cost of my EZ Ai Assist plan?
No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.
Are all of these capabilities available in the app?
Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.