Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Lightweight high-speed vision tier with its own API ID
128K context and a 32K vision-series output ceiling
Images, video, files and text as documented inputs
Native function calling in the GLM-4.6V family
Before you choose
Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
Visual and file inputs support understanding with text output, not native image, video or audio generation. File format, size and ingestion requirements depend on the endpoint and app integration.
Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.
Extended thinking
enabled · defaultdisabled
Use thinking.type enabled or disabled. Thinking is enabled by default; the exact decision to reason depends on the model and request. Evaluate both modes using the same acceptance criteria. Do not assume GLM-5.3's low/high/max effort controls apply to this model.
The cached-input rate is $0.004 per million tokens. Keep that precision; $0.00 would incorrectly imply free cached input.
FlashX is the paid lightweight vision tier. GLM-4.6V-Flash, listed separately as free, is not the requested model and is not interchangeable.
Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Check a batch of receipt images
Keep uncertain extraction out of automated actions.
Extract merchant, date, currency and total from each supplied receipt image. Include the image ID and visible supporting text, and mark blurred or conflicting fields as unknown. Flag possible duplicates without deleting anything. Return a review queue; do not approve expenses or infer missing totals.
Workflow 02
Triage visual content quality
Use a repeatable checklist for many images.
Evaluate these product images against the supplied quality checklist. For each image list visible failures, uncertain checks and the next reviewer action. Reference the affected region and avoid inferring facts outside the frame. Do not alter, publish or remove any image.
Workflow 03
Verify a visual extraction pipeline
Separate model errors from integration errors.
Compare these source images, extracted fields and downstream validation results. Identify whether each mismatch comes from unreadable evidence, incorrect extraction or a schema mapping problem. Cite the source, propose focused test cases and require human review for unresolved records before any tool action.
Developer reference
Z.ai API pricing
These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.
Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.
Common questions
A few things worth knowing.
Is FlashX the free vision tier?
No. The official table lists paid rates for GLM-4.6V-FlashX. The similarly named GLM-4.6V-Flash is a different, free-priced entry.
Why does cached input show $0.004?
That is the published per-million-token rate. Rounding to two decimals would make a paid rate look like zero.
Does it retain native function calling?
The API includes the GLM-4.6V family in its native function-tool support and lists the FlashX ID. The integration must still validate and execute calls.
Are its limits the same as GLM-4.5V?
No. The overview lists 128K context for this model and the API gives the GLM-4.6V series a 32K output ceiling. GLM-4.5V has 64K context and 16K output.
Does the context window guarantee complete recall?
No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.
Are these prices the cost of my EZ Ai Assist plan?
No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.
Are all of these capabilities available in the app?
Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.