Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Text-based engineering work over a published 1M-token context
Always-on reasoning with three effort levels
Function calling for supported coding and agent integrations
Structured JSON output and streaming in the documented API
Before you choose
Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.
Choose the reasoning effort
lowhighmax · default
Thinking is always enabled and cannot be disabled. Set reasoning_effort to low, high or max; max is the default. A disabled thinking.type request is not supported. Lower effort changes the reasoning budget, not the model into a non-thinking endpoint.
Use reasoning_effort low, high or max. Migrating a disabled-thinking request requires changing thinking.type to enabled; lowering effort does not turn reasoning off.
The model is text-only. GLM-5.3-Flash and FlashX are the native multimodal choices in this family, not aliases for this ID.
Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Plan a safe repository migration
Turn a broad change into reviewable increments.
Use the repository map, interface definitions and migration requirements below. Identify affected callers with file references, propose the smallest ordered set of changes, and attach a regression test and rollback checkpoint to each step. Separate confirmed dependencies from assumptions. Do not claim to have inspected files or run commands that I have not supplied.
Workflow 02
Resolve conflicting technical evidence
Make an architecture decision traceable.
Compare the supplied design proposals, incident notes and measured constraints. Create a decision table with supporting document references, contradictions and unanswered questions. Recommend one next experiment that would distinguish the leading options. Keep measurements separate from estimates and do not invent evidence to complete the table.
Workflow 03
Review a difficult patch
Prioritize concrete failures over broad rewrites.
Review this patch against the acceptance criteria and surrounding code. Identify reproducible bugs, data-loss risks and missing failure handling, citing the affected function for each. Put uncertain concerns in a separate questions section. Suggest a minimal fix and tests without rewriting unrelated code or claiming those tests have run.
Developer reference
Z.ai API pricing
These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.
Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.
Common questions
A few things worth knowing.
Can GLM-5.3 turn thinking off?
No. Thinking is always enabled. Its effort settings are low, high and max, defaulting to max; disabled thinking is unsupported.
Can it inspect a screenshot directly?
The official GLM-5.3 guide specifies text-only input. Use a separately documented vision model or provide a textual description; do not inherit Flash's image support.
How large are its documented limits?
Z.ai publishes 1M context and 128K maximum output. This guide preserves those units instead of assuming a binary conversion or promising a 128K final answer.
Does a benchmark prove it will solve my coding task?
No. Provider evaluations help identify candidates, but repository-specific tests, review and measured rework determine whether a model meets your requirements.
Does the context window guarantee complete recall?
No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.
Are these prices the cost of my EZ Ai Assist plan?
No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.
Are all of these capabilities available in the app?
Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.