Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
128K text context in the dedicated model guide
16K maximum output, distinct from later GLM families
Instruction following and code-related text tasks
Function calling and structured output documented for integrations
Before you choose
This guide describes Z.ai-hosted API behavior. It does not establish license terms, self-hosted performance or identical limits for separately distributed weights.
Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.
Non-reasoning model
This model does not expose the later GLM-4.5-and-newer thinking switch in the reviewed API schema. It can perform reasoning tasks in text, but do not send later-family thinking or effort controls without documented support.
The full dated API ID is glm-4-32b-0414-128k. It is not an alias for a later GLM-4.5 or GLM-5 model.
The pricing table lists $0.10 input and $0.10 output per million tokens, with no published cached-input rate for this row.
Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Extract a bounded record set
Check an older model with explicit source fields.
Extract the fields in this schema from the supplied text records. Preserve record IDs, include source references and use null for missing values. Do not infer dates, names or amounts from neighboring records. Return a separate list of ambiguous cases that need review before database entry.
Workflow 02
Review a small utility function
Keep code assistance focused and testable.
Inspect this utility function and its stated contract. Identify concrete input cases that violate the contract, explain the relevant code path and propose the smallest fix. Provide boundary tests and expected outcomes, separating predicted behavior from results actually supplied in this conversation.
Workflow 03
Compare instruction-following candidates
Use one scorecard across model generations.
Build an evaluation from these representative writing and extraction requests. Define exact constraints, acceptable answers, error categories and a scoring method. Keep the test prompts identical across candidate models, and include latency and usage fields without filling in invented scores.
Developer reference
Z.ai API pricing
These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.
Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
No separate permanent storage entitlement is implied. A missing cache rate is not a zero-priced cache feature; confirm the current provider terms.
The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.
Common questions
A few things worth knowing.
What does the long model name identify?
It is the documented dated model identifier. Preserve the complete API name rather than shortening it to an ambiguous GLM-4 label.
Does it use the newer thinking switch?
The reviewed API restricts that switch to GLM-4.5 and newer families. This guide does not assign those controls to the older 32B model.
Is cached input free?
No free cache rate is established. The official pricing row shows no cached-input rate; missing data must not be converted to zero.
Do these API limits apply to self-hosted weights?
Not necessarily. This guide covers the hosted endpoint. Licensing, hardware requirements, runtime behavior and limits for self-hosting need separate verification.
Does the context window guarantee complete recall?
No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.
Are these prices the cost of my EZ Ai Assist plan?
No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.
Are all of these capabilities available in the app?
Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.