Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Free token rates in the current direct API table
200K context and 128K maximum output in the Flash tab
Lightweight text, coding and translation workflows
Thinking enabled by default with an explicit off switch
Before you choose
Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.
Extended thinking
enabled · defaultdisabled
The model reasons whenever thinking is enabled, rather than dynamically skipping reasoning for a simple request. Use thinking.type disabled for the non-thinking path. Thinking is enabled by default; compare both modes on your own tasks and do not copy GLM-5.3 effort levels into this model.
A zero token price is a published commercial term for this API model, not unlimited capacity or free EZ Ai Assist access.
Keep the exact glm-4.7-flash ID distinct from the paid flashx ID. Recheck quotas and optional service costs before scaling.
Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Build a lightweight evaluation set
Test a free API tier before depending on it.
Create a small evaluation set from these representative text tasks. Include easy, ambiguous and failure cases, expected answer criteria and a human-review flag. Define how to record correctness, latency and rate-limit errors. Do not invent model scores or assume unlimited service availability.
Workflow 02
Rewrite with strict constraints
Check that a simple edit preserves facts.
Rewrite the supplied product description in plain language within the requested length. Preserve every factual claim and remove unsupported superlatives. Return the rewritten text and a short checklist showing how it satisfies the constraints; flag any source ambiguity rather than adding missing facts.
Workflow 03
Compare translations against a glossary
Keep meaning and terminology reviewable.
Compare these source passages and draft translations using the supplied glossary. Identify omissions, changed meaning and inconsistent terms with passage references. Suggest focused corrections and flag phrases needing a native-language review. Do not add context that is absent from the source.
Developer reference
Z.ai API pricing
These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.
The provider currently lists free input, cached input and output for this API model. Free token rates do not remove rate limits, quotas, account requirements or separate tool charges, and do not make EZ Ai Assist subscriptions free.
No separate permanent storage entitlement is implied. A missing cache rate is not a zero-priced cache feature; confirm the current provider terms.
The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.
Common questions
A few things worth knowing.
What exactly is free?
The current direct API table lists input, cached input and output tokens as free for GLM-4.7-Flash. Account requirements, rate limits and separately used tools still matter.
Does free API pricing make the app free?
No. EZ Ai Assist subscriptions and provider API billing are separate. Check your workspace and plan for model access.
Is it the same endpoint as FlashX?
No. FlashX has its own identifier and paid rates. Verify the selected ID before comparing cost or behavior.
What should a production evaluation include?
Measure task quality, retries, latency, quotas and failure handling. A zero token price does not establish reliability or suitability for consequential work.
Does the context window guarantee complete recall?
No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.
Are these prices the cost of my EZ Ai Assist plan?
No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.
Are all of these capabilities available in the app?
Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.