Back to models
Chat Models

Z.ai

GLM-4.7-FlashX

Best Value

Paid lightweight GLM-4.7 tier for fast text and coding workflows with 200K context and 128K output.

High throughputLow cost

At a glance

Know the model before you prompt.

Z.ai API specifications
Context window
200K
Maximum output
128K
Inputs → output
Text → Text
Knowledge cutoff
Not verified

API model ID: glm-4.7-flashx

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Lightweight high-speed tier in the GLM-4.7 family
  • 200K text context and 128K maximum output in its own tab
  • Thinking control for simple and more involved requests
  • Coding, translation and long-text workflows described by the family guide

Before you choose

  • Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
  • This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
  • Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.

Extended thinking

enabled · defaultdisabled

The model reasons whenever thinking is enabled, rather than dynamically skipping reasoning for a simple request. Use thinking.type disabled for the non-thinking path. Thinking is enabled by default; compare both modes on your own tasks and do not copy GLM-5.3 effort levels into this model.

  • FlashX is paid; do not copy the free GLM-4.7-Flash pricing row. Its input, cached-input and output rates are distinct.
  • The shared guide provides separate tabs for the flagship, FlashX and Flash. Use the exact FlashX ID in API requests.
  • Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
  • For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Triage small code changes

Keep a high-throughput review concrete.

Review these small diffs against their linked issue descriptions. For each, return confirmed defects, missing tests and unresolved questions with file references. Prioritize actionable findings and omit style preferences outside the requirements. Do not approve or merge any change.

Workflow 02

Normalize a document collection

Apply one schema without fabricating fields.

Extract the requested fields from each supplied text document into the schema below. Preserve document IDs and a supporting quotation or section reference for each value. Use null for absent fields and flag contradictions. Do not infer missing dates or entities from unrelated documents.

Workflow 03

Evaluate a low-cost processing route

Measure the whole successful task.

Design an evaluation for routing this text-processing queue to a lightweight model. Define labeled test cases, error severity, latency, usage and retry measures. Specify when to escalate an item to a stronger model or a human, and return an empty scoring sheet without invented results.

Developer reference

Z.ai API pricing

These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.07
Cached input$0.01
Output$0.40
  • Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
  • Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
  • The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.

Common questions

A few things worth knowing.

Is FlashX free like Flash?

No. FlashX has paid token rates, while the current table lists Flash tokens as free. Similar names do not imply identical pricing.

Are its limits borrowed from base GLM-4.7?

The model guide has a separate FlashX tab listing 200K context and 128K maximum output, corroborated by the model overview.

Does high-speed mean a fixed throughput guarantee?

No fixed throughput guarantee is established here. Test latency under your actual input sizes, concurrency and integration overhead.

Can it replace human review for code?

It can help triage and suggest tests, but a generated review is not proof of correctness. Run the relevant tests and inspect consequential changes.

Does the context window guarantee complete recall?

No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.

Are these prices the cost of my EZ Ai Assist plan?

No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.

Are all of these capabilities available in the app?

Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider