Back to models
Chat Models

Z.ai

GLM-4.6

Best for Coding

Text model with 200K context, hybrid thinking and explicitly enabled streaming tool-call arguments.

200K contextCoding

At a glance

Know the model before you prompt.

Z.ai API specifications
Context window
200K
Maximum output
128K
Inputs → output
Text → Text
Knowledge cutoff
Not verified

API model ID: glm-4.6

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • 200K context and 128K maximum output
  • Hybrid thinking for text-based reasoning and coding
  • Streaming tool-call arguments with explicit API settings
  • Function calling and structured task responses

Before you choose

  • Context and output limits use the provider's published K/M units. They are capacity ceilings, not a guarantee of complete recall or a final answer of that length. A model-specific knowledge cutoff was not verified.
  • This is a text-input guide. Supply extracted text for document analysis; do not assume the model can directly inspect images, videos or attached files just because other GLM models can.
  • Answers and proposed code still need validation. A tool call is a request for an integration to execute an action, not proof it ran. Keep approvals around consequential changes and check results against source evidence.

Extended thinking

enabled · defaultdisabled

GLM-4.6 uses hybrid thinking, enabled by default, with thinking.type enabled or disabled. The model can decide when to reason in the enabled path. This differs from GLM-4.7's enabled-thinking behavior and GLM-5.3's always-on reasoning.

  • Tool-argument streaming requires both stream true and tool_stream true. tool_stream defaults to false; assemble argument fragments before validation and execution.
  • The migration guide recommends temperature 1.0 and top_p 0.95. Change one sampling control at a time when testing rather than treating defaults as quality guarantees.
  • Select the exact API identifier shown here. Website slugs use hyphens for URLs and are not substitutes for dotted API model names. API access and the GLM Coding Plan endpoint have different entitlements.
  • For supported interleaved-thinking tool loops, retain the returned reasoning_content with the tool history. Preserved thinking uses thinking.clear_thinking false and unmodified history; its documented defaults differ between the standard API and Coding Plan endpoints. Check model support before enabling it.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Plan a version migration test

Compare behavior before changing production traffic.

Using this existing integration and representative requests, design a migration test plan for GLM-4.6. Cover message history, reasoning settings, structured output, tool calls, latency and usage. Define pass criteria and rollback triggers, and identify assumptions that require a live test without inventing its outcome.

Workflow 02

Audit a tool-stream parser

Check partial arguments before execution.

Review this streaming parser and sample event sequence. Identify cases where fragmented tool arguments, multiple calls or interrupted streams could produce invalid actions. Propose the smallest parser changes and regression cases. Require complete validation before execution and do not claim the proposed tests have run.

Workflow 03

Reconcile a long requirements set

Turn a large context into traceable checks.

Read the supplied requirements and issue notes. Group them by behavior, identify duplicates and contradictions, and write acceptance tests linked to source sections. Flag unresolved product decisions instead of making them. Keep the result focused on verifiable requirements rather than a broad redesign.

Developer reference

Z.ai API pricing

These are Z.ai direct API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.60
Cached input$0.11
Output$2.20
  • Input and output are billed separately per 1,000,000 tokens. Compare actual task usage, retries and latency rather than the input rate alone. These rates are not guaranteed reseller or app prices.
  • Only eligible cache hits receive the cached-input rate. Cached-input storage is listed as limited-time free, not permanently free. Recheck the provider table for storage terms and future changes.
  • The pricing page separately lists built-in Web Search at $0.01 per use. This is a service fee when that tool is used, not a charge on every prompt or a promise that this model or your workspace has automatic search.

Common questions

A few things worth knowing.

What enables streamed tool arguments?

Set both stream and tool_stream to true. The migration guide says tool_stream is disabled by default. Your client must combine and validate argument fragments.

How does its thinking differ from GLM-4.7?

GLM-4.6 is documented as hybrid: the enabled path can decide whether to think. GLM-4.7's enabled path uses thinking, while both support disabling it.

Is 128K its context window?

No. Its context is 200K and maximum output is 128K, according to the migration guide.

Should temperature and top_p both be changed at once?

The migration guide recommends adjusting one at a time. Keep other conditions stable so a test can identify what caused a behavior change.

Does the context window guarantee complete recall?

No. A large context is a capacity limit, not an accuracy guarantee. Label sources, split unrelated material, ask for evidence references and test whether important details were omitted. Output also has its own ceiling.

Are these prices the cost of my EZ Ai Assist plan?

No. This is a dated reference to direct Z.ai API token pricing. EZ Ai Assist subscriptions, Z.ai's GLM Coding Plan, optional tools and third-party hosting are separate products with their own terms.

Are all of these capabilities available in the app?

Not necessarily. The guide describes provider documentation, not workspace entitlements or an integration test. Check your model picker, accepted inputs and available controls. None of these examples runs tools or changes external systems by itself.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider