Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
A 1,048,576-token context for large document sets and repository material
Native text, image and video understanding with text responses
Always-on thinking with low, high and max effort for different task budgets
Strict final-answer JSON Schema, required tool choice and dynamic tool loading in the documented K3 API
Before you choose
The API accepts max_completion_tokens up to 1,048,576, but input plus that budget must fit the context window. This is a request ceiling, not a promise of a million-token final answer. The default budget is 131,072.
Vision and video understanding produce text; they do not imply native image, video or audio generation. Check what the app accepts before planning a multimodal workflow.
Public image URLs are not accepted by the documented API workflow. Use supported encoded content or uploaded file references, and follow the model-specific image and video guidance.
Tools require an integration to execute them. Validate arguments and results, and require approval before writes, purchases or other consequential actions. No model benchmark guarantees success on your workflow.
Choose the reasoning effort
lowhighmax · default
K3 always uses thinking. Set the top-level reasoning_effort to low, high or max; max is the default. Thinking cannot be disabled. Evaluate latency, billed tokens and correctness before increasing effort, and do not substitute another provider's effort names.
Use kimi-k3 with max_completion_tokens; the current API deprecates max_tokens in favor of that field. Preserve complete assistant messages, including returned reasoning and tool history, in multi-turn workflows.
K3 documents strict JSON Schema output via response_format.json_schema and strict true. The schema constrains the final content, not reasoning text. It also supports required tool choice and dynamic tool loading.
K3 fixes temperature at 1.0, top_p at 0.95, n at 1 and penalties at zero. Omit unsupported sampling changes; a prompt does not override these API constraints.
Older quickstarts warn about the built-in web-search workflow. The later platform changelog introduces separate Search and Search Pro APIs; these are distinct services, not automatic search access inside a chat prompt or an EZ Ai Assist feature guarantee.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Plan a cross-repository change
Connect a long code context to a small, reviewable plan.
Use the repository map, files and issue description supplied here to plan this change. Identify the interfaces affected, cite the relevant files, and separate confirmed dependencies from assumptions. Propose an incremental implementation with a test and rollback checkpoint for each step. Do not claim you ran commands or inspected anything outside the supplied material.
Workflow 02
Reconcile conflicting project evidence
Make a decision brief traceable to its sources.
Compare the project documents below and produce a decision brief. For each recommendation, cite the document and section, show any conflicting evidence, and list the unresolved question. Keep approved decisions separate from proposals. Finish with the smallest set of facts a person must verify before acting; do not fill missing evidence with plausible details.
Workflow 03
Review a recorded product flow
Use visual evidence without guessing hidden behavior.
Review this supplied product walkthrough and the acceptance criteria. Identify visible usability problems, reference the relevant frame or timestamp where available, and distinguish observation from inference. Prioritize three small fixes that preserve the current design. Describe how to test each one interactively without claiming to have performed the test.
Developer reference
Kimi API pricing
These are Kimi API reference prices, not EZ Ai Assist subscription prices.
Prices are per 1,000,000 tokens and exclude applicable taxes. Input and output are billed separately; inspect actual usage rather than estimating tokens from character counts.
K3 cache writes are separate charges by TTL. Five minutes is the default; a cache hit uses only the cached-input rate, refreshes the lifetime and does not incur another write charge. Do not sum every row as though all tokens incur every rate.
These are direct Kimi platform rates, not universal reseller prices. Recheck the provider's current table and account limits before budgeting a deployment.
Common questions
A few things worth knowing.
How is K3 thinking controlled?
Thinking stays on. The top-level reasoning_effort accepts low, high and max, defaulting to max. Compare quality, response time and actual usage on representative tasks before selecting a setting.
Can K3 return a million-token answer?
The documented completion request ceiling is 1,048,576, with a 131,072 default, but input plus the requested completion budget must fit the context window. This is not a guaranteed final-answer length; reasoning and other request constraints matter.
Why are there two cache-write prices?
K3 publishes separate write charges for five-minute and one-hour cache lifetimes. Eligible hits use the cached-input rate and refresh the lifetime without another write fee. These are different billing events, not five mandatory charges on every token.
Does structured output constrain the entire response?
K3's strict JSON Schema feature constrains final content, not its reasoning text. Validate the parsed result and its factual claims in your application. A correct schema does not prove the values are correct.
Does the context window guarantee complete recall?
No. Context is a capacity limit, not a retrieval-accuracy guarantee. Organize documents, label evidence and ask for references to the supplied material. Evaluate omissions and contradictions before trusting a long answer.
Can it create images or videos?
The documented model accepts text, image and video inputs and returns text. That supports visual analysis, not native image or video generation. Input support in your app workspace must be checked separately.
Are these API capabilities included in my EZ Ai Assist plan?
This guide describes the provider API, not subscription entitlements or a promise that every setting is exposed in the app. Check your workspace for model access and supported inputs, and use the EZ Ai Assist pricing page for plan details.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.