Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Coding-focused work over a 262,144-token context
Always-on thinking for multi-step debugging and implementation planning
Text, image and video input for code and visual context
Tool-call workflows when supported by the surrounding integration
Before you choose
The model quickstart's 32,768 max_tokens value is a default, not a verified maximum output. A model-specific output ceiling and knowledge cutoff were not established; neither is borrowed from K3.
Vision and video understanding produce text; they do not imply native image, video or audio generation. Check what the app accepts before planning a multimodal workflow.
Public image URLs are not accepted by the documented API workflow. Use supported encoded content or uploaded file references, and follow the model-specific image and video guidance.
Tools require an integration to execute them. Validate arguments and results, and require approval before writes, purchases or other consequential actions. No model benchmark guarantees success on your workflow.
Extended thinking
enabled · default
K2.7 Code thinking is always on and cannot be disabled; a disabled setting returns an error. HighSpeed uses the same model behavior. Do not copy K3's low/high/max controls into these IDs unless the provider documents support.
The model quickstart lists 32,768 as the max_tokens default, not a maximum. The current API reference favors max_completion_tokens; follow its model-specific request schema instead of treating an older example as the output ceiling.
In thinking workflows, tool_choice supports auto or none. Preserve returned reasoning_content with assistant history as the provider recommends; omission can reduce performance. Do not assume K3's required tool choice, dynamic loading or strict schema features apply.
K2.7 Code fixes temperature at 1.0; K2.6 uses its mode-specific values. Both document top_p 0.95, n 1 and zero penalties. Follow the specific model rather than copying a generic sampling configuration.
Older quickstarts warn about the built-in web-search workflow. The later platform changelog introduces separate Search and Search Pro APIs; these are distinct services, not automatic search access inside a chat prompt or an EZ Ai Assist feature guarantee.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Debug from a reproducible failure
Follow the evidence before proposing a patch.
Using this failing test, stack trace and related code, identify the most likely cause. Cite the lines that support your conclusion and list any missing evidence. Propose the smallest patch and a regression test that would fail before the fix. Keep unrelated code unchanged and distinguish a predicted result from a test that has actually been run.
Workflow 02
Audit a tool-driven implementation
Check the boundary between a plan and executed work.
Review this coding-agent transcript, tool results and final diff against the original request. Find unsupported completion claims, missing validation and changes outside scope. Link each finding to the supplied evidence. End with a focused verification checklist and identify actions that require human approval before they are attempted.
Workflow 03
Translate a screen into acceptance tests
Turn visual requirements into concrete checks.
Compare this screenshot, component code and requested behavior. Write acceptance tests for layout, keyboard interaction, loading and failure states. Separate checks that can be automated from those requiring visual review. Preserve existing colors and spacing in your recommendations, and flag any behavior the screenshot alone cannot establish.
Developer reference
Kimi API pricing
These are Kimi API reference prices, not EZ Ai Assist subscription prices.
Prices are per 1,000,000 tokens and exclude applicable taxes. Input and output are billed separately; inspect actual usage rather than estimating tokens from character counts.
Only eligible cached input receives the cache-hit rate. Uncached input and output keep their own rates. K3's separately published cache-write prices must not be assumed for K2 models.
These are direct Kimi platform rates, not universal reseller prices. Recheck the provider's current table and account limits before budgeting a deployment.
Common questions
A few things worth knowing.
Can thinking be turned off?
No. K2.7 Code's documented thinking is always enabled; a disabled setting returns an error. It does not share K2.6's optional-thinking behavior.
Is HighSpeed a different-quality model?
The provider describes HighSpeed as the same model with a faster serving option and different rates. Measure your actual latency and output quality; the name does not guarantee a fixed speed for every request.
Is 32,768 the maximum output?
The quickstart lists that value as a default. The reviewed sources do not establish a model-specific maximum, so this guide leaves the output ceiling unverified rather than converting a default into a limit.
What history should a tool loop preserve?
Keep the full assistant history, including returned reasoning_content, as the provider recommends. In thinking mode, the documented tool_choice values are auto and none; do not assume K3-only controls work here.
Does the context window guarantee complete recall?
No. Context is a capacity limit, not a retrieval-accuracy guarantee. Organize documents, label evidence and ask for references to the supplied material. Evaluate omissions and contradictions before trusting a long answer.
Can it create images or videos?
The documented model accepts text, image and video inputs and returns text. That supports visual analysis, not native image or video generation. Input support in your app workspace must be checked separately.
Are these API capabilities included in my EZ Ai Assist plan?
This guide describes the provider API, not subscription entitlements or a promise that every setting is exposed in the app. Check your workspace for model access and supported inputs, and use the EZ Ai Assist pricing page for plan details.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.