Back to models
NewChat Models

Kimi

Kimi K2.7 Code

Best for Coding

Coding-focused Kimi with 256K context, visual inputs and always-on thinking for multi-step software work.

256K contextCoding

At a glance

Know the model before you prompt.

Kimi API specifications
Context window
262,144 tokens
Maximum output
Not verified
Inputs → output
Text + Images + Video → Text
Knowledge cutoff
Not verified

API model ID: kimi-k2.7-code

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Coding-focused work over a 262,144-token context
  • Always-on thinking for multi-step debugging and implementation planning
  • Text, image and video input for code and visual context
  • Tool-call workflows when supported by the surrounding integration

Before you choose

  • The model quickstart's 32,768 max_tokens value is a default, not a verified maximum output. A model-specific output ceiling and knowledge cutoff were not established; neither is borrowed from K3.
  • Vision and video understanding produce text; they do not imply native image, video or audio generation. Check what the app accepts before planning a multimodal workflow.
  • Public image URLs are not accepted by the documented API workflow. Use supported encoded content or uploaded file references, and follow the model-specific image and video guidance.
  • Tools require an integration to execute them. Validate arguments and results, and require approval before writes, purchases or other consequential actions. No model benchmark guarantees success on your workflow.

Extended thinking

enabled · default

K2.7 Code thinking is always on and cannot be disabled; a disabled setting returns an error. HighSpeed uses the same model behavior. Do not copy K3's low/high/max controls into these IDs unless the provider documents support.

  • The model quickstart lists 32,768 as the max_tokens default, not a maximum. The current API reference favors max_completion_tokens; follow its model-specific request schema instead of treating an older example as the output ceiling.
  • In thinking workflows, tool_choice supports auto or none. Preserve returned reasoning_content with assistant history as the provider recommends; omission can reduce performance. Do not assume K3's required tool choice, dynamic loading or strict schema features apply.
  • K2.7 Code fixes temperature at 1.0; K2.6 uses its mode-specific values. Both document top_p 0.95, n 1 and zero penalties. Follow the specific model rather than copying a generic sampling configuration.
  • Older quickstarts warn about the built-in web-search workflow. The later platform changelog introduces separate Search and Search Pro APIs; these are distinct services, not automatic search access inside a chat prompt or an EZ Ai Assist feature guarantee.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Debug from a reproducible failure

Follow the evidence before proposing a patch.

Using this failing test, stack trace and related code, identify the most likely cause. Cite the lines that support your conclusion and list any missing evidence. Propose the smallest patch and a regression test that would fail before the fix. Keep unrelated code unchanged and distinguish a predicted result from a test that has actually been run.

Workflow 02

Audit a tool-driven implementation

Check the boundary between a plan and executed work.

Review this coding-agent transcript, tool results and final diff against the original request. Find unsupported completion claims, missing validation and changes outside scope. Link each finding to the supplied evidence. End with a focused verification checklist and identify actions that require human approval before they are attempted.

Workflow 03

Translate a screen into acceptance tests

Turn visual requirements into concrete checks.

Compare this screenshot, component code and requested behavior. Write acceptance tests for layout, keyboard interaction, loading and failure states. Separate checks that can be automated from those requiring visual review. Preserve existing colors and spacing in your recommendations, and flag any behavior the screenshot alone cannot establish.

Developer reference

Kimi API pricing

These are Kimi API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input · cache miss$0.95
Input · cache hit$0.19
Output$4.00
  • Prices are per 1,000,000 tokens and exclude applicable taxes. Input and output are billed separately; inspect actual usage rather than estimating tokens from character counts.
  • Only eligible cached input receives the cache-hit rate. Uncached input and output keep their own rates. K3's separately published cache-write prices must not be assumed for K2 models.
  • These are direct Kimi platform rates, not universal reseller prices. Recheck the provider's current table and account limits before budgeting a deployment.

Common questions

A few things worth knowing.

Can thinking be turned off?

No. K2.7 Code's documented thinking is always enabled; a disabled setting returns an error. It does not share K2.6's optional-thinking behavior.

Is HighSpeed a different-quality model?

The provider describes HighSpeed as the same model with a faster serving option and different rates. Measure your actual latency and output quality; the name does not guarantee a fixed speed for every request.

Is 32,768 the maximum output?

The quickstart lists that value as a default. The reviewed sources do not establish a model-specific maximum, so this guide leaves the output ceiling unverified rather than converting a default into a limit.

What history should a tool loop preserve?

Keep the full assistant history, including returned reasoning_content, as the provider recommends. In thinking mode, the documented tool_choice values are auto and none; do not assume K3-only controls work here.

Does the context window guarantee complete recall?

No. Context is a capacity limit, not a retrieval-accuracy guarantee. Organize documents, label evidence and ask for references to the supplied material. Evaluate omissions and contradictions before trusting a long answer.

Can it create images or videos?

The documented model accepts text, image and video inputs and returns text. That supports visual analysis, not native image or video generation. Input support in your app workspace must be checked separately.

Are these API capabilities included in my EZ Ai Assist plan?

This guide describes the provider API, not subscription entitlements or a promise that every setting is exposed in the app. Check your workspace for model access and supported inputs, and use the EZ Ai Assist pricing page for plan details.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider