Back to models
Chat Models

Kimi

Kimi K2.7 Code HighSpeed

Best for Coding

Faster-serving K2.7 Code endpoint with the same model behavior, 256K context and separately priced API usage.

High speedCoding

At a glance

Know the model before you prompt.

Kimi API specifications
Context window
262,144 tokens
Maximum output
Not verified
Inputs → output
Text + Images + Video → Text
Knowledge cutoff
Not verified

API model ID: kimi-k2.7-code-highspeed

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • The same K2.7 Code model served through a separately priced HighSpeed endpoint
  • A 262,144-token context and text, image and video inputs
  • Always-on thinking and integrated tool-call workflows
  • Provider-reported throughput around 180 tokens/second, up to 260 for short contexts; actual serving conditions vary

Before you choose

  • The model quickstart's 32,768 max_tokens value is a default, not a verified maximum output. A model-specific output ceiling and knowledge cutoff were not established; neither is borrowed from K3.
  • Vision and video understanding produce text; they do not imply native image, video or audio generation. Check what the app accepts before planning a multimodal workflow.
  • Public image URLs are not accepted by the documented API workflow. Use supported encoded content or uploaded file references, and follow the model-specific image and video guidance.
  • Tools require an integration to execute them. Validate arguments and results, and require approval before writes, purchases or other consequential actions. No model benchmark guarantees success on your workflow.

Extended thinking

enabled · default

K2.7 Code thinking is always on and cannot be disabled; a disabled setting returns an error. HighSpeed uses the same model behavior. Do not copy K3's low/high/max controls into these IDs unless the provider documents support.

  • The model quickstart lists 32,768 as the max_tokens default, not a maximum. The current API reference favors max_completion_tokens; follow its model-specific request schema instead of treating an older example as the output ceiling.
  • In thinking workflows, tool_choice supports auto or none. Preserve returned reasoning_content with assistant history as the provider recommends; omission can reduce performance. Do not assume K3's required tool choice, dynamic loading or strict schema features apply.
  • K2.7 Code fixes temperature at 1.0; K2.6 uses its mode-specific values. Both document top_p 0.95, n 1 and zero penalties. Follow the specific model rather than copying a generic sampling configuration.
  • Older quickstarts warn about the built-in web-search workflow. The later platform changelog introduces separate Search and Search Pro APIs; these are distinct services, not automatic search access inside a chat prompt or an EZ Ai Assist feature guarantee.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Compare two serving options fairly

Measure speed without losing correctness checks.

Design a comparison of standard and HighSpeed serving for this set of coding tasks. Keep prompts, context and acceptance criteria identical. Define how to record time to first response, total latency, usage, correctness and rework. Provide a results template without invented measurements and explain how to decide whether the faster option earns its extra cost.

Workflow 02

Triage a pull-request queue

Keep fast reviews focused on actionable defects.

Triage these supplied pull-request diffs against their acceptance criteria. For each, list only concrete bugs, risk level, supporting code and the next test to run. Put uncertain findings in a separate questions section. Do not approve, merge or change anything; return a prioritized review queue that a developer can verify.

Workflow 03

Shorten a debugging feedback loop

Choose the next discriminating test instead of guessing.

Read this bug report and the tests already attempted. Propose the next three diagnostic steps, ordered by how well they distinguish the remaining hypotheses. For each step, give the expected observations and what each would mean. Keep commands scoped to the issue and flag destructive actions. Do not claim a hypothesis is confirmed without supplied results.

Developer reference

Kimi API pricing

These are Kimi API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input · cache miss$1.90
Input · cache hit$0.38
Output$8.00
  • Prices are per 1,000,000 tokens and exclude applicable taxes. Input and output are billed separately; inspect actual usage rather than estimating tokens from character counts.
  • Only eligible cached input receives the cache-hit rate. Uncached input and output keep their own rates. K3's separately published cache-write prices must not be assumed for K2 models.
  • These are direct Kimi platform rates, not universal reseller prices. Recheck the provider's current table and account limits before budgeting a deployment.

Common questions

A few things worth knowing.

Does HighSpeed change the model weights?

Kimi describes it as the same K2.7 Code model with faster serving. Treat differences in latency, concurrency and cost as endpoint properties to evaluate, not proof of a different quality tier.

Is 180 or 260 tokens per second guaranteed?

No. Those are provider-reported throughput figures, with the higher figure associated with short contexts. Request size, load and integration overhead can change the experienced speed, including inside EZ Ai Assist.

How does the price differ from standard Code?

The current direct API table lists twice the standard Code rate for uncached input, cached input and output. Compare the total cost of successful tasks, including retries, rather than latency alone.

Do the thinking and output controls change?

Thinking remains always on. The quickstart's 32,768 completion setting is a default, not an established maximum. Use the documented schema for this endpoint rather than copying K3 effort levels.

Does the context window guarantee complete recall?

No. Context is a capacity limit, not a retrieval-accuracy guarantee. Organize documents, label evidence and ask for references to the supplied material. Evaluate omissions and contradictions before trusting a long answer.

Can it create images or videos?

The documented model accepts text, image and video inputs and returns text. That supports visual analysis, not native image or video generation. Input support in your app workspace must be checked separately.

Are these API capabilities included in my EZ Ai Assist plan?

This guide describes the provider API, not subscription entitlements or a promise that every setting is exposed in the app. Check your workspace for model access and supported inputs, and use the EZ Ai Assist pricing page for plan details.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider