Back to models
NewChat Models

DeepSeek

DeepSeek V4 Flash

Best Overall

V4 Flash 0731 on OpenRouter, with long-context text workflows and hosting limits distinct from newer direct API aliases.

1M contextVersioned model

At a glance

Know the model before you prompt.

OpenRouter API specifications
Context window
1,048,576 tokens
Maximum output
Not verified
Inputs → output
Text → Text
Knowledge cutoff
Not verified

API model ID: deepseek/deepseek-v4-flash-0731

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Version-specific V4 Flash 0731 listing on OpenRouter, separate from the newer DeepSeek direct alias
  • A published 1,048,576-token context for large text and code inputs
  • Reasoning-oriented generation and tool-call support on compatible serving endpoints
  • Multiple hosting routes whose cost, availability and supported parameters can differ

Before you choose

  • OpenRouter's aggregate completion limit and individual provider limits may differ; no single verified per-endpoint maximum output is claimed here. Confirm the selected provider's completion limit.
  • The original direct-API 384K output description should not be treated as a universal guarantee for every routed endpoint.
  • This is not the V4 Flash Vision experimental variant or V4.1 Flash. Do not infer vision support or newer reasoning controls from their documentation.
  • No knowledge cutoff was established in the reviewed version-specific source. Published model or benchmark claims do not replace testing on your own inputs.

Reasoning behavior

The listing supports reasoning, but this guide does not assert one default or effort mapping across all hosting providers. Inspect the selected endpoint's accepted parameters; do not transplant V4.1 Flash's direct-API controls into this older routed version.

  • Use deepseek/deepseek-v4-flash-0731 for this OpenRouter listing; record the selected hosting provider when evaluating results.
  • Function-call output is a request for an integration to execute a tool, not proof that browsing or code execution actually happened.
  • For reproducible comparisons, keep the prompt, host, model version and evaluation criteria fixed. Check any fallback route before assuming it shares the same limits.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Review a large change in bounded passes

Make a repository review manageable and verifiable.

Review the code change and acceptance criteria I provide in two passes: first map affected interfaces and risks, then inspect only the highest-risk paths. Cite file names and concrete evidence for each finding. Separate confirmed bugs from questions, propose minimal fixes, and list tests without claiming to have run them.

Workflow 02

Compare versioned model outputs

Evaluate a pinned older model against a candidate replacement.

Compare these anonymized outputs from two model versions against the same rubric. Score factual support, instruction following, completeness and unnecessary changes. Quote the evidence behind each score, record ties, and identify cases requiring a human decision. Do not infer the provider or choose a winner based on the model name.

Workflow 03

Plan an alias migration

Check the difference between a model name and its serving route.

Using the supplied integration configuration and provider notices, create a migration checklist for a legacy model alias. Identify where IDs are stored, what must be pinned, which parameters require retesting and how to detect fallback changes. Include a small regression suite and rollback criteria. Mark anything not established by the evidence as unknown.

Developer reference

OpenRouter API pricing

These are named-provider OpenRouter reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Relace route reference on OpenRouter · USD per 1,000,000 tokens
Token typePrice
Input$0.0157
Output$1.28
  • These are the Relace route's displayed input/output rates checked October 6, 2026, not rates guaranteed across all OpenRouter providers.
  • The aggregate listing contains other hosts and promotional prices. Provider selection, availability, caching and routing can change the actual bill; check the selected endpoint.
  • These are not DeepSeek direct-API peak/off-peak rates. No cross-provider cache price or discount is assumed.

Common questions

A few things worth knowing.

Which V4 Flash version does this page describe?

The OpenRouter V4 Flash 0731 listing supplied for this guide. It is a versioned model reference, not a claim that the old deepseek-v4-flash direct-API alias still serves those weights.

Has V4 Flash been retired everywhere?

The retirement notice concerns DeepSeek's direct API, where old Flash IDs temporarily route to V4.1 Flash. Other hosts can serve different versions. Check the actual route rather than applying one provider's lifecycle notice globally.

Why is the output limit not verified?

A model-level aggregate limit is not enough to establish the maximum completion length of a selected provider endpoint. This guide keeps that distinction instead of presenting the old 384K figure as a universal hosted limit.

Can I assume it accepts images?

No. This text-focused 0731 guide is separate from the Vision experimental variant and V4.1 Flash. Use a verified vision-capable model and endpoint for screenshots or image analysis.

Does the listed price apply to every host?

No. The table names one serving route, Relace. OpenRouter lists other providers with different rates and capabilities; inspect your routing choice before comparing costs.

How should I compare it with V4.1 Flash?

Use the same representative tasks and score the results blind where practical. Record latency, billed usage, host and model version. Check failure cases and tool behavior, not only a single impressive answer.

Can these prompts enable tools in EZ Ai Assist?

No. Prompts describe tasks; tools require an available integration and appropriate permissions. Verify model access and tool controls in your app workspace before relying on them.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider