Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Native visual understanding alongside text input in DeepSeek's V4.1 Flash release
Thinking and non-thinking modes, with thinking enabled by default
Tool calls and JSON output through an appropriately configured API integration
OpenAI-compatible, Anthropic-compatible and Responses API request formats documented by DeepSeek
Before you choose
DeepSeek publishes rounded 1M context and 384K output labels; this guide preserves those labels instead of inventing exact token counts.
Vision input does not imply native image generation, audio generation or video output. The reviewed sources do not establish a knowledge cutoff.
In thinking mode, temperature, presence_penalty and frequency_penalty have no effect. top_p is clamped to 0.95–1.0; in non-thinking mode it is fixed at 1.0.
Tool results need validation and permission checks. A prompt cannot enable an API tool or grant access to a private document on its own.
Choose the reasoning effort
lowhigh · defaultmax
Thinking is on by default. For OpenAI-format requests, use thinking.type enabled/disabled and reasoning_effort. The effective levels are low, high and max: minimal maps to low; medium and xhigh map to high; ultra maps to max. Compare quality and latency on your own task before increasing effort.
With the OpenAI SDK, put the DeepSeek-specific thinking object in extra_body. Do not assume another provider's reasoning parameter names are interchangeable.
For requests carrying tools, DeepSeek requires returned reasoning_content from prior turns to be preserved in subsequent requests; missing content can cause a 400 error. Without tools, prior reasoning_content is ignored.
The September release notice says legacy V4 Flash IDs temporarily route to V4.1 Flash. Pin and verify your actual serving model rather than assuming an old ID still means old weights.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Review a screenshot against requirements
Turn visible evidence into a short, testable review.
Compare this screenshot with the requirements I provide. List visible mismatches, cite the relevant requirement, and describe a small test for each finding. Separate what the image proves from what would need an interactive check. Finish with the three most useful changes; do not assume hidden screens or behavior.
Workflow 02
Extract a document decision log
Keep a long document task tied to quoted evidence.
Read the project documents below and create a decision log with decision, owner if stated, supporting section, unresolved question and next verification step. Mark missing values as not stated. Do not turn proposals into approved decisions. End with any contradictions that a project owner should resolve.
Workflow 03
Evaluate a reasoning budget
Prepare a repeatable quality and latency comparison.
Using these representative tasks and acceptance criteria, design a small evaluation for low, high and max reasoning effort. Define a scoring rubric, failure categories and a results table for measured latency, token use and correctness. Do not invent results or prices. Recommend how we should choose the least expensive setting that passes.
Developer reference
DeepSeek API pricing
These are DeepSeek direct API reference prices, not EZ Ai Assist subscription prices.
Published direct API time-based rates · USD per 1,000,000 tokens
Token type
Off-peak
Peak
Input · cache miss
$0.15
$0.30
Input · cache hit
$0.003
$0.006
Output
$0.60
$1.20
Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours, including weekends and those holidays, are off-peak.
Off-peak rates are 50% of peak rates. A cache hit uses the cache-hit input rate; it is not a discount on output tokens.
These rates apply to deepseek-flash on DeepSeek's direct API. OpenRouter and other hosting services can publish different rates and billing rules. Recheck the provider before estimating a bill.
Common questions
A few things worth knowing.
Which API ID represents V4.1 Flash?
DeepSeek's current model table identifies deepseek-flash as DeepSeek V4.1 Flash. Older V4 Flash IDs can route to newer weights, so record the provider and resolved version when reproducibility matters.
Can it work with screenshots?
The release and current pricing table document vision input. Ask for observations grounded in the supplied image, and verify interaction claims separately. Image input is not image generation.
Can thinking be disabled?
Yes, the direct API supports enabled and disabled thinking. Thinking defaults to high effort; low, high and max are the distinct mapped levels. Availability of these controls in EZ Ai Assist must be checked in the app.
When do the lower prices apply?
Outside the published weekday UTC peak windows, including weekends and Chinese public holidays. Use UTC when scheduling and check the provider's current table rather than assuming your local evening is off-peak.
Does a long context guarantee complete recall?
No. Organize the input, identify the relevant sections and ask for evidence references. Evaluate retrieval and answer quality on your own documents instead of treating the maximum context size as a quality guarantee.
Why might a multi-turn tool request fail?
DeepSeek requires prior reasoning_content to be returned when requests carry tools, even for turns without a tool call. Follow its thinking-mode documentation and inspect the error before retrying.
Are these the prices of an EZ Ai Assist plan?
No. The table is a developer reference for DeepSeek's direct API. Your EZ Ai Assist subscription, available models and exposed controls are separate; use the pricing page and app for those details.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.