Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Hybrid reasoning that can be enabled or disabled
A 256,000-token context with a 32,000-token output ceiling
Text-based tool and agent workflows across 23 supported languages
Optional thinking token budgets and separate thinking/text response blocks
Before you choose
This is a text-only reasoning model, not Command A Vision or A+.
Thinking consumes a budget that must leave room for the final response; a longer reasoning trace is not evidence of a better answer.
The model card's free usage stops at rate limits and directs production deployments to Model Vault.
Context capacity is not a guarantee of complete recall. Keep source IDs, evaluate omissions and verify citations against the underlying documents.
A generated plan or tool call does not authorize or execute an external action. Use a configured integration, validate results and keep consequential changes behind human approval.
Extended thinking
enabled · defaultdisabled
Cohere's hybrid reasoning uses the thinking object and is enabled by default. Set type disabled to turn it off. An optional token_budget limits thinking; leave room for the final response, with the provider recommending at least 1K answer tokens. Do not mistake the guide's 31K example for a universal limit across all reasoning models.
Use command-a-reasoning-08-2025 for the model covered here. The website slug is a navigation label, not necessarily the API model ID.
The default is thinking enabled. To disable it use thinking.type disabled, and parse the final text separately from thinking blocks.
When setting token_budget, leave at least 1K tokens for the answer as the provider recommends. The guide's 31K example relates to a 32K output envelope, not an extra allowance beyond it.
Keep tool permissions narrow, store the results actually returned and distinguish proposed actions from completed actions.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Compare proposals with explicit uncertainty
Keep a complex decision grounded in supplied facts.
Compare these proposals against the evaluation criteria below. Build an evidence matrix with source references, unresolved assumptions and disqualifying gaps. Explain the tradeoffs in a concise decision summary, then list what must be verified before selection. Do not invent vendor capabilities or treat a missing answer as a positive result.
Workflow 02
Budget a multi-step analysis
Test thinking depth against an answer-quality rubric.
Create an evaluation plan for these analysis tasks with thinking enabled, disabled and constrained by a budget. Define the same output requirements for each run and a rubric for correctness, evidence coverage and concise conclusions. Include a table for measured usage and latency without filling it with invented results.
Workflow 03
Audit a research tool sequence
Check whether each step is justified by evidence.
Review this research workflow and the actual tool outputs provided. Identify steps that lack evidence, conclusions that overreach and redundant calls. Propose a shorter sequence with a clear stop condition and a human review point. Do not assume an external source was read unless its contents are included here.
Developer reference
Cohere API pricing
These are Cohere API reference terms, not EZ Ai Assist subscription prices.
The current model card offers trial and production keys free usage until rate limits are reached. This is not unlimited free production. Cohere directs production deployments to Model Vault. No public per-token production rate is established here.
Confirm limits for your exact model and key type; a production key alone is not evidence of unrestricted capacity.
Obtain deployment-specific commercial terms before budgeting. Do not substitute another Command model's token rates or show unverified prices as $0.
Common questions
A few things worth knowing.
Is thinking always required?
No. This is a hybrid model: thinking is enabled by default but can be disabled. Evaluate both settings rather than making every task use the deepest possible analysis.
What does token_budget control?
It caps thinking tokens. The model proceeds to its final response after the budget is exceeded. Reserve answer space; the provider's 31K thinking example leaves about 1K within the 32K output envelope.
Can it reason about uploaded images?
This ID is text-only. Do not inherit vision from Command A+ or Command A Vision. Extract authorized text first or select a model with documented image input.
Why is no per-token rate shown?
The current card describes free usage within rate limits and production through Model Vault. It does not establish a public per-token production price, so this page does not borrow the standard Command A rate.
Can it see current information automatically?
No. Model knowledge and a long context are not a live search service. Provide current evidence or connect an authorized retrieval workflow where supported, and verify the sources before acting on an answer.
How should I test it on my own work?
Use representative examples, explicit pass criteria and difficult counterexamples. Score factual support, omitted requirements and invalid outputs as well as latency and cost. Keep a human escalation path for uncertain or consequential results.
Are these API capabilities included in my EZ Ai Assist plan?
This guide describes the provider API, not subscription entitlements or a promise that every setting is exposed in the app. Check your workspace for model access and supported inputs, and use the EZ Ai Assist pricing page for plan details.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.