Back to models
LegacyChat Models

Google

Gemini 3.6 Flash

Best Overall

Google’s earlier Gemini 3.6 Flash supports multimodal workflows and four thinking levels; the recommended replacement for Gemini 3 Flash Preview.

1M contextAgentic

At a glance

Know the model before you prompt.

Google API specifications
Input token limit
1,048,576 tokens
Maximum output
65,536 tokens
Inputs → output
Text + Images + Video + Audio + PDF → Text
Knowledge cutoff
March 2026

API model ID: gemini-3.6-flash

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Streaming responses and multimodal understanding
  • Function calling and structured outputs
  • Context caching and Batch API
  • Code execution, Google Search grounding, and URL context
  • File search and Google Maps grounding
  • Computer use (Preview), plus Flex and Priority processing

Before you choose

  • No native image or audio generation; these models return text.
  • Live API is not supported; audio input is not a real-time voice session.
  • Tool access, upload limits, and permissions depend on the integration. Validate outputs against the original evidence.
  • Google’s model card lists a March 2026 cutoff, with some domains still at January 2025. Use current evidence for changing facts.
  • This guide does not describe Gemini 3.5 Flash-Lite or Flash Cyber, which appeared in the same announcement but are separate models.

Gemini thinking levels

minimallowmedium · defaulthigh

Medium is the documented default. Minimal is supported here but does not guarantee that thinking is completely off. Compare lower levels for bounded tasks and higher levels for dependent reasoning steps. Unlike 3.6 Flash, the newer 3.7 and 3.8 Flash models reject minimal, so review configurations when migrating.

  • The exact model ID is gemini-3.6-flash. The lifecycle page lists no announced shutdown date for this model.
  • Google recommends this model as the replacement for gemini-3-flash-preview, which likewise has no shutdown date listed in the reviewed table. A replacement recommendation is not itself a retirement notice.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Check a Flash migration

Find configuration differences before changing production traffic.

Compare these existing Gemini 3 Flash prompts and API settings with the documented Gemini 3.6 Flash configuration below. Identify assumptions about thinking defaults, tools, schemas, token limits, and billing. Propose a small side-by-side evaluation with explicit pass criteria. Do not change production settings or assume matching outputs.

Workflow 02

Extract a visual procedure

Turn provided frames into a checkable sequence.

Use the supplied screenshots or video frames to describe the visible procedure step by step. Cite frame labels or timestamps and list prerequisites only when the source supports them. Mark hidden actions as unknown. Finish with a concise checklist and the questions a reviewer should resolve before following the procedure.

Workflow 03

Summarize a code investigation

Keep conclusions linked to test evidence.

Read the bug report, patch, and test output below. Summarize the intended fix, the evidence that supports it, and any untested paths. Identify one focused regression test for each remaining risk. Distinguish a passing test from proof that all cases work. Do not invent execution results or propose unrelated code changes.

Developer reference

Google API pricing

These are Google Gemini Developer API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.75
Cached input$0.075
Output (including thinking)$3.75
  • The table shows promotional Standard rates through December 31, 2026. From January 1, 2027, the published rates are $1.50 input, $0.15 cached input, and $7.50 output per million tokens.
  • Context-cache storage is billed separately: $0.50 per million token-hours through December 31, 2026, then $1.00 from January 1, 2027.
  • Thinking tokens are included in billed output. Batch, Flex, and Priority have separate rates and availability; grounding and other tools can add charges.
  • Prices and quotas can change. Confirm the model, input modality, processing tier, and current Google rate before estimating API spend.

Common questions

A few things worth knowing.

Why does this card say Legacy?

It distinguishes an earlier Flash generation from newer catalog models. It does not mean the API has already shut down. Google’s reviewed lifecycle page lists no announced shutdown date for Gemini 3.6 Flash.

Is 3.6 Flash the recommended replacement for 3 Flash Preview?

Yes, that is the replacement listed in Google’s lifecycle table. Test prompts, tool schemas, thinking defaults, and cost before migration; a recommendation does not guarantee identical behavior.

Does minimal turn thinking off?

Not necessarily. Google documents minimal as a supported level, but it is not a guarantee of zero thinking. Medium is the default for this model, and the model can adapt its work to the request.

Does the announcement include other models?

Yes. Flash-Lite and Flash Cyber appear alongside 3.6 Flash in the announcement. Their capabilities, access rules, and pricing must not be copied into this model’s guide.

Can I use these prompts in EZ Ai Assist?

Yes—adapt these original editorial examples to the inputs and controls available in your workspace. Share only material you are authorized to provide. They are starting points, not benchmarks or a promise of a particular result.

Does the input limit guarantee accurate long-document answers?

No. The 1,048,576-token input limit is capacity, not a recall guarantee. Organize sources with names and sections, request citations, and check them. The separate maximum output is 65,536 tokens; actual requests must also fit integration limits.

Are these the prices of my EZ Ai Assist plan?

No. The table describes direct Google API token usage. EZ Ai Assist subscriptions and app capabilities are separate. Use our pricing page for current plans and check the app for enabled models and controls.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider