Back to models
PreviewChat Models

Google

Gemini 3 Flash

Best Overall

Google’s Gemini 3 Flash Preview supports multimodal inputs and tool workflows, with separate audio rates and a Gemini 3.6 Flash migration path.

FlashValue

At a glance

Know the model before you prompt.

Google API specifications
Input token limit
1,048,576 tokens
Maximum output
65,536 tokens
Inputs → output
Text + Images + Video + Audio + PDF → Text
Knowledge cutoff
January 2025

API model ID: gemini-3-flash-preview

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Streaming responses and multimodal understanding
  • Function calling and structured outputs
  • Context caching and Batch API
  • Code execution, Google Search grounding, and URL context
  • File search and Google Maps grounding
  • Computer use, plus Flex and Priority processing

Before you choose

  • No native image or audio generation; these models return text.
  • Live API is not supported; audio input is not a real-time voice session.
  • Tool access, upload limits, and permissions depend on the integration. Validate outputs against the original evidence.
  • The API endpoint is preview. Google recommends Gemini 3.6 Flash as a replacement, but the reviewed lifecycle table gives no shutdown date.
  • Audio input and cached audio have different token rates from text, images, and video. Estimate costs by input modality.

Gemini thinking levels

minimallowmediumhigh · default

High is the documented default for Gemini 3 Flash Preview. Minimal, low, medium, and high are supported; minimal does not guarantee zero thinking. Compare explicit settings when evaluating a replacement because later Flash defaults differ.

  • The exact API ID is gemini-3-flash-preview, not gemini-3-flash.
  • Google’s lifecycle table lists no shutdown date and recommends Gemini 3.6 Flash as the replacement. Recheck the table before scheduling a migration.
  • Google’s Gemini 3 developer guide lists a January 2025 cutoff. Grounding or supplied sources are needed for reliable discussion of newer facts.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Build a timestamped recap

Keep a recording summary anchored to what was actually supplied.

Summarize this supplied recording or transcript for a teammate who missed it. Identify decisions, assigned actions, and unresolved questions, with timestamps or line references when available. Distinguish explicit commitments from suggestions. Mark unclear speech as uncertain and do not infer names, deadlines, or facts that are not in the source.

Workflow 02

Check extracted records

Find mistakes before structured data enters another system.

Compare these extracted records with the source pages and schema below. Identify missing required fields, unsupported values, and formatting mismatches. Cite the source location for each correction and leave unknown values empty according to the schema. Return a validation summary without inventing records or submitting changes.

Workflow 03

Prepare a replacement trial

Test defaults and costs as well as answer quality.

Design a small comparison of this existing Gemini 3 Flash workflow and Gemini 3.6 Flash. Use the supplied examples to define correctness checks, failure cases, and latency measurements. Include thinking-default differences and audio-input costs where relevant. State what must be measured rather than inventing scores or declaring a winner.

Developer reference

Google API pricing

These are Google Gemini Developer API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input: text, image, video$0.50
Input: audio$1.00
Cached input: text, image, video$0.05
Cached input: audio$0.10
Output (including thinking)$3.00
  • Audio input is $1.00 per million tokens, with cached audio at $0.10; do not apply the text/image/video input rate to audio.
  • Context-cache storage is billed separately at $1.00 per million token-hours. Thinking tokens are included in output charges.
  • Batch, Flex, and Priority have separate rates and availability. Grounding and other tools can add costs. Preview pricing and limits may change.

Common questions

A few things worth knowing.

Is Gemini 3 Flash already retired?

The reviewed lifecycle table does not list a shutdown date for gemini-3-flash-preview. It recommends Gemini 3.6 Flash as the replacement. A recommended successor is not the same as an announced shutdown.

Which thinking default should I expect?

Google documents high as the default for this preview model, with minimal, low, medium, and high available. Some later Flash models default to medium. Specify and compare settings when your integration allows it.

Why does the table have separate audio rows?

Google prices audio input differently from text, image, and video input. Cached audio also has its own rate. Native output remains text, so accepting audio does not mean this endpoint generates speech.

How should I evaluate a migration to 3.6 Flash?

Use the same representative inputs, compare supported controls and tools, and measure quality, latency, and total token cost. Include audio cases if your workflow uses them. Review preview and lifecycle notices before changing an integration.

Can I use these prompts in EZ Ai Assist?

Yes—adapt these original editorial examples to the inputs and controls available in your workspace. Share only material you are authorized to provide. They are starting points, not benchmarks or a promise of a particular result.

Does the input limit guarantee accurate long-document answers?

No. The 1,048,576-token input limit is capacity, not a recall guarantee. Organize sources with names and sections, request citations, and check them. The separate maximum output is 65,536 tokens; actual requests must also fit integration limits.

Are these the prices of my EZ Ai Assist plan?

No. The table describes direct Google API token usage. EZ Ai Assist subscriptions and app capabilities are separate. Use our pricing page for current plans and check the app for enabled models and controls.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider