Back to models
Chat Models

Google

Gemini 3.7 Flash

Google’s Gemini 3.7 Flash combines multimodal inputs, configurable thinking, and structured tool workflows for repeatable everyday tasks.

Chat

At a glance

Know the model before you prompt.

Google API specifications
Input token limit
1,048,576 tokens
Maximum output
65,536 tokens
Inputs → output
Text + Images + Video + Audio + PDF → Text
Knowledge cutoff
March 2026

API model ID: gemini-3.7-flash

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Streaming responses and multimodal understanding
  • Function calling and structured outputs
  • Context caching and Batch API
  • Code execution, Google Search grounding, and URL context
  • File search and Google Maps grounding
  • Computer use (Preview), plus Flex and Priority processing

Before you choose

  • No native image or audio generation; these models return text.
  • Live API is not supported; audio input is not a real-time voice session.
  • Tool access, upload limits, and permissions depend on the integration. Validate outputs against the original evidence.
  • The model card lists March 2026 as the knowledge cutoff while noting that some domains remain at January 2025. Provide up-to-date sources when needed.
  • Do not treat an announcement’s benchmark scores as expected performance on your own data.

Gemini thinking levels

lowmedium · defaulthigh

Medium is the documented default. Try low for narrow extraction and high for tasks with interdependent constraints, then measure quality and latency. The minimal setting returns an error. Thinking levels do not set a strict token budget; higher thinking can increase billed output.

  • The API ID is gemini-3.7-flash. No shutdown date is listed in the reviewed lifecycle table.
  • Use documented tool schemas and preserve conversation state across tool turns. Search grounding and code execution only operate when the integration enables them.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Build a document intake table

Extract consistent fields without filling evidence gaps.

Read these project intake documents and create one row per project with its goal, deadline, owner, dependencies, and unresolved requirements. Cite the document and section for each field. Use unknown when a value is absent. Flag conflicting dates and names for review, and do not silently combine separate projects.

Workflow 02

Analyze a dashboard screenshot

Separate observed values from explanations that need data.

Analyze the dashboard screenshot and supporting notes below. List the visible trends and anomalies, quoting the displayed labels and values. Separate observations from possible explanations. Identify the source data needed to verify each explanation. Finish with three focused follow-up questions; do not invent hidden metrics.

Workflow 03

Review a tool result

Turn tool output into a bounded next-step recommendation.

Given the task, allowed tools, and tool results below, determine whether the evidence is sufficient to answer the user. Cite the fields you used, flag failures or stale data, and propose at most one additional read-only lookup if necessary. Do not claim that a proposed lookup has run. Return a concise answer with explicit uncertainty.

Developer reference

Google API pricing

These are Google Gemini Developer API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.75
Cached input$0.075
Output (including thinking)$3.75
  • The table shows promotional Standard rates through December 31, 2026. From January 1, 2027, the published rates are $1.50 input, $0.15 cached input, and $7.50 output per million tokens.
  • Context-cache storage is billed separately: $0.50 per million token-hours through December 31, 2026, then $1.00 from January 1, 2027.
  • Thinking tokens are included in billed output. Batch, Flex, and Priority have separate rates and availability; grounding and other tools can add charges.
  • Prices and quotas can change. Confirm the model, input modality, processing tier, and current Google rate before estimating API spend.

Common questions

A few things worth knowing.

Should I move every 3.7 Flash workflow to 3.8 Flash?

Not automatically. Evaluate both against your quality, latency, and cost requirements, including regression cases. Keep model IDs explicit when reproducibility matters and review the lifecycle page independently of release announcements.

What is the default thinking level?

Medium. This model supports low, medium, and high; minimal is not accepted. If your integration hides the control, the API documentation alone does not tell you which setting that integration uses.

Does it know everything published before March 2026?

No. A cutoff is not a complete knowledge inventory. Google notes that some domains remain at January 2025, and the model can still be wrong about older material. Supply sources and ask it to distinguish evidence from inference.

Are caching and thinking included in the displayed rate?

The table separates cached input from ordinary input, and output includes thinking tokens. Stored-cache time is an additional charge. Processing tiers and tools can change total cost, so a single token rate is not a complete bill.

Can I use these prompts in EZ Ai Assist?

Yes—adapt these original editorial examples to the inputs and controls available in your workspace. Share only material you are authorized to provide. They are starting points, not benchmarks or a promise of a particular result.

Does the input limit guarantee accurate long-document answers?

No. The 1,048,576-token input limit is capacity, not a recall guarantee. Organize sources with names and sections, request citations, and check them. The separate maximum output is 65,536 tokens; actual requests must also fit integration limits.

Are these the prices of my EZ Ai Assist plan?

No. The table describes direct Google API token usage. EZ Ai Assist subscriptions and app capabilities are separate. Use our pricing page for current plans and check the app for enabled models and controls.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider