Back to models
NewChat Models

Google

Gemini 3.8 Flash

Google’s Gemini 3.8 Flash supports multimodal understanding, three thinking levels, and tool-assisted workflows with promotional API pricing.

FastMultimodalHigh throughput

At a glance

Know the model before you prompt.

Google API specifications
Input token limit
1,048,576 tokens
Maximum output
65,536 tokens
Inputs → output
Text + Images + Video + Audio + PDF → Text
Knowledge cutoff
March 2026

API model ID: gemini-3.8-flash

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Streaming responses and multimodal understanding
  • Function calling and structured outputs
  • Context caching and Batch API
  • Code execution, Google Search grounding, and URL context
  • File search and Google Maps grounding
  • Computer use (Preview), plus Flex and Priority processing

Before you choose

  • No native image or audio generation; these models return text.
  • Live API is not supported; audio input is not a real-time voice session.
  • Tool access, upload limits, and permissions depend on the integration. Validate outputs against the original evidence.
  • Google’s model card gives a March 2026 knowledge cutoff, but some domains remain at January 2025. Supply current evidence for time-sensitive questions.
  • This guide covers regular Gemini 3.8 Flash, not the separate Flash Cyber or Live models.

Gemini thinking levels

lowmedium · defaulthigh

Medium is the documented default. Start there, then compare low for straightforward tasks and high for more complex reasoning. The minimal setting returns an error for this model. Levels are relative thinking controls, not fixed token budgets or guarantees of correctness.

  • Use the exact ID gemini-3.8-flash. No shutdown date is listed in the reviewed Gemini Developer API lifecycle page.
  • For multi-turn tool work, preserve the required conversation state and thought signatures through the official SDK or documented API flow. Tool actions still require an integration and appropriate permissions.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Triage mixed-format requests

Combine supplied text and screenshots into a compact review queue.

Review the support requests and screenshots I provide. For each request, identify the observed issue, relevant evidence, likely owner, and one safe next step. Return a table with an uncertainty column. Do not infer missing account details or claim to have tested the product. Flag any item that needs a human decision.

Workflow 02

Check a focused change

Ask for concrete defects, not a broad rewrite.

Review this small code change against the behavior described below. Trace the affected inputs and error paths, identify concrete defects, and cite the relevant lines. Separate findings from assumptions. Propose minimal tests and explain what each would catch. Do not refactor unrelated code or execute external actions.

Workflow 03

Summarize a supplied clip

Keep visual and audio claims tied to the provided material.

Using the clip or transcript supplied here, summarize the sequence of events and extract the decisions and open questions. Include timestamps where the source supports them. Distinguish visible actions from spoken claims and uncertain interpretations. Do not invent missing speech, people, or events. End with a short verification checklist.

Developer reference

Google API pricing

These are Google Gemini Developer API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.75
Cached input$0.075
Output (including thinking)$3.75
  • The table shows promotional Standard rates through December 31, 2026. From January 1, 2027, the published rates are $1.50 input, $0.15 cached input, and $7.50 output per million tokens.
  • Context-cache storage is billed separately: $0.50 per million token-hours through December 31, 2026, then $1.00 from January 1, 2027.
  • Thinking tokens are included in billed output. Batch, Flex, and Priority have separate rates and availability; grounding and other tools can add charges.
  • Prices and quotas can change. Confirm the model, input modality, processing tier, and current Google rate before estimating API spend.

Common questions

A few things worth knowing.

When is 3.8 Flash worth comparing with an earlier Flash model?

Compare it when a repeatable workflow is sensitive to latency, cost, or errors. Run the same inputs with a shared scoring rubric and include difficult cases. The newer release name does not establish that it is the best choice for every task.

Is minimal thinking supported?

No. Google documents low, medium, and high for this model, with medium as the default. Sending minimal produces an error; do not reuse an older Flash configuration without checking it.

Can it generate voice or video?

This model understands supported audio and video inputs but returns text. Native media generation and the Live API are not supported by this endpoint. Separate Google models and products have different capabilities.

Will the API prices change after the promotion?

Google currently schedules the promotional rates to end December 31, 2026. From January 1, 2027, Standard input, cached input, and output are listed at $1.50, $0.15, and $7.50 per million tokens. Recheck the official table before budgeting.

Can I use these prompts in EZ Ai Assist?

Yes—adapt these original editorial examples to the inputs and controls available in your workspace. Share only material you are authorized to provide. They are starting points, not benchmarks or a promise of a particular result.

Does the input limit guarantee accurate long-document answers?

No. The 1,048,576-token input limit is capacity, not a recall guarantee. Organize sources with names and sections, request citations, and check them. The separate maximum output is 65,536 tokens; actual requests must also fit integration limits.

Are these the prices of my EZ Ai Assist plan?

No. The table describes direct Google API token usage. EZ Ai Assist subscriptions and app capabilities are separate. Use our pricing page for current plans and check the app for enabled models and controls.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider