Back to models
NewChat Models

Google

Gemini 3.1 Flash-Lite

Best Value

Google’s Gemini 3.1 Flash-Lite supports translation, extraction, and lightweight multimodal workflows, with minimal thinking by default.

ValueHigh volume

At a glance

Know the model before you prompt.

Google API specifications
Input token limit
1,048,576 tokens
Maximum output
65,536 tokens
Inputs → output
Text + Images + Video + Audio + PDF → Text
Knowledge cutoff
Not verified

API model ID: gemini-3.1-flash-lite

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Multimodal understanding with text output
  • Function calling and structured outputs
  • Search grounding, Maps grounding, and URL context
  • Code execution and file search
  • Context caching and Batch API

Before you choose

  • No native image or audio generation; use a dedicated generation model.
  • The Live API is not supported by this model.
  • Large input capacity does not guarantee that every detail will be recovered correctly.
  • Computer use is not supported.

Gemini thinking levels

minimal · defaultlowmediumhigh

Start with minimal for bounded extraction and classification. Increase the thinking level only if tests show a quality improvement. Minimal does not guarantee zero thinking; difficult inputs can still require some reasoning.

  • This page documents the stable API ID shown above, not an earlier preview mentioned in the launch article.
  • The model-specific card does not establish a training cutoff; no cutoff is inferred from its release date.
  • Tool execution, permissions, and validation belong to the application integrating the API. Always inspect citations and generated structured data.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Translate customer messages

Preserve customer-specific details while changing the language.

Translate these customer messages into Spanish. Keep message IDs, URLs, order numbers, prices, and product names unchanged. Preserve the original tone without adding promises. Return a JSON array with id and translation. Put unclear source phrases in a separate review list rather than guessing.

Workflow 02

Tag customer feedback

Use a closed vocabulary for consistent downstream handling.

For each customer review below, select up to two tags from delivery, fit, durability, usability, and value. Return the review ID, tags, sentiment, and a supporting phrase copied from that review. Use unclassified if none fit. Do not infer the customer's identity or invent an experience they did not describe.

Workflow 03

Summarize a release note

Turn supplied material into a compact, checkable update.

Summarize the release notes I supply into three sections: user-visible changes, required actions, and known limitations. Keep version numbers and dates exact. Link each bullet to a heading in the source. If no action is required, say so; do not create upgrade instructions that are absent from the notes.

Developer reference

Google API pricing

These are Google API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input (text / image / video)$0.25
Input (audio)$0.50
Cached input (text / image / video)$0.025
Cached input (audio)$0.05
Output (including thinking)$1.50
Additional API fees · USD
UsageRate and unit
Cache storage$1.00per 1M tokens per hour
Google Search grounding (after allowance)$14.00per 1,000 search queries
Google Maps grounding (after allowance)$14.00per 1,000 search queries
  • Prices are standard paid-tier API reference rates checked October 6, 2026. Free-tier terms and Batch, Flex, or Priority rates differ; verify the selected service tier before estimating spend.
  • Output pricing includes thinking tokens. Cached reads and cache storage are separate costs.
  • Google lists a 5,000-per-month Search allowance shared across Gemini 3.x models, with a separate shared Gemini 3 Maps allowance. One prompt can trigger multiple billable searches.

Common questions

A few things worth knowing.

What is Gemini 3.1 Flash-Lite useful for?

Consider it for translation, classification, and short summaries. Supply clear constraints and assess correctness, latency, and total usage on your own examples before making it a default.

How should I configure thinking?

This Flash-Lite model supports minimal, low, medium, and high thinking levels, with minimal as the default. More thinking can increase response time and output-token cost; it is not a substitute for validation.

Can I send images, audio, or video?

The Google model card supports these inputs and text output. Upload and tool availability depend on your integration. This does not make the model an image generator, voice generator, or Live API model.

How does this compare with the other Flash-Lite version?

3.1 Flash-Lite remains a distinct stable API ID. It has separate audio pricing and does not support computer use according to its model card. Evaluate it against 3.5 Flash-Lite using your own quality and cost targets.

What costs are additional to input and output tokens?

Cache storage and grounding can add charges beyond ordinary input/output tokens. Audio input has a separate rate where shown. Review the units and free allowances in the provider pricing page.

Does the large input limit guarantee a correct answer?

No. Keep evidence relevant, specify the required output, request source references, and check the result. The 1,048,576 input-token limit and 65,536 output-token limit are different constraints.

Are these the prices and limits of my EZ Ai Assist plan?

No. This guide separates Google’s direct API reference from EZ Ai Assist subscriptions. Use the pricing page for plans and the app for current model access. The example prompts are starting points, not guaranteed results.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider