Google’s stable Gemini 2.5 Flash-Lite supports budget-sensitive multimodal tasks with optional thinking. Direct API access is restricted to prior users.
Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Multimodal understanding with text output
Function calling and structured outputs
Search grounding, Maps grounding, and URL context
Code execution and file search
Context caching and Batch API
Before you choose
No native image or audio generation; use a dedicated generation model.
The Live API is not supported by this model.
Large input capacity does not guarantee that every detail will be recovered correctly.
Stable-model access is restricted to prior Google API users; old preview IDs have separate retirement dates.
Extended thinking
Thinking is off by default. Unlike Gemini 3 thinking levels, this model uses a token budget: enable 512–24,576 thinking tokens or dynamic thinking with -1; 0 disables thinking.
This page documents the stable API ID shown above, not an earlier preview mentioned in the launch article.
Gemini 2.5 uses thinkingBudget rather than Gemini 3 thinkingLevel; do not copy reasoning defaults between generations.
Tool execution, permissions, and validation belong to the application integrating the API. Always inspect citations and generated structured data.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Clean a contact import
Limit the operation to formatting supplied data.
Validate these contact-import rows against the schema below. Trim surrounding whitespace and flag missing required fields, duplicate IDs, and malformed email fields. Do not enrich contacts, infer personal details, or change a supplied address. Return accepted rows and a separate error list keyed by row number.
Workflow 02
Extract receipt line items
Keep a low-complexity extraction task auditable.
Read the receipt I supply and return merchant, date, currency, and line items with quantity and amount. Mark unreadable text as null. Keep discounts and tax separate. Compare the sum with the printed total and report a mismatch without silently correcting any printed value.
Workflow 03
Classify incoming requests
Make the decision rules explicit before processing a batch.
Apply the routing rules below to this batch of requests. For each request return its ID, the matching rule number, destination queue, and whether manual review is needed. Do not create new queues. If multiple rules conflict, flag the conflict and leave the destination unresolved.
Developer reference
Google API pricing
These are Google API reference prices, not EZ Ai Assist subscription prices.
Prices are standard paid-tier API reference rates checked October 6, 2026. Free-tier terms and Batch, Flex, or Priority rates differ; verify the selected service tier before estimating spend.
Output pricing includes thinking tokens. Cached reads and cache storage are separate costs.
Search grounding includes 1,500 free requests per day on the paid tier, shared between 2.5 Flash and Flash-Lite; Maps has its own 1,500-per-day paid-tier allowance.
Common questions
A few things worth knowing.
What is Gemini 2.5 Flash-Lite useful for?
Consider it for bounded extraction and classification in established workflows. Supply clear constraints and assess correctness, latency, and total usage on your own examples before making it a default.
How should I configure thinking?
Thinking is off by default. Unlike Gemini 3 thinking levels, this model uses a token budget: enable 512–24,576 thinking tokens or dynamic thinking with -1; 0 disables thinking.
Can I send images, audio, or video?
The Google model card supports these inputs and text output. Upload and tool availability depend on your integration. This does not make the model an image generator, voice generator, or Live API model.
Is this model retired?
No. Google says the stable 2.5 models remain served, with access limited to users who actively used them before. Retired preview IDs are different. Check current provider and EZ Ai Assist availability independently.
What costs are additional to input and output tokens?
Cache storage and grounding can add charges beyond ordinary input/output tokens. Audio input has a separate rate where shown. Review the units and free allowances in the provider pricing page.
Does the large input limit guarantee a correct answer?
No. Keep evidence relevant, specify the required output, request source references, and check the result. The 1,048,576 input-token limit and 65,536 output-token limit are different constraints.
Are these the prices and limits of my EZ Ai Assist plan?
No. This guide separates Google’s direct API reference from EZ Ai Assist subscriptions. Use the pricing page for plans and the app for current model access. The example prompts are starting points, not guaranteed results.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.