Back to models
PreviewChat Models

Google

Gemini 3.1 Pro

Best for Coding

Google’s Gemini 3.1 Pro Preview supports complex synthesis, coding, and multimodal analysis with a 1M-token input limit.

ReasoningCoding

At a glance

Know the model before you prompt.

Google API specifications
Input token limit
1,048,576 tokens
Maximum output
65,536 tokens
Inputs → output
Text + Images + Video + Audio + PDF → Text
Knowledge cutoff
January 2025

API model ID: gemini-3.1-pro-preview

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Streaming responses and multimodal understanding
  • Function calling and structured outputs
  • Context caching and Batch API
  • Code execution, Google Search grounding, and URL context
  • Google Maps grounding; File search (AI Studio only)
  • Flex and Priority processing

Before you choose

  • No native image or audio generation; these models return text.
  • Live API is not supported; audio input is not a real-time voice session.
  • Tool access, upload limits, and permissions depend on the integration. Validate outputs against the original evidence.
  • This is a preview API model. Availability and behavior can change; test updates against representative cases.
  • File search is documented for AI Studio only. Do not assume every Google product or EZ Ai Assist offers the same tools.

Gemini thinking levels

lowmediumhigh · default

High is the default. Google documents low, medium, and high for 3.1 Pro. Try lower thinking for simpler requests and compare it with high for multi-step analysis. These levels are relative, not precise token budgets. Preserve required thought signatures when continuing tool calls.

  • Use gemini-3.1-pro-preview for the standard model. The customtools variant is a separate ID aimed at different tool-selection behavior.
  • The reviewed lifecycle table lists no announced shutdown date. Preview status is not a promise of indefinite availability.
  • The January 2025 cutoff is listed in Google’s Gemini 3 developer guide. Use current supplied or grounded sources for later events.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Reconcile a research packet

Evaluate conflicting evidence before making a recommendation.

Compare the supplied research on this product decision. For each major claim, cite its source, date, method, and relevant limitations. Explain conflicts rather than averaging them away. Build a decision matrix using the criteria below and identify which missing evidence could change the recommendation. Do not browse or invent sources.

Workflow 02

Trace a system-level defect

Follow a failure across the code and logs you provide.

Investigate this cross-service bug using only the architecture notes, code, and logs below. Reconstruct the event sequence and identify where the observed behavior diverges from the contract. Rank hypotheses by evidence, propose a minimal diagnostic test for each, and distinguish confirmed defects from speculation. Do not deploy changes.

Workflow 03

Plan a constrained rollout

Make dependencies and rollback decisions explicit.

Create a phased rollout plan from the requirements, constraints, and current system description below. Identify dependencies, acceptance checks, responsible roles, and rollback triggers. Show which decisions must be resolved before implementation. Compare two feasible approaches and explain their tradeoffs without assuming unavailable budget or staffing.

Developer reference

Google API pricing

These are Google Gemini Developer API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token type≤ 200,000input tokens> 200,000input tokens
Input$2.00$4.00
Cached input$0.20$0.40
Output (including thinking)$12.00$18.00
  • The two columns depend on total input length: up to 200,000 tokens versus more than 200,000 tokens. Use the corresponding input, cached-input, and output rate for the request.
  • Google lists gemini-3.1-pro-preview and gemini-3.1-pro-preview-customtools together under these prices.
  • Context-cache storage costs $4.50 per million token-hours and is separate from cached-input token charges. Thinking tokens count as output.
  • Batch, Flex, and Priority have separate pricing. Grounding and other tools can add fees; preview access, quotas, and prices may change.

Common questions

A few things worth knowing.

Is Gemini 3.1 Pro a stable API endpoint?

The documented ID is gemini-3.1-pro-preview. Treat it as preview, review Google’s lifecycle notices, and validate important workflows when behavior changes. The marketing name does not remove that preview status.

When do the higher long-context prices apply?

The pricing table distinguishes requests with up to 200,000 input tokens from those above that level. The higher column includes a different output rate as well as input and cached-input rates; do not price only the excess tokens at the higher rate.

How does it differ from Custom Tools?

The customtools endpoint is tuned to favor developer-defined tools over bash in relevant workflows. It is not a blanket upgrade or a promise of better quality on every task. Evaluate tool selection and answer quality separately.

Can it search my files directly in the app?

Google documents File search for this model as AI Studio only. Other integrations may implement their own retrieval. This specification does not establish which file-search features EZ Ai Assist currently enables.

Can I use these prompts in EZ Ai Assist?

Yes—adapt these original editorial examples to the inputs and controls available in your workspace. Share only material you are authorized to provide. They are starting points, not benchmarks or a promise of a particular result.

Does the input limit guarantee accurate long-document answers?

No. The 1,048,576-token input limit is capacity, not a recall guarantee. Organize sources with names and sections, request citations, and check them. The separate maximum output is 65,536 tokens; actual requests must also fit integration limits.

Are these the prices of my EZ Ai Assist plan?

No. The table describes direct Google API token usage. EZ Ai Assist subscriptions and app capabilities are separate. Use our pricing page for current plans and check the app for enabled models and controls.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider