Back to models
Enterprise accessChat Models

Groq

LLaMA 3.3 70B Versatile

Meta's 70B text model on Groq with 131K context. Free/developer access ended; committed-spend enterprise contracts are exempt.

Text workflows131K context

At a glance

Know the model before you prompt.

Groq API specifications
Context window
131,072 tokens
Maximum output
32,768 tokens
Inputs → output
Text → Text
Knowledge cutoff
Not verified

API model ID: llama-3.3-70b-versatile

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Meta's LLaMA 3.3 70B text model hosted on Groq, separate from other LLaMA serving services
  • A documented 131,072-token context and 32,768-token maximum output on this endpoint
  • Tool use and function-call workflows when supported by the caller's integration
  • JSON object mode for structured text tasks; validate the contents and expected schema yourself

Before you choose

  • Groq's free/developer endpoint access ended; only the documented enterprise committed-spend exception is unaffected by that notice.
  • This is a text-only model on the reviewed endpoint. It is not a vision model and has no verified dedicated reasoning-effort control in this guide.
  • JSON object mode is not a promise of strict JSON Schema enforcement. Do not borrow GPT-OSS feature support.
  • Groq's approximately 280 tokens/second estimate is not a service-level guarantee. The reviewed hosting sources do not establish a knowledge cutoff here.

Non-reasoning model

No dedicated thinking toggle or configurable reasoning-effort levels are established for this endpoint. You can ask for an explanation and verification steps, but that does not enable a hidden API mode or guarantee correct reasoning.

  • Keep the hosted ID llama-3.3-70b-versatile and its Groq-specific limits distinct from model names used by other platforms.
  • For a migration, compare GPT-OSS 120B on actual prompts and test tool schemas, error handling and output parsing. It is not a guaranteed drop-in behavior match.
  • Groq's original replacement list also names Qwen3.6 27B, but that model later lost free/developer access too. Consult the latest lifecycle notice before choosing a replacement.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Compare policy revisions

Use a large text window for a traceable document comparison.

Compare the two policy versions I provide. List material changes by section, explain their practical effect in plain language and quote the supporting passage for each. Separate confirmed differences from ambiguous wording. Do not give legal conclusions; identify the questions an authorized reviewer should resolve before publication.

Workflow 02

Validate a text extraction

Check a structured result against its original evidence.

Audit this extracted JSON against the source text and the field definitions below. Mark each field as supported, contradicted or not stated, with a source location. Return a corrected object only where the evidence is clear, and list unresolved fields separately. Do not invent missing values or assume that valid JSON means accurate data.

Workflow 03

Evaluate a replacement for a 70B workflow

Make migration decisions from behavior rather than labels.

Using these saved prompts, expected outputs and application constraints, design a regression evaluation for replacing our current 70B text endpoint. Include instruction following, JSON parsing, tool-call arguments, latency and error behavior. Define pass thresholds and rollback criteria. Do not assert compatibility or benchmark results without measurements.

Developer reference

Groq API pricing

These are Groq API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans

Groq lists enterprise access with Contact Sales pricing. Use the rates in your eligible contract; this guide does not present former public token rates as a current offer.

  • The free/developer shutdown is not a change to every hosting service or to the underlying model's existence.
  • When comparing a replacement, include current model rates, tool costs and measured failure/retry behavior, not only the former endpoint's token speed.

Common questions

A few things worth knowing.

Can free or developer-tier Groq accounts still use it?

Groq's notice ended that access on August 16, 2026. Enterprise customers with a committed-spend contract are not affected. Verify your account's eligibility rather than assuming an old example will still run.

Is LLaMA 3.3 retired everywhere?

No such global claim is made here. This notice is specific to Groq's access tiers. Other hosting services and deployment arrangements have their own availability and terms.

Why is it listed under Groq instead of Meta?

Meta developed LLaMA, while this catalog entry describes Groq hosting. Retaining the route keeps the original provider context and ensures limits and prices are not mixed with another service.

Does it have a thinking-effort selector?

The reviewed Groq reference does not establish one for this endpoint. Asking for a reasoned explanation is not the same as configuring the API reasoning controls documented for GPT-OSS or Qwen.

Can I rely on JSON Schema enforcement?

This guide documents JSON object mode, not strict schema enforcement. Parse and validate the response against your requirements and handle missing, invalid or unsupported values.

Which replacement should I evaluate?

Groq's notice includes GPT-OSS 120B. Test it on your real workflows, especially tool calling and structured output. Qwen3.6 also appears in the older notice but has a later free/developer shutdown of its own.

Where are the current token prices?

Groq's current listing directs eligible enterprise users to Contact Sales. Use the actual contract rate; no former public price is represented as a current self-service tariff.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider