Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
Meta's LLaMA 3.3 70B text model hosted on Groq, separate from other LLaMA serving services
A documented 131,072-token context and 32,768-token maximum output on this endpoint
Tool use and function-call workflows when supported by the caller's integration
JSON object mode for structured text tasks; validate the contents and expected schema yourself
Before you choose
Groq's free/developer endpoint access ended; only the documented enterprise committed-spend exception is unaffected by that notice.
This is a text-only model on the reviewed endpoint. It is not a vision model and has no verified dedicated reasoning-effort control in this guide.
JSON object mode is not a promise of strict JSON Schema enforcement. Do not borrow GPT-OSS feature support.
Groq's approximately 280 tokens/second estimate is not a service-level guarantee. The reviewed hosting sources do not establish a knowledge cutoff here.
Non-reasoning model
No dedicated thinking toggle or configurable reasoning-effort levels are established for this endpoint. You can ask for an explanation and verification steps, but that does not enable a hidden API mode or guarantee correct reasoning.
Keep the hosted ID llama-3.3-70b-versatile and its Groq-specific limits distinct from model names used by other platforms.
For a migration, compare GPT-OSS 120B on actual prompts and test tool schemas, error handling and output parsing. It is not a guaranteed drop-in behavior match.
Groq's original replacement list also names Qwen3.6 27B, but that model later lost free/developer access too. Consult the latest lifecycle notice before choosing a replacement.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Compare policy revisions
Use a large text window for a traceable document comparison.
Compare the two policy versions I provide. List material changes by section, explain their practical effect in plain language and quote the supporting passage for each. Separate confirmed differences from ambiguous wording. Do not give legal conclusions; identify the questions an authorized reviewer should resolve before publication.
Workflow 02
Validate a text extraction
Check a structured result against its original evidence.
Audit this extracted JSON against the source text and the field definitions below. Mark each field as supported, contradicted or not stated, with a source location. Return a corrected object only where the evidence is clear, and list unresolved fields separately. Do not invent missing values or assume that valid JSON means accurate data.
Workflow 03
Evaluate a replacement for a 70B workflow
Make migration decisions from behavior rather than labels.
Using these saved prompts, expected outputs and application constraints, design a regression evaluation for replacing our current 70B text endpoint. Include instruction following, JSON parsing, tool-call arguments, latency and error behavior. Define pass thresholds and rollback criteria. Do not assert compatibility or benchmark results without measurements.
Developer reference
Groq API pricing
These are Groq API reference prices, not EZ Ai Assist subscription prices.
Groq lists enterprise access with Contact Sales pricing. Use the rates in your eligible contract; this guide does not present former public token rates as a current offer.
The free/developer shutdown is not a change to every hosting service or to the underlying model's existence.
When comparing a replacement, include current model rates, tool costs and measured failure/retry behavior, not only the former endpoint's token speed.
Common questions
A few things worth knowing.
Can free or developer-tier Groq accounts still use it?
Groq's notice ended that access on August 16, 2026. Enterprise customers with a committed-spend contract are not affected. Verify your account's eligibility rather than assuming an old example will still run.
Is LLaMA 3.3 retired everywhere?
No such global claim is made here. This notice is specific to Groq's access tiers. Other hosting services and deployment arrangements have their own availability and terms.
Why is it listed under Groq instead of Meta?
Meta developed LLaMA, while this catalog entry describes Groq hosting. Retaining the route keeps the original provider context and ensures limits and prices are not mixed with another service.
Does it have a thinking-effort selector?
The reviewed Groq reference does not establish one for this endpoint. Asking for a reasoned explanation is not the same as configuring the API reasoning controls documented for GPT-OSS or Qwen.
Can I rely on JSON Schema enforcement?
This guide documents JSON object mode, not strict schema enforcement. Parse and validate the response against your requirements and handle missing, invalid or unsupported values.
Which replacement should I evaluate?
Groq's notice includes GPT-OSS 120B. Test it on your real workflows, especially tool calling and structured output. Qwen3.6 also appears in the older notice but has a later free/developer shutdown of its own.
Where are the current token prices?
Groq's current listing directs eligible enterprise users to Contact Sales. Use the actual contract rate; no former public price is represented as a current self-service tariff.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.