Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
OpenAI's compact open-weight mixture-of-experts model on Groq's hosted infrastructure
A 131,072-token text context with up to 65,536 output tokens on this serving endpoint
Low, medium and high effort for reasoning tasks, with function/tool-call support
JSON object mode and supported strict JSON Schema outputs for structured workflows
Automatic prompt caching for eligible prefixes and optional supported search/execution tools
Before you choose
Text-only input/output; this endpoint is not a replacement for a verified vision or audio model.
The 65,536 output ceiling is Groq-specific. Do not copy OpenAI's higher underlying-model output limit into an endpoint configuration.
The approximately 1,000 tokens/second figure is a provider estimate, not a guarantee. Input size, reasoning and tool round trips affect real response time.
Compact does not mean equally capable on every difficult task. Evaluate failure cases and provide a review or escalation path for uncertain outputs.
Choose the reasoning effort
lowmediumhigh
Select low, medium or high reasoning_effort on Groq. For routine extraction, start with the least expensive measured configuration that passes your quality checks. include_reasoning changes response visibility; it does not document an off switch for the model's reasoning.
reasoning_format is not supported for GPT-OSS on Groq. Use include_reasoning when controlling whether the separate reasoning field is returned.
OpenAI's model reference describes small-footprint open-weight deployment under Apache 2.0; this page's limits and prices concern Groq hosting, not local hardware requirements.
Use strict: true only with a schema that meets Groq's supported requirements. Validate meaning and handle API errors even when output structure is constrained.
Groq's model card rounds cached input to $0.037; its caching policy specifies a 50% discount. The table shows the calculated $0.0375 rate transparently, rather than silently rounding a low price.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Classify support messages
Define an escalation path instead of forcing every message into a category.
Classify each support message below using only the allowed categories I provide. Return the message ID, category, a short supporting excerpt and whether human review is needed. Use the review path for ambiguous or conflicting requests. Do not infer sensitive personal attributes or add details that are absent from the message.
Workflow 02
Normalize a small dataset
Make low-cost extraction easy to audit.
Normalize the supplied product records into the requested JSON schema. Preserve original IDs and source values, standardize only the fields covered by these rules, and report missing or conflicting values separately. Do not guess prices, units or dates. Include a concise validation checklist that can be applied to every output row.
Workflow 03
Patch a focused utility function
Keep a small coding task bounded by explicit tests.
Given this utility function, its intended behavior and the failing examples, propose the smallest fix. Explain the specific defect, preserve the public interface and list edge-case tests. Do not refactor surrounding code or add dependencies. If the expected behavior is ambiguous, ask a targeted question before choosing an implementation.
Developer reference
Groq API pricing
These are Groq API reference prices, not EZ Ai Assist subscription prices.
Cached input is calculated from Groq's 50% discount on $0.075 input: $0.0375 per million tokens. The model card displays the rounded $0.037 figure; confirm the billing rate with Groq for precise estimates.
Cache hits depend on exact matching prefixes and are not guaranteed. The discount applies to eligible input, not output tokens.
These are hosted inference rates, independent of the model's open-weight license. Tools and other service features may have separate charges or terms.
Common questions
A few things worth knowing.
When should I try 20B instead of 120B?
Start with 20B for well-defined tasks that have clear acceptance checks. Compare representative outputs and escalate cases that fail. A lower token price is only useful when the result is good enough for the intended job.
Is this an OpenAI or Groq model?
OpenAI supplies the model weights and architecture. Groq serves the API endpoint in this catalog entry, so Groq's context/output limits, feature support and charges are the relevant hosting reference.
Why show four decimal places for cached input?
Groq's caching guide promises a 50% input discount; half of $0.075 is $0.0375. Its model card displays $0.037. The guide identifies the calculation and the rounding difference instead of hiding it.
Can I turn it into a vision model with a prompt?
No. The documented Groq endpoint accepts text. A prompt cannot add an unsupported input modality or enable a tool integration that is not configured.
Does strict JSON remove the need for validation?
No. It constrains output structure for supported schemas, not the truth or suitability of values. Check source evidence, business rules and missing information before acting on an output.
What reasoning levels are available?
Groq documents low, medium and high for GPT-OSS. The separate include_reasoning option affects response visibility; reasoning_format is not supported on these endpoints.
Can I rely on the advertised token speed?
Treat approximately 1,000 tokens per second as a provider reference, not a service-level guarantee. Measure actual end-to-end latency and cost with your input length, settings and tools.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.