Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.
Supported API features and tools
OpenAI's open-weight mixture-of-experts reasoning model, served here by Groq
Text input/output with a 131,072-token context and a 65,536-token maximum output on Groq
Low, medium and high reasoning effort, plus tool use through supported integrations
JSON object output and strict JSON Schema output on Groq's supported structured-output path
Automatic prompt caching on eligible repeated inputs; browser search and code execution when explicitly enabled and supported
Before you choose
This hosted endpoint is text-only; do not infer image, audio or video input from other GPT models.
OpenAI's model reference lists a different output ceiling. This page uses Groq's 65,536 serving limit, not the underlying model reference's 131,072.
Groq's approximately 500 tokens/second figure is a published estimate, not a latency or throughput guarantee for your workload.
Schema compliance does not guarantee factual correctness. Validate values, authorize tool calls and treat external content as untrusted.
Choose the reasoning effort
lowmediumhigh
Use reasoning_effort to choose low, medium or high on the Groq endpoint. Start with a bounded task and measure whether extra effort improves the result. Hiding returned reasoning is not the same as selecting a non-reasoning model.
Groq's reasoning_format parameter is not supported for GPT-OSS. Returned reasoning uses a separate reasoning field; include_reasoning controls whether it is included in the response.
OpenAI distributes the weights under Apache 2.0. That license and local deployment options do not mean hosted Groq inference is free.
For strict structured output, follow Groq's supported JSON Schema subset, including required fields and additionalProperties: false. Check the current integration's tool/streaming restrictions before combining features.
Prompt caching is automatic and exact-prefix dependent. Cache hits are not guaranteed; inspect usage rather than assuming every repeated request receives the discount.
Put it to work
Start with a more useful prompt.
Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.
Workflow 01
Extract a review-ready risk register
Produce structured data with evidence attached to each entry.
From the project notes below, create a JSON risk register using the schema I supply. Each risk must include its evidence location, likelihood rationale, impact, mitigation and owner if stated. Use null for missing information if the schema permits it; otherwise report the schema conflict. Do not invent owners or claim that schema-valid output is factually verified.
Workflow 02
Review code with a tool boundary
Make proposed tool actions auditable before execution.
Review these source files for the reported defect. First list the evidence you need and any proposed read-only tool calls. Do not modify files or run destructive commands without approval. Then identify the smallest supported fix, cite affected functions and propose regression tests. Clearly separate inspected evidence from steps that still need to be executed.
Workflow 03
Compare a long set of requirements
Keep a large-context comparison focused on traceable differences.
Compare the old and new requirements documents I provide. Return a change matrix containing requirement ID, old intent, new intent, downstream impact and source location. Flag ambiguities and incompatible requirements instead of resolving them by assumption. Finish with the five changes that most urgently need a product or engineering decision.
Developer reference
Groq API pricing
These are Groq API reference prices, not EZ Ai Assist subscription prices.
Groq publishes these input/output rates for openai/gpt-oss-120b. The hosted service's pricing is separate from the open-weight license.
Eligible cached input receives a 50% discount. Caching is automatic, depends on matching input prefixes and is not guaranteed.
Enabled search or execution tools may have separate terms or charges. Compare actual billed usage and latency for your workload, not just a published token-speed estimate.
Common questions
A few things worth knowing.
Is GPT-OSS 120B made by Groq?
OpenAI created the open-weight model; Groq is the serving provider for this catalog route. The guide therefore uses Groq's hosted ID, limits and prices while linking OpenAI's original model documentation.
Why does the output limit differ from OpenAI's page?
A hosting provider can impose its own serving limit. Groq lists 65,536 output tokens for this endpoint; OpenAI's underlying model reference lists 131,072. This page does not substitute the larger number for Groq's limit.
Can it inspect an image?
Not on the text-only endpoint documented here. Use a verified vision model or provide an independently extracted text description, recognizing that transcription can lose visual information.
Can I request strict JSON?
Groq documents strict structured outputs for GPT-OSS 120B. Supply a supported schema and validate the returned values. Valid JSON and schema compliance do not establish that an answer is correct.
How do I control reasoning?
Use low, medium or high reasoning_effort. For response visibility, Groq uses include_reasoning; reasoning_format is not supported for these GPT-OSS endpoints.
Is 500 tokens per second guaranteed?
No. It is Groq's approximate published token-speed figure. Measure end-to-end time, including input processing, reasoning and tools, using realistic prompts and concurrency.
Does the open-weight license make API calls free?
No. OpenAI's weight license and Groq's hosted inference charges are different things. The displayed token rates are provider API references, not EZ Ai Assist subscription charges.
Check the source
Official documentation
Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.