Back to models
NewChat Models

Groq

GPT-OSS 120B

Best Overall

OpenAI's larger open-weight reasoning model on Groq, with tool use, structured outputs, and a 131K context window.

500 tok/sOpen model

At a glance

Know the model before you prompt.

Groq API specifications
Context window
131,072 tokens
Maximum output
65,536 tokens
Inputs → output
Text → Text
Knowledge cutoff
June 1, 2024

API model ID: openai/gpt-oss-120b

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • OpenAI's open-weight mixture-of-experts reasoning model, served here by Groq
  • Text input/output with a 131,072-token context and a 65,536-token maximum output on Groq
  • Low, medium and high reasoning effort, plus tool use through supported integrations
  • JSON object output and strict JSON Schema output on Groq's supported structured-output path
  • Automatic prompt caching on eligible repeated inputs; browser search and code execution when explicitly enabled and supported

Before you choose

  • This hosted endpoint is text-only; do not infer image, audio or video input from other GPT models.
  • OpenAI's model reference lists a different output ceiling. This page uses Groq's 65,536 serving limit, not the underlying model reference's 131,072.
  • Groq's approximately 500 tokens/second figure is a published estimate, not a latency or throughput guarantee for your workload.
  • Schema compliance does not guarantee factual correctness. Validate values, authorize tool calls and treat external content as untrusted.

Choose the reasoning effort

lowmediumhigh

Use reasoning_effort to choose low, medium or high on the Groq endpoint. Start with a bounded task and measure whether extra effort improves the result. Hiding returned reasoning is not the same as selecting a non-reasoning model.

  • Groq's reasoning_format parameter is not supported for GPT-OSS. Returned reasoning uses a separate reasoning field; include_reasoning controls whether it is included in the response.
  • OpenAI distributes the weights under Apache 2.0. That license and local deployment options do not mean hosted Groq inference is free.
  • For strict structured output, follow Groq's supported JSON Schema subset, including required fields and additionalProperties: false. Check the current integration's tool/streaming restrictions before combining features.
  • Prompt caching is automatic and exact-prefix dependent. Cache hits are not guaranteed; inspect usage rather than assuming every repeated request receives the discount.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Extract a review-ready risk register

Produce structured data with evidence attached to each entry.

From the project notes below, create a JSON risk register using the schema I supply. Each risk must include its evidence location, likelihood rationale, impact, mitigation and owner if stated. Use null for missing information if the schema permits it; otherwise report the schema conflict. Do not invent owners or claim that schema-valid output is factually verified.

Workflow 02

Review code with a tool boundary

Make proposed tool actions auditable before execution.

Review these source files for the reported defect. First list the evidence you need and any proposed read-only tool calls. Do not modify files or run destructive commands without approval. Then identify the smallest supported fix, cite affected functions and propose regression tests. Clearly separate inspected evidence from steps that still need to be executed.

Workflow 03

Compare a long set of requirements

Keep a large-context comparison focused on traceable differences.

Compare the old and new requirements documents I provide. Return a change matrix containing requirement ID, old intent, new intent, downstream impact and source location. Flag ambiguities and incompatible requirements instead of resolving them by assumption. Finish with the five changes that most urgently need a product or engineering decision.

Developer reference

Groq API pricing

These are Groq API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans
Standard processing · USD per 1,000,000 tokens
Token typePrice
Input$0.15
Cached input$0.075
Output$0.60
  • Groq publishes these input/output rates for openai/gpt-oss-120b. The hosted service's pricing is separate from the open-weight license.
  • Eligible cached input receives a 50% discount. Caching is automatic, depends on matching input prefixes and is not guaranteed.
  • Enabled search or execution tools may have separate terms or charges. Compare actual billed usage and latency for your workload, not just a published token-speed estimate.

Common questions

A few things worth knowing.

Is GPT-OSS 120B made by Groq?

OpenAI created the open-weight model; Groq is the serving provider for this catalog route. The guide therefore uses Groq's hosted ID, limits and prices while linking OpenAI's original model documentation.

Why does the output limit differ from OpenAI's page?

A hosting provider can impose its own serving limit. Groq lists 65,536 output tokens for this endpoint; OpenAI's underlying model reference lists 131,072. This page does not substitute the larger number for Groq's limit.

Can it inspect an image?

Not on the text-only endpoint documented here. Use a verified vision model or provide an independently extracted text description, recognizing that transcription can lose visual information.

Can I request strict JSON?

Groq documents strict structured outputs for GPT-OSS 120B. Supply a supported schema and validate the returned values. Valid JSON and schema compliance do not establish that an answer is correct.

How do I control reasoning?

Use low, medium or high reasoning_effort. For response visibility, Groq uses include_reasoning; reasoning_format is not supported for these GPT-OSS endpoints.

Is 500 tokens per second guaranteed?

No. It is Groq's approximate published token-speed figure. Measure end-to-end time, including input processing, reasoning and tools, using realistic prompts and concurrency.

Does the open-weight license make API calls free?

No. OpenAI's weight license and Groq's hosted inference charges are different things. The displayed token rates are provider API references, not EZ Ai Assist subscription charges.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider