Back to models
Enterprise accessChat Models

Groq

LLaMA 3.1 8B Instant

Compact LLaMA text model on Groq. Free/developer access ended; eligible enterprise contracts retain access under Groq's exception.

CompactText workflows

At a glance

Know the model before you prompt.

Groq API specifications
Context window
131,072 tokens
Maximum output
131,072 tokens
Inputs → output
Text → Text
Knowledge cutoff
Not verified

API model ID: llama-3.1-8b-instant

Capabilities & boundaries

What it supports. Where the limits are.

Tool support requires the appropriate API integration; a supported tool is not automatically active in every chat.

Supported API features and tools

  • Meta's compact LLaMA 3.1 8B model served through the named Groq text endpoint
  • Published context and maximum completion limits of 131,072 tokens; actual requests must still fit the endpoint's constraints
  • Tool use through a supported function-call integration with caller-controlled execution
  • JSON object mode for text extraction and classification workflows with downstream validation

Before you choose

  • Free/developer Groq access has ended; the committed-spend enterprise exception does not establish access for every account or app.
  • Text-only input/output. No dedicated reasoning-effort control or strict JSON Schema guarantee is inferred from other Groq models.
  • Equal published context and output ceilings do not mean you can supply a full context and also generate that many additional tokens. Check the accepted request budget.
  • The approximately 560 tokens/second reference is not a latency guarantee. Verify quality on difficult or ambiguous inputs rather than assuming speed implies suitability.

Non-reasoning model

This guide does not establish a dedicated thinking mode for LLaMA 3.1 8B Instant. Give clear instructions, include a small validation rubric and route uncertain cases for review; do not send unsupported reasoning-effort settings borrowed from another model.

  • Use llama-3.1-8b-instant only if the Groq account has appropriate current access. Check scheduled tasks and fallback routes as well as interactive requests.
  • Groq recommends GPT-OSS 20B as a replacement, but its reasoning and structured-output behavior differ. Re-test prompt budgets, parsing, latency and errors rather than only changing the ID.
  • A verified knowledge cutoff is not established in this guide's hosting sources. Ground time-sensitive answers in supplied evidence or a supported retrieval tool.

Put it to work

Start with a more useful prompt.

Original examples from EZ Ai Assist. Adapt these to your task and the features available in your workspace.

Workflow 01

Route short requests safely

Use a small model for an explicit, bounded classification task.

Route each short request below to one of the teams in the supplied routing table. Return the request ID, chosen team and a brief evidence-based reason. If the rules conflict or no team matches, select manual review. Preserve the original wording and do not infer missing customer details or commitments.

Workflow 02

Rewrite without changing the facts

Constrain a fast writing task with a fidelity check.

Rewrite this status update in clear, concise language for a nontechnical reader. Preserve every date, number, owner and commitment exactly. Do not add promises or remove uncertainty. After the rewrite, list any statements that were ambiguous in the source and confirm which factual details you preserved.

Workflow 03

Test a compact-model migration

Capture behavior that might change when moving to a reasoning model.

Create a migration test matrix from these examples for replacing a compact text model with GPT-OSS 20B. Cover response formatting, classification consistency, tool arguments, timeout handling and output-token budgets. Define what we must measure and how to compare failures. Do not claim that the new model is faster or more accurate before testing.

Developer reference

Groq API pricing

These are Groq API reference prices, not EZ Ai Assist subscription prices.

View EZ Ai Assist plans

Current Groq access is contract-dependent, and the model listing directs enterprise customers to Contact Sales. No public self-service token price is asserted for this restricted endpoint.

  • Confirm eligible contract rates and access with Groq before estimating new usage.
  • A replacement can change reasoning-token use and latency even when its per-token rate looks attractive. Compare complete request costs and acceptance-test results.

Common questions

A few things worth knowing.

What happened to self-service Groq access?

Free and developer-tier access ended August 16, 2026. Committed-spend enterprise contracts are explicitly unaffected by the notice. Check your account rather than relying on an old working configuration.

Is the underlying LLaMA model gone?

No. This guide describes Groq's hosting policy, not a global withdrawal of Meta's model. Availability on other services or private deployments must be checked separately.

Can this model handle screenshots?

The reviewed Groq endpoint is text-only. Use a verified vision endpoint for images; changing the wording of a prompt cannot add an unsupported modality.

Are context and output limits additive?

Do not assume so. Both published ceilings are 131,072 tokens, but a real request must obey the provider's combined budgeting and serving constraints. Leave room for the desired response and validate the accepted limits.

What should I migrate to?

Groq recommends GPT-OSS 20B. Evaluate it with your actual prompts, parser and tools. Its reasoning behavior and structured-output support differ, so an ID change alone is not a complete migration test.

Does Instant guarantee a particular response time?

No. Groq's approximate 560 tokens/second figure is a reference measurement, not an end-to-end promise. Input processing, concurrency and external tools affect user-visible latency.

Why not display the old public prices?

The current access model is enterprise/contract-dependent. Quoting a historical self-service rate as a current offer would be misleading; consult the actual contract or compare a currently available alternative.

Check the source

Official documentation

Specifications and API prices checked on . Example prompts and workflow advice are editorial guidance from EZ Ai Assist.

Same provider