ApiaryActiveLive
Try: pause · settings · learn · wipe
← Community / Reading Room
GF
craft · 18 min read

Groq Free Tier Limits Explained (and When to Go Local)

Search mixes them up constantly.

By Austin Little

Groq is the fast free-cloud option a lot of operators actually use — until a 429 error shows up mid-draft. This page explains how Groq's free-tier limits work in plain English, what to check in the Console before you build a habit, and when you should stop fighting the meter and go local with Ollama instead.

AI disclosure. This page was drafted with AI assistance and edited for Apiary. We don't invent quotes, stats, people, or events. If something looks off, tell Austin — that's the point of a living hive.

Groq vs Grok (clear this up first)

Search mixes them up constantly.

  • Groq (with a q) — inference company. Console at console.groq.com. Fast API access to hosted models.
  • Grok (with a k) — xAI's chatbot product. Different company, different product.

This article is about Groq. If you landed here looking for the other one, you are in the wrong hive room.

Why people search "Groq free tier limits explained"

Because the marketing story is simple ("fast," "free tier") and the operating reality is a set of ceilings that stack. You can be "well under" one limit and still get blocked by another. People also want to know whether they need a credit card, whether creating more API keys helps (it does not multiply org limits), and when local AI is the calmer path.

If you want the wider free-AI map, start with How to Use AI Without Paying. For private drafts that should never leave the machine, see Private ChatGPT Alternative: Local AI That Stays on Your Machine. For another fast-cloud evaluation, see Is Cerebras Free AI Worth It for Writing and Research?. For the Apiary meaning of tending knowledge with AI, see What Is an AI Beekeeper?.

What "free tier" means at Groq (as of October 2026)

From Groq's own materials and billing docs:

  • There is a Free plan aimed at getting started with the APIs.
  • Groq's billing FAQs describe upgrading from Free to Developer by adding a payment method, and downgrading back to Free from account Settings under Billing.
  • That strongly implies Free is a real ongoing tier, not only a vanishing trial — but product packaging changes.

Apiary house rule for this lane: we treat free tiers as useful fuel, never as a promise of unlimited production capacity.

The four ceilings you must understand

Groq's rate-limit documentation measures usage in several ways. For text chat work, the ones you will feel are:

AbbreviationMeaningWhat it feels like when you hit it
RPMRequests per minuteYou sent too many calls too fast
RPDRequests per dayYou sent too many calls today
TPMTokens per minuteYour prompts+answers were too fat too fast
TPDTokens per dayYou burned through today's token budget

There are also audio-oriented measures (ASH/ASD) for speech models, and some orgs may see separate input/output token caps (ITPM/OTPM). Groq notes that cached tokens do not count toward rate limits — useful if you repeat long system prompts.

Critical operating rule from Groq's docs: limits apply at the organization level, not per API key. Extra keys are for hygiene and rotation, not for multiplying quota.

Second critical rule: you hit whichever limit arrives first. A burst of tiny requests can exhaust RPM while TPM still looks fine. A few huge prompts can exhaust TPM/TPD while request count looks fine.

When you exceed a limit, the API returns HTTP 429 Too Many Requests. Response headers can show remaining requests/tokens and reset timing. Build retries with backoff if you are scripting; if you are a human in a chat UI, take the hint and slow down or switch models/tools.

Where to read your numbers (do this, don't trust a blog table)

Groq's rate-limit docs say the public table is a high-level summary and that you should view the current, exact rate limits for your organization on the limits page in account settings.

That sentence is the whole article in miniature:

  1. Open the Groq Console.
  2. Find Settings → Limits (wording can shift — look for Limits).
  3. Read the row for the model you actually call.

Why this page will not hard-code a giant free-tier spreadsheet

Third-party blogs publish neat free-tier tables (RPM/RPD/TPM/TPD per model). Some of those numbers may match what you see in Console on a given day. Groq's own rate-limit page also shows per-model numbers, while noting Developer plan base limits in the upgrade callout — packaging and which table applies to Free vs Developer is easy to misread.

Apiary rule: we refuse to invent or over-confidently freeze specific free-tier RPM/RPD/token numbers in a evergreen article unless they were verified same-day from the official Limits page for a Free org. For this page, treat any specific number you see elsewhere as unverified until you open Console.

What we will say, grounded in official docs:

  • Limits are multi-dimensional (requests and tokens, minute and day).
  • They are org-wide.
  • 429 means stop and respect reset headers.
  • Developer plan exists for higher limits (paid path) — not required for learning and light use if Free covers you.
  • Console Limits page is authoritative for your account.

A practical mental model for writers and small operators

Think of Free Groq as a fast shared workshop, not as your private garage.

Good fits:

  • Short drafting bursts when your laptop model is slow
  • Trying an OpenAI-compatible endpoint in a BYO-LLM app
  • Quick rewrites of non-sensitive text
  • Learning how API keys and base URLs work without paying

Bad fits:

  • All-day autonomous agents hammering the API
  • Paste of customer PII, medical charts, kids' school records, or bank details
  • Production bots that must not fail at 2pm when the day budget is gone
  • Anything you would not put on a postcard

Privacy plain English: free cloud can see your prompts. Local Ollama cannot, unless you copy the text elsewhere. Dignity first.

Setup without drama (free path)

High-level steps — always prefer Groq's current quickstart if UI labels moved:

  1. Create an account at the official Console signup.
  2. Verify email if asked.
  3. Create an API key. Copy it once; store it like a password.
  4. Check Limits for your org before you build a workflow.

Do not paste your key into random "free AI wrapper" sites. Prefer apps where the key stays in your browser/local config (BYO-LLM), matching Apiary's design philosophy: the brain is yours.

How tokens burn faster than people expect

A "request" is not one polite sentence. Each call includes:

  • System prompt (if any)
  • Conversation history you re-send
  • Your new user message
  • The model's completion

Long chat threads in naive apps re-send history and chew TPD quietly. Practical habits:

  • Start a fresh thread for a new job.
  • Keep system prompts short.
  • Ask for shorter answers when you only need a punch list.
  • Set a sensible max output length in APIs when you script.
  • Prefer local for long iterative drafting; use Groq for the spike.

When to go local (the decision tree)

Go local (Ollama) when:

  1. The text is private or awkward if leaked.
  2. You hit 429s during the hours you actually write.
  3. You want zero account dependency for basic drafting.
  4. You are offline or on flaky internet.
  5. You are teaching a relative and want fewer moving parts after install — see AI for Grandparents: Free, Safe, and Actually Useful.

Stay on Groq Free when:

  1. You need speed on a weak laptop for non-private work.
  2. You are testing BYO-LLM wiring.
  3. Your daily volume is light and Limits page shows comfortable headroom.
  4. You accept that the meter is part of the product.

Pay for Developer (or another paid path) only when Free reliably blocks real work you care about and local cannot cover the job. Paying for throughput on purpose is operator judgment. Paying because a modal scared you is not.

Groq Free vs local vs other free cloud

NeedBetter defaultNotes
Private journal / client notesLocalNo prompt leaving the box
Fast non-private rewriteGroq FreeCheck Limits; watch 429
Forever-$0 after installLocalElectricity + disk
Phone convenienceOfficial free web chatsStill cloud; still not for secrets
Trial credits that expireOther vendors (e.g. some trial-credit clouds)Read the fine print — see Cerebras sibling
Apiary-style app designBYO-LLM to local or free keyNo baked-in landlord key

Handling 429s like an adult

If you are a human using a chat UI:

  • Pause.
  • Switch to a smaller/faster model row if Limits show more headroom there — after checking Console, not after a tweet.
  • Finish the job in Ollama.
  • Come back tomorrow if you burned the day budget on toys.

If you are scripting:

  • Read retry-after when present.
  • Exponential backoff.
  • Cap concurrency hard.
  • Do not create five keys to "bypass" org limits — that is not how org limits work, and trying to evade platform rules is a bad habit.

Security hygiene for API keys

  • One key per environment (laptop vs experiment) is fine; it does not raise Free caps.
  • Revoke keys you pasted into a ticket or chat log.
  • Never commit keys to git.
  • Prefer OS keychain / env vars over keys sitting in a desktop screenshot.
  • If a "helpful" Discord bot asks for your Groq key, that is a theft attempt.

Writing workflow that respects the meter

Combine this page with How to Write With AI Without Sounding Fake:

  1. Ugly human notes offline.
  2. Outline with local model.
  3. Optional: one hard section on Groq Free if you need speed and the text is non-private.
  4. Voice pass and truth pass by you.
  5. Publish only what you verified.

That order keeps you from burning TPD on brochure fluff you will delete anyway.

Common traps specific to Groq Free

  • Building a customer-facing bot on Free and acting shocked at throttles. Free is for start and light use.
  • Assuming blog RPM tables match your org forever. Console Limits wins.
  • Pasting secrets because "it's just an API." APIs are still someone else's computer.
  • Confusing Groq with Grok and following the wrong docs.
  • Leaving a runaway script in a loop overnight. Wake up to a spent day budget and a lesson.
  • Thinking Developer is mandatory to learn. Learn on Free + local first.

Keeping the stack from rotting

Once a quarter:

  1. Open Console → Limits. Note what changed.
  2. Rotate the API key if it might have leaked.
  3. Delete unused experiments.
  4. Confirm you still know how to run Ollama without any cloud key.
  5. Re-read billing settings if any payment method is on file — cancel cleanly if you upgraded for a test and do not need it.

Joe-Google test

People type:

  • "groq free tier limits"
  • "groq free tier explained"
  • "groq rate limits"
  • "groq API free"
  • "groq 429"
  • "groq vs ollama"

They rarely type "organization-scoped token bucket replenishment." Explain the second idea in first-list words.

Walkthrough: first hour on Groq free (operator version)

  1. Create account on the official console host — type the domain carefully.
  2. Turn on MFA.
  3. Open Settings → find Limits / Organization limits (labels can move; follow live UI).
  4. Copy your limits into a note: model id, RPM, RPD, TPM, TPD if shown.
  5. Create one API key; store in password manager; name it "laptop-dev".
  6. Send one tiny hello-world request from a BYO client or curl-equivalent you already understand.
  7. Read response headers if visible; note remaining.
  8. Stop. Do not burn the day quota on a poorly written eval loop.

Walkthrough: connecting a BYO app

  1. Find the app's custom provider / OpenAI-compatible settings.
  2. Paste base URL from Groq docs (verify live).
  3. Paste key.
  4. Paste a model id that appears on your Limits page.
  5. Test with non-sensitive text.
  6. Write a personal rule: fallback to Ollama when 429 appears twice in ten minutes.

Failure story patterns (composite)

  • Eval hammer: A script compared 40 prompts × 15 models overnight on free tier. Hit RPD. Author thought Groq was "down." Fix: local eval for bulk; cloud for spot checks.
  • Key in screenshot: Key posted in a Discord debug channel. Fix: revoke immediately; assume abuse.
  • Shared family org: Teen's homework bot and parent's marketing tool shared one org bucket. Fix: separate orgs/accounts if appropriate, or schedule usage.

Tokens: rough intuition without fake precision

A short email rewrite might be hundreds of tokens. A pasted chapter might be thousands. Free TPM/TPD ceilings disappear faster on long-context habits. Chunk documents. Summarize locally first. Send cloud only the paragraph that needs the stronger model.

Caching note

Groq docs state cached tokens do not count toward rate limits. That can reward reusable prefixes in some architectures. It is not a reason to put secrets in a reusable prefix. Design caches like you design logs: assume they are sensitive.

Groq vs free ChatGPT vs Ollama (roles)

NeedPrefer
Grandma-friendly UIFree official chat apps
Private draftsOllama
Fast API for a personal toolGroq free (within limits)
Production SLAPaid plans / proper architecture
Zero networkOllama offline after pull

Billing FAQs mindset

Groq publishes billing FAQs covering upgrades from free to developer tiers, payment methods, and downgrades. Read primary docs before you assume free forever or assume a card is required.

Monitoring without a datacenter

For personal use, a spreadsheet tab is enough: date, model, rough requests, whether you hit 429, whether you fell back to local. Review monthly. If 429s cluster on the same workflow, change the workflow.

Ethical load on free tiers

Free tiers are shared resources. Using them to spam, scrape personal data, or brute-force is how free tiers shrink for everyone. Stay inside the provider's terms. Build polite clients.

Local fallback recipe (conceptual)

When your BYO client catches 429:

  1. Show user-facing message: "Cloud busy — trying local."
  2. Point at localhost model.
  3. If local missing, queue the job.
  4. Do not silently drop user text.

This single UX choice makes free cloud viable for real people.

Hardware reality vs API reality

Weak laptop + free Groq = smart pairing for non-private work. Strong laptop + Ollama = often enough to never open Groq. Strong laptop + Groq = convenience when you want speed and already trust the cloud for that prompt. There is no moral ranking — only fit.

Apiary angle again

We teach free paths because paywalls on knowledge are the wrong default. Groq free is one path. It is not the hive. If Groq changed free tiers tomorrow, your Ollama weights and your articles would still exist. Architect for that world.

Checklist before you recommend Groq to a friend

  • [ ] They can store a key safely.
  • [ ] They understand org-level limits.
  • [ ] They have a local fallback.
  • [ ] They will not paste customer PII.
  • [ ] They know where the official docs live.

If not, start them on Ollama chat only.

Extended FAQ

Does Groq train on my API prompts? Read their current data/privacy statements — do not guess from a meme. When unsure, keep sensitive text local.

Is Llama on Groq the same as Llama in Ollama? Related family, different hosting, quantization, and tooling. Test quality on your tasks; do not assume identical behavior.

Can I use Groq in commercial products on free tier? Read current terms and limits. Free tiers may be unsuitable for customer-facing reliability even when allowed.

What does ASH/ASD mean? Audio seconds per hour/day on speech models per Groq's unit list — relevant if you do transcription, not if you only chat text.

Sources / further reading

  • Groq rate limits docs: https://console.groq.com/docs/rate-limits
  • Groq billing FAQs: https://console.groq.com/docs/billing-faqs
  • GroqCloud marketing/plans: https://groq.com/groqcloud
  • Ollama: https://ollama.com
  • Sibling drafts: Use AI Without Paying; Write With AI Without Sounding Fake; Cerebras Free AI Worth It; Private ChatGPT Alternative; AI for Grandparents; What Is an AI Beekeeper?

Reading rate-limit headers without becoming a platform engineer

When you call the API directly, Groq documents headers such as:

  • retry-after (seconds) — often present when you already hit 429
  • x-ratelimit-limit-requests / remaining-requests / reset-requests — tied to RPD in their header notes
  • x-ratelimit-limit-tokens / remaining-tokens / reset-tokens — tied to TPM in their header notes

You do not need to memorize header names to write blog posts. You need them if you are wiring a script. Practical translation:

  • If remaining-requests crashes toward zero, you are request-heavy.
  • If remaining-tokens crashes toward zero, your prompts or completions are fat.
  • If reset is a few seconds, you bumped a per-minute ceiling; wait.

Choosing models without chasing the logo

Model catalogs churn. Names appear, rename, and deprecate. Operator rules that stay stable:

  1. Pick the smallest model that does the writing job.
  2. Check that model's row on your Limits page.
  3. Prefer production models for anything you rely on; treat preview models as temporary.
  4. Do not hard-code a model ID into a tutorial you will not maintain — or do, and accept you must update it.

For writing and light research summaries of text you provide, a mid-small chat model is usually enough. Save giant reasoning models for hard problems — and expect tighter daily budgets on bigger rows when Free is constrained.

BYO-LLM wiring checklist (Groq key)

When an app says it speaks OpenAI-compatible APIs:

  • Base URL points at Groq's documented OpenAI-compatible endpoint (verify current docs).
  • Model field uses a model ID your key can call.
  • API key field gets your Groq key — not your bank password, not a fake key unless the app is pointed at Ollama instead.
  • Network calls leave your machine — so still no secrets in the prompt.
  • If the app offers a local Ollama mode, use that for private work and keep Groq as a profile you switch on deliberately.

Apiary's philosophy matches this: the app coordinates; it should not own your inference landlord relationship.

Cost psychology: free is not "no constraints"

People treat paid tiers as serious and free tiers as toys. Flip that. Free tiers are serious about fairness limits. The constraint is how you learn good architecture:

  • Cache and reuse thoughtfully.
  • Don't poll in tight loops.
  • Separate "explore" from "deadline" workloads.
  • Keep a local fallback so a 429 is an inconvenience, not a halt.

That discipline still helps if you later pay.

Team and household sharing

Because limits are org-scoped:

  • A household sharing one Groq org shares one budget.
  • A teammate's runaway notebook can empty your day.
  • For teaching, consider separate orgs when possible, or schedule usage so experiments do not collide with real deadlines.

For grandparents and careful shared computers, local may be kinder than a shared cloud key — see the grandparents guide.

Myths to drop

  • "Free means I can scrape the whole internet through the API." No. Limits, terms, and decency say otherwise.
  • "If I use the playground only, limits don't apply." Assume usage still counts somehow until docs say otherwise; check your usage views.
  • "Zero-data retention marketing means I can paste SSNs." Product retention features are not a substitute for judgment. Local is still the private path.
  • "I read a March blog with exact RPM numbers, so I'm fine in September." Re-open Limits.

Mini playbook: deadline day

You have a draft due tonight. Laptop is hot. Groq returns 429.

  1. Save everything locally.
  2. Switch Ollama to a small model.
  3. Continue section-by-section.
  4. Use Groq later only if Limits recover and text is non-private.
  5. Do not create a second paid account in a panic without reading cancel paths.

Deadlines reward boring backups.

FAQ

Do I need a credit card for Groq Free? Groq's materials describe a Free plan and require a payment method to upgrade to Developer. Re-verify on the official signup/billing pages the day you set up — packaging can change.

Do more API keys give more quota? No. Docs say limits are organization-level.

Is Groq free unlimited? No. Rate limits are the product constraint on Free.

Can I use Groq inside Apiary-style BYO-LLM apps? If the app accepts an OpenAI-compatible base URL + key, often yes. Keep the key local to the app config. Prefer local endpoint for private work.

What does 429 mean? You hit a rate limit. Slow down, switch tool, or wait for reset. Check response headers when scripting.

Should I put Groq on a production checkout flow? Not on Free if downtime matters. Design for paid capacity or local fallback on purpose.

Is Groq a local model? No. Inference runs on Groq's side. Local means Ollama (or similar) on your machine.

How is this different from Cerebras? Different company, different hardware story, different free/trial packaging. Cerebras' own docs currently describe a time-bounded free trial credits model with payment method requirements — not the same as "always Free tier." Read the Cerebras sibling before you assume they are interchangeable.

What if the Limits page looks different than this article? Trust the Limits page. Flag the article for update. That is what living hive docs are for.

A one-hour lab (learn the limits safely)

Minutes 0–10: Create Console account. Generate key. Open Limits. Write down the model you will try and the limit labels you see (do not publish those numbers if they might be account-specific and sensitive — they are your operating notes).

Minutes 10–25: Call a tiny non-private prompt from a BYO chat UI or curl. Confirm it works.

Minutes 25–40: Install/open Ollama if needed. Run the same prompt locally. Compare latency and privacy posture — not as a benchmark white paper, as a feel check.

Minutes 40–50: Intentionally send a few rapid tiny requests (still non-private) only if you are okay burning a little RPM learning — or skip and just read 429 docs. Do not attack the API; learn politely.

Minutes 50–60: Write a sticky note: "Groq for fast public-ish bursts. Local for private. Console Limits is truth."

Operator scenarios

Freelance writer on a weak laptop: outline locally, polish non-sensitive sections on Groq Free when Limits allow, final voice pass by hand.

Shop owner drafting invoices: local only for customer lines; never send account numbers to any free cloud.

Student learning APIs: Groq Free is a good classroom. Still follow school rules on AI-assisted work.

Agent experimenter: Free will teach you about 429s quickly. That lesson is cheaper than a surprise bill — as long as you did not put a card down "just in case" without reading cancel paths.

Dignity and misuse

Fast free inference is still a tool. Do not use it to spam, to harass, to generate scam scripts, or to automate abuse. Platforms rate-limit partly for fairness and safety. Apiary content assumes lawful, decent use.

Frequently asked
Do I need a credit card for Groq Free?
Groq's materials describe a Free plan and require a payment method to upgrade to Developer. Re-verify on the official signup/billing pages the day you set up — packaging can change.
Do more API keys give more quota?
No. Docs say limits are organization-level.
Is Groq free unlimited?
No. Rate limits are the product constraint on Free.
Can I use Groq inside Apiary-style BYO-LLM apps?
If the app accepts an OpenAI-compatible base URL + key, often yes. Keep the key local to the app config. Prefer local endpoint for private work.
What does 429 mean?
You hit a rate limit. Slow down, switch tool, or wait for reset. Check response headers when scripting.
References & sources
  1. Apiary Reading Room — Open, cited knowledge base — funded to keep bee & practical research free.
From the Apiary Reading Room. Opinion & editorial — not financial advice. We don't overclaim.
More from the Reading Room