Verified provider facts · Reviewed 2026-08-22

Free LLM APIs with direct API key links

Compare permanent free tiers, no-card options, models and published limits. Every claim links to an official source.

13permanent free tiers
13need no credit card
13free tiers OpenAI compatible
2026-08-22sources last reviewed

Data-driven shortcuts

Strong starting points

Selected by published catalog facts, never by sponsorship.

Re-checked 2026-08-22

What changed since the last review

First re-check since publication. Every source was read again: GitHub Models is gone, Novita has no free rows left, and Groq dropped both Llama models from the free table. Two entries could not be re-read and keep their old check date.

  • Lifecycle GitHub Models Retired on 2026-07-30 as announced: playground, model catalog, inference API and BYOK endpoints are all gone. Kept as a tombstone entry.
  • Lifecycle Novita AI inclusionai/Ling-3.0-flash and Mind Lab Macaron V1 Venti are no longer priced at zero; no model on the pricing table is free any more. Moved to metered access.
  • Limit changed GroqCloud llama-3.3-70b-versatile and llama-3.1-8b-instant left the free plan table. openai/gpt-oss-safeguard-20b, qwen/qwen3.6-27b and groq/compound-mini are in it. The RPM/RPD columns now quote openai/gpt-oss-120b at 30 RPM / 1K RPD.
  • Limit changed Cerebras Inference zai-glm-4.7 left the Free Trial table; only gpt-oss-120b and gemma-4-31b remain at 5 RPM / 30K TPM / 1M TPH / 1M TPD.
  • Added Z.AI Open Platform GLM-4.6V-Flash, a vision model, is now listed at $0 across every pricing column alongside GLM-4.7-Flash and GLM-4.5-Flash.
  • Added SambaNova Cloud DeepSeek-V3.2 and gemma-4-31B-it appear in the free table as preview models, on the same 20 RPM / 20 RPD / 200K TPD numbers.
  • Correction Pollinations.AI The API docs do publish limits after all, as a minimum interval per tier: Anonymous one request every 15s, Seed 5s, Flower 3s, Nectar none. The entry previously said no figures were published.
  • Correction Nebius Token Factory The 60 RPM / 400,000 TPM pair carried here as a baseline appears on the source page only inside a table labelled a visual example. No account default is published; the entry no longer quotes a number.
  • Correction Vercel AI Gateway The $5 monthly free credit is no longer published on the pricing page. The free tier still exists but no amount is stated, so none is quoted here.
  • Correction Cloudflare Workers AI Some models sit outside the 10,000 Neuron allowance and need a paid billing method regardless: @cf/moonshotai/kimi-k2.6, @cf/zai-org/glm-5.2 and @cf/deepseek-ai/deepseek-v4-pro-0813.
  • Lifecycle Moonshot AI (Kimi) Docs moved to platform.kimi.ai. The lowest tier still requires a $1 recharge before any call goes through, so the entry moved out of the permanent free tier list into metered access.

A free tier closing turns into a bill. What the paid ones actually cost, per million tokens, with the sentence each figure came from: the field notes cost table.

Every week on record, as JSON

Catalog

Compare the full directory

Permanent free tiers, aggregators, trials and metered access stay visibly separate.

26 matches

ProviderFree accessCardOpenAILimitsSample probeSources checkedAccess
Google Gemini API https://generativelanguage.googleapis.com/v1beta/openai/ Provider free tier Not required Yes Dynamic / unknown Rate limits apply per project, not per API key, and depend on the project's usage tier; the Free tier is the entry tier and moving up requires enabling billing. Google publishes the tier structure but directs you to AI Studio for the exact RPM, TPM and RPD active on your project, so no single free-tier number can be quoted here. Not checked No authenticated probe has been published. Gemini API rate limitsGemini API billingGemini OpenAI compatibility Get API access
GroqCloud https://api.groq.com/openai/v1 Provider free tier Not required Yes 30 RPM · 1000 requests/day Free plan limits are per model and enforced per organization. openai/gpt-oss-120b, openai/gpt-oss-20b, openai/gpt-oss-safeguard-20b and qwen/qwen3.6-27b: 30 RPM, 1K RPD, 8K TPM, 200K TPD. groq/compound and groq/compound-mini: 30 RPM, 250 RPD, 70K TPM, no daily token cap listed. meta-llama/llama-prompt-guard-2-22m and -86m: 30 RPM, 14.4K RPD, 15K TPM, 500K TPD. Whisper large-v3 and turbo are metered in audio seconds: 20 RPM, 2K RPD, 7.2K ASH, 28.8K ASD. llama-3.3-70b-versatile and llama-3.1-8b-instant no longer appear in the free table. The RPM/RPD columns here quote openai/gpt-oss-120b. Not checked No authenticated probe has been published. GroqCloud rate limitsGroq OpenAI compatibilityGroq billing FAQ Get API access
SambaNova Cloud https://api.sambanova.ai/v1 Provider free tier Not required Yes 20 RPM · 20 requests/day The Free Tier applies whenever no payment method is linked to the account. Every model in the free table shares the same numbers: 20 RPM, 20 RPD and 200,000 tokens per day. Production models are DeepSeek-V3.1, Meta-Llama-3.3-70B-Instruct and gpt-oss-120b; DeepSeek-V3.2 and gemma-4-31B-it are listed as preview. SambaNova publishes no tokens-per-minute figure for the free tier, only RPM, RPD and TPD. MiniMax-M2.7 appears in the Developer tier tables only. Linking a payment method moves the account to the Developer Tier. Not checked No authenticated probe has been published. SambaNova model rate limitsSambaNova developer guideSambaNova OpenAI compatibility Get API access
Cohere https://api.cohere.ai/compatibility/v1 Provider free tier Not required Yes 20 RPM Trial keys are free, rate limited and not licensed for production. Chat is limited to 20 requests/minute, Rerank to 10/minute, Audio Transcriptions and EmbedJob to 5/minute, Tokenize to 100/minute, and Embed to 2,000 inputs/minute. Every trial key is additionally capped at 1,000 API calls per month across the account. Not checked No authenticated probe has been published. Cohere API rate limitsCohere Compatibility APICohere pricing Get API access
Cloudflare Workers AI https://api.cloudflare.com/client/v4/accounts/ACCOUNT_ID/ai/v1 Provider free tier Not required Yes Dynamic / unknown Both the Workers Free and Workers Paid plans include 10,000 Neurons per day at no charge, resetting daily at 00:00 UTC. Neurons are a compute unit rather than a request count, so the number of requests you get depends on the model and prompt size. On the Free plan there is no overage: once the allocation is spent, requests fail until reset. Workers Paid bills overage at $0.011 per 1,000 Neurons. A handful of models sit outside the free allowance entirely and require a paid billing method regardless — @cf/moonshotai/kimi-k2.6, @cf/zai-org/glm-5.2 and @cf/deepseek-ai/deepseek-v4-pro-0813 — reachable through Workers Paid or prepaid AI Gateway credits. Not checked No authenticated probe has been published. Workers AI pricingWorkers AI OpenAI compatible endpointsWorkers AI models Get API access
Hugging Face Inference Providers https://router.huggingface.co/v1 Provider free tier Not required Yes Dynamic / unknown Signed-in free users receive $0.10 in monthly Inference Provider credits (Hugging Face notes this is subject to change); PRO users receive $2.00 and Team or Enterprise organizations $2.00 per seat. Credits only apply to requests routed by Hugging Face, not to requests made with your own provider key. Once the monthly credits are spent you must purchase credits to continue. Not checked No authenticated probe has been published. Inference Providers pricing and billingInference Providers OpenAI compatibilityInference Providers settings Get API access
SiliconFlow https://api.siliconflow.com/v1 Provider free tier Not required Yes 1000 RPM Rate limits for free models are fixed; paid models are tiered by monthly spend and start at tier L0 with 1,000 RPM and 40,000 TPM. Limits are enforced per user account rather than per API key, and each model is limited separately. deepseek-ai/DeepSeek-R1 and deepseek-ai/DeepSeek-V3 carry an extra cap of 30 requests/hour and 100 requests/day. Not checked No authenticated probe has been published. SiliconFlow rate limitsSiliconFlow quick startSiliconFlow models Get API access
Fireworks AI https://api.fireworks.ai/inference/v1 Provider free tier Not required Yes 10 RPM An account with no payment method and no credits is limited to 10 requests per minute across the entire account. Adding a payment method and active credits raises the ceiling to a maximum of 6,000 RPM. Fireworks is pre-paid, so the 10 RPM envelope is the only no-cost path and there is no published daily request cap on it. Not checked No authenticated probe has been published. Fireworks account quotasFireworks serverless rate limitsFireworks pricing Get API access
Z.AI Open Platform https://api.z.ai/api/paas/v4 Provider free tier Not required Yes Dynamic / unknown GLM-4.7-Flash, GLM-4.5-Flash and the vision model GLM-4.6V-Flash are listed at $0 for input, cached input, cached input storage and output on the official pricing table, making them free to call rather than merely discounted. Z.AI does not publish per-model RPM or TPM for these models on the pricing page; the rate limit reference is a separate page and the console is authoritative. Not checked No authenticated probe has been published. Z.AI pricingZ.AI rate limitsOpenAI Python SDK with Z.AI Get API access
Novita AI https://api.novita.ai/openai Metered access Not required Yes Dynamic / unknown No model on the pricing table is priced at zero any more. The two entries this catalog previously listed as Free — inclusionai/Ling-3.0-flash and Mind Lab Macaron V1 Venti — are gone from the table. The cheapest rows are now Llama 3.1 8B Instruct at $0.02 input / $0.05 output per million and BAAI:BGE-M3 embeddings at $0.01 per million. No signup credit is published, and the page states no request-rate numbers. Not checked No authenticated probe has been published. Novita pricingNovita LLM API referenceNovita model library Get API access
Mistral La Plateforme https://api.mistral.ai/v1 Provider free tier Not required Yes Dynamic / unknown Mistral documents a Free mode (also referred to as the Experiment plan) that lets you activate a workspace and generate an API key without paying, but the public documentation does not currently expose a reachable rate-limit page for it: the tier URL listed in Mistral's own documentation index returns 404 as of the check date. The console is the only authoritative source for the Experiment plan's requests-per-second and tokens-per-minute ceilings. Not checked No authenticated probe has been published. Mistral La Plateforme documentationMistral pricingMistral API documentation index Get API access
Alibaba Cloud Model Studio https://dashscope-intl.aliyuncs.com/compatible-mode/v1 Provider free tier Not required Yes Dynamic / unknown Rate limiting is applied at the Alibaba Cloud root account level and aggregates usage across all RAM users, workspaces and API keys under that account. Each model carries its own RPM and TPM limit published in per-model tables, and free quota is granted per model rather than as one account-wide allowance, so no single free-tier number applies. Not checked No authenticated probe has been published. Model Studio rate limitingModel Studio OpenAI compatibilityModel Studio models Get API access
Moonshot AI (Kimi) https://api.moonshot.ai/v1 Metered access Not required Yes Dynamic / unknown The tier table is keyed to cumulative recharge rather than to a standing free allowance: the docs state you must recharge at least $1 before you can start using the platform, and Tier0 then gives 1 concurrent request, 3 RPM, 500K TPM and 1.5M TPD. Tier1 at $10 cumulative gives 50 concurrent, 200 RPM, 2M TPM and unlimited TPD, rising to Tier5 at $3,000. A $5 voucher arrives only once cumulative payments reach $5, and vouchers do not count toward that cumulative total. The docs moved to platform.kimi.ai and the old platform.moonshot.ai paths now 301 there; the page also carries a notice that the tier rules were slated for revision in August because of abusive traffic. Not checked No authenticated probe has been published. Kimi rate limitsKimi API overviewKimi pricing Get API access
Pollinations.AI https://text.pollinations.ai/openai Provider free tier Not required Yes 4 RPM The API docs now publish four tiers by minimum interval between requests: Anonymous, no signup, basic models, one request every 15 seconds; Seed, free registration, standard models, one request every 5 seconds; Flower, paid, advanced models, one request every 3 seconds; Nectar, enterprise, all models, no limit. The RPM column here is the Anonymous tier expressed per minute (one per 15s = 4/min). Since 2025-03-31 free-tier images may carry a watermark, and removing it with the nologo parameter needs an account. Not checked No authenticated probe has been published. Pollinations API documentationPollinations project repositoryPollinations.AI home Get API access
Ollama Cloud https://ollama.com/v1 Provider free tier Not required Yes Dynamic / unknown Cloud models run on Ollama's servers while keeping the local CLI workflow, and require only an ollama.com account to start. The Cloud documentation page describes access and model retirement policy but does not publish hourly or daily request limits, so the account page is the authoritative source for the active quota. Not checked No authenticated probe has been published. Ollama Cloud documentationOllama Cloud API accessOllama model library Get API access
Cerebras Inference https://api.cerebras.ai/v1 Free trial credit Required Yes 5 RPM Free Trial limits are 5 RPM and 30K TPM per model, capped at 1M tokens per hour and 1M tokens per day, for gpt-oss-120b and gemma-4-31b — zai-glm-4.7 is no longer in that table. gemma-4-31b additionally caps images at 2 per request and 4 MB per payload. New accounts receive $5 in credits that expire 30 days after being granted, and Cerebras states plainly that it offers no always-free per-model allowance. Skipping the payment method at sign-up leaves Playground and API access inactive. Not checked No authenticated probe has been published. Cerebras rate limitsCerebras pricingCerebras OpenAI compatibility Get API access
Vercel AI Gateway https://ai-gateway.vercel.sh/v1 Free trial credit Not required Yes Dynamic / unknown Every Vercel team account gets a free tier, but it covers a subset of models rather than the full catalog, and free-tier requests are rate limited per model at lower limits than the paid tier — those per-model numbers are not published. The free credits start counting from your first Gateway request. The pricing page no longer states a dollar figure for the monthly credit: the $5 per month this catalog previously quoted is gone from the page, so no amount is quoted here. Not checked No authenticated probe has been published. AI Gateway pricingAI Gateway getting startedAI Gateway models Get API access
IBM watsonx.ai https://us-south.ml.cloud.ibm.com/ml/v1 Free trial credit Not required No Dynamic / unknown IBM offers a no-cost trial of watsonx.ai alongside the paid Essentials and Standard plans, and the pricing page invites you to start building at no cost. The page states plan pricing per resource unit but does not publish request-rate limits for the trial, so the console is authoritative for the trial's ceilings. Not checked No authenticated probe has been published. watsonx.ai pricingwatsonx.ai API referencewatsonx.ai foundation models Get API access
OpenRouter https://openrouter.ai/api/v1 Free model aggregator Not required Yes 20 RPM · 50 requests/day Model variants whose ID ends in :free are capped at 20 requests/minute regardless of account status. The daily cap depends on lifetime credit purchases: under 10 credits gives 50 requests/day, and 10 or more credits raises it to 1,000 requests/day. OpenRouter governs capacity globally, so extra accounts or extra API keys do not raise these limits. Not checked No authenticated probe has been published. OpenRouter API rate limitsOpenRouter free model variantsOpenRouter quickstart Get API access
GitHub Models https://models.github.ai/inference Retired 2026-07-30 Retired free tier Not required Yes Dynamic / unknown Retired on 2026-07-30. GitHub shut down the playground, the model catalog, the inference API and the BYOK endpoints; the changelog entry dated that day reads “GitHub Models is now retired.” There is no free tier left to quote. GitHub points existing users at Microsoft Foundry for a broad model catalog and at GitHub Copilot for AI workflows. Not checked No authenticated probe has been published. GitHub Models is being fully retired on July 30, 2026 New access closed
Together AI https://api.together.xyz/v1 Metered access Not required Yes Dynamic / unknown Together applies dynamic per-model rate limits that rise with sustained successful traffic and fall when traffic drops, and states plainly that there are no fixed per-model limits published. Requests above your dynamic rate return 429 with x-ratelimit-reset; requests at or below it that still fail return 503. No standing free allowance is documented on the rate-limits page. Not checked No authenticated probe has been published. Together serverless rate limitsTogether OpenAI compatibilityTogether pricing Get API access
Nebius Token Factory https://api.tokenfactory.nebius.com/v1 Metered access Not required Yes Dynamic / unknown Limits are dynamic and the docs do not publish an account default: they tell you to read the Rate Limits page inside Token Factory. The 60 RPM / 400,000 TPM pair this catalog previously carried as a baseline appears on the page only inside a table labelled a visual example, so it is not quoted as a limit here. The published rules that do apply: usage is evaluated in rolling 15-minute windows, averaging at or above 80% of the current limit raises it by 20% for the next window, averaging at or below 50% divides it by 1.5, and the ceiling is 20x the base allocation before an Enterprise plan is required. No standing free allowance is documented. Not checked No authenticated probe has been published. Nebius rate limits and scalingNebius Token Factory quickstartNebius billing and consumption Get API access
Perplexity API https://api.perplexity.ai Metered access Not required Yes 50 RPM Usage tiers are set by cumulative API credit purchases and never downgrade. Tier 0 ($0 purchased) allows 1 query/second and 50 requests/minute on the Agent API; Tier 1 ($50+) allows 3 QPS and 150/min, rising to 33 QPS and 2,000/min at Tier 4. The Search API is separately limited to 50 requests/second at every tier. Tier 0 sets a rate ceiling but is not a free token grant, so credits are still required to make calls. Not checked No authenticated probe has been published. Perplexity rate limits and usage tiersPerplexity pricingPerplexity OpenAI compatibility Get API access
DeepInfra https://api.deepinfra.com/v1/openai Metered access Not required Yes Dynamic / unknown The official pricing page lists a per-million-token input and output price for every serving model and does not describe a free allowance or publish request-rate limits. Treat DeepInfra as pay-as-you-go and check the dashboard for any promotional credit attached to a new account. Not checked No authenticated probe has been published. DeepInfra pricingDeepInfra OpenAI compatibilityDeepInfra models Get API access
Chutes https://llm.chutes.ai/v1 Metered access Not required Yes Dynamic / unknown Chutes prices inference per token and sells subscription plans that bundle a daily quota: Plus at $10/month with a bundled daily quota and 6% off pay-as-you-go rates beyond it, and Pro at $20/month with a larger daily quota and 10% off. The pricing page documents no zero-cost tier, and per-plan request-rate numbers are shown on the plan limits page rather than in pricing. Not checked No authenticated probe has been published. Chutes pricingChutes documentationChutes app Get API access
Scaleway Generative APIs https://api.scaleway.ai/v1 Metered access Required Yes Dynamic / unknown Every model served through Generative APIs - Serverless is limited by tokens per minute, queries per minute and concurrent requests. Scaleway states that base limits apply only if you have registered a valid payment method, and that they increase automatically if you also verify your identity. The exact numbers live in the Organization quotas page rather than in the rate-limits documentation. Not checked No authenticated probe has been published. Scaleway Generative APIs rate limitsScaleway Generative APIs quickstartScaleway Generative APIs pricing Get API access

Copy, configure, run

One SDK, a different Base URL

Use your own Groq API key from an environment variable. The same pattern works for other OpenAI-compatible providers in the directory.

Generate a coding-agent config
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["GROQ_API_KEY"],
    base_url="https://api.groq.com/openai/v1",
)

response = client.chat.completions.create(
    model="llama-3.3-70b-versatile",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

Why trust this directory

Facts are useful only while they stay traceable

Official sources

Every limit and lifecycle claim links to the provider page it came from.

Review dates

Each record carries the date its sources were last checked.

Honest boundaries

Trials and metered access are separated; a sampled request is never presented as uptime.

Keep the list current

Found a stale limit or missing provider?

Open a correction with the official source and the date you checked it. English or Simplified Chinese is welcome. Read the contribution guide.

No source at hand? Name the provider that changed on you and leave the page for somebody else to find: one line, in a form with one field.

Related reading: field notes on what AI coding agents cost and where they break.

Need one stable hosted endpoint?