Verified provider facts · Reviewed 2026-07-25

Free LLM APIs with direct API key links

Compare permanent free tiers, no-card options, models and published limits. Every claim links to an official source.

15permanent free tiers
15need no credit card
15free tiers OpenAI compatible
2026-07-25sources last reviewed

Data-driven shortcuts

Strong starting points

Selected by published catalog facts, never by sponsorship.

Catalog

Compare the full directory

Permanent free tiers, aggregators, trials and metered access stay visibly separate.

26 matches

ProviderFree accessCardOpenAILimitsSample probeSources checkedAccess
Google Gemini API https://generativelanguage.googleapis.com/v1beta/openai/ Provider free tier Not required Yes Dynamic / unknown Rate limits apply per project, not per API key, and depend on the project's usage tier; the Free tier is the entry tier and moving up requires enabling billing. Google publishes the tier structure but directs you to AI Studio for the exact RPM, TPM and RPD active on your project, so no single free-tier number can be quoted here. Not checked No authenticated probe has been published. Gemini API rate limitsGemini API billingGemini OpenAI compatibility Get API access
GroqCloud https://api.groq.com/openai/v1 Provider free tier Not required Yes 30 RPM · 1000 requests/day Free plan limits are per model and enforced per organization. llama-3.3-70b-versatile: 30 RPM, 1,000 RPD, 12K TPM, 100K TPD. llama-3.1-8b-instant: 30 RPM, 14,400 RPD, 6K TPM, 500K TPD. openai/gpt-oss-120b and openai/gpt-oss-20b: 30 RPM, 1,000 RPD, 8K TPM, 200K TPD. groq/compound: 30 RPM, 250 RPD, 70K TPM. Cached tokens do not count toward token limits. The RPM/RPD columns here quote llama-3.3-70b-versatile. Not checked No authenticated probe has been published. GroqCloud rate limitsGroq OpenAI compatibilityGroq billing FAQ Get API access
SambaNova Cloud https://api.sambanova.ai/v1 Provider free tier Not required Yes 20 RPM · 20 requests/day The Free Tier applies whenever no payment method is linked to the account. For DeepSeek-V3.1, Meta-Llama-3.3-70B-Instruct and gpt-oss-120b the free limits are 20 RPM, 20 RPD and 200,000 tokens per day. Linking a payment method moves the account to the Developer Tier (60-240 RPM, 12,000-48,000 RPD, 20M tokens/day across all models). Not checked No authenticated probe has been published. SambaNova model rate limitsSambaNova developer guideSambaNova OpenAI compatibility Get API access
Cohere https://api.cohere.ai/compatibility/v1 Provider free tier Not required Yes 20 RPM Trial keys are free, rate limited and not licensed for production. Chat is limited to 20 requests/minute, Rerank to 10/minute, Audio Transcriptions and EmbedJob to 5/minute, Tokenize to 100/minute, and Embed to 2,000 inputs/minute. Every trial key is additionally capped at 1,000 API calls per month across the account. Not checked No authenticated probe has been published. Cohere API rate limitsCohere Compatibility APICohere pricing Get API access
Cloudflare Workers AI https://api.cloudflare.com/client/v4/accounts/ACCOUNT_ID/ai/v1 Provider free tier Not required Yes Dynamic / unknown Both the Workers Free and Workers Paid plans include 10,000 Neurons per day at no charge, resetting daily at 00:00 UTC. Neurons are a compute unit rather than a request count, so the number of requests you get depends on the model and prompt size. On the Free plan there is no overage: once the allocation is spent, requests fail until reset. Workers Paid bills overage at $0.011 per 1,000 Neurons. Not checked No authenticated probe has been published. Workers AI pricingWorkers AI OpenAI compatible endpointsWorkers AI models Get API access
Hugging Face Inference Providers https://router.huggingface.co/v1 Provider free tier Not required Yes Dynamic / unknown Signed-in free users receive $0.10 in monthly Inference Provider credits (Hugging Face notes this is subject to change); PRO users receive $2.00 and Team or Enterprise organizations $2.00 per seat. Credits only apply to requests routed by Hugging Face, not to requests made with your own provider key. Once the monthly credits are spent you must purchase credits to continue. Not checked No authenticated probe has been published. Inference Providers pricing and billingInference Providers OpenAI compatibilityInference Providers settings Get API access
SiliconFlow https://api.siliconflow.com/v1 Provider free tier Not required Yes 1000 RPM Rate limits for free models are fixed; paid models are tiered by monthly spend and start at tier L0 with 1,000 RPM and 40,000 TPM. Limits are enforced per user account rather than per API key, and each model is limited separately. deepseek-ai/DeepSeek-R1 and deepseek-ai/DeepSeek-V3 carry an extra cap of 30 requests/hour and 100 requests/day. Not checked No authenticated probe has been published. SiliconFlow rate limitsSiliconFlow quick startSiliconFlow models Get API access
Fireworks AI https://api.fireworks.ai/inference/v1 Provider free tier Not required Yes 10 RPM An account with no payment method and no credits is limited to 10 requests per minute across the entire account. Adding a payment method and active credits raises the ceiling to a maximum of 6,000 RPM. Fireworks is pre-paid, so the 10 RPM envelope is the only no-cost path and there is no published daily request cap on it. Not checked No authenticated probe has been published. Fireworks account quotasFireworks serverless rate limitsFireworks pricing Get API access
Z.AI Open Platform https://api.z.ai/api/paas/v4 Provider free tier Not required Yes Dynamic / unknown GLM-4.7-Flash and GLM-4.5-Flash are listed at $0 for input, cached input and output on the official pricing table, making them free to call rather than merely discounted. Z.AI does not publish per-model RPM or TPM for these models on the pricing page; the rate limit reference is a separate page and the console is authoritative. Not checked No authenticated probe has been published. Z.AI pricingZ.AI rate limitsOpenAI Python SDK with Z.AI Get API access
Novita AI https://api.novita.ai/openai Provider free tier Not required Yes Dynamic / unknown The official pricing table lists inclusionai/Ling-3.0-flash and Mind Lab's Macaron V1 Venti with both input and output priced Free; every other listed model is billed per million tokens. Novita does not publish RPM or RPD for the free models on the pricing page. Not checked No authenticated probe has been published. Novita pricingNovita LLM API referenceNovita model library Get API access
Mistral La Plateforme https://api.mistral.ai/v1 Provider free tier Not required Yes Dynamic / unknown Mistral documents a Free mode (also referred to as the Experiment plan) that lets you activate a workspace and generate an API key without paying, but the public documentation does not currently expose a reachable rate-limit page for it: the tier URL listed in Mistral's own documentation index returns 404 as of the check date. The console is the only authoritative source for the Experiment plan's requests-per-second and tokens-per-minute ceilings. Not checked No authenticated probe has been published. Mistral La Plateforme documentationMistral pricingMistral API documentation index Get API access
Alibaba Cloud Model Studio https://dashscope-intl.aliyuncs.com/compatible-mode/v1 Provider free tier Not required Yes Dynamic / unknown Rate limiting is applied at the Alibaba Cloud root account level and aggregates usage across all RAM users, workspaces and API keys under that account. Each model carries its own RPM and TPM limit published in per-model tables, and free quota is granted per model rather than as one account-wide allowance, so no single free-tier number applies. Not checked No authenticated probe has been published. Model Studio rate limitingModel Studio OpenAI compatibilityModel Studio models Get API access
Moonshot AI (Kimi) https://api.moonshot.ai/v1 Provider free tier Not required Yes Dynamic / unknown Limits are expressed as concurrency, RPM, TPM and TPD, and the published tier table is keyed to cumulative recharge rather than to a standing free allowance; reaching $5 of cumulative recharge grants a $5 voucher. New accounts start at the lowest tier, so treat the console tier page as authoritative before planning volume. Not checked No authenticated probe has been published. Kimi rate limitsKimi API overviewKimi pricing Get API access
Pollinations.AI https://text.pollinations.ai/openai Provider free tier Not required Yes Dynamic / unknown Pollinations exposes an OpenAI-compatible text endpoint that can be called without a key for basic use, and the official API docs do not publish RPM or RPD figures. The docs do state that since 2025-03-31 free-tier images may include watermarks, so free-tier output is not identical to paid output. Not checked No authenticated probe has been published. Pollinations API documentationPollinations project repositoryPollinations.AI home Get API access
Ollama Cloud https://ollama.com/v1 Provider free tier Not required Yes Dynamic / unknown Cloud models run on Ollama's servers while keeping the local CLI workflow, and require only an ollama.com account to start. The Cloud documentation page describes access and model retirement policy but does not publish hourly or daily request limits, so the account page is the authoritative source for the active quota. Not checked No authenticated probe has been published. Ollama Cloud documentationOllama Cloud API accessOllama model library Get API access
Cerebras Inference https://api.cerebras.ai/v1 Free trial credit Required Yes 5 RPM Free Trial limits are 5 RPM and 30K TPM per model, capped at 1M tokens per hour and 1M tokens per day, for gpt-oss-120b, zai-glm-4.7 and gemma-4-31b. New accounts receive $5 in credits that expire 30 days after being granted. Cerebras states that if you skip adding a payment method at sign-up, Playground and API access remain inactive until you add one. Not checked No authenticated probe has been published. Cerebras rate limitsCerebras pricingCerebras OpenAI compatibility Get API access
Vercel AI Gateway https://ai-gateway.vercel.sh/v1 Free trial credit Not required Yes Dynamic / unknown Every Vercel team account receives $5 of AI Gateway credits per month on the free tier, usable only against the Free Tier model subset rather than the full catalogue. Free tier requests are additionally rate limited per model with lower limits than the paid tier, and Vercel does not publish those per-model numbers. Credits start counting from your first Gateway request. Not checked No authenticated probe has been published. AI Gateway pricingAI Gateway getting startedAI Gateway models Get API access
IBM watsonx.ai https://us-south.ml.cloud.ibm.com/ml/v1 Free trial credit Not required No Dynamic / unknown IBM offers a no-cost trial of watsonx.ai alongside the paid Essentials and Standard plans, and the pricing page invites you to start building at no cost. The page states plan pricing per resource unit but does not publish request-rate limits for the trial, so the console is authoritative for the trial's ceilings. Not checked No authenticated probe has been published. watsonx.ai pricingwatsonx.ai API referencewatsonx.ai foundation models Get API access
OpenRouter https://openrouter.ai/api/v1 Free model aggregator Not required Yes 20 RPM · 50 requests/day Model variants whose ID ends in :free are capped at 20 requests/minute regardless of account status. The daily cap depends on lifetime credit purchases: under 10 credits gives 50 requests/day, and 10 or more credits raises it to 1,000 requests/day. OpenRouter governs capacity globally, so extra accounts or extra API keys do not raise these limits. Not checked No authenticated probe has been published. OpenRouter API rate limitsOpenRouter free model variantsOpenRouter quickstart Get API access
GitHub Models https://models.github.ai/inference Retires 2026-07-30 Retiring free tier Not required Yes Dynamic / unknown Limits vary by model and Copilot plan. New customers are no longer accepted, and the service retires on 2026-07-30. Not checked No authenticated probe has been published. GitHub Models retirement announcementGitHub Models catalog APIGitHub Models billing New access closed
Together AI https://api.together.xyz/v1 Metered access Not required Yes Dynamic / unknown Together applies dynamic per-model rate limits that rise with sustained successful traffic and fall when traffic drops, and states plainly that there are no fixed per-model limits published. Requests above your dynamic rate return 429 with x-ratelimit-reset; requests at or below it that still fail return 503. No standing free allowance is documented on the rate-limits page. Not checked No authenticated probe has been published. Together serverless rate limitsTogether OpenAI compatibilityTogether pricing Get API access
Nebius Token Factory https://api.tokenfactory.nebius.com/v1 Metered access Not required Yes 60 RPM Limits are dynamic with a published baseline of 60 RPM and 400,000 TPM. Usage is evaluated in rolling 15-minute windows: averaging at or above 80% of the current limit raises it by 20% for the next window, averaging at or below 50% divides it by 1.5, and the ceiling is 20x the base allocation before an Enterprise plan is required. No standing free allowance is documented on this page. Not checked No authenticated probe has been published. Nebius rate limits and scalingNebius Token Factory quickstartNebius billing and consumption Get API access
Perplexity API https://api.perplexity.ai Metered access Not required Yes 50 RPM Usage tiers are set by cumulative API credit purchases and never downgrade. Tier 0 ($0 purchased) allows 1 query/second and 50 requests/minute on the Agent API; Tier 1 ($50+) allows 3 QPS and 150/min, rising to 33 QPS and 2,000/min at Tier 4. The Search API is separately limited to 50 requests/second at every tier. Tier 0 sets a rate ceiling but is not a free token grant, so credits are still required to make calls. Not checked No authenticated probe has been published. Perplexity rate limits and usage tiersPerplexity pricingPerplexity OpenAI compatibility Get API access
DeepInfra https://api.deepinfra.com/v1/openai Metered access Not required Yes Dynamic / unknown The official pricing page lists a per-million-token input and output price for every serving model and does not describe a free allowance or publish request-rate limits. Treat DeepInfra as pay-as-you-go and check the dashboard for any promotional credit attached to a new account. Not checked No authenticated probe has been published. DeepInfra pricingDeepInfra OpenAI compatibilityDeepInfra models Get API access
Chutes https://llm.chutes.ai/v1 Metered access Not required Yes Dynamic / unknown Chutes prices inference per token and sells subscription plans that bundle a daily quota: Plus at $10/month with a bundled daily quota and 6% off pay-as-you-go rates beyond it, and Pro at $20/month with a larger daily quota and 10% off. The pricing page documents no zero-cost tier, and per-plan request-rate numbers are shown on the plan limits page rather than in pricing. Not checked No authenticated probe has been published. Chutes pricingChutes documentationChutes app Get API access
Scaleway Generative APIs https://api.scaleway.ai/v1 Metered access Required Yes Dynamic / unknown Every model served through Generative APIs - Serverless is limited by tokens per minute, queries per minute and concurrent requests. Scaleway states that base limits apply only if you have registered a valid payment method, and that they increase automatically if you also verify your identity. The exact numbers live in the Organization quotas page rather than in the rate-limits documentation. Not checked No authenticated probe has been published. Scaleway Generative APIs rate limitsScaleway Generative APIs quickstartScaleway Generative APIs pricing Get API access

Copy, configure, run

One SDK, a different Base URL

Use your own Groq API key from an environment variable. The same pattern works for other OpenAI-compatible providers in the directory.

Generate a coding-agent config
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["GROQ_API_KEY"],
    base_url="https://api.groq.com/openai/v1",
)

response = client.chat.completions.create(
    model="llama-3.3-70b-versatile",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

Why trust this directory

Facts are useful only while they stay traceable

Official sources

Every limit and lifecycle claim links to the provider page it came from.

Review dates

Each record carries the date its sources were last checked.

Honest boundaries

Trials and metered access are separated; a sampled request is never presented as uptime.

Keep the list current

Found a stale limit or missing provider?

Open a correction with the official source and the date you checked it. English or Simplified Chinese is welcome. Read the contribution guide.

Related reading: field notes on what AI coding agents cost and where they break.

Need one stable hosted endpoint?