Alibaba · 9 providers in this catalog

Free Qwen API access

Every provider in this catalog that lists a Qwen model on free terms, with the limit each one publishes.

What the Qwen family is

Alibaba's Qwen family covers general chat and coding variants, and it is the family most consistently offered by providers headquartered in mainland China.

Providers offering Qwen on free terms

9 of the 25 providers detailed in this catalog list at least one Qwen model. Because Alibaba publishes the weights rather than the service, the model id can be identical across all of them while the free allowance is not.

ProviderFree accessPublished limitsCardMatching models
Cloudflare Workers AI Provider free tier Published in compute units Not required @cf/qwen/qwen2.5-coder-32b-instruct
SiliconFlow Provider free tier 1000 requests per minute Not required Qwen/Qwen3-8B
Alibaba Cloud Model Studio Provider free tier Published per model Not required qwen-plus, qwen-turbo, qwen3-coder-plus
Ollama Cloud Provider free tier Enforced but not published Not required qwen3-coder:480b-cloud
Together AI Metered access Dynamic, no fixed numbers Not required Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo
Nebius Token Factory Metered access 60 requests per minute Not required Qwen/Qwen3-235B-A22B
DeepInfra Metered access Enforced but not published Not required Qwen/Qwen3-Next-80B-A3B-Instruct
Chutes Metered access Published per paid plan Not required Qwen/Qwen3-32B
Scaleway Generative APIs Metered access Enforced but not published Required qwen3-coder-30b-a3b-instruct

How the free terms actually differ

  • Cloudflare Workers AI — Both the Workers Free and Workers Paid plans include 10,000 Neurons per day at no charge, resetting daily at 00:00 UTC.
  • SiliconFlow — Rate limits for free models are fixed; paid models are tiered by monthly spend and start at tier L0 with 1,000 RPM and 40,000 TPM.
  • Alibaba Cloud Model Studio — Rate limiting is applied at the Alibaba Cloud root account level and aggregates usage across all RAM users, workspaces and API keys under that account.
  • Ollama Cloud — Cloud models run on Ollama's servers while keeping the local CLI workflow, and require only an ollama.com account to start.
  • Together AI — Together applies dynamic per-model rate limits that rise with sustained successful traffic and fall when traffic drops, and states plainly that there are no fixed per-model limits published.
  • Nebius Token Factory — Limits are dynamic with a published baseline of 60 RPM and 400,000 TPM.
  • DeepInfra — The official pricing page lists a per-million-token input and output price for every serving model and does not describe a free allowance or publish request-rate limits.
  • Chutes — Chutes prices inference per token and sells subscription plans that bundle a daily quota: Plus at $10/month with a bundled daily quota and 6% off pay-as-you-go rates beyond it, and Pro at $20/month with a larger daily quota and 10% off.
  • Scaleway Generative APIs — Every model served through Generative APIs - Serverless is limited by tokens per minute, queries per minute and concurrent requests.

Picking one

2 of these 9 publish a fixed request count; the rest state a tier or a credit balance instead, which means the only honest answer to "how much is free" is the one in their console. 8 take no credit card. 7 answer a cross-origin browser request, so a key for those can be checked from the browser checker with nothing installed.

A model id that matches Qwen is not a promise that the free tier includes it. Providers move individual models between paid and free without changing the id, which is why each row links to the provider page carrying that provider's own wording and the date it was read.