Qwen API Pricing, Models & Rate Limits (2026)
Qwen is Alibaba Cloud's model family, served directly from Alibaba Cloud's Model Studio (DashScope) platform rather than through a third-party host. The Max tier targets frontier-class reasoning and long context; Plus trades a little quality for a lower price.
Also known as: Alibaba Cloud, Model Studio, DashScope.
How much does the Qwen API cost?
Qwen API pricing is usage-based and varies by model and token volume. Use the current model table and provider documentation to estimate a workload before committing to production.
For cross-provider quota units and fixed-workload capacity, see the LLM API rate-limit comparison; this page remains the authoritative owner for Qwen provider facts.
Qwen product vs API
Qwen consumer access and API billing are separate surfaces; check the provider documentation for current account terms.
Three decisions unique to Qwen
Qwen current-model price mechanics
| Current model | Input | Cached input | Output | Batch | Verified |
|---|
| Qwen 3.7 Plus | $0.800/M | Unavailable — no model cache rate | $2.000/M | Unavailable | 2026-07-23 |
|---|
| Qwen 3.8 Max | $1.600/M | Unavailable — no model cache rate | $6.400/M | Unavailable | 2026-07-10 |
|---|
| Qwen 3.7 Max | $1.600/M | Unavailable — no model cache rate | $6.400/M | Unavailable | 2026-07-23 |
|---|
Messages API vs OpenAI compatibility map
| Choice | Decision rule | Evidence |
|---|
| Billing | ChatGPT plan never includes API credits | Separate metered API account |
|---|
| Input/cached/output | Token prices are model rows | Use calculator for workload totals |
|---|
| Limits/auth | Per-model QPS/QPM caps by account tier · Bearer API key | Verify before production |
|---|
Adoption map: what is documented versus unavailable
| Dimension | Recorded value | Decision consequence |
|---|
Try Qwen side by side →Verified 2026-08-14. dated provider pricing/source →
Price range /M
$1.10–$2.80
Qwen model pricing
Compare Qwen models by input, output, and blended token cost below.
Pricing values are registry-backed and were most recently verified on 2026-07-23. Sources: https://www.alibabacloud.com/help/en/model-studio/pricing. Model detail pages preserve each model's own title and verification date.
* Blended comparison assumes 3 input tokens for every output token; it is not the provider's billing unit.
Speed
Best for
Build with Qwen
Qwen implementation details
Verified 2026-08-14 against source.
Qwen publishes its current authentication, limits, and data-handling details in the linked documentation.
| OpenAI-compatible | Partial |
| API base URL | https://dashscope-intl.aliyuncs.com/compatible-mode/v1 |
| Auth model | Bearer API key |
| Prompt caching | Not documented |
| Batch discount | Not documented |
| Free tier | Free quota for new Alibaba Cloud accounts |
| Free-tier limits | Quota is limited to eligible new accounts and varies by model, region, and account. |
| Free-tier expiry | Not published |
| Rate-limit model | Per-model QPS/QPM caps by account tier |
| Data residency | Singapore/international region via DashScope Intl; mainland China served from a separate region |
| Trains on API data | Not documented |
| SLA published | No |
Switching to and from Qwen
The closest parity-aware alternative to
Qwen 3.8 Max ($2.80/M) outside Qwen is
Grok 4.6 ($3.00/M, +7.1%) — a
config migration. Biggest gap: loses documented data-residency options.
The closest parity-aware alternative to
Qwen 3.7 Max ($2.80/M) outside Qwen is
Grok 4.6 ($3.00/M, +7.1%) — a
config migration. Biggest gap: loses documented data-residency options.
Full Qwen alternatives comparison →Calling Qwen through All AI Ask
Calling Qwen directly means adapting client code away from the plain OpenAI request shape — our gateway removes that: every model, including Qwen's, is called the same way.
curl https://allaiask.com/api/v1/chat \
-H "Authorization: Bearer $ALLAIASK_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "qwen3.7-plus", "messages": [{"role": "user", "content": "Hello"}]}'FAQ
Is Qwen OpenAI-compatible?
Partially. Qwen publishes an OpenAI-compatible mode for some endpoints, but not a full 1:1 replacement for its native API — check https://www.alibabacloud.com/help/en/model-studio/models before relying on it for every feature you use.
Does Qwen support prompt caching?
Not documented as of 2026-08-14 — we did not find a published prompt-caching feature for Qwen. If that changes, this page updates.
Does Qwen have a free tier?
Yes — Free quota for new Alibaba Cloud accounts. Quota is limited to eligible new accounts and varies by model, region, and account.
How much does the Qwen API cost?
Current Qwen models range from $1.10 to $2.80 per million blended tokens (3:1 input:output). Full per-model pricing is in the table below.
Where is Qwen API data hosted?
Singapore/international region via DashScope Intl; mainland China served from a separate region
Batch 48 · qwen provider adoption evidence. Surface verification: 2026-08-14. Missing joins are deliberately published as Unavailable.
Qwen product–endpoint–entitlement map
Frozen Batch 48 qwen fixture: DashScope international, mainland, resource-package, and unsupported credential surfaces. First-party evidence checked 2026-08-14; unresolved joins fail closed. Formula / decision rule: eligible = product ∧ endpoint ∧ key class ∧ allowance owner ∧ model visibility
| Frozen fixture | Inputs and observation | Formula / boundary | State |
|---|
batch48-qwen-m1-r1 general pay-as-you-go · international | product=Model Studio; endpoint=dashscope-intl; protocol=OpenAI-compatible; key=API; balance=prepaid; workload=chat; model=Qwen-Max The international general endpoint and prepaid balance join; the model is visible to this credential. | join=product+endpoint+key+balance+model → PASS Do not transfer this entitlement to mainland credentials. | PASS — general surface qualified. |
batch48-qwen-m1-r2 coding plan · resource package | product=coding plan/resource package; endpoint=coding.aliyun; protocol=agent; key=subscription; allowance=coding; model=Unavailable Coding allowance is owned by the subscription surface; no general API model visibility is asserted. | visible model = key class ∧ product entitlement ∧ endpoint join Subscription allowance is not prepaid API balance. | UNAVAILABLE — model visibility not joined. |
batch48-qwen-m1-r3 Anthropic-compatible · alternate region · unsupported pairing | product=Model Studio; endpoint=anthropic-compatible; region=mainland; key=international; correction=region/key review The endpoint exists, but the international key and mainland region do not form a supported pair in the fixture. | eligible only when region(key)=region(endpoint) ∧ product owns protocol A compatible protocol cannot waive regional entitlement. | FAIL CLOSED — correct region and credential. |
Provenance: Frozen Batch 48 qwen fixture: module 1 observation and calculation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Surface verification date: 2026-08-14. Alibaba Cloud Model Studio docs (first-party source)
Qwen model-mode contract registry
Frozen Batch 48 qwen fixture: General, reasoning, coding-agent, vision, long-output, and tool-heavy mode fixtures. First-party evidence checked 2026-08-14; unresolved joins fail closed. Formula / decision rule: mode pass = endpoint/model/version ∧ accepted mode = effective mode ∧ capability evidence
| Frozen fixture | Inputs and observation | Formula / boundary | State |
|---|
batch48-qwen-m2-r1 general chat · always-on reasoning | endpoint=chat; model=Qwen3-Max; version=2026-08; requested=reasoning; effective=reasoning; context=262K; output=32K; tools=joined The frozen general model returns the requested reasoning mode with joined context and tool evidence. | pass = requested mode = effective mode ∧ context/output/tool joins This observation is not evidence for a later alias revision. | PASS — revision-scoped contract. |
batch48-qwen-m2-r2 optional thinking · coding agent · vision | endpoint=coding; model=Qwen3-Coder; version=2026-07; thinking=optional; vision=image hash joined; tools=schema joined Optional thinking and image/tool inputs are observed on the coding endpoint; account entitlement remains required. | effective capability = submitted control ∧ endpoint support ∧ account entitlement Vision support cannot be copied to a text-only revision. | PASS WITH GATE — account entitlement required. |
batch48-qwen-m2-r3 long output · tool-heavy | endpoint=chat; model=Qwen-Max; output request=48K; tools=6; schema=sha256:5d20; stream=joined; lifecycle=Unavailable The schema and stream join, but lifecycle status for the alias is not published in this fixture. | promotion requires output evidence ∧ all tool IDs ∧ lifecycle state Do not silently transfer limits or modes across GLM/Qwen hosts. | UNAVAILABLE — lifecycle gate missing. |
Provenance: Frozen Batch 48 qwen fixture: module 2 observation and calculation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Surface verification date: 2026-08-14. Alibaba Cloud Model Studio docs (first-party source)
Qwen subscription-versus-general-API isolation canary
Frozen Batch 48 qwen fixture: Same-key, wrong-key, balance, endpoint-swap, concurrency, and coding-surface isolation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Formula / decision rule: route = product ∧ credential class ∧ endpoint ∧ debit source; missing join → Unavailable
| Frozen fixture | Inputs and observation | Formula / boundary | State |
|---|
batch48-qwen-m3-r1 same key · endpoint swap | account=acct-q1; key hash=sha256:91a; general endpoint=joined; coding endpoint=joined; model=Qwen3; debit=prepaid The same account can call the general surface, but debit ownership changes at the coding endpoint. | safe retry = same product ∧ same debit source ∧ idempotent request A successful general request does not prove coding access. | PASS WITH SPLIT — preserve debit owner. |
batch48-qwen-m3-r2 exhausted coding allowance · exhausted prepaid balance | coding error=allowance_exhausted; general error=balance_exhausted; request IDs=req-c7/req-g7; retry=not transient The two failures have distinct product owners and neither should be retried as a network throttle. | retry eligible = transient ∧ debit not exhausted ∧ request idempotent Never route a billing/allowance error into a generic retry loop. | FAIL CLOSED — replenish the named owner. |
batch48-qwen-m3-r3 concurrent coding sessions · general workload on coding surface | sessions=4; concurrency cap=Unavailable; coding key hash=sha256:44b; general workload=misrouted; correction=general endpoint Concurrency capacity is not published and the general workload is on the wrong product surface. | admit only if cap known ∧ workload surface = credential product Unknown concurrency is neither zero nor unlimited. | UNAVAILABLE — route correction and live cap required. |
Provenance: Frozen Batch 48 qwen fixture: module 3 observation and calculation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Surface verification date: 2026-08-14. Alibaba Cloud Model Studio docs (first-party source)
Try qwen through the Batch 48 route →