← All providers

Qwen API Pricing, Models & Rate Limits (2026)

Qwen is Alibaba Cloud's model family, served directly from Alibaba Cloud's Model Studio (DashScope) platform rather than through a third-party host. The Max tier targets frontier-class reasoning and long context; Plus trades a little quality for a lower price.

Also known as: Alibaba Cloud, Model Studio, DashScope.

How much does the Qwen API cost?

Qwen API pricing is usage-based and varies by model and token volume. Use the current model table and provider documentation to estimate a workload before committing to production.

Verified 2026-08-14 source

For cross-provider quota units and fixed-workload capacity, see the LLM API rate-limit comparison; this page remains the authoritative owner for Qwen provider facts.

Qwen product vs API

Qwen consumer access and API billing are separate surfaces; check the provider documentation for current account terms.

Three decisions unique to Qwen

Qwen current-model price mechanics

Current modelInputCached inputOutputBatchVerified
Qwen 3.7 Plus$0.800/MUnavailable — no model cache rate$2.000/MUnavailable2026-07-23
Qwen 3.8 Max$1.600/MUnavailable — no model cache rate$6.400/MUnavailable2026-07-10
Qwen 3.7 Max$1.600/MUnavailable — no model cache rate$6.400/MUnavailable2026-07-23

Messages API vs OpenAI compatibility map

ChoiceDecision ruleEvidence
BillingChatGPT plan never includes API creditsSeparate metered API account
Input/cached/outputToken prices are model rowsUse calculator for workload totals
Limits/authPer-model QPS/QPM caps by account tier · Bearer API keyVerify before production

Adoption map: what is documented versus unavailable

DimensionRecorded valueDecision consequence
Try Qwen side by side →

Verified 2026-08-14. dated provider pricing/source

Current models
3
Legacy models
0
Price range /M
$1.10–$2.80
Max context
256K
Median tok/s
49
Next retirement

Qwen model pricing

Compare Qwen models by input, output, and blended token cost below.

ModelInput /MOutput /MBlended /M
Qwen 3.7 Plus$0.80$2.00$1.10
Qwen 3.8 Max$1.60$6.40$2.80
Qwen 3.7 Max$1.60$6.40$2.80

Pricing values are registry-backed and were most recently verified on 2026-07-23. Sources: https://www.alibabacloud.com/help/en/model-studio/pricing. Model detail pages preserve each model's own title and verification date.

* Blended comparison assumes 3 input tokens for every output token; it is not the provider's billing unit.

Speed

Fastest measured Qwen model is Qwen 3.7 Plus at 84 tokens/sec (340ms TTFT), median across measured Qwen models is 49 tokens/sec. See the full speed benchmark methodology.

Best for

Muse Spark 1.3 Contributor is our pick for CodingMuse Spark 1.3 Contributor is our pick for Structured Data ExtractionMuse Spark 1.3 Contributor is our pick for Writing & ContentGLM-5.2 is our pick for Math & ReasoningMuse Spark 1.3 Contributor is our pick for Agents & Tool Use
What it will cost →
Qwen's 3 priced models, ranked by verbosity-adjusted monthly cost, not list rate.

Related Qwen pages

Qwen alternatives →All LLM API pricing →Qwen speed benchmarks →Qwen cost calculator →

Build with Qwen

Qwen rate limits →

Qwen implementation details

Verified 2026-08-14 against source.

Qwen publishes its current authentication, limits, and data-handling details in the linked documentation.

OpenAI-compatiblePartial
API base URLhttps://dashscope-intl.aliyuncs.com/compatible-mode/v1
Auth modelBearer API key
Prompt cachingNot documented
Batch discountNot documented
Free tierFree quota for new Alibaba Cloud accounts
Free-tier limitsQuota is limited to eligible new accounts and varies by model, region, and account.
Free-tier expiryNot published
Rate-limit modelPer-model QPS/QPM caps by account tier
Data residencySingapore/international region via DashScope Intl; mainland China served from a separate region
Trains on API dataNot documented
SLA publishedNo
DocsOfficial pricingFree-tier terms

Switching to and from Qwen

The closest parity-aware alternative to Qwen 3.8 Max ($2.80/M) outside Qwen is Grok 4.6 ($3.00/M, +7.1%) — a config migration. Biggest gap: loses documented data-residency options.
The closest parity-aware alternative to Qwen 3.7 Max ($2.80/M) outside Qwen is Grok 4.6 ($3.00/M, +7.1%) — a config migration. Biggest gap: loses documented data-residency options.
Full Qwen alternatives comparison →

Calling Qwen through All AI Ask

Calling Qwen directly means adapting client code away from the plain OpenAI request shape — our gateway removes that: every model, including Qwen's, is called the same way.

curl https://allaiask.com/api/v1/chat \
  -H "Authorization: Bearer $ALLAIASK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen3.7-plus", "messages": [{"role": "user", "content": "Hello"}]}'

FAQ

Is Qwen OpenAI-compatible?

Partially. Qwen publishes an OpenAI-compatible mode for some endpoints, but not a full 1:1 replacement for its native API — check https://www.alibabacloud.com/help/en/model-studio/models before relying on it for every feature you use.

Does Qwen support prompt caching?

Not documented as of 2026-08-14 — we did not find a published prompt-caching feature for Qwen. If that changes, this page updates.

Does Qwen have a free tier?

Yes — Free quota for new Alibaba Cloud accounts. Quota is limited to eligible new accounts and varies by model, region, and account.

How much does the Qwen API cost?

Current Qwen models range from $1.10 to $2.80 per million blended tokens (3:1 input:output). Full per-model pricing is in the table below.

Where is Qwen API data hosted?

Singapore/international region via DashScope Intl; mainland China served from a separate region

Batch 48 · qwen provider adoption evidence. Surface verification: 2026-08-14. Missing joins are deliberately published as Unavailable.

Qwen product–endpoint–entitlement map

Frozen Batch 48 qwen fixture: DashScope international, mainland, resource-package, and unsupported credential surfaces. First-party evidence checked 2026-08-14; unresolved joins fail closed. Formula / decision rule: eligible = product ∧ endpoint ∧ key class ∧ allowance owner ∧ model visibility

Frozen fixtureInputs and observationFormula / boundaryState
batch48-qwen-m1-r1
general pay-as-you-go · international
product=Model Studio; endpoint=dashscope-intl; protocol=OpenAI-compatible; key=API; balance=prepaid; workload=chat; model=Qwen-Max
The international general endpoint and prepaid balance join; the model is visible to this credential.
join=product+endpoint+key+balance+model → PASS
Do not transfer this entitlement to mainland credentials.
PASS — general surface qualified.
batch48-qwen-m1-r2
coding plan · resource package
product=coding plan/resource package; endpoint=coding.aliyun; protocol=agent; key=subscription; allowance=coding; model=Unavailable
Coding allowance is owned by the subscription surface; no general API model visibility is asserted.
visible model = key class ∧ product entitlement ∧ endpoint join
Subscription allowance is not prepaid API balance.
UNAVAILABLE — model visibility not joined.
batch48-qwen-m1-r3
Anthropic-compatible · alternate region · unsupported pairing
product=Model Studio; endpoint=anthropic-compatible; region=mainland; key=international; correction=region/key review
The endpoint exists, but the international key and mainland region do not form a supported pair in the fixture.
eligible only when region(key)=region(endpoint) ∧ product owns protocol
A compatible protocol cannot waive regional entitlement.
FAIL CLOSED — correct region and credential.

Provenance: Frozen Batch 48 qwen fixture: module 1 observation and calculation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Surface verification date: 2026-08-14. Alibaba Cloud Model Studio docs (first-party source)

Qwen model-mode contract registry

Frozen Batch 48 qwen fixture: General, reasoning, coding-agent, vision, long-output, and tool-heavy mode fixtures. First-party evidence checked 2026-08-14; unresolved joins fail closed. Formula / decision rule: mode pass = endpoint/model/version ∧ accepted mode = effective mode ∧ capability evidence

Frozen fixtureInputs and observationFormula / boundaryState
batch48-qwen-m2-r1
general chat · always-on reasoning
endpoint=chat; model=Qwen3-Max; version=2026-08; requested=reasoning; effective=reasoning; context=262K; output=32K; tools=joined
The frozen general model returns the requested reasoning mode with joined context and tool evidence.
pass = requested mode = effective mode ∧ context/output/tool joins
This observation is not evidence for a later alias revision.
PASS — revision-scoped contract.
batch48-qwen-m2-r2
optional thinking · coding agent · vision
endpoint=coding; model=Qwen3-Coder; version=2026-07; thinking=optional; vision=image hash joined; tools=schema joined
Optional thinking and image/tool inputs are observed on the coding endpoint; account entitlement remains required.
effective capability = submitted control ∧ endpoint support ∧ account entitlement
Vision support cannot be copied to a text-only revision.
PASS WITH GATE — account entitlement required.
batch48-qwen-m2-r3
long output · tool-heavy
endpoint=chat; model=Qwen-Max; output request=48K; tools=6; schema=sha256:5d20; stream=joined; lifecycle=Unavailable
The schema and stream join, but lifecycle status for the alias is not published in this fixture.
promotion requires output evidence ∧ all tool IDs ∧ lifecycle state
Do not silently transfer limits or modes across GLM/Qwen hosts.
UNAVAILABLE — lifecycle gate missing.

Provenance: Frozen Batch 48 qwen fixture: module 2 observation and calculation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Surface verification date: 2026-08-14. Alibaba Cloud Model Studio docs (first-party source)

Qwen subscription-versus-general-API isolation canary

Frozen Batch 48 qwen fixture: Same-key, wrong-key, balance, endpoint-swap, concurrency, and coding-surface isolation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Formula / decision rule: route = product ∧ credential class ∧ endpoint ∧ debit source; missing join → Unavailable

Frozen fixtureInputs and observationFormula / boundaryState
batch48-qwen-m3-r1
same key · endpoint swap
account=acct-q1; key hash=sha256:91a; general endpoint=joined; coding endpoint=joined; model=Qwen3; debit=prepaid
The same account can call the general surface, but debit ownership changes at the coding endpoint.
safe retry = same product ∧ same debit source ∧ idempotent request
A successful general request does not prove coding access.
PASS WITH SPLIT — preserve debit owner.
batch48-qwen-m3-r2
exhausted coding allowance · exhausted prepaid balance
coding error=allowance_exhausted; general error=balance_exhausted; request IDs=req-c7/req-g7; retry=not transient
The two failures have distinct product owners and neither should be retried as a network throttle.
retry eligible = transient ∧ debit not exhausted ∧ request idempotent
Never route a billing/allowance error into a generic retry loop.
FAIL CLOSED — replenish the named owner.
batch48-qwen-m3-r3
concurrent coding sessions · general workload on coding surface
sessions=4; concurrency cap=Unavailable; coding key hash=sha256:44b; general workload=misrouted; correction=general endpoint
Concurrency capacity is not published and the general workload is on the wrong product surface.
admit only if cap known ∧ workload surface = credential product
Unknown concurrency is neither zero nor unlimited.
UNAVAILABLE — route correction and live cap required.

Provenance: Frozen Batch 48 qwen fixture: module 3 observation and calculation. First-party evidence checked 2026-08-14; unresolved joins fail closed. Surface verification date: 2026-08-14. Alibaba Cloud Model Studio docs (first-party source)

Try qwen through the Batch 48 route →

Try Qwen for free

Run real prompts against every current Qwen model, and every other provider on this site, in one workspace.

Try It Free