Qwen API Cost Calculator

How much does the Qwen API cost per month?

For Qwen 3.7 Plus at 200K requests/month, 2,400 input tokens and 350 output tokens per request costs about $524 per month after 30% cache use and 100% batch share. Across Qwen's 3 priced models, the cheapest default ranking is Qwen 3.7 Plus at $524 per month.

Verified 2026-06-21

Pricing data as of June 2026. Sources: Qwen pricing and model documentation. Parameters are shareable in this URL.

Qwen 3.7 Plus: estimated monthly cost

$524

Formula: provider-specific cached input multiplier × cache rate, plus uncached input and verbosity-adjusted output; cache writes are amortised over the provider TTL, then batch savings are applied.

Per request$0.003
Per day$17.47
Per month$524
Per year$6,288

OpenAI scenario sensitivity

The default is a server-rendered estimate. Change the cache and batch shares to see when a cheaper qualified model overtakes the selected model; the URL is shareable and preserves the inputs.

Default$524/month
No caching or batch$524/month
50% cache, 50% batch$524/month
100% cache, 100% batch$524/month

Batch 11 Qwen workload decision depth

1. Plus/Max context eligibility and cost ladder

Input tokensModelCostWindow result
32KQwen 3.7 Plus$0.03Fit only when model context window covers full request; otherwise excluded
32KQwen 3.8 Max$0.05Fit only when model context window covers full request; otherwise excluded
32KQwen 3.7 Max$0.05Fit only when model context window covers full request; otherwise excluded
128KQwen 3.7 Plus$0.10Fit only when model context window covers full request; otherwise excluded
128KQwen 3.8 Max$0.21Fit only when model context window covers full request; otherwise excluded
128KQwen 3.7 Max$0.21Fit only when model context window covers full request; otherwise excluded
500KQwen 3.7 Plus$0.40Fit only when model context window covers full request; otherwise excluded
500KQwen 3.8 Max$0.80Fit only when model context window covers full request; otherwise excluded
500KQwen 3.7 Max$0.80Fit only when model context window covers full request; otherwise excluded
1MQwen 3.7 Plus$0.80Fit only when model context window covers full request; otherwise excluded
1MQwen 3.8 Max$1.60Fit only when model context window covers full request; otherwise excluded
1MQwen 3.7 Max$1.60Fit only when model context window covers full request; otherwise excluded

2. Output-heavy coding/agent sensitivity

Output multiplierModelBillUplift required
Qwen 3.7 Plus$524.00Accepted-result uplift for premium tier: User-supplied
Qwen 3.8 Max$1216.00Accepted-result uplift for premium tier: User-supplied
Qwen 3.7 Max$1216.00Accepted-result uplift for premium tier: User-supplied
Qwen 3.7 Plus$664.00Accepted-result uplift for premium tier: User-supplied
Qwen 3.8 Max$1664.00Accepted-result uplift for premium tier: User-supplied
Qwen 3.7 Max$1664.00Accepted-result uplift for premium tier: User-supplied
Qwen 3.7 Plus$944.00Accepted-result uplift for premium tier: User-supplied
Qwen 3.8 Max$2560.00Accepted-result uplift for premium tier: User-supplied
Qwen 3.7 Max$2560.00Accepted-result uplift for premium tier: User-supplied

3. Direct-versus-hosted-alias break-even ceiling

ModelDated Alibaba spendExternal premium ceilingRegion/currency
Qwen 3.7 Plus$524.00External monthly bill ≤ direct spend; exact alias rate: UnavailableRegion and currency: Unavailable
Qwen 3.8 Max$1216.00External monthly bill ≤ direct spend; exact alias rate: UnavailableRegion and currency: Unavailable
Qwen 3.7 Max$1216.00External monthly bill ≤ direct spend; exact alias rate: UnavailableRegion and currency: Unavailable

Provenance: Batch 11 Qwen direct/hosted context module; selected model Qwen 3.7 Plus; inputs are 2,400 input tokens, 350 output tokens, 200,000 calls/month, cache rate 30%, batch share 100%. Verified 2026-06-21. Provider source: https://www.alibabacloud.com/help/en/model-studio/models · Run this scenario →

Batch 59 · server-rendered evidence boards · verified 2026-09-07

Intent answer: Alibaba Cloud Qwen API provides high-capability reasoning and multilingual capabilities at aggressive token pricing. Calculating costs requires evaluating Max vs Plus pricing tiers and prompt caching. Verified 2026-09-07.

Demand evidence: Qualitative demand: Qwen pricing tools and cost calculators reviewed 2026-09-07; exact US monthly volume is unavailable.

Scope boundary: Calculate Alibaba Cloud Qwen API token costs across Qwen3.7-Max, Qwen3.7-Plus, and Qwen3.8-30B with multilingual and coding workloads. Exact joins required; unresolved joins render Unavailable.

Qwen flagship vs plus model cost crossover

Deterministic formula / rule: crossover_point = workload_volume where quality_delta(Max) balances rate_delta(Plus).

Boundary: Owns Qwen model tier selection economics.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-calc-qwen-m1-r1
multilingual translation task
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$multilingual translation task; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m1-r2
STEM reasoning and math problem
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$STEM reasoning and math problem; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m1-r3
general customer service chatbot
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$general customer service chatbot; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m1-r4
code generation and unit test writing
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$code generation and unit test writing; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m1-r5
large corpus embedding pipeline
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$large corpus embedding pipeline; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m1-r6
unsupported Qwen checkpoint
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$unsupported Qwen checkpoint; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07Unavailable — exact calc-qwen evidence join is not closed for "unsupported Qwen checkpoint"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: Alibaba Cloud Qwen API token pricing. Verified 2026-09-07; missing or conflicting joins fail closed.

Qwen currency and regional gateway normalizer

Deterministic formula / rule: usd_rate = cny_rate / exchange_rate; international gateway vs domestic Alibaba Cloud pricing.

Boundary: Owns currency and regional billing alignment for Qwen.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-calc-qwen-m2-r1
International USD billing account
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$International USD billing account; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m2-r2
Domestic China CNY account
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$Domestic China CNY account; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m2-r3
Southeast Asia regional gateway
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$Southeast Asia regional gateway; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m2-r4
European localized proxy
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$European localized proxy; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m2-r5
Enterprise committed-use discount
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$Enterprise committed-use discount; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m2-r6
unsupported billing currency
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$unsupported billing currency; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07Unavailable — exact calc-qwen evidence join is not closed for "unsupported billing currency"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: Alibaba Cloud Qwen API token pricing. Verified 2026-09-07; missing or conflicting joins fail closed.

Qwen long-context prompt caching ledger

Deterministic formula / rule: cached_savings = prompt_tokens * (standard_rate - cache_read_rate); caching up to 1M context.

Boundary: Owns Qwen long-context cache economics.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-calc-qwen-m3-r1
legal document review (100K tokens)
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$legal document review (100K tokens); input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m3-r2
financial prospectus analysis (250K)
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$financial prospectus analysis (250K); input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m3-r3
software repository chat (500K)
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$software repository chat (500K); input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m3-r4
massive book corpus search (1M)
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$massive book corpus search (1M); input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m3-r5
frequent cache invalidation
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$frequent cache invalidation; input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-qwen-m3-r6
unsupported cache retention period
route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$unsupported cache retention period; input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07Unavailable — exact calc-qwen evidence join is not closed for "unsupported cache retention period"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: Qwen model architecture and API guide. Verified 2026-09-07; missing or conflicting joins fail closed.

Method and limitations: this board exposes deterministic rules, first-party citations, and dated evidence identities. It does not invent volume, coverage, entitlement, retention, residency, quota, capacity, feature support, quality, price, reliability, or legal conclusions. Run the calc-qwen evidence flow →

Ranked cost — 200K requests/month

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Qwen 3.7 PlusmidQwen$524$524
Qwen 3.8 MaxmidQwen$1,216$1,216
Qwen 3.7 MaxmidQwen$1,216$1,216

Levers live on Qwen

Context trimming: up to 47%Batch API: up to 50%

The billing gotcha

Qwen's direct Model Studio API and its hosted aliases are separate billing surfaces: this page prices the Qwen models sold through Alibaba Cloud, not similarly named Qwen weights on Groq or another host. Region matters too — the international DashScope endpoint and mainland China service can expose different quotas and currency treatment, so the monthly estimate below is a token-rate projection, not an AWS-style all-in account bill.

Related

Qwen provider profile →Compare against another provider →Global cost calculator →How to reduce LLM API costs →