Qwen API Cost Calculator
How much does the Qwen API cost per month?
For Qwen 3.7 Plus at 200K requests/month, 2,400 input tokens and 350 output tokens per request costs about $524 per month after 30% cache use and 100% batch share. Across Qwen's 3 priced models, the cheapest default ranking is Qwen 3.7 Plus at $524 per month.
Pricing data as of June 2026. Sources: Qwen pricing and model documentation. Parameters are shareable in this URL.
Qwen 3.7 Plus: estimated monthly cost
Formula: provider-specific cached input multiplier × cache rate, plus uncached input and verbosity-adjusted output; cache writes are amortised over the provider TTL, then batch savings are applied.
| Per request | $0.003 |
|---|---|
| Per day | $17.47 |
| Per month | $524 |
| Per year | $6,288 |
OpenAI scenario sensitivity
The default is a server-rendered estimate. Change the cache and batch shares to see when a cheaper qualified model overtakes the selected model; the URL is shareable and preserves the inputs.
| Default | $524/month |
|---|---|
| No caching or batch | $524/month |
| 50% cache, 50% batch | $524/month |
| 100% cache, 100% batch | $524/month |
Batch 11 Qwen workload decision depth
1. Plus/Max context eligibility and cost ladder
| Input tokens | Model | Cost | Window result |
|---|---|---|---|
| 32K | Qwen 3.7 Plus | $0.03 | Fit only when model context window covers full request; otherwise excluded |
| 32K | Qwen 3.8 Max | $0.05 | Fit only when model context window covers full request; otherwise excluded |
| 32K | Qwen 3.7 Max | $0.05 | Fit only when model context window covers full request; otherwise excluded |
| 128K | Qwen 3.7 Plus | $0.10 | Fit only when model context window covers full request; otherwise excluded |
| 128K | Qwen 3.8 Max | $0.21 | Fit only when model context window covers full request; otherwise excluded |
| 128K | Qwen 3.7 Max | $0.21 | Fit only when model context window covers full request; otherwise excluded |
| 500K | Qwen 3.7 Plus | $0.40 | Fit only when model context window covers full request; otherwise excluded |
| 500K | Qwen 3.8 Max | $0.80 | Fit only when model context window covers full request; otherwise excluded |
| 500K | Qwen 3.7 Max | $0.80 | Fit only when model context window covers full request; otherwise excluded |
| 1M | Qwen 3.7 Plus | $0.80 | Fit only when model context window covers full request; otherwise excluded |
| 1M | Qwen 3.8 Max | $1.60 | Fit only when model context window covers full request; otherwise excluded |
| 1M | Qwen 3.7 Max | $1.60 | Fit only when model context window covers full request; otherwise excluded |
2. Output-heavy coding/agent sensitivity
| Output multiplier | Model | Bill | Uplift required |
|---|---|---|---|
| 1× | Qwen 3.7 Plus | $524.00 | Accepted-result uplift for premium tier: User-supplied |
| 1× | Qwen 3.8 Max | $1216.00 | Accepted-result uplift for premium tier: User-supplied |
| 1× | Qwen 3.7 Max | $1216.00 | Accepted-result uplift for premium tier: User-supplied |
| 2× | Qwen 3.7 Plus | $664.00 | Accepted-result uplift for premium tier: User-supplied |
| 2× | Qwen 3.8 Max | $1664.00 | Accepted-result uplift for premium tier: User-supplied |
| 2× | Qwen 3.7 Max | $1664.00 | Accepted-result uplift for premium tier: User-supplied |
| 4× | Qwen 3.7 Plus | $944.00 | Accepted-result uplift for premium tier: User-supplied |
| 4× | Qwen 3.8 Max | $2560.00 | Accepted-result uplift for premium tier: User-supplied |
| 4× | Qwen 3.7 Max | $2560.00 | Accepted-result uplift for premium tier: User-supplied |
3. Direct-versus-hosted-alias break-even ceiling
| Model | Dated Alibaba spend | External premium ceiling | Region/currency |
|---|---|---|---|
| Qwen 3.7 Plus | $524.00 | External monthly bill ≤ direct spend; exact alias rate: Unavailable | Region and currency: Unavailable |
| Qwen 3.8 Max | $1216.00 | External monthly bill ≤ direct spend; exact alias rate: Unavailable | Region and currency: Unavailable |
| Qwen 3.7 Max | $1216.00 | External monthly bill ≤ direct spend; exact alias rate: Unavailable | Region and currency: Unavailable |
Provenance: Batch 11 Qwen direct/hosted context module; selected model Qwen 3.7 Plus; inputs are 2,400 input tokens, 350 output tokens, 200,000 calls/month, cache rate 30%, batch share 100%. Verified 2026-06-21. Provider source: https://www.alibabacloud.com/help/en/model-studio/models · Run this scenario →
Batch 59 · server-rendered evidence boards · verified 2026-09-07
Intent answer: Alibaba Cloud Qwen API provides high-capability reasoning and multilingual capabilities at aggressive token pricing. Calculating costs requires evaluating Max vs Plus pricing tiers and prompt caching. Verified 2026-09-07.
Demand evidence: Qualitative demand: Qwen pricing tools and cost calculators reviewed 2026-09-07; exact US monthly volume is unavailable.
Scope boundary: Calculate Alibaba Cloud Qwen API token costs across Qwen3.7-Max, Qwen3.7-Plus, and Qwen3.8-30B with multilingual and coding workloads. Exact joins required; unresolved joins render Unavailable.
Qwen flagship vs plus model cost crossover
Deterministic formula / rule: crossover_point = workload_volume where quality_delta(Max) balances rate_delta(Plus).
Boundary: Owns Qwen model tier selection economics.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch59-calc-qwen-m1-r1multilingual translation task | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$multilingual translation task; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m1-r2STEM reasoning and math problem | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$STEM reasoning and math problem; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m1-r3general customer service chatbot | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$general customer service chatbot; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m1-r4code generation and unit test writing | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$code generation and unit test writing; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m1-r5large corpus embedding pipeline | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$large corpus embedding pipeline; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m1-r6unsupported Qwen checkpoint | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$unsupported Qwen checkpoint; workload shape; Qwen3.7-Max cost; Qwen3.7-Plus cost; price ratio; quality benchmark spread; tier selection recommendation; verified=2026-09-07 | Unavailable — exact calc-qwen evidence join is not closed for "unsupported Qwen checkpoint" | FAIL CLOSED — manual, probe, or source evidence required |
First-party citation: Alibaba Cloud Qwen API token pricing. Verified 2026-09-07; missing or conflicting joins fail closed.
Qwen currency and regional gateway normalizer
Deterministic formula / rule: usd_rate = cny_rate / exchange_rate; international gateway vs domestic Alibaba Cloud pricing.
Boundary: Owns currency and regional billing alignment for Qwen.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch59-calc-qwen-m2-r1International USD billing account | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$International USD billing account; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m2-r2Domestic China CNY account | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$Domestic China CNY account; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m2-r3Southeast Asia regional gateway | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$Southeast Asia regional gateway; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m2-r4European localized proxy | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$European localized proxy; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m2-r5Enterprise committed-use discount | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$Enterprise committed-use discount; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m2-r6unsupported billing currency | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$unsupported billing currency; account region; currency; base rate; conversion formula; effective USD/M input; effective USD/M output; verified=2026-09-07 | Unavailable — exact calc-qwen evidence join is not closed for "unsupported billing currency" | FAIL CLOSED — manual, probe, or source evidence required |
First-party citation: Alibaba Cloud Qwen API token pricing. Verified 2026-09-07; missing or conflicting joins fail closed.
Qwen long-context prompt caching ledger
Deterministic formula / rule: cached_savings = prompt_tokens * (standard_rate - cache_read_rate); caching up to 1M context.
Boundary: Owns Qwen long-context cache economics.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch59-calc-qwen-m3-r1legal document review (100K tokens) | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$legal document review (100K tokens); input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m3-r2financial prospectus analysis (250K) | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$financial prospectus analysis (250K); input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m3-r3software repository chat (500K) | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$software repository chat (500K); input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m3-r4massive book corpus search (1M) | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$massive book corpus search (1M); input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m3-r5frequent cache invalidation | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$frequent cache invalidation; input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07 | Unavailable — frozen calc-qwen fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-qwen-m3-r6unsupported cache retention period | route=$/llm-cost-calculator/qwen; owner=$calc-qwen; scenario=$unsupported cache retention period; input document size; query frequency; standard input expense; cached input expense; monthly dollar savings; cache efficiency %; verified=2026-09-07 | Unavailable — exact calc-qwen evidence join is not closed for "unsupported cache retention period" | FAIL CLOSED — manual, probe, or source evidence required |
First-party citation: Qwen model architecture and API guide. Verified 2026-09-07; missing or conflicting joins fail closed.
Method and limitations: this board exposes deterministic rules, first-party citations, and dated evidence identities. It does not invent volume, coverage, entitlement, retention, residency, quota, capacity, feature support, quality, price, reliability, or legal conclusions. Run the calc-qwen evidence flow →
Ranked cost — 200K requests/month
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Qwen 3.7 Plusmid | Qwen | $524 | $524 | — | — |
| Qwen 3.8 Maxmid | Qwen | $1,216 | $1,216 | — | — |
| Qwen 3.7 Maxmid | Qwen | $1,216 | $1,216 | — | — |
Levers live on Qwen
The billing gotcha
Qwen's direct Model Studio API and its hosted aliases are separate billing surfaces: this page prices the Qwen models sold through Alibaba Cloud, not similarly named Qwen weights on Groq or another host. Region matters too — the international DashScope endpoint and mainland China service can expose different quotas and currency treatment, so the monthly estimate below is a token-rate projection, not an AWS-style all-in account bill.
