← Back to all pricing

Qwen 3.6 27B on Groq API Pricing: Extreme LPU Speed for Bilingual AI

Comprehensive Qwen 3.6 27B on Groq API pricing analysis ($0.60/M input, $3.00/M output), LPU hardware acceleration (600+ tps), bilingual reasoning, and open-weights ROI.

Legacy — superseded by Qwen 3.8 30B See Qwen 3.8 30B pricing.
No announced shutdown date. Source · Full retirement tracker

How much does Qwen 3.6 27B cost per million tokens?

Qwen 3.6 27B on Groq costs $0.60 per million input tokens and $3.00 per million output tokens ($1.20/M blended at 3:1). Combines elite Chinese/English bilingual reasoning with extreme inference speed (600+ tps) powered by Groq LPUs. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$0.60/M
Output
$3.00/M
Blended
$1.20/M
Provider
Verified 2026-06-19source

How much does Qwen 3.6 27B cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.2100
Medium1,000500$2.1000
Long4,0002,000$8.4000

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.

Volume ladder, output-cost sensitivity, and migration ledger for Qwen 3.6 27B

1. Fixed-shape monthly cost ladder

WorkloadInput / output tokens100K requests1M requests10M requests
Short chat500 / 150$75.00$750.00$7500.00
Code review4,000 / 800$480.00$4800.00$48000.00
Document summary16,000 / 2,000$1560.00$15600.00$156000.00

Cost = requests × (input tokens × $0.60/M + output tokens × $3.00/M) ÷ 1,000,000, at three fixed text-token workload shapes. No volume discount, cache, or batch rate is inferred — Qwen 3.6 27B's registry entry does not price those tiers.

2. Output-token cost-sensitivity band

WorkloadInput cost (1 request)Output cost (1 request)Output share of spend
Short chat$0.0003$0.000460.0%
Code review$0.0024$0.002450.0%
Document summary$0.0096$0.006038.5%

Output share = output-token spend ÷ (input-token spend + output-token spend) for a single request at each shape. Across these three shapes, Qwen 3.6 27B's output share spans 38.5% to 60.0% — a 21.5%-point swing driven entirely by workload shape, not by any change in the $0.60/$3.00 per-1M rates.

3. Successor rate delta and lifecycle ledger

RateQwen 3.6 27BQwen 3.8 30BDelta ($ / %)
Input $/M$0.60$0.60$0.0000 (0.0%)
Output $/M$3.00$3.00$0.0000 (0.0%)

Delta = successor rate − Qwen 3.6 27B rate, both taken from the current dated pricing registry. Cache, batch, and context-tier pricing are not substituted across models; each figure is model-specific or marked Unavailable.

FieldRecorded valueNote
Lifecycle statuslegacyVerified 2026-08-14
Deprecation announcedUnavailableNo inference beyond the dated record
Shutdown dateUnavailableNull/unavailable is not a promise of indefinite availability
SuccessorQwen 3.8 30Bqwen3-8-30b priced in this registry
Context windowUnavailableUnavailable — no model-specific spec sourced
Compare Qwen 3.6 27B against its successor →

Price verified 2026-06-19; lifecycle verified 2026-08-14. Luna is the data owner. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated price source · lifecycle source · Test Qwen 3.6 27B in All AI Ask.

Continuous SEO Builder · Batch 71 Audit · 2026-09-08Owner: qwen3-6-27b

Qwen 3.6 27B on Groq API Pricing: Extreme LPU Speed for Bilingual AI

Qwen 3.6 27B on Groq costs $0.60 per million input tokens and $3.00 per million output tokens ($1.20/M blended at 3:1). Combines elite Chinese/English bilingual reasoning with extreme inference speed (600+ tps) powered by Groq LPUs. Verified 2026-09-08.

Module 1 · Qwen 3.6 27B on Groq LPU Token Economics
Blended Cost = (Input Tokens × $0.60 + Output Tokens × $3.00) / 1,000,000

Groq LPU hardware powers Qwen 3.6 27B at extreme velocity and competitive token costs.

Boundary: Standard pay-as-you-go rate card on Groq Cloud; includes dedicated LPU acceleration.
ScenarioRendered Evidence & Bounds
Scenario 1Cross-border bilingual document translation (4K in, 2K out): $0.00840 per document
Scenario 2Mandarin/English customer service turn (1.5K in, 300 out): $0.00180 per message
Scenario 3Bilingual coding snippet refactor (3K in, 1K out): $0.00480 per snippet
Scenario 4Real-time live voice translation turn (800 in, 200 out): $0.00108 per turn
Scenario 5Technical manual localization review (10K in, 3K out): $0.01500 per section
Scenario 6Monthly 50M token bilingual enterprise workload: $60.00 infrastructure budget
Module 2 · Groq LPU Hardware Acceleration Speed & SLA Advantages
Throughput Advantage = 600+ TPS (Groq LPU) vs 80-120 TPS (Traditional GPU Clusters)

LPU hardware delivers 5x to 8x faster streaming generation than conventional GPU clusters.

Boundary: Evaluates user experience and operational advantages of ultra-fast inference speed.
ScenarioRendered Evidence & Bounds
Scenario 1Generates 600+ tokens per second: renders a 500-token response in under 0.9 seconds
Scenario 2Time-to-first-token latency under 120ms enables instant interactive application feedback
Scenario 3Deterministic LPU architecture guarantees consistent latency without GPU thermal throttling
Scenario 4High-speed bilingual synthesis makes real-time conversational voice translation viable
Scenario 5Eliminates developer waiting time during iterative code generation and validation loops
Scenario 6Delivers superior cost-performance ratio for time-sensitive production user interfaces
Module 3 · Qwen 3.6 27B Managed API vs Self-Hosted GPU Cluster Break-Even
Self-Hosted TCO = 2× NVIDIA A100/H100 Monthly Lease ($3,200) + DevOps Overhead ($1,500)

Managed LPU inference provides massive TCO savings over self-hosted infrastructure below 3.9B tokens/mo.

Boundary: Calculates monthly token volume required before self-hosting an unquantized 27B model breaks even.
ScenarioRendered Evidence & Bounds
Scenario 1Managed Groq API spend at 50M tokens/mo: $60.00 total infrastructure spend
Scenario 2Managed Groq API spend at 500M tokens/mo: $600.00 total infrastructure spend
Scenario 3Managed Groq API spend at 2B tokens/mo: $2,400.00 total infrastructure spend
Scenario 4Self-hosted cluster break-even threshold: >3.9 billion tokens per month
Scenario 5Managed API eliminates GPU cluster maintenance, cold starts, and DevOps operational overhead
Scenario 6Strong recommendation: utilize Groq managed LPU API for all workloads under 3.5B tokens/mo
Explore Related Analyses:Groq provider profileCompare vs Qwen 3.8 30BCompare vs DeepSeek V4 FlashFastest AI models comparison

How fast is Qwen 3.6 27B?

Not yet measured — see the speed benchmark leaderboard for models we do track.

How much does Qwen 3.6 27B cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.12
1,000,000$1.20
10,000,000$12.00
100,000,000$120.00

How does Qwen 3.6 27B compare with other models?

GPT-OSS 20B$0.13/MGPT-OSS 120B$0.26/MLlama 4 Maverick$0.30/MQwen 3.8 30B$1.20/MQwen 3.8 30B$1.20/MQwen 3.7 Plus$1.10/MGLM-5.1$1.00/M
See all Groq models →

What are common questions about Qwen 3.6 27B?

Is Qwen 3.6 27B cheaper than Qwen 3.8 30B?

Qwen 3.6 27B costs $1.20/M blended tokens, Qwen 3.8 30B costs $1.20/M — Qwen 3.8 30B is cheaper.

How much does 1 million tokens cost with Qwen 3.6 27B?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $1.20. Pure input costs $0.60/M; pure output costs $3.00/M.

What does Qwen 3.6 27B cost at high volume?

At 100 million blended tokens a month, Qwen 3.6 27B costs approximately $120.00. See the cost-at-scale table below for other volumes.

Try Qwen 3.6 27B for free

Run real prompts against Qwen 3.6 27B and every other model on this page in one workspace.

Try Qwen 3.6 27B Free