← Back to all pricing

Llama 4 Maverick on Groq API Pricing: Blazing LPU Hardware Speed

Comprehensive Llama 4 Maverick on Groq API pricing analysis ($0.20/M input, $0.60/M output), 800+ tps LPU hardware inference, agent tool calling, and open-weights ROI.

Deprecated — use Muse Spark 1.3 (Meta) See Muse Spark 1.3 pricing.
No announced shutdown date. Source · Full retirement tracker

How much does Llama 4 Maverick cost per million tokens?

Llama 4 Maverick on Groq costs $0.20 per million input tokens and $0.60 per million output tokens ($0.30/M blended at 3:1). Delivers next-generation open-weights reasoning with extreme 800+ tps inference velocity on Groq LPUs. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$0.20/M
Output
$0.60/M
Blended
$0.30/M
Provider
Verified 2026-07-10source

How much does Llama 4 Maverick cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.0500
Medium1,000500$0.5000
Long4,0002,000$2.0000

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.

Batch 9 continuity economics for Llama 4 Maverick

1. Text-only workload bills

ShapeInput / output100K calls1M calls10M calls
Chat1,000 / 300$38.00$380.00$3800.00
Coding agent4,000 / 1,000$140.00$1400.00$14000.00
Long input20,000 / 2,000$520.00$5200.00$52000.00

Formula: calls × (input tokens × $0.20/M + output tokens × $0.60/M) ÷ 1,000,000. Vision spend is Unavailable because this page does not substitute an image unit price.

2. Traffic-drain budget (fallback, not successor)

DrainLlama 4 Maverick baselineGPT-OSS 20B fallbackBlended monthlySavings
0%$1400.00$600.00$1400.00$0.0000
25%$1400.00$600.00$1200.00$200.00
50%$1400.00$600.00$1000.00$400.00
100%$1400.00$600.00$600.00$800.00

The 1M-call coding-agent shape is held constant. GPT-OSS 20B is labelled only as a same-provider fallback; it is not a successor.

3. Continuity evidence ledger

FieldValueEvidence / boundary
successorIdmuse-spark-1-3Verified 2026-08-14
Lifecycle / shutdownlegacyUnavailable
Price verified2026-07-10https://groq.com/pricing
Muse Spark gatewayUnavailable
Required retest fieldsquality, latency, context fit, compatibilityMeasured evidence required before migration claim

Successor rate audit

RateLlama 4 MaverickMuse Spark 1.3Delta
Input $/M$0.20$1.25$1.05 / 525.0%
Output $/M$0.60$4.25$3.65 / 608.3%
Input/output mix at one call:
ShapeInput spendOutput spendOutput share
Chat$0.0002$0.000247.4%
Coding agent$0.0008$0.000642.9%
Long input$0.0040$0.001223.1%
A shutdown date, cache/batch terms, and gateway availability are not inferred when the registry is silent.

Verified 2026-07-10. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never treated as zero or an inferred successor. Dated price source · Run this evidence in All AI Ask.

Continuous SEO Builder · Batch 72 Audit · 2026-09-08Owner: llama-4-maverick

Llama 4 Maverick on Groq API Pricing: Blazing LPU Hardware Speed

Llama 4 Maverick on Groq costs $0.20 per million input tokens and $0.60 per million output tokens ($0.30/M blended at 3:1). Delivers next-generation open-weights reasoning with extreme 800+ tps inference velocity on Groq LPUs. Verified 2026-09-08.

Module 1 · Llama 4 Maverick on Groq LPU Token Economics
Blended Cost = (Input Tokens × $0.20 + Output Tokens × $0.60) / 1,000,000

Groq LPUs power Llama 4 Maverick at $0.30/M blended tokens with unprecedented generation velocity.

Boundary: Standard pay-as-you-go rate card on Groq Cloud; includes dedicated LPU acceleration.
ScenarioRendered Evidence & Bounds
Scenario 1Interactive voice bot turn (400 in, 100 out): $0.000140 per conversational turn
Scenario 2Customer inquiry classification (800 in, 50 out): $0.000190 per ticket
Scenario 3Python script synthesis and test (2K in, 500 out): $0.000700 per code snippet
Scenario 4Real-time document summarization (8K in, 800 out): $0.002080 per document
Scenario 5Full agent multi-turn tool calling loop (16K in, 2K out): $0.004400 per full loop
Scenario 6Monthly 100M token production tier: $30.00 total infrastructure spend
Module 2 · Groq LPU Hardware Acceleration Speed & Latency Benchmarks
Throughput Advantage = 800+ TPS (Groq LPU) vs 90-130 TPS (GPU Cloud Providers)

Groq LPUs deliver 6x to 9x faster token throughput than conventional GPU cloud instances.

Boundary: Evaluates user responsiveness and concurrency advantages of custom tensor streaming architecture.
ScenarioRendered Evidence & Bounds
Scenario 1Generates 800+ tokens per second: a 400-token response streams in under 0.5 seconds
Scenario 2Time-to-first-token under 90ms enables truly instantaneous voice conversational agents
Scenario 3Deterministic LPU silicon architecture ensures zero jitter or queue queuing spikes
Scenario 4High-speed agent loops complete 5-turn reasoning traces in under 2.5 seconds total elapsed time
Scenario 5Dramatically improves human developer productivity during real-time code autocomplete
Scenario 6Industry-leading price-to-speed ratio across all open-weights hosting alternatives
Module 3 · Llama 4 Maverick Managed Groq API vs Self-Hosting TCO
Self-Hosted TCO = 4× NVIDIA H100 Lease ($8,000/mo) + Kubernetes SRE Overhead ($2,500/mo)

Managed LPU inference provides overwhelming cost savings over self-hosting below 35B tokens/mo.

Boundary: Calculates the monthly token volume threshold before on-premise hardware deployment breaks even.
ScenarioRendered Evidence & Bounds
Scenario 1Managed Groq API spend at 100M tokens/mo: $30.00 monthly cost
Scenario 2Managed Groq API spend at 1B tokens/mo: $300.00 monthly cost
Scenario 3Managed Groq API spend at 10B tokens/mo: $3,000.00 monthly cost
Scenario 4Self-hosted cluster break-even point: requires >35 billion tokens per month
Scenario 5Managed API eliminates GPU reservation commitments, thermal maintenance, and cluster idle waste
Scenario 6Recommended architectural strategy: utilize Groq managed LPUs for all production workloads
Explore Related Analyses:Groq provider profileCompare vs Llama 3.3 70BCompare vs Grok 3 MiniFastest AI models comparison

How fast is Llama 4 Maverick?

Not yet measured — see the speed benchmark leaderboard for models we do track.

How much does Llama 4 Maverick cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.03
1,000,000$0.30
10,000,000$3.00
100,000,000$30.00

How does Llama 4 Maverick compare with other models?

GPT-OSS 20B$0.13/MGPT-OSS 120B$0.26/MQwen 3.8 30B$1.20/MQwen 3.6 27B$1.20/MGPT-4o Mini$0.26/MGrok-3 Mini$0.26/MGPT-OSS 120B$0.26/M
See all Groq models →

What are common questions about Llama 4 Maverick?

Is Llama 4 Maverick cheaper than GPT-4o Mini?

Llama 4 Maverick costs $0.30/M blended tokens, GPT-4o Mini costs $0.26/M — GPT-4o Mini is cheaper.

How much does 1 million tokens cost with Llama 4 Maverick?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.30. Pure input costs $0.20/M; pure output costs $0.60/M.

What does Llama 4 Maverick cost at high volume?

At 100 million blended tokens a month, Llama 4 Maverick costs approximately $30.00. See the cost-at-scale table below for other volumes.

Try Llama 4 Maverick for free

Run real prompts against Llama 4 Maverick and every other model on this page in one workspace.

Try Llama 4 Maverick Free