How Much Does a Coding Agent Cost per Month?
At production volume (20,000 calls/month), the cheapest effective option is Amazon Nova Micro at $18.26/month. The most expensive frontier option, GPT-5.4 Pro, runs $19,488/month — A coding agent reads large amounts of file context and writes substantial diffs — both sides of the bill are large compared to a chat turn.
How much does coding agent cost per month?
At production volume (20,000 calls/month), the cheapest effective option for coding agent is Amazon Nova Micro at $18.26 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $19,488 per month for the same workload.
Token shape
| Shape | large file context in, large diff out |
| Input / output tokens per call | 20K in / 2K out |
| Cacheable input | 60% |
| Batch-eligible | No |
Input is the files and instructions relevant to one change; output is a multi-file diff plus explanation. Repeated reads of the same files across a session make a majority of input cacheable.
What drives this workload's cost?
The main token-volume driver here is input tokens: each call sends 20K input tokens and requests up to 2K output tokens, at 20,000 calls per month in the default volume. 60% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is not marked batch-eligible because the profile assumes a synchronous response.
Volume
Ranked cost — Production volume
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Amazon Nova Microbudget | Amazon | $19.60 | $18.26 | 0.76× | — |
| Amazon Nova Litebudget | Amazon | $33.60 | $32.74 | 0.91× | — |
| GPT-5 Nanobudgetlegacy | OpenAI | $36.00 | $26.19 | — | — |
| Gemini 2.5 Flash Litebudgetlegacy | $56.00 | $35.18 | — | ▲2 | |
| GPT-OSS 20Bbudget | Groq | $42.00 | $65.16 | 2.93× | ▼1 |
| Ministral 8Bbudget | Mistral | $66.00 | $65.58 | 0.93× | ▲1 |
| Mistral Small 3.1budget | Mistral | $84.00 | $80.40 | 0.85× | ▲4 |
| GPT-4o Minibudgetlegacy | OpenAI | $84.00 | $54.58 | — | — |
| Grok-3 Minibudgetlegacy | xAI | $84.00 | $84.00 | — | — |
| GPT-OSS 120Bbudget | Groq | $84.00 | $104 | 1.82× | — |
| Llama 4 Maverickbudgetlegacy | Groq | $104 | $104 | — | ▲1 |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $130 | $79.77 | 0.78× | ▲1 |
| Muse Spark 1.3 Contributorbudget | Meta | $48.00 | $144 | 12.95× | ▼8 |
| Codestralbudget | Mistral | $156 | $148 | 0.79× | — |
| Gemini 3.1 Flash Litebudgetlegacy | $160 | $100 | 0.87× | — |
Show all 68 models
| GPT-5 Minibudgetlegacy | OpenAI | $180 | $131 | — | ▲1 |
| GPT-OSS 120B (Cerebras)budget | Cerebras | $170 | $210 | 2.32× | ▼1 |
| Gemini 3.5 Flash Litebudget | $220 | $158 | — | — | |
| Gemini 2.5 Flashbudgetlegacy | $220 | $158 | — | — | |
| Mistral Large 3budget | Mistral | $260 | $260 | — | ▲1 |
| DeepSeek V4 Flashbudget | DeepSeek | $229 | $313 | 2.59× | ▼1 |
| GLM-5.1midlegacy | Z.ai | $328 | $328 | — | — |
| Qwen 3.8 30Bmid | Groq | $360 | $360 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $360 | $360 | — | — |
| Qwen 3.7 Plusmid | Qwen | $400 | $400 | — | — |
| Amazon Nova Promid | Amazon | $448 | $448 | — | — |
| Gemini 3.7 Flashmid | $450 | $294 | — | — | |
| GPT-5.4 Minimidlegacy | OpenAI | $480 | $333 | — | — |
| Gemini 3.1 Flashmidlegacy | $480 | $324 | — | — | |
| Claude Haiku 4.5mid | Anthropic | $600 | $503 | — | ▲1 |
| o3-Minimidlegacy | OpenAI | $616 | $400 | — | ▲1 |
| GPT-5.6 Lunamid | OpenAI | $640 | $444 | — | ▲1 |
| Grok 4.3mid | xAI | $600 | $641 | 1.41× | ▼3 |
| Muse Spark 1.3mid | Meta | $670 | $670 | — | — |
| Mistral Medium 3mid | Mistral | $900 | $852 | 0.84× | ▲5 |
| Qwen 3.8 Maxmid | Qwen | $896 | $896 | — | ▲1 |
| Qwen 3.7 Maxmid | Qwen | $896 | $896 | — | ▲1 |
| Gemini 3.6 Flashmid | $900 | $588 | — | ▲1 | |
| GPT-5midlegacy | OpenAI | $900 | $655 | — | ▲2 |
| Grok-3midlegacy | xAI | $960 | $960 | — | ▲2 |
| Gemini 3.5 Flashmidlegacy | $960 | $648 | — | ▲2 | |
| Grok-4.20 Reasoningmid | xAI | $1,040 | $1,040 | — | ▲3 |
| Grok-4.20mid | xAI | $1,040 | $1,040 | — | ▲3 |
| Grok 4.6mid | xAI | $1,040 | $1,040 | — | ▲3 |
| Grok 4.5mid | xAI | $1,040 | $1,040 | — | ▲3 |
| DeepSeek V4 Promid | DeepSeek | $686 | $1,051 | 3.30× | ▼11 |
| Gemini 3.1 Promid | $1,280 | $686 | 0.63× | ▲4 | |
| GPT-4.1midlegacy | OpenAI | $1,120 | $728 | — | ▲1 |
| Claude Sonnet 5mid | Anthropic | $1,200 | $1,006 | — | ▲1 |
| GLM-5.2mid | Z.ai | $736 | $1,241 | 3.87× | ▼14 |
| GPT-4omidlegacy | OpenAI | $1,400 | $910 | — | ▲1 |
| GPT-5.6 Terramid | OpenAI | $1,600 | $1,110 | — | ▲1 |
| GPT-5.4midlegacy | OpenAI | $1,600 | $1,110 | — | ▲1 |
| GLM 4.7 (Cerebras)mid | Cerebras | $1,010 | $1,728 | 7.53× | ▼10 |
| Claude Sonnet 4.6mid | Anthropic | $1,800 | $1,510 | — | — |
| Claude Sonnet 4.5midlegacy | Anthropic | $1,800 | $1,510 | — | — |
| Claude Sonnet 4midlegacy | Anthropic | $1,800 | $1,510 | — | — |
| GPT-5.6 Solmid | OpenAI | $2,400 | $1,615 | — | — |
| Claude Opus 4.8mid | Anthropic | $3,000 | $2,476 | 0.96× | — |
| Claude Opus 4.7midlegacy | Anthropic | $3,000 | $2,516 | — | — |
| Claude Opus 4.6midlegacy | Anthropic | $3,000 | $2,516 | — | — |
| Claude Opus 4.5midlegacy | Anthropic | $3,000 | $2,516 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $5,200 | $3,239 | — | — |
| Claude Fable 5frontier | Anthropic | $6,000 | $5,032 | — | — |
| Claude Opus 5frontier | Anthropic | $9,000 | $7,548 | — | — |
| Claude Opus 4.1frontierlegacy | Anthropic | $9,000 | $7,548 | — | — |
| Claude Opus 4frontierlegacy | Anthropic | $9,000 | $7,548 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $19,200 | $13,604 | 1.04× | — |
Batch 39 · server-rendered decision evidence · verified 2026-08-27
Coding-agent workload cost mechanics
Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.
Repository-context fan-out ledger
Formula / rubric: Task input = initial scan + retrieved files + repeated state + tool results + cache writes/reads + compaction summaries.
Provenance: Frozen small/medium/large repository shapes at 5, 15, and 30 turns; rates use the dated registry. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
small / 5 turnsbatch39-coding-agent-m1-r1 | 180K repo tokens; 6 retrieved files; 5× tool results 3K; 40% cache hit | uncached 252K, cached 168K, output 18K; base bill $1.1100 = (168000×$5.00 + 18000×$15.00)/1M | Cache-read rate must be provider-specific; this is an OpenAI-shaped example. | CALCULATED — context fan-out closed. |
medium / 15 turnsbatch39-coding-agent-m1-r2 | 620K repo; 18 files; 15 tool results; 2 compactions; 55% cache hit | compaction summary 24K; repeated state 410K; cache write/read split unavailable. | Do not treat repeated context as free. | Unavailable — cache write/read split is not returned |
large / 30 turnsbatch39-coding-agent-m1-r3 | 2.4M repo; 44 files; 30 tool results; 5 compactions; 3 parallel workers | retrieval and tool token counts close; per-worker duplicate context not measured. | Budget parallel agents only with observed duplicate-context usage. | Unavailable — per-worker duplicate context is not measured |
Module citations: OpenAI API pricing. All AI Ask evidence registry (verified 2026-08-27).
Test-and-repair loop cost tree
Formula / rubric: Expected spend = firstPass + pFail×repair + pReview×feedback + pCancel×restart; probabilities remain user inputs.
Provenance: Bug-fix, feature, and refactor trees expose probability variables instead of inventing success rates. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
bug fixbatch39-coding-agent-m2-r1 | first pass $0.09; repair $0.06; reviewer $0.03; pFail 0.25; pReview 0.10 | Expected = .09 + .25×.06 + .10×.03 = $0.108/task. | Use accepted-task denominator; do not substitute benchmark pass rate. | CALCULATED — expected path. |
featurebatch39-coding-agent-m2-r2 | first pass $0.22; two repairs $0.11/$0.08; pFail and pCancel user-supplied | Best case $0.22; bounded observed path $0.41; expected is Unavailable — pFail and pCancel are user-supplied | No expected value until probabilities are entered. | Unavailable — pFail and pCancel are user-supplied |
refactor/restartbatch39-coding-agent-m2-r3 | tests fail after cancellation; restart context 140K; reviewer feedback 2 turns | Restart token debit exists; cancellation credit is not documented. | Worst-case bound includes full restart; no credit assumed. | Unavailable — cancellation credit is not documented |
Module citations: Anthropic Messages API pricing. All AI Ask evidence registry (verified 2026-08-27).
Parallel-agent budget and SLO envelope
Formula / rubric: Wall clock ≈ ceil(dependency depth/workers)×step time + merge/review + rate-limit wait; spend = workers×duplicated context + accepted work.
Provenance: Frozen 1/3/10-worker envelopes; provider concurrency facts are measured or named Unavailable. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
1 workerbatch39-coding-agent-m3-r1 | dependency depth 8; 12K context/step; 8×42 s; merge 0 | critical path 336 s; duplicated context 0; spend 8×step bill. | Use when ordering dominates and concurrency is not needed. | CALCULATED — serial baseline. |
3 workersbatch39-coding-agent-m3-r2 | depth 8; 3 branches; 12K context each; merge/review 70 s | wall-clock lower bound 182 s; duplicate context 36K; concurrency limit Unavailable — provider limit is not sourced | Do not promise 182 s until limit and queue wait are observed. | Unavailable — provider limit is not sourced |
10 workersbatch39-coding-agent-m3-r3 | 10 branches; 4 shared dependencies; merge 180 s; retries 5% | critical path and queue wait cannot be closed; 10× context budget is explicit. | Run only with a measured provider concurrency limit. | Unavailable — critical path and queue wait cannot be closed |
Module citations: Google Gemini API quotas. All AI Ask evidence registry (verified 2026-08-27).
Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →
