How Much Does a Coding Agent Cost per Month?

At production volume (20,000 calls/month), the cheapest effective option is Amazon Nova Micro at $18.26/month. The most expensive frontier option, GPT-5.4 Pro, runs $19,488/month — A coding agent reads large amounts of file context and writes substantial diffs — both sides of the bill are large compared to a chat turn.

How much does coding agent cost per month?

At production volume (20,000 calls/month), the cheapest effective option for coding agent is Amazon Nova Micro at $18.26 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $19,488 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapelarge file context in, large diff out
Input / output tokens per call20K in / 2K out
Cacheable input60%
Batch-eligibleNo

Input is the files and instructions relevant to one change; output is a multi-file diff plus explanation. Repeated reads of the same files across a session make a majority of input cacheable.

What drives this workload's cost?

The main token-volume driver here is input tokens: each call sends 20K input tokens and requests up to 2K output tokens, at 20,000 calls per month in the default volume. 60% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is not marked batch-eligible because the profile assumes a synchronous response.

Volume

Side project
2,000 calls/mo
$1.83/mo cheapest
Production
20,000 calls/mo
$18.26/mo cheapest
Scale
200,000 calls/mo
$183/mo cheapest

Ranked cost — Production volume

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$19.60$18.260.76×
Amazon Nova LitebudgetAmazon$33.60$32.740.91×
GPT-5 NanobudgetlegacyOpenAI$36.00$26.19
Gemini 2.5 Flash LitebudgetlegacyGoogle$56.00$35.182
GPT-OSS 20BbudgetGroq$42.00$65.162.93×1
Ministral 8BbudgetMistral$66.00$65.580.93×1
Mistral Small 3.1budgetMistral$84.00$80.400.85×4
GPT-4o MinibudgetlegacyOpenAI$84.00$54.58
Grok-3 MinibudgetlegacyxAI$84.00$84.00
GPT-OSS 120BbudgetGroq$84.00$1041.82×
Llama 4 MaverickbudgetlegacyGroq$104$1041
GPT-5.4 NanobudgetlegacyOpenAI$130$79.770.78×1
Muse Spark 1.3 ContributorbudgetMeta$48.00$14412.95×8
CodestralbudgetMistral$156$1480.79×
Gemini 3.1 Flash LitebudgetlegacyGoogle$160$1000.87×
Show all 68 models
GPT-5 MinibudgetlegacyOpenAI$180$1311
GPT-OSS 120B (Cerebras)budgetCerebras$170$2102.32×1
Gemini 3.5 Flash LitebudgetGoogle$220$158
Gemini 2.5 FlashbudgetlegacyGoogle$220$158
Mistral Large 3budgetMistral$260$2601
DeepSeek V4 FlashbudgetDeepSeek$229$3132.59×1
GLM-5.1midlegacyZ.ai$328$328
Qwen 3.8 30BmidGroq$360$360
Qwen 3.6 27BmidlegacyGroq$360$360
Qwen 3.7 PlusmidQwen$400$400
Amazon Nova PromidAmazon$448$448
Gemini 3.7 FlashmidGoogle$450$294
GPT-5.4 MinimidlegacyOpenAI$480$333
Gemini 3.1 FlashmidlegacyGoogle$480$324
Claude Haiku 4.5midAnthropic$600$5031
o3-MinimidlegacyOpenAI$616$4001
GPT-5.6 LunamidOpenAI$640$4441
Grok 4.3midxAI$600$6411.41×3
Muse Spark 1.3midMeta$670$670
Mistral Medium 3midMistral$900$8520.84×5
Qwen 3.8 MaxmidQwen$896$8961
Qwen 3.7 MaxmidQwen$896$8961
Gemini 3.6 FlashmidGoogle$900$5881
GPT-5midlegacyOpenAI$900$6552
Grok-3midlegacyxAI$960$9602
Gemini 3.5 FlashmidlegacyGoogle$960$6482
Grok-4.20 ReasoningmidxAI$1,040$1,0403
Grok-4.20midxAI$1,040$1,0403
Grok 4.6midxAI$1,040$1,0403
Grok 4.5midxAI$1,040$1,0403
DeepSeek V4 PromidDeepSeek$686$1,0513.30×11
Gemini 3.1 PromidGoogle$1,280$6860.63×4
GPT-4.1midlegacyOpenAI$1,120$7281
Claude Sonnet 5midAnthropic$1,200$1,0061
GLM-5.2midZ.ai$736$1,2413.87×14
GPT-4omidlegacyOpenAI$1,400$9101
GPT-5.6 TerramidOpenAI$1,600$1,1101
GPT-5.4midlegacyOpenAI$1,600$1,1101
GLM 4.7 (Cerebras)midCerebras$1,010$1,7287.53×10
Claude Sonnet 4.6midAnthropic$1,800$1,510
Claude Sonnet 4.5midlegacyAnthropic$1,800$1,510
Claude Sonnet 4midlegacyAnthropic$1,800$1,510
GPT-5.6 SolmidOpenAI$2,400$1,615
Claude Opus 4.8midAnthropic$3,000$2,4760.96×
Claude Opus 4.7midlegacyAnthropic$3,000$2,516
Claude Opus 4.6midlegacyAnthropic$3,000$2,516
Claude Opus 4.5midlegacyAnthropic$3,000$2,516
GPT-4 TurbofrontierlegacyOpenAI$5,200$3,239
Claude Fable 5frontierAnthropic$6,000$5,032
Claude Opus 5frontierAnthropic$9,000$7,548
Claude Opus 4.1frontierlegacyAnthropic$9,000$7,548
Claude Opus 4frontierlegacyAnthropic$9,000$7,548
GPT-5.4 ProfrontierlegacyOpenAI$19,200$13,6041.04×

Batch 39 · server-rendered decision evidence · verified 2026-08-27

Coding-agent workload cost mechanics

Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.

Repository-context fan-out ledger

Formula / rubric: Task input = initial scan + retrieved files + repeated state + tool results + cache writes/reads + compaction summaries.

Provenance: Frozen small/medium/large repository shapes at 5, 15, and 30 turns; rates use the dated registry. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
small / 5 turns
batch39-coding-agent-m1-r1
180K repo tokens; 6 retrieved files; 5× tool results 3K; 40% cache hituncached 252K, cached 168K, output 18K; base bill $1.1100 = (168000×$5.00 + 18000×$15.00)/1MCache-read rate must be provider-specific; this is an OpenAI-shaped example.CALCULATED — context fan-out closed.
medium / 15 turns
batch39-coding-agent-m1-r2
620K repo; 18 files; 15 tool results; 2 compactions; 55% cache hitcompaction summary 24K; repeated state 410K; cache write/read split unavailable.Do not treat repeated context as free.Unavailable — cache write/read split is not returned
large / 30 turns
batch39-coding-agent-m1-r3
2.4M repo; 44 files; 30 tool results; 5 compactions; 3 parallel workersretrieval and tool token counts close; per-worker duplicate context not measured.Budget parallel agents only with observed duplicate-context usage.Unavailable — per-worker duplicate context is not measured

Module citations: OpenAI API pricing. All AI Ask evidence registry (verified 2026-08-27).

Test-and-repair loop cost tree

Formula / rubric: Expected spend = firstPass + pFail×repair + pReview×feedback + pCancel×restart; probabilities remain user inputs.

Provenance: Bug-fix, feature, and refactor trees expose probability variables instead of inventing success rates. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
bug fix
batch39-coding-agent-m2-r1
first pass $0.09; repair $0.06; reviewer $0.03; pFail 0.25; pReview 0.10Expected = .09 + .25×.06 + .10×.03 = $0.108/task.Use accepted-task denominator; do not substitute benchmark pass rate.CALCULATED — expected path.
feature
batch39-coding-agent-m2-r2
first pass $0.22; two repairs $0.11/$0.08; pFail and pCancel user-suppliedBest case $0.22; bounded observed path $0.41; expected is Unavailable — pFail and pCancel are user-suppliedNo expected value until probabilities are entered.Unavailable — pFail and pCancel are user-supplied
refactor/restart
batch39-coding-agent-m2-r3
tests fail after cancellation; restart context 140K; reviewer feedback 2 turnsRestart token debit exists; cancellation credit is not documented.Worst-case bound includes full restart; no credit assumed.Unavailable — cancellation credit is not documented

Module citations: Anthropic Messages API pricing. All AI Ask evidence registry (verified 2026-08-27).

Parallel-agent budget and SLO envelope

Formula / rubric: Wall clock ≈ ceil(dependency depth/workers)×step time + merge/review + rate-limit wait; spend = workers×duplicated context + accepted work.

Provenance: Frozen 1/3/10-worker envelopes; provider concurrency facts are measured or named Unavailable. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
1 worker
batch39-coding-agent-m3-r1
dependency depth 8; 12K context/step; 8×42 s; merge 0critical path 336 s; duplicated context 0; spend 8×step bill.Use when ordering dominates and concurrency is not needed.CALCULATED — serial baseline.
3 workers
batch39-coding-agent-m3-r2
depth 8; 3 branches; 12K context each; merge/review 70 swall-clock lower bound 182 s; duplicate context 36K; concurrency limit Unavailable — provider limit is not sourcedDo not promise 182 s until limit and queue wait are observed.Unavailable — provider limit is not sourced
10 workers
batch39-coding-agent-m3-r3
10 branches; 4 shared dependencies; merge 180 s; retries 5%critical path and queue wait cannot be closed; 10× context budget is explicit.Run only with a measured provider concurrency limit.Unavailable — critical path and queue wait cannot be closed

Module citations: Google Gemini API quotas. All AI Ask evidence registry (verified 2026-08-27).

Calculate your repository-agent budget

Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroAmazon cost calculator →LLM Chatbot cost →RAG Question Answering cost →

FAQ

Why is output so much larger here than a chatbot?
A code change often touches multiple functions or files, and a coding model typically explains what it changed — both add up to far more output tokens than a chat reply.
Does verbosity matter more for coding agents?
Yes — a model that pads its diffs with unrequested commentary bills for tokens that add no value, and that shows up directly in the effective-cost ranking below.