How Much Does an Agentic Tool-Use Loop Cost per Month?
At production volume (25,000 calls/month), the cheapest effective option is Amazon Nova Micro at $27.38/month. The most expensive frontier option, GPT-5.4 Pro, runs $29,232/month — An agent task is not one call — it is a chain of several small round-trips (plan, call a tool, read the result, decide the next step), and the per-task cost is the sum of the whole chain.
How much does agentic tool loop cost per month?
At production volume (25,000 calls/month), the cheapest effective option for agentic tool loop is Amazon Nova Micro at $27.38 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $29,232 per month for the same workload.
Token shape
| Shape | many small round-trips per completed task |
| Input / output tokens per call | 3K in / 0K out × 8 turns |
| Cacheable input | 70% |
| Batch-eligible | No |
"callsPerMonth" here counts completed tasks, each made of 8 model round-trips (plan → tool call → observe, repeated). Tool schemas and growing scratchpad context make most of each turn's input cacheable.
What drives this workload's cost?
The main token-volume driver here is input tokens: each call sends 3K input tokens and requests up to 0K output tokens, at 25,000 calls per month in the default volume. 70% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is not marked batch-eligible because the profile assumes a synchronous response.
Volume
Ranked cost — Production volume
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Amazon Nova Microbudget | Amazon | $29.40 | $27.38 | 0.76× | — |
| Amazon Nova Litebudget | Amazon | $50.40 | $49.10 | 0.91× | — |
| GPT-5 Nanobudgetlegacy | OpenAI | $54.00 | $36.49 | — | — |
| Gemini 2.5 Flash Litebudgetlegacy | $84.00 | $47.29 | — | ▲2 | |
| GPT-OSS 20Bbudget | Groq | $63.00 | $97.74 | 2.93× | ▼1 |
| Ministral 8Bbudget | Mistral | $99.00 | $98.37 | 0.93× | ▲1 |
| Mistral Small 3.1budget | Mistral | $126 | $121 | 0.85× | ▲4 |
| GPT-4o Minibudgetlegacy | OpenAI | $126 | $73.47 | — | — |
| Grok-3 Minibudgetlegacy | xAI | $126 | $126 | — | — |
| GPT-OSS 120Bbudget | Groq | $126 | $156 | 1.82× | — |
| Llama 4 Maverickbudgetlegacy | Groq | $156 | $156 | — | ▲1 |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $195 | $108 | 0.78× | ▲1 |
| Muse Spark 1.3 Contributorbudget | Meta | $72.00 | $215 | 12.95× | ▼8 |
| Codestralbudget | Mistral | $234 | $223 | 0.79× | — |
| Gemini 3.1 Flash Litebudgetlegacy | $240 | $137 | 0.87× | — |
Show all 68 models
| GPT-5 Minibudgetlegacy | OpenAI | $270 | $182 | — | ▲1 |
| GPT-OSS 120B (Cerebras)budget | Cerebras | $255 | $314 | 2.32× | ▼1 |
| Gemini 3.5 Flash Litebudget | $330 | $220 | — | — | |
| Gemini 2.5 Flashbudgetlegacy | $330 | $220 | — | — | |
| Mistral Large 3budget | Mistral | $390 | $390 | — | ▲1 |
| DeepSeek V4 Flashbudget | DeepSeek | $343 | $469 | 2.59× | ▼1 |
| GLM-5.1midlegacy | Z.ai | $492 | $492 | — | — |
| Qwen 3.8 30Bmid | Groq | $540 | $540 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $540 | $540 | — | — |
| Qwen 3.7 Plusmid | Qwen | $600 | $600 | — | — |
| Amazon Nova Promid | Amazon | $672 | $672 | — | — |
| Gemini 3.7 Flashmid | $675 | $400 | — | — | |
| GPT-5.4 Minimidlegacy | OpenAI | $720 | $457 | — | — |
| Gemini 3.1 Flashmidlegacy | $720 | $445 | — | — | |
| Claude Haiku 4.5mid | Anthropic | $900 | $689 | — | ▲1 |
| o3-Minimidlegacy | OpenAI | $924 | $539 | — | ▲1 |
| GPT-5.6 Lunamid | OpenAI | $960 | $610 | — | ▲1 |
| Grok 4.3mid | xAI | $900 | $962 | 1.41× | ▼3 |
| Muse Spark 1.3mid | Meta | $1,005 | $1,005 | — | — |
| Mistral Medium 3mid | Mistral | $1,350 | $1,278 | 0.84× | ▲5 |
| Qwen 3.8 Maxmid | Qwen | $1,344 | $1,344 | — | ▲1 |
| Qwen 3.7 Maxmid | Qwen | $1,344 | $1,344 | — | ▲1 |
| Gemini 3.6 Flashmid | $1,350 | $799 | — | ▲1 | |
| GPT-5midlegacy | OpenAI | $1,350 | $912 | — | ▲2 |
| Grok-3midlegacy | xAI | $1,440 | $1,440 | — | ▲2 |
| Gemini 3.5 Flashmidlegacy | $1,440 | $889 | — | ▲2 | |
| Grok-4.20 Reasoningmid | xAI | $1,560 | $1,560 | — | ▲3 |
| Grok-4.20mid | xAI | $1,560 | $1,560 | — | ▲3 |
| Grok 4.6mid | xAI | $1,560 | $1,560 | — | ▲3 |
| Grok 4.5mid | xAI | $1,560 | $1,560 | — | ▲3 |
| DeepSeek V4 Promid | DeepSeek | $1,030 | $1,576 | 3.30× | ▼11 |
| Gemini 3.1 Promid | $1,920 | $919 | 0.63× | ▲4 | |
| GPT-4.1midlegacy | OpenAI | $1,680 | $980 | — | ▲1 |
| Claude Sonnet 5mid | Anthropic | $1,800 | $1,378 | — | ▲1 |
| GLM-5.2mid | Z.ai | $1,104 | $1,862 | 3.87× | ▼14 |
| GPT-4omidlegacy | OpenAI | $2,100 | $1,225 | — | ▲1 |
| GPT-5.6 Terramid | OpenAI | $2,400 | $1,525 | — | ▲1 |
| GPT-5.4midlegacy | OpenAI | $2,400 | $1,525 | — | ▲1 |
| GLM 4.7 (Cerebras)mid | Cerebras | $1,515 | $2,592 | 7.53× | ▼10 |
| Claude Sonnet 4.6mid | Anthropic | $2,700 | $2,067 | — | — |
| Claude Sonnet 4.5midlegacy | Anthropic | $2,700 | $2,067 | — | — |
| Claude Sonnet 4midlegacy | Anthropic | $2,700 | $2,067 | — | — |
| GPT-5.6 Solmid | OpenAI | $3,600 | $2,199 | — | — |
| Claude Opus 4.8mid | Anthropic | $4,500 | $3,385 | 0.96× | — |
| Claude Opus 4.7midlegacy | Anthropic | $4,500 | $3,445 | — | — |
| Claude Opus 4.6midlegacy | Anthropic | $4,500 | $3,445 | — | — |
| Claude Opus 4.5midlegacy | Anthropic | $4,500 | $3,445 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $7,800 | $4,298 | — | — |
| Claude Fable 5frontier | Anthropic | $9,000 | $6,889 | — | — |
| Claude Opus 5frontier | Anthropic | $13,500 | $10,334 | — | — |
| Claude Opus 4.1frontierlegacy | Anthropic | $13,500 | $10,334 | — | — |
| Claude Opus 4frontierlegacy | Anthropic | $13,500 | $10,334 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $28,800 | $18,727 | 1.04× | — |
Batch 40 · server-rendered decision evidence · verified 2026-08-27
Agentic tool-loop cost evidence
Frozen inputs, formulas, provenance, and decision boundaries are visible in the initial HTML. Unsupported evidence fails closed as Unavailable.
Recurrence-based context ledger
Formula / scoring rule: Cₙ = system + userₙ + Σ(previous assistant/tool) + tool schema + scratchpad; total bill = Σ(inputₙ×rate + outputₙ×rate).
Provenance: Frozen 8-step loop: 3K initial input, 300 output/step, 1.2K tool result/step; cache hit is explicit.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
step 1batch40-agentic-tool-loop-m1-r1 | system 600 + user 3,000 + tools 800; output 300 | Input 4,400; output 300; bill = (4,400×5 + 300×15)/1M = $0.0265. | Constant-per-step shortcut is invalid after history grows. | CALCULATED — recurrence row. |
step 4batch40-agentic-tool-loop-m1-r2 | prior history 6,000; current tool result 1,200; schema 800; output 300 | Input 11,000; cache-hit prefix Unavailable — provider returned cache usage is absent | Do not apply a cache discount from prompt similarity. | Unavailable — provider returned cache usage is absent |
step 8batch40-agentic-tool-loop-m1-r3 | history 14,400; tool result 1,200; output 300; stop=success | Input 17,000; cumulative input 85,600; cumulative output 2,400. | Cost is task-shaped, not 8 × first-step cost. | CALCULATED — recurrence row. |
Module citation: All AI Ask evidence registry (verified 2026-08-27).
Heterogeneous tool settlement tree
Formula / scoring rule: Tool-loop cost = model bill + Σ(tool fee + retry fee + compensating action); unsupported fee remains Unavailable.
Provenance: Frozen search, browser, code, database, and write-action calls; write action is side-effect gated.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
search / browserbatch40-agentic-tool-loop-m2-r1 | 2 searches; 1 browser fetch; 3,100 payload tokens; 1 timeout | Model payload bill calculated; external search fee Unavailable — user tool tariff is not supplied | Do not treat tool fee as zero. | Unavailable — user tool tariff is not supplied |
code / databasebatch40-agentic-tool-loop-m2-r2 | 1 sandbox run; 2 DB reads; 1 retry; 4,800 payload tokens | Retry is counted once; database egress Unavailable — not priced in fixture | External costs remain separate from model tokens. | Unavailable — not priced in fixture |
write actionbatch40-agentic-tool-loop-m2-r3 | 1 side-effecting write; idempotency key w-40; timeout before acknowledgement | Compensating action required; duplicate-effect risk Unavailable — not observable | No successful completion credit until acknowledgement closes. | Unavailable — not observable |
Module citation: OpenAI function calling documentation.
Completed-task budget controller
Formula / scoring rule: Cost/completed = total path cost / completed tasks; stopped and escalated tasks remain in numerator.
Provenance: 5/10/20-step caps; planner/executor routing and completion rates are declared scenarios only.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
5-step capbatch40-agentic-tool-loop-m3-r1 | 25K tasks; 5 steps; completion scenario 82%; budget $0.08 | Stopped 4,500; completed 20,500; cost per completed Unavailable — observed completion rate is absent | Scenario completion is not reliability evidence. | Unavailable — observed completion rate is absent |
10-step capbatch40-agentic-tool-loop-m3-r2 | branch factor 1.3; retry ceiling 2; budget $0.12 | Escalations and over-budget count Unavailable — not measured on a live run | Do not publish a completion winner from assumptions. | Unavailable — not measured on a live run |
planner/executorbatch40-agentic-tool-loop-m3-r3 | cheap planner $0.002/step; premium executor $0.02/step; 70/30 split | Illustrative blended step = .7×.002 + .3×.02 = $0.0074. | Routing cost excludes tool fees and acceptance. | CALCULATED — scenario only. |
Module citation: All AI Ask agent workload registry.
Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →
