Writing a Code Snippet: Which Cheap AI Model Wins?
A focused coding task: produce a correct, efficient, 0-indexed iterative Fibonacci function in Python — and nothing but the code. We sent this exact task to all 20 budget models (each under $3 per million output tokens) through the live API and measured real speed, cost, and accuracy. The full output of every model is included below — click any row to read it.
The prompt
Every model received exactly this, with default generation settings:
Write a Python function called `fib(n)` that returns the n-th Fibonacci number (0-indexed, so fib(0)=0 and fib(1)=1). Use an efficient iterative approach and handle n=0 correctly. Return only the code inside a single code block, with no explanation.
How accuracy was graded
- Algorithm is correct (fib(0)=0, fib(1)=1, fib(n)=fib(n-1)+fib(n-2)).
- Uses an efficient O(n) iterative approach (no naive recursion).
- Handles the n=0 edge case correctly.
- Returns only the code, with no surrounding explanation or stray reasoning.
Results — sortable by speed, cost & accuracy
| # | Model | Accuracy ↓ | Speed ↕ | Cost ↕ | Output |
|---|---|---|---|---|---|
| 1 | GPT-5.4 NanoOpenAI · gpt-5.4-nano | 100★ | 44.2 t/s1742 ms | $0.000109$1.25/M out | |
Agent grade — 100/100. Correct iterative algorithm; fib(0)=0 and fib(1)=1 verified. Clean, code-only, and adds a sensible negative-input guard. Verbatim output (66 in / 77 out tokens, 1742 ms, $0.000109) ```python
def fib(n):
if n < 0:
raise ValueError("n must be non-negative")
if n == 0:
return 0
a, b = 0, 1
for _ in range(1, n):
a, b = b, a + b
return b
``` | |||||
| 2 | Muse Spark 1.3 ContributorMeta · muse-spark-1.3-contributor | 100★ | 190.9 t/s3703 ms | $0.000148$0.2/M out | |
Agent grade — 100/100. Correct iterative algorithm; fib(0)=0 and fib(1)=1 verified. Clean, code-only. A heavy internal reasoner — most of the 707 output tokens were hidden reasoning, making it slower and pricier per answer than its low per-token rate suggests. Verbatim output (67 in / 707 out tokens, 3703 ms, $0.000148) ```python
def fib(n):
if n <= 1:
return n
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 3 | Llama 3.3 70BGroq · llama-3.3-70b | 99 | 198 t/s298 ms | $0.000103$0.79/M out | |
Agent grade — 99/100. Correct and elegant (returns n for n<=1). Clean and code-only. Verbatim output (95 in / 59 out tokens, 298 ms, $0.000103) ```python
def fib(n):
if n <= 1:
return n
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 4 | Llama 3.1 8BGroq · llama-3.1-8b | 99 | 409.7 t/s144 ms | $0.0000095$0.08/M out | |
Agent grade — 99/100. Correct, clean, code-only — and the fastest response in this test. Verbatim output (95 in / 59 out tokens, 144 ms, $0.0000095) ```python
def fib(n):
if n <= 1:
return n
a, b = 0, 1
for i in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 5 | Mistral Medium 3Mistral · mistral-medium | 99 | 36.8 t/s1629 ms | $0.00015$2/M out | |
Agent grade — 99/100. Correct, concise, and strictly code-only. No reasoning leakage. Verbatim output (76 in / 60 out tokens, 1629 ms, $0.00015) ```python
def fib(n):
if n == 0:
return 0
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 6 | Mistral Small 3.1Mistral · mistral-small | 99 | 95.4 t/s629 ms | $0.0000474$0.6/M out | |
Agent grade — 99/100. Correct, concise, code-only — identical quality to the larger Mistral at a fraction of the cost. Verbatim output (76 in / 60 out tokens, 629 ms, $0.0000474) ```python
def fib(n):
if n == 0:
return 0
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 7 | Ministral 8BMistral · ministral-8b | 99 | 43.9 t/s1368 ms | $0.0000186$0.15/M out | |
Agent grade — 99/100. Correct, code-only, and one of the cheapest runs in the test. Verbatim output (64 in / 60 out tokens, 1368 ms, $0.0000186) ```python
def fib(n):
if n == 0:
return 0
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 8 | CodestralMistral · codestral | 99 | 91.9 t/s653 ms | $0.0000732$0.9/M out | |
Agent grade — 99/100. Correct, code-only — exactly what you want from a code-specialised model. Verbatim output (64 in / 60 out tokens, 653 ms, $0.0000732) ```python
def fib(n):
if n == 0:
return 0
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 9 | Llama 4 ScoutGroq · llama-4-scout | 98 | 218 t/s344 ms | $0.0000332$0.34/M out | |
Agent grade — 98/100. Correct with explicit base cases, code-only. No issues. Verbatim output (70 in / 75 out tokens, 344 ms, $0.0000332) ```python
def fib(n):
if n <= 0:
return 0
elif n == 1:
return 1
else:
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 10 | Amazon Nova LiteAmazon · nova-lite | 98 | 133.5 t/s562 ms | $0.0000219$0.24/M out | |
Agent grade — 98/100. Correct, code-only, clean explicit base cases. Verbatim output (65 in / 75 out tokens, 562 ms, $0.0000219) ```python
def fib(n):
if n == 0:
return 0
elif n == 1:
return 1
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 11 | Gemini 3.1 Flash LiteGoogle · gemini-3.1-flash-lite | 97 | 76.4 t/s995 ms | $0.00013$1.5/M out | |
Agent grade — 97/100. Correct and code-only. Returns 0 for negative n silently rather than raising, otherwise spot-on. Verbatim output (63 in / 76 out tokens, 995 ms, $0.00013) ```python
def fib(n):
if n <= 0:
return 0
elif n == 1:
return 1
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 12 | Amazon Nova MicroAmazon · nova-micro | 97 | 137.7 t/s530 ms | $0.0000125$0.14/M out | |
Agent grade — 97/100. Correct and code-only; used a bare ``` fence without a language tag — a trivial nit. Verbatim output (65 in / 73 out tokens, 530 ms, $0.0000125) ```
def fib(n):
if n == 0:
return 0
elif n == 1:
return 1
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 13 | GPT-OSS 120BGroq · gpt-oss-120b | 82 | 297 t/s394 ms | $0.0000899$0.6/M out | |
Agent grade — 82/100. Correct canonical solution with a negative guard, but a <think> block leaked in ahead of the code block. Verbatim output (131 in / 117 out tokens, 394 ms, $0.0000899) <think>
The user wants a Python function fib(n) that returns n-th Fibonacci number using an efficient iterative approach, handling n=0 correctly. Provide only code inside a single code block, no explanation.
We must comply. Provide code.
</think>
```python
def fib(n):
if n < 0:
raise ValueError("n must be a non-negative integer")
a, b = 0, 1
for _ in range(n):
a, b = b, a + b
return a
``` | |||||
| 14 | DeepSeek V4 ProDeepSeek · deepseek-v4-pro | 82 | 88.6 t/s2336 ms | $0.000208$0.87/M out | |
Agent grade — 82/100. Correct with explicit base cases, but preceded by a leaked <think> block. Verbatim output (64 in / 207 out tokens, 2336 ms, $0.000208) <think>
We need to provide a Python function `fib(n)` that returns the n-th Fibonacci number (0-indexed, fib(0)=0, fib(1)=1). Use iterative approach, efficient, handle n=0. Only code in a single code block, no explanation.
So output:
```python
def fib(n):
if n == 0:
return 0
elif n == 1:
return 1
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
```
That's it.
</think>
```python
def fib(n):
if n == 0:
return 0
if n == 1:
return 1
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 15 | Grok 4.3xAI · grok-4.3 | 80 | 28.7 t/s3305 ms | $0.000316$2.5/M out | |
Agent grade — 80/100. The canonical solution is correct, but a (truncated) <think> reasoning block leaked into the output despite the "only the code" instruction. Verbatim output (63 in / 95 out tokens, 3305 ms, $0.000316) <think>
The task is: Write a Python function called `fib(n)` that returns the n-th Fibonacci number (0-indexed, so fib(0)=0 and fib(1)=1). Use an efficient iterative approach and handle n=0 correctly. Return ...
</think>
```python
def fib(n):
a, b = 0, 1
for _ in range(n):
a, b = b, a + b
return a
``` | |||||
| 16 | GPT-OSS 20BGroq · gpt-oss-20b | 80 | 355.9 t/s576 ms | $0.0000713$0.3/M out | |
Agent grade — 80/100. Final code is correct, but the model exposed its <think> reasoning before the code block. Verbatim output (131 in / 205 out tokens, 576 ms, $0.0000713) <think>
User wants a Python function fib(n) that returns nth Fibonacci number 0-indexed. Use efficient iterative approach, handle n=0 correctly. Return only code inside a single code block, no explanation.
We need to output just the code inside a code block. So something like:
```
def fib(n):
a, b = 0, 1
for _ in range(n):
a, b = b, a + b
return a
```
Edge cases: fib(0)=0; works: loop 0 times, a=0. fib(1)=1: loop once, a,b -> a=1,b=1, return a=1. Works. Also negative? Not required.
So output code.
</think>
```python
def fib(n):
a, b = 0, 1
for _ in range(n):
a, b = b, a + b
return a
``` | |||||
| 17 | DeepSeek V4 FlashDeepSeek · deepseek-v4-flash | 80 | 88.8 t/s2860 ms | $0.0000801$0.28/M out | |
Agent grade — 80/100. Correct canonical solution, but leaked a <think> reasoning block before the code. Verbatim output (64 in / 254 out tokens, 2860 ms, $0.0000801) <think>
We need to write a Python function fib(n) that returns the n-th Fibonacci number, 0-indexed. Efficient iterative approach, handle n=0 correctly. Return only code in a single code block, no explanation.
We'll write a simple iterative solution: start a=0, b=1, loop n times, update a,b. For n=0, return a (0). For n=1, after 1 iteration a becomes 1, b becomes 1? Actually careful: typical iterative:
def fib(n):
a, b = 0, 1
for _ in range(n):
a, b = b, a + b
return a
Test: n=0: loop runs 0 times, return a=0. n=1: loop runs 1: a=1, b=1; return a=1. Good. That's efficient O(n), O(1) space. Also handle negative? Not required.
We'll output the code in a code block.
</think>
```python
def fib(n):
a, b = 0, 1
for _ in range(n):
a, b = b, a + b
return a
``` | |||||
| 18 | GPT-OSS 120B (Cerebras)Cerebras · cerebras-gpt-oss-120b | 80 | 863.8 t/s213 ms | $0.000182$0.75/M out | |
Agent grade — 80/100. Correct canonical solution at blistering speed, but a <think> block leaked into the output. Verbatim output (127 in / 184 out tokens, 213 ms, $0.000182) <think>
The user wants a Python function called `fib(n)` that returns the n-th Fibonacci number with 0-indexed. Use efficient iterative approach and handle n=0 correctly. Return only the code inside a single code block, with no explanation. So I need to provide just a code block with the function definition. Simple. Ensure no extra text. Provide something like:
```python
def fib(n):
a, b = 0, 1
for _ in range(n):
a, b = b, a + b
return a
```
That handles n=0 returns 0. That's correct. Provide only code block.
</think>
```python
def fib(n):
a, b = 0, 1
for _ in range(n):
a, b = b, a + b
return a
``` | |||||
| 19 | GLM 4.7 (Cerebras)Cerebras · cerebras-glm-4.7 | 80 | 718.8 t/s377 ms | $0.000892$2.75/M out | |
Agent grade — 80/100. Correct solution, but the model exposed its <think> reasoning before the final code block. Verbatim output (65 in / 271 out tokens, 377 ms, $0.000892) <think>
We need to write a Python function called fib(n) that returns the n-th Fibonacci number (0-indexed). Use an efficient iterative approach and handle n=0 correctly. Return only the code inside a single code block, with no explanation.
So the solution: define function fib(n). Use iterative method: if n <= 1: return n. else: a, b = 0, 1; for i in range(2, n+1): a, b = b, a+b; return b.
Make sure to handle n=0 correctly: returns 0. Also handle n=1 returns 1.
We need to output only the code inside a single code block. So we'll output:
```python
def fib(n):
if n <= 1:
return n
a, b = 0, 1
for _ in range(2, n+1):
a, b = b, a + b
return b
```
That's it. No explanation.
</think>
```python
def fib(n):
if n <= 1:
return n
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
| 20 | Qwen 3 32BGroq · qwen3-32b | 68 | 353.2 t/s4658 ms | $0.00099$0.59/M out | |
Agent grade — 68/100. The final code is correct, but the model dumped a ~1,650-token <think> monologue into the output — a major breach of the code-only instruction and the priciest run here. Verbatim output (68 in / 1645 out tokens, 4658 ms, $0.00099) <think>
Okay, I need to write a Python function called fib(n) that returns the n-th Fibonacci number. Let me think about how to approach this.
First, the problem says it's 0-indexed. So fib(0) is 0, fib(1) is 1. For example, fib(2) should be 1, fib(3) is 2, and so on. Got it.
The user wants an efficient iterative approach. So recursion isn't the way to go here because for larger n, recursion would be inefficient and cause stack overflows. Iterative method is better for both time and space complexity.
Let me plan the steps. For n=0, return 0. For n=1, return 1. For higher numbers, calculate iteratively.
Maybe start with two variables, a and b, representing fib(0) and fib(1). Then for each step from 2 up to n, compute the next Fibonacci number by adding the previous two. Let me think:
Initialize a = 0 (fib(0)), b = 1 (fib(1)). Then for i in range(2, n+1), compute c = a + b. Then update a and b: a becomes the old b, and b becomes c. Wait, but how does this work for each step?
Wait, let's take an example. Let's say n is 2. For i=2, we compute 0+1=1. So the result is 1. Which is correct. For n=3, the loop runs once (i=2 and i=3?), depends on the loop setup. Wait, maybe I need to loop from 2 to n, inclusive?
Alternatively, maybe the loop should run n-1 times. Let me think: if n is 0, return 0. If n is 1, return 1. Else, for n >= 2, we can iterate (n-1) times. Wait, let's think of it as starting with a=0, b=1. For each step, we calculate the next number. So for n=2, we need one step: a + b = 0 + 1 = 1. For n=3, two steps: first step gives 1 (for n=2), second step gives 1 + 1 = 2 (for n=3). So if n is 4, three steps: 0, 1 -> 1, 1 -> 1, 2 -> 2, 3. The result is 3. Which is correct.
So for the general case where n >= 1, we need to perform (n-1) iterations. Because for each iteration, we're moving from the i-th to (i+1)-th Fibonacci number.
Let me structure this:
if n is 0: return 0
else:
a, b = 0, 1
for i in range(2, n+1):
c = a + b
a = b
b = c
return b
Wait, let's check for n=1. If n is 1, the loop would be range(2, 2), which doesn't execute. So returns b which is 1. Correct.
n=2: loop runs once (i=2). c=0+1=1. a becomes 1, b becomes 1. Return b=1. Correct.
n=3: loop runs for i=2 and 3? Wait, range(2, n+1) when n=3 is range(2,4), which is 2 and 3. Wait, but for n=3, how many steps? Let me see. Starting with a=0, b=1. Then for i in 2 to n (inclusive), which for n=3 is two iterations?
Wait wait, for n=3:
Initially a=0, b=1 (for 0 and 1). Then for i=2, which would compute fib(2) as a + b = 1. Then a becomes 1 (previous b) and b becomes 1 (new value). Then i=3, compute a + b = 1 + 1 = 2. a becomes 1, b becomes 2. So return b is 2. Which is correct for fib(3) = 2. So the loop for n=3 runs for two steps (i=2 and i=3). So the number of iterations is (n - 1) steps. Because for n >= 1, the loop runs from 2 to n inclusive. So for n=1, no steps. For n=2, one step. For n=3, two steps. Which seems right.
So the code structure would be:
def fib(n):
if n == 0:
return 0
a, b = 0, 1
for i in range(2, n + 1):
c = a + b
a = b
b = c
return b
Testing this for various values. Let's see:
fib(0) → 0. Correct.
fib(1) → 1. Correct.
fib(2): loop runs once. a=0, b=1 → c=1 → a becomes 1, b becomes 1 → return 1.
fib(3): loop runs twice. First iteration (i=2): same as before. a=1, b=1. Second iteration (i=3): c=1+1=2. a becomes 1, b becomes 2. Return 2. Correct.
fib(4): returns 3. Let's check. After first iteration (i=2): a=1, b=1. i=3: a=1, b=2. i=4: c=1+2=3. a=2, b=3. return 3. Correct.
This seems to work.
What about n=5? Let's compute. The sequence is 0,1,1,2,3,5. For n=5, should return 5. Let's simulate.
Loop runs from 2 to 5: i=2,3,4,5 (4 iterations).
Start with a=0, b=1.
i=2:
c=0+1=1 → a=1, b=1.
i=3:
c=1+1=2 → a=1, b=2.
i=4:
c=1+2=3 → a=2, b=3.
i=5:
c=2+3=5 → a=3, b=5. return 5. Correct.
Looks good.
Now, what if n is 0? The code handles that. So the code seems correct.
Another test case: n=0 returns 0. What if someone passes a negative number? The problem says to handle n=0 correctly, but perhaps the function should assume n is non-negative. Since the problem says nothing about invalid inputs, perhaps no need to handle them. So assuming that n is a non-negative integer.
So the code is correct. The code handles all cases.
The code is efficient: O(n) time, O(1) space. Which is the best possible for this problem without using closed-form formulas (which would be O(1) time but might have precision issues for large n).
Therefore, the Python function as written should work.
</think>
```python
def fib(n):
if n == 0:
return 0
a, b = 0, 1
for _ in range(2, n + 1):
a, b = b, a + b
return b
``` | |||||
Batch 51 · cheap-model-tests-code-snippet-fibonacci evidence contributions. Every board is server-rendered from frozen fixtures; historical outputs are not presented as new runs. Verification date: 2026-09-01.
Deterministic test receipts
These values are read from the exported test records. Prompt hashes identify the immutable test input; raw-output hashes identify each verbatim model response. Version and validation fields are recorded metadata, not a new run or a re-grade.
| Test / model | Prompt SHA-256 | Raw output SHA-256 | Grader | Solver / harness | Validation flags |
|---|---|---|---|---|---|
code-snippet-fibonaccigpt-5.4-nano | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | a57327357e981f19bce34cfb0c23297cc50ded8b330f716c24f57ccfa80bd655 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccigemini-3.1-flash-lite | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 412402373211028756fd80bb8d4dcb36eccff6f30f720459c6d8b76ce0baf11d | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccigrok-4.3 | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | cb0caeb9645e1bae32af74ffcd103c8181c0a774deae8d40a2b64051c67ea769 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccillama-4-scout | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 4f153c75da4a20c233cecae0c093c4443c29afbcddbfea4cfbff72aa26a10986 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccigpt-oss-120b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 9bfe8306c29b98bdce17c9ed9ba9c4cf1049a28bab74a0b5100ff111d73bd7f3 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccigpt-oss-20b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 0677c7b2d9fb1de2c1a8337499c71c8635ae1c8b03730c41a502f4e6d50f3b9c | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccillama-3.3-70b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | d25a76102a8f210d19b5e4a03fd019813491443e11f3afebd03780c35e682936 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccillama-3.1-8b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 116ae97fa49c8db2ecc0a76bf31769aef837266ce264e4bc80ff2a30b6543f3e | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonacciqwen3-32b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | dde27b8a27cec3bbf94b0170dccb6a9b6a685bf200e497a6cbbc5099627535cb | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccideepseek-v4-flash | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 115b101e08808e34b5e082f1075d331d63d5f542516c47b4f7f472b7d06c37d9 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccideepseek-v4-pro | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | c0e8b6d5774ef5d5da18b4b9c9de94b1f71289f1d9ca1cd66ecabfaae344605c | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccimistral-medium | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4b | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccimistral-small | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4b | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonacciministral-8b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4b | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccicodestral | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4b | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccicerebras-gpt-oss-120b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | a1d43cb0157cd3d04bde4fb07d1e3152e86036c295e9832872c4513e58674bb1 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccicerebras-glm-4.7 | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 5b09a6ecae0b0c4ebeaf0fed639ffa7d38aa30e8baea93c2c861070b29a3c3aa | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccinova-lite | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 3e1f8b3777e1bd772c4a101c70e97e395ed1f59eca5dfe703ebc3a2d80427572 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccinova-micro | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 1306c9cd67e804a40617e532eae72b59247ae26c25cadafdbdfbaf2e2ed6783e | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccimuse-spark-1.3-contributor | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | d25a76102a8f210d19b5e4a03fd019813491443e11f3afebd03780c35e682936 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
Code-extraction and instruction-compliance receipt
Frozen Batch 51 fixture board. Formula / decision rule: code-only pass = exactly one valid artifact + no forbidden wrapper/leak + syntax pass Boundary: Raw output is preserved; a failing answer is never silently cleaned before validation.
| Frozen fixture / field ID | Identity keys | Deterministic rule | Output / bounded state | Validation |
|---|---|---|---|---|
batch51-cheap-model-tests-code-snippet-fibonacci-m1-r1clean Python fence | suite=cheap; slug=code-snippet-fibonacci; model/run=exact; raw hash=joined; fence count=1; artifact hash=joined | one parseable Python artifact and no non-code text | parse=pass; code-only=pass | PASS — clean artifact |
batch51-cheap-model-tests-code-snippet-fibonacci-m1-r2explanation-before-code | suite=cheap; slug=code-snippet-fibonacci; model/run=exact; raw hash=joined; non-code text=present; fence count=1 | answer-only rule fails when explanation precedes code | preserve raw text; code-only=fail even if artifact parses | FAIL — wrapper text |
batch51-cheap-model-tests-code-snippet-fibonacci-m1-r3leaked <think> | suite=cheap; slug=code-snippet-fibonacci; model/run=exact; raw hash=joined; reasoning marker=<think> | reasoning marker is a forbidden wrapper for this test | leak=detected; do not remove marker before grading | FAIL — reasoning leak |
batch51-cheap-model-tests-code-snippet-fibonacci-m1-r4multiple code blocks | suite=cheap; slug=code-snippet-fibonacci; model/run=exact; raw hash=joined; fence count=2 | multiple candidate artifacts make deterministic extraction ambiguous | artifact=Unresolved; code-only=fail | UNRESOLVED — multiple artifacts |
batch51-cheap-model-tests-code-snippet-fibonacci-m1-r5syntax error | suite=cheap; slug=code-snippet-fibonacci; model/run=exact; raw hash=joined; fence count=1; parser=error | syntax must pass before execution | parse=fail; preserve parser diagnostic | FAIL — syntax |
batch51-cheap-model-tests-code-snippet-fibonacci-m1-r6no-code fixtures | suite=cheap; slug=code-snippet-fibonacci; model/run=exact; raw hash=joined; fence count=0 | no executable artifact cannot satisfy code test | artifact=missing; code-only=fail | FAIL — no artifact |
Provenance: Batch 51 cheap-model-tests-code-snippet-fibonacci module 1; historical run 2026-06-16; audit verification 2026-09-01. First-party Fibonacci test evidence (JSON). Missing or conflicting joins fail closed.
Fibonacci executable edge-case harness
Frozen Batch 51 fixture board. Formula / decision rule: correctness pass = every in-scope case passes + no side effect + iterative/O(n) gate passes Boundary: Non-integer and negative behavior is out of scope unless the frozen policy says otherwise.
| Frozen fixture / field ID | Identity keys | Deterministic rule | Output / bounded state | Validation |
|---|---|---|---|---|
batch51-cheap-model-tests-code-snippet-fibonacci-m2-r1n=0 | suite=cheap; slug=code-snippet-fibonacci; exact model/run; harness=FIB-v1; input=0; expected=0 | execute extracted function against deterministic expected value | observed=Unavailable until executable receipt joins; side effect=Unknown | UNAVAILABLE — execution receipt |
batch51-cheap-model-tests-code-snippet-fibonacci-m2-r2n=1 | suite=cheap; slug=code-snippet-fibonacci; exact model/run; harness=FIB-v1; input=1; expected=1 | same fixed-case rule; no cross-case inference | pass/fail=Unavailable until run artifact joins | UNAVAILABLE — execution receipt |
batch51-cheap-model-tests-code-snippet-fibonacci-m2-r3n=2 | suite=cheap; slug=code-snippet-fibonacci; exact model/run; harness=FIB-v1; input=2; expected=1 | same fixed-case rule | observed value=Unavailable; preserve error/timeout if present | UNAVAILABLE — execution receipt |
batch51-cheap-model-tests-code-snippet-fibonacci-m2-r4n=10 | suite=cheap; slug=code-snippet-fibonacci; exact model/run; harness=FIB-v1; input=10; expected=55 | execute and record observed value, timeout, and print/mutation side effect | pass/fail=Unavailable until exact run joins | UNAVAILABLE — execution receipt |
batch51-cheap-model-tests-code-snippet-fibonacci-m2-r5n=30 | suite=cheap; slug=code-snippet-fibonacci; exact model/run; harness=FIB-v1; input=30; expected=832040 | execution and iterative gate both required | pass/fail=Unavailable; timeout=preserved if observed | UNAVAILABLE — execution receipt |
batch51-cheap-model-tests-code-snippet-fibonacci-m2-r6non-integer/negative-policy fixtures | suite=cheap; slug=code-snippet-fibonacci; exact model/run; harness=FIB-v1; inputs=-1 and 2.5; policy=out-of-scope | out-of-scope inputs are not correctness cases | state=Out of scope; do not penalize or credit | BOUNDED — policy exclusion |
Provenance: Batch 51 cheap-model-tests-code-snippet-fibonacci module 2; historical run 2026-06-16; audit verification 2026-09-01. First-party Fibonacci test evidence (JSON). Missing or conflicting joins fail closed.
Verified-code observed frontier
Frozen Batch 51 fixture board. Formula / decision rule: eligible = strict-100 + executable pass + iterative pass + code-only pass; missing cost never wins Boundary: This frontier covers one small function, not repository-scale coding.
| Frozen fixture / field ID | Identity keys | Deterministic rule | Output / bounded state | Validation |
|---|---|---|---|---|
batch51-cheap-model-tests-code-snippet-fibonacci-m3-r1strict-100 score | suite=cheap; slug=code-snippet-fibonacci; exact model/run; score=100 gate; all cases=joined; cost/latency=recorded fields | filter to strict score and complete validation receipt | winner/tie=Unavailable until exact eligible set joins | UNAVAILABLE — eligibility |
batch51-cheap-model-tests-code-snippet-fibonacci-m3-r2executable-pass | suite=cheap; slug=code-snippet-fibonacci; exact model/run; harness=FIB-v1; execution=pass; complexity=joined | execution pass alone is insufficient without iterative gate | eligible=Unavailable until complexity receipt joins | UNAVAILABLE — complexity gate |
batch51-cheap-model-tests-code-snippet-fibonacci-m3-r3code-only-pass | suite=cheap; slug=code-snippet-fibonacci; exact model/run; extraction=pass; leak=absent; raw hash=joined | code-only pass is one gate in the frontier | eligible=Unavailable until executable receipt joins | UNAVAILABLE — execution gate |
batch51-cheap-model-tests-code-snippet-fibonacci-m3-r4fastest verified | suite=cheap; slug=code-snippet-fibonacci; verified runs=joined; latency=recorded; correctness=pass | select lowest recorded latency among verified runs only | winner/tie=Unavailable; no live latency claim | UNAVAILABLE — latency vector |
batch51-cheap-model-tests-code-snippet-fibonacci-m3-r5cheapest verified | suite=cheap; slug=code-snippet-fibonacci; verified runs=joined; recorded cost=present | select lowest recorded cost among verified runs | winner/tie=Unavailable; current prices out of scope | UNAVAILABLE — cost vector |
batch51-cheap-model-tests-code-snippet-fibonacci-m3-r6missing-cost policies | suite=cheap; slug=code-snippet-fibonacci; verified run=joined; recorded cost=missing | missing cost is not zero and cannot win cheapest policy | cost frontier=Unavailable; correctness evidence remains separate | BOUNDED — missing cost |
Provenance: Batch 51 cheap-model-tests-code-snippet-fibonacci module 3; historical run 2026-06-16; audit verification 2026-09-01. First-party Fibonacci test evidence (JSON). Missing or conflicting joins fail closed.
Verdict
For code generation on a budget, GPT-5.4 Nano took the top accuracy score (100/100) in this run. If raw throughput matters most, GPT-OSS 120B (Cerebras) was fastest at 863.8 tokens/sec, and Llama 3.1 8B was the cheapest to run. Every model here costs under $3 per million output tokens, so the real differentiator is how cleanly each one follows instructions — expand the rows above to judge for yourself.
