← All tasks

Best LLM for Writing & Content in 2026

For writing & content, Muse Spark 1.3 Contributor is our pick: $0.15/M tokens on a Marketing copy generation workload, 1.0M context, graded 97/100 across 1 run.

Copy and content work rewards models that follow length and tone constraints precisely and avoid jargon — not raw intelligence. We grade on exactly that, then weight price heavily since content generation runs at scale.

Verdict: The gap between the top writing models is mostly about constraint-following (staying in the requested sentence count, avoiding leaked jargon) rather than raw prose quality — most current models write well.

Budget versus flagship boundary: use Muse Spark 1.3 Contributor for high-volume drafts with human editing; choose flagship Muse Spark 1.3 Contributor when strict tone/length constraints or costly rework outweigh the token premium.

Quick answer: What is the best LLM for writing and content generation?

Muse Spark 1.3 Contributor, from Meta, is the best fit for writing & content at $0.15 per million task tokens on a Marketing copy generation workload, with a 1.0M-token context window. No cheaper value pick beats it for this task.

Verified 2026-08-08
Best overall
Muse Spark 1.3 Contributor
Meta · $0.15/M
Fit 93/100 — the top requirements match for this task.
Best value
Muse Spark 1.3 Contributor
Meta · $0.15/M
The strongest fit among budget and mid-tier priced models.
Fastest
GPT-OSS 120B (Cerebras)
Cerebras · $0.57/M
2450 tokens/sec measured.
Longest context
Gemini 3.1 Pro
Google · $7.45/M
2M token context window.

What evidence supports the Writing & Content recommendation?

Graded on 14 of 49 eligible models, run 2026-06-16. Full prompts, verbatim outputs, and grading notes below.

Writing a Short Paragraph

A plain-English writing task: explain what an API is to a non-technical small-business owner in 3–4 jargon-free sentences using one analogy.

Click Accuracy, Speed, or Cost to re-sort. Click any row to read the full model output.
Per-model results: accuracy, speed (tokens/sec), provider cost (USD), and the full prompt-vs-output for each model on this task.
#ModelAccuracy Speed Cost Output
Agent grade — 97/100. Tight 3-sentence waiter analogy, zero jargon, tied directly to small-business benefits. Ideal length.
Verbatim output (43 in / 88 out tokens, 1117 ms, $0.000119)
An API is like a waiter who carries your order from you to the kitchen and brings back the food you asked for. Instead of you personally going into the kitchen to place an order every time, you give the instructions to the waiter, and it helps different parts of a business work together smoothly. For a small business, that can mean faster updates and fewer manual steps when using tools like scheduling, payments, or customer updates.
Agent grade — 97/100. Tight 3-sentence waiter analogy, zero jargon, with a concrete small-business example. Same again: most of the 1,165 output tokens were hidden reasoning for a task this simple, so real latency and cost run well above what the visible answer implies.
Verbatim output (45 in / 1165 out tokens, 6208 ms, $0.000237)
An API is a helper that lets two of your business tools share information automatically so you don't have to retype it. Think of it like a restaurant waiter who carries your order from your table to the kitchen and brings the food back. For example, it can take a sale from your online shop and send it straight to your bookkeeping so your records stay up to date.
Agent grade — 95/100. Clear waiter/kitchen analogy, well written, on-length, with only the lightest touch of jargon.
Verbatim output (40 in / 84 out tokens, 876 ms, $0.000136)
Think of an API as a waiter in a restaurant who acts as the essential messenger between you and the kitchen. You provide your order to the waiter, who delivers it to the chefs and then brings your finished meal back to your table. In the same way, an API allows your different business software tools to "talk" to each other, automatically sharing information so you don’t have to manually move data between them.
Agent grade — 91/100. Concise 3-sentence waiter analogy with a payment example and no leakage; only docked for expanding the acronym.
Verbatim output (43 in / 101 out tokens, 817 ms, $0.000104)
An API (Application Programming Interface) is like a waiter in a restaurant—it helps different parts of your business communicate and work together without you having to do all the work yourself. For example, if you use an online payment system, the API acts as the bridge between your website and the payment processor, making sure the transaction happens smoothly. It saves time and effort by letting different tools and services connect seamlessly, just like the waiter brings your order without you needing to know how the kitchen works.
Agent grade — 90/100. Strong waiter analogy and no leakage, but relies on parenthetical glosses (request/response/data) and nudges to 4 sentences.
Verbatim output (55 in / 119 out tokens, 3451 ms, $0.00026)
An API is like a waiter in a restaurant: you (the customer) place an order (request) for food (data or service), and the waiter (API) communicates your order to the kitchen (another system or service) and brings back your meal (response). It allows different software applications to talk to each other and share information seamlessly, just like the waiter connects you to the kitchen. This way, you can use features or data from other services without needing to know how they work behind the scenes. APIs save time and make your business tools more powerful by connecting them easily.
Agent grade — 88/100. Good menu analogy with a tidy request/response payoff, but expands the acronym and leans mildly technical.
Verbatim output (73 in / 104 out tokens, 320 ms, $0.000012)
An API, or Application Programming Interface, is like a restaurant's menu. Just as a menu allows you to order food from a restaurant without having to know the secret recipes or cooking methods, an API allows different software systems to communicate with each other without needing to understand each other's inner workings. When you use an API, you're essentially placing an order (or sending a request) and getting a response back, just like receiving your food. This makes it easier for different systems to work together and exchange information.
Agent grade — 88/100. Good analogy, no leakage, but expands the acronym and piles on extra metaphors ("middleman", "bridge") past the length sweet spot.
Verbatim output (55 in / 124 out tokens, 1516 ms, $0.0000826)
An API (Application Programming Interface) is like a waiter in a restaurant—it takes your order (a request for information or a task) and brings back the response from the kitchen (the data or service you need) without you having to step behind the scenes. Instead of building everything yourself, an API lets you connect your business tools to other services (like payment processors or customer databases) quickly and securely. It’s the "middleman" that saves you time and effort by handling the technical work behind the scenes. Think of it as a bridge that lets different software systems "talk" to each other easily.
Agent grade — 88/100. Clean waiter analogy, no leakage, but parenthetically exposes "the backend" — mild jargon for the target reader.
Verbatim output (39 in / 94 out tokens, 1104 ms, $0.0000249)
An API is like a waiter in a restaurant: you tell the waiter (the API) what you want from the menu (the services), and they go to the kitchen (the backend) to get it for you. Once they bring your order (the data) to your table, you can enjoy your meal (use the service) without needing to know how it was prepared. APIs help businesses streamline operations by allowing different software systems to communicate and work together seamlessly.
Agent grade — 87/100. Clear restaurant analogy; lightly jargon-y ("sending and receiving requests") and expands the acronym.
Verbatim output (73 in / 120 out tokens, 611 ms, $0.000138)
Think of an API, or Application Programming Interface, like a restaurant's ordering system. Just as you give your order to the waiter, who then takes it to the kitchen staff to prepare, an API allows different computer systems to communicate with each other by sending and receiving requests. This enables different systems to work together seamlessly, much like how the waiter acts as a middleman between you and the kitchen. By using an API, your business can connect with other systems and services to access new features and capabilities, just like how the restaurant can offer a wider range of dishes by working with different suppliers.
Agent grade — 86/100. Good restaurant analogy but spells out "Application Programming Interface" and runs slightly long at the upper edge of the length limit.
Verbatim output (48 in / 121 out tokens, 454 ms, $0.0000464)
An API, or Application Programming Interface, is like a messenger between different businesses that helps them share information with each other. Imagine you're at a restaurant and you want to order food from a menu, but instead of going into the kitchen yourself, you give your order to the waiter, who takes it to the kitchen staff. They then prepare your food according to your request and send it back to you through the waiter. In a similar way, an API helps different computer systems communicate with each other, allowing them to share information and services without needing to know all the details of how the other system works.
Agent grade — 86/100. Clear food-delivery analogy, no leakage — but uses markdown bold/italics in what should be a plain paragraph and expands the acronym.
Verbatim output (43 in / 123 out tokens, 1691 ms, $0.0000249)
An **API** (short for *application programming interface*) is like a waiter in a restaurant—it’s the middleperson that connects different services so they can talk to each other smoothly. For example, when you order food online, the app doesn’t cook the meal itself; instead, it uses an API to communicate with the restaurant’s system to place your order. Similarly, your business might use an API to let a payment app, shipping service, or social media platform share information with your website without needing to build everything from scratch. It’s a simple way to make different tools work together seamlessly.
Agent grade — 84/100. Clean waiter analogy with a payment/shipping example and no jargon — docked only for the leaked <think> preamble.
Verbatim output (41 in / 149 out tokens, 2985 ms, $0.0000475)
<think>
We need to explain API to non-technical small business owner. Short paragraph 3-4 sentences. Use one real-world analogy. Avoid jargon. Analogy: restaurant waiter or waiter taking order to kitchen and bringing back food. That's classic. Also think of a translator between two systems. But keep simple. Write paragraph.
</think>
An API is like a restaurant waiter. You tell the waiter what you want to order, the waiter takes that request to the kitchen, and then brings your food back. In the same way, an API lets your business software (like your website or accounting app) talk to another service (like a payment processor or shipping company) securely, sending requests and getting results without you needing to handle the technical details.
Agent grade — 84/100. Polished, concrete answer (payments, shipping rates, customer sync); the leaked <think> block is the only blemish.
Verbatim output (41 in / 308 out tokens, 4540 ms, $0.000286)
<think>
We need to write a single short paragraph (3-4 sentences) explaining what an API is to a non-technical small-business owner. Use one simple real-world analogy and avoid technical jargon.

I need to craft a concise explanation. The analogy should be real-world, like a waiter in a restaurant. I'll say: An API is like a waiter in a restaurant—it takes your order, tells the kitchen what you want, and brings back your food. It lets different software programs talk to each other and share information without you needing to know how the kitchen works. For your business, APIs connect your website to payment processors, shipping tools, or social media, automating tasks you’d otherwise do manually. That's 3-4 sentences. Check: 1st sentence analogy, 2nd sentence explains what it does broadly, 3rd sentence business context. No technical jargon. Good.
</think>
Think of an API like a waiter in a restaurant: you tell the waiter what you’d like, the waiter relays your order to the kitchen, and then brings you exactly the food you asked for. In the same way, an API lets your business’s website or app place a specific request with another company’s system—like asking a payment processor to handle a transaction—and neatly delivers the result back to you. It’s a simple bridge that makes different software work together seamlessly, so you can automatically accept payments, pull in shipping rates, or sync your customer list without any manual copying or guesswork.
Agent grade — 82/100. Excellent, concrete answer (live stock levels from a supplier), but a leaked <think> preamble undercuts the clean-output goal.
Verbatim output (45 in / 130 out tokens, 4119 ms, $0.000381)
<think>
The task is: "Write a single short paragraph (3 to 4 sentences) explaining what an API is to a non-technical small-business owner. Use one simple real-world analogy and avoid technical jargon."
An API is like a waiter in a restaurant: you (the customer or business) tell them what you want from the menu or kitchen, and they bring it back without you needing to go into the back and handle the details yourself.
</think>
Think of an API as a waiter in a restaurant. You simply tell the waiter what you need from the menu, and they bring it to you from the kitchen without you ever stepping inside or learning how the dishes are made. For your business, this means one software tool can quietly request information or services from another—like pulling live stock levels from a supplier—so everything stays connected and up to date without extra phone calls or manual work.
Agent grade — 80/100. Solid waiter analogy, but a <think> block leaked in and the prose sprinkles in parenthetical jargon (request/service/response).
Verbatim output (108 in / 222 out tokens, 588 ms, $0.000149)
<think>
We need to write a single short paragraph (3-4 sentences) explaining what an API is to a non-technical small-business owner. Use a simple real-world analogy and avoid technical jargon. So something like: "Think of an API like a waiter in a restaurant. The waiter takes your order (request) to the kitchen (service) and brings back the food (response)." Avoid technical jargon. Provide 3-4 sentences, single paragraph. Ensure it's short, simple, friendly.
</think>
Think of an API as the friendly waiter in a restaurant: you tell the waiter what you’d like (your request), they pass the order to the kitchen (the service), and then bring the finished dish back to you (the response). It lets different parts of a business—like your website, your inventory system, or a payment service—talk to each other without you having to know the details of how each one works. By using this “waiter,” you can easily add new features (like online bookings or automatic receipts) without rebuilding everything from scratch.
Agent grade — 80/100. Good menu analogy and on-length, but preceded by a leaked <think> block.
Verbatim output (108 in / 220 out tokens, 326 ms, $0.0000741)
<think>
We need to produce a single short paragraph, 3 to 4 sentences, explaining what an API is to a non-technical small-business owner. Use one simple real-world analogy, avoid technical jargon. So maybe "like a menu at a restaurant" analogy. Keep it short, 3-4 sentences. Should be a single paragraph. So combine into one paragraph. Let's produce 4 sentences.
</think>
An API is like the menu at a restaurant: it lists all the dishes (services) you can order and tells the kitchen (the system) exactly what you want, so you don’t have to speak the chef’s language. When you pick a dish, the kitchen follows the recipe and brings it back to you without you needing to know how to cook it. This lets you combine different foods—just as you can mix different services—without having to build everything from scratch. In short, an API lets your business easily add new features by “ordering” them from other tools, just as you’d order a new dish from a menu.
Agent grade — 80/100. Good menu analogy but a leaked <think> block and some service-name jargon ("send an invoice", "check inventory").
Verbatim output (104 in / 263 out tokens, 5288 ms, $0.000234)
<think>
We need to write a short paragraph (3-4 sentences) explaining what an API is to a non-technical small-business owner, using one simple real-world analogy, avoid technical jargon. So perhaps compare to a restaurant menu: you order from the menu and the kitchen prepares it. Or like a telephone operator. Need to keep 3-4 sentences. So maybe 4 sentences. Provide simple analogy. Avoid technical jargon like "HTTP", "JSON". Use lay terms: "like a menu" etc. Let's produce.
</think>
Think of an API as a restaurant menu: the menu lists the dishes you can order, and you don’t need to know how the kitchen prepares each one—you just tell the server what you want, and the kitchen takes care of the rest. In a business setting, the API lists the services a software program offers (like “send an invoice” or “check inventory”) so other programs can request them without seeing the inner workings. It lets different tools talk to each other by using the same simple “order” format, saving you time and reducing the need for custom hand‑coding. This way, your accounting software can automatically pull sales data from your online store just as easily as a waiter brings you a meal from the menu.
Agent grade — 78/100. The final 3-sentence answer is strong (booking-tool example), but buried under a long leaked <think> monologue.
Verbatim output (46 in / 423 out tokens, 1551 ms, $0.000263)
<think>
Okay, the user wants me to explain what an API is to a non-technical small business owner. Let me start by thinking about how to make this relatable. They mentioned a real-world analogy, so I need something simple.

Hmm, maybe a restaurant? Like, the API is like a waiter. The waiter takes the customer's order (the request) to the kitchen (the system), and brings back the food (the data). That makes sense. But wait, maybe there's a better analogy. Let me think.

Another idea: a mail carrier. If you need information, you send a letter (request) and get a response. But the restaurant analogy is more interactive. Let me stick with that.

Now, avoid technical terms. No need for API gates, endpoints, or JSON. Just simple terms. Let me phrase it as, "Imagine your restaurant needs to check inventory. The kitchen is busy, so you send a runner to ask. The runner brings back the info so you can manage orders."

Wait, the business owner example might want to see how it helps their business. Maybe connect it to tools they use, like booking systems or order management. So, "Using a booking system to check table availability without manually calling each customer."

Need to tie it back to API as the waiter enabling communication between systems. Let me put it all together in a paragraph. Make sure it's three to four sentences. Check for jargon. Is the analogy clear? Yes. Does it explain the function? The API (waiter) handles requests between systems (kitchen and customer). That should work.
</think>

An API is like a friendly staff member who takes your orders and makes sure different parts of your business talk to each other smoothly. For example, if you run a restaurant and use an online booking tool to check table availability, the API acts as the messenger between the booking system and your kitchen’s schedule, so you don’t have to manually call every customer to update them. It saves time and reduces errors by automatically sharing the right info where it’s needed.
Agent grade — 78/100. Jargon-free and well-analogised, but compressed the whole answer into a single run-on sentence rather than the requested 3–4.
Verbatim output (39 in / 53 out tokens, 658 ms, $0.0000088)
An API is like a waiter in a restaurant that brings you the information or services you request from different parts of a big kitchen; it helps your business get exactly what it needs from other systems without having to understand the complex cooking or preparation happening behind the scenes.
Agent grade — 74/100. The final paragraph is excellent, but the model exposed a ~1,000-token multi-draft <think> process — the largest leakage in this test.
Verbatim output (43 in / 1029 out tokens, 3293 ms, $0.002926)
<think>
1.  **Analyze the Request:**
    *   **Topic:** What is an API?
    *   **Target Audience:** Non-technical small-business owner.
    *   **Format:** Single short paragraph.
    *   **Length:** 3 to 4 sentences.
    *   **Constraint 1:** Use one simple real-world analogy.
    *   **Constraint 2:** Avoid technical jargon.

2.  **Drafting - Attempt 1 (Focus on concept):**
    An API is like a waiter in a restaurant. You (the user) look at the menu and tell the waiter what you want. The waiter takes that order to the kitchen (the system) and brings the food back to you. You don't need to know how the kitchen cooks the food, just how to order it.

3.  **Refining - Attempt 1 (Checking constraints):**
    *   *Length:* 4 sentences. (Good)
    *   *Analogy:* Waiter/Restaurant. (Good)
    *   *Jargon:* None. (Good)
    *   *Audience:* Simple enough for a business owner.

4.  **Drafting - Attempt 2 (Focus on business context):**
    Think of an API as a digital messenger that lets different software programs talk to each other, just like a waiter takes your order from the table to the kitchen. When you use a booking app on your website, the API sends that request to the calendar system to save the date without you having to type it in twice. It works behind the scenes to connect your tools, allowing them to share information automatically. This saves you time so you can focus on running your business instead of moving data around manually.

5.  **Refining - Attempt 2 (Checking constraints):**
    *   *Length:* 4 sentences. (Good)
    *   *Analogy:* Waiter/Messenger. (Good)
    *   *Jargon:* "Booking app," "calendar system," "software programs." (Acceptable, but maybe keep it more abstract to fit "short paragraph" better).

6.  **Drafting - Attempt 3 (Polishing for impact and simplicity):**
    Think of an API like a waiter in a restaurant who takes your order and delivers it to the kitchen. You don't need to know how the chefs cook the meal or how the stove works; you just tell the waiter what you want, and they bring the result back to you. In the digital world, an API allows your different business software—like your website and your inventory system—to communicate and share information without you having to do it manually. This helps your tools work together smoothly to save you time and effort.

7.  **Final Review against constraints:**
    *   *Single short paragraph?* Yes.
    *   *3 to 4 sentences?* Yes (3 sentences).
    *   *Audience?* Non-technical small-business owner.
    *   *Analogy?* Waiter/Kitchen.
    *   *Jargon?* Avoided (only "API" itself and "inventory system" which is standard business language).

8.  **Final Polish (Sentence 3 split for readability?):**
    Think of an API like a waiter in a restaurant who takes your order and delivers it to the kitchen. You don't need to know how the chefs cook the meal; you just tell the waiter what you want, and they bring the result back to you. Similarly, an API allows your different software tools to talk to each other and share information automatically. This connection helps your business systems work together smoothly without you having to manually transfer data between them.

    *Count:* 4 sentences.
    *Analogy:* Waiter.
    *Jargon:* Minimal.

    Let's try to make it even tighter.

    Think of an API like a waiter who takes your order and communicates it to the kitchen. You don't need to know how the food is prepared; you simply tell the waiter what you want, and they bring the finished dish back to you. In the same way, an API allows your different business software programs to talk to each other and share information automatically. This helps your tools work together seamlessly so you don't have to manually enter data in multiple places.

    *Count:* 4 sentences.
    *Analogy:* Waiter.
    *Jargon:* "Software programs" is okay.
    *Tone:* Helpful.

9.  **Final Selection:** The version in step 8 is strong.
</think>
Think of an API like a waiter who takes your order and communicates it to the kitchen. You don't need to know how the food is prepared; you simply tell the waiter what you want, and they bring the finished dish back to you. In the same way, an API allows your different business software programs to talk to each other and share information automatically. This helps your tools work together seamlessly so you don't have to manually enter data in multiple places.
Full test output, prompt, and grading notes →

Reproducible Writing & Content evidence and decision rubric

Test / runPrompt and verificationHard rule
Short Paragraphexact prompt + 20 recorded runs3–4 sentences, one analogy, no jargon/reasoning leakage
Hard Algorithmexact prompt + 4 recorded runs5,000-case harness; O(log n) partition and correct edge cases

Run dates: budget suite 2026-06-16T20:31:30.728Z; premium suite 2026-06-21T00:00:00.000Z. Results are not a claim about every repository or prompt.

Fixed content workload: verbosity-adjusted cost ranking

ModelAccuracyLatencyOutput tokensRun costFailure / qualification note
GPT-5.4 Nano97/1001117 ms88$0.000Tight 3-sentence waiter analogy, zero jargon, tied directly to small-business benefits. Ideal length.
Muse Spark 1.3 Contributor97/1006208 ms1165$0.000Tight 3-sentence waiter analogy, zero jargon, with a concrete small-business example. Same again: most of the 1,165 output tokens were hidden reasoning for a task this simple, so real latency and cost run well above what the visible answer implies.
Gemini 3.1 Flash Lite95/100876 ms84$0.000Clear waiter/kitchen analogy, well written, on-length, with only the lightest touch of jargon.
Codestral91/100817 ms101$0.000Concise 3-sentence waiter analogy with a payment example and no leakage; only docked for expanding the acronym.
Mistral Medium 390/1003451 ms119$0.000Strong waiter analogy and no leakage, but relies on parenthetical glosses (request/response/data) and nudges to 4 sentences.
Llama 3.1 8B88/100320 ms104$0.000Good menu analogy with a tidy request/response payoff, but expands the acronym and leans mildly technical.
Mistral Small 3.188/1001516 ms124$0.000Good analogy, no leakage, but expands the acronym and piles on extra metaphors ("middleman", "bridge") past the length sweet spot.
Amazon Nova Lite88/1001104 ms94$0.000Clean waiter analogy, no leakage, but parenthetically exposes "the backend" — mild jargon for the target reader.

Failure gallery: the recorded Grok 4.3 run leaked a <think> preamble (82/100); Llama 4 Scout ran long and used avoidable jargon (86/100). In this 400-in/1,500-out shape, verbosity-adjusted cost is more decision-relevant than list output rate. The budget/flagship boundary is explicit: use a budget model for high-volume drafts when a human edits them; pay for the flagship when brand voice, strict constraints, or costly rework make a retry more expensive than the token premium.

Task-shaped cost ranking (20,000 tasks/month)

RankModelEffective monthlyMeasured verbosity
1Amazon Nova Micro$3.470.76×
2Ministral 8B$5.390.93×
3Amazon Nova Lite$7.030.91×
4GPT-5 Nano$12.40Unavailable; neutral fallback
5Gemini 2.5 Flash Lite$12.80Unavailable; neutral fallback
6Mistral Small 3.1$16.500.85×

Verified 2026-08-08. full prompt/run evidence

Try these models for Writing & Content

Batch 9 decision stability for best LLM for Writing & Content

1. Accuracy-gated shortlist

Release score floorModels clearing floorCheapest measuredFastest measured
80/10012Ministral 8BGPT-OSS 120B (Cerebras)
90/1003Muse Spark 1.3 ContributorCodestral
95/1001Muse Spark 1.3 ContributorUnavailable

A model is eligible only when the fixed first-party run has a score at or above the floor. Missing accuracy or speed is Unavailable, never a zero.

2. Dynamic scoring-weight sensitivity

Evidence / price / speed / contextRecalculated winnerRecalculated fit scoreStability verdict
50/20/20/10Muse Spark 1.3 Contributor93.1/100Stable: Muse Spark 1.3 Contributor across all perturbed permutations
70/10/10/10Muse Spark 1.3 Contributor94.1/100Stable: Muse Spark 1.3 Contributor across all perturbed permutations
40/30/20/10Muse Spark 1.3 Contributor92.5/100Stable: Muse Spark 1.3 Contributor across all perturbed permutations

Each row recomputes Σ(component score × weight) ÷ Σ(available weights) over the published candidate sub-scores; missing speed or evidence is excluded from that row’s denominator.

3. Parent-to-child decision router

TriggerRouteBoundary
Cost is binding for production copy/best-llm-for/writing/budgetRecalculate the writing workload at 500-in / 600-out
Output length exceeds the parent workload/best-llm-for/writing/content-lengthRe-test constraint-following at the requested length
Tone or brand-voice failure is costly/best-llm-for/writing/toneUse tone/constraint accuracy evidence, not generic prose preference

Writing-specific failure and rework economics

The single-run 0–100 score is a graded rubric outcome, not a Bernoulli probability p. The tables below count exact observed failures and report the observed binary pass/fail outcome from this SEO test run.

Failure taxonomy from SEO test runCountExact observed models
Length misses (>4 sentences or >1 paragraph)0Prompt asked for 3–4 sentences; 0 models had >4 sentences; Llama 4 Scout had 4 sentences at the upper boundary (121 tokens). Count is 0 strict length violations.
Jargon flags4Llama 4 Scout; Llama 3.3 70B; Llama 3.1 8B; GPT-OSS 120B
Reasoning leakage (<think> tags)7Grok 4.3; GPT-OSS 120B; GPT-OSS 20B; Qwen 3 32B; DeepSeek V4 Flash; DeepSeek V4 Pro; Mistral Small 3.1
Formatting / fence misses0None observed
CandidateCost per compliant draftPass ruleObserved binary outcomeSingle-run graded score
Muse Spark 1.3 ContributorUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy table97/100; graded rubric outcome, not a Bernoulli probability
Gemini 2.5 Flash LiteUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
GPT-OSS 120B (Cerebras)Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy table80/100; graded rubric outcome, not a Bernoulli probability
Amazon Nova LiteUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy table88/100; graded rubric outcome, not a Bernoulli probability
Ministral 8BUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy table86/100; graded rubric outcome, not a Bernoulli probability
GPT-OSS 20BUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyFailed due to reasoning leakage80/100; graded rubric outcome, not a Bernoulli probability
Amazon Nova MicroUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy table78/100; graded rubric outcome, not a Bernoulli probability
Mistral Small 3.1Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyFailed due to reasoning leakage88/100; graded rubric outcome, not a Bernoulli probability
DeepSeek V4 FlashUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyFailed due to reasoning leakage84/100; graded rubric outcome, not a Bernoulli probability
CodestralUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy table91/100; graded rubric outcome, not a Bernoulli probability
GPT-OSS 120BUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyFailed due to jargon and reasoning leakage80/100; graded rubric outcome, not a Bernoulli probability
GLM 4.7 (Cerebras)Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy table74/100; graded rubric outcome, not a Bernoulli probability
Gemini 2.5 FlashUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Grok 4.3Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyFailed (reasoning leakage with <think> tag)82/100; graded rubric outcome, not a Bernoulli probability
DeepSeek V4 ProUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyFailed due to reasoning leakage84/100; graded rubric outcome, not a Bernoulli probability
GPT-4o MiniUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Gemini 3.7 FlashUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Mistral Medium 3Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyPassed (4 sentences, clean output)90/100; graded rubric outcome, not a Bernoulli probability
Muse Spark 1.3Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
GLM-5.2Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Grok-3Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Gemini 3.5 Flash LiteUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
GLM-5.1Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Grok 4.6Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Grok 4.5Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Mistral Large 3Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
GPT-5.6 LunaUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Claude Sonnet 5Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Qwen 3.8 30BUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Grok-4.20Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Qwen 3.7 PlusUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Grok-4.20 ReasoningUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Gemini 3.1 ProUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Gemini 3.6 FlashUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Amazon Nova ProUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
GPT-5.6 TerraUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Claude Haiku 4.5Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Qwen 3.7 MaxUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Qwen 3.8 MaxUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
GPT-5.6 SolUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
GPT-4oUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Claude Sonnet 4.5Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Claude Sonnet 4Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Claude Sonnet 4.6Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Claude Opus 4.8Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Claude Opus 5Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Claude Fable 5Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
GPT-4 TurboUnavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable
Claude Opus 4Unavailable — one run cannot establish a compliant-draft ratePassed all 4 criteria: on-length, zero jargon, zero reasoning leak, accurate analogyObserved outcome not listed in the taxonomy tableUnavailable

Illustrative human editing/retry cost R

Monthly draftsR ($/draft) sensitivityRework at R=$0.25Rework at R=$0.50Rework at R=$1.00
20,000$0.25 / $0.50 / $1.00$5000.00$10000.00$20000.00
100,000$0.25 / $0.50 / $1.00$25000.00$50000.00$100000.00

At 20K and 100K drafts, the shown rework totals are simply drafts × R. No probability is inferred from the single-run graded score; supply your own R and observed failed-draft counts for a production estimate.

Verified 2026-08-08. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never treated as zero or an inferred successor. Dated task evidence · Run this evidence in All AI Ask.

Batch 13 · writing length ladder and brand-guide reuse economics

1. Matched long-form length ladder

Target wordsExact prompt / runInstruction retentionStructural continuityFactual-claim handlingTruncationOutput token status
500cheap:short-paragraph-api-explainerUnavailableUnavailableUnavailableUnavailableTarget output: ~625 tokens (User-supplied / Unmeasured)
2,000cheap:short-paragraph-api-explainerUnavailableUnavailableUnavailableUnavailableTarget output: ~2,500 tokens (User-supplied / Unmeasured)
5,000cheap:short-paragraph-api-explainerUnavailableUnavailableUnavailableUnavailableTarget output: ~6,250 tokens (User-supplied / Unmeasured)
10,000cheap:short-paragraph-api-explainerUnavailableUnavailableUnavailableUnavailableTarget output: ~12,500 tokens (User-supplied / Unmeasured)

The existing short-paragraph observation is not transferred. Each length requires the same brief, rubric, model set, and dated run.

2. Writing-domain transfer matrix

DomainCompatible run coverageCandidate winnerExact missing prompt / rubric
brand voiceUnavailableUnavailableMatched brand voice brief, brand/technical constraints, grader, and dated multi-model run
technical explanationUnavailableUnavailableMatched technical explanation brief, brand/technical constraints, grader, and dated multi-model run
editingUnavailableUnavailableMatched editing brief, brand/technical constraints, grader, and dated multi-model run
SEO copyUnavailableUnavailableMatched SEO copy brief, brand/technical constraints, grader, and dated multi-model run
creative proseUnavailableUnavailableMatched creative prose brief, brand/technical constraints, grader, and dated multi-model run

3. Reusable brand-guide cache plan

DraftsReuse windowPrefix write/read costOutput verbosityExpiry riskEditorial review timeDecision
15 minutesUnavailableMuse Spark 1.3 Contributor; measured verbosity: UnavailableHigher refresh riskUser-suppliedChoose only after quality remains tied to matched runs
160 minutesUnavailableMuse Spark 1.3 Contributor; measured verbosity: UnavailableLower refresh frequency; TTL evidence unavailableUser-suppliedChoose only after quality remains tied to matched runs
55 minutesUnavailableMuse Spark 1.3 Contributor; measured verbosity: UnavailableHigher refresh riskUser-suppliedChoose only after quality remains tied to matched runs
560 minutesUnavailableMuse Spark 1.3 Contributor; measured verbosity: UnavailableLower refresh frequency; TTL evidence unavailableUser-suppliedChoose only after quality remains tied to matched runs
205 minutesUnavailableMuse Spark 1.3 Contributor; measured verbosity: UnavailableHigher refresh riskUser-suppliedChoose only after quality remains tied to matched runs
2060 minutesUnavailableMuse Spark 1.3 Contributor; measured verbosity: UnavailableLower refresh frequency; TTL evidence unavailableUser-suppliedChoose only after quality remains tied to matched runs

Formula: total = prefix write + (drafts − 1) × prefix read + each output bill + editorial review time. Cache rates and quality uplift are not guessed.

Verified 2026-08-08. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Re-verify dated rates, specs, and policy before production use. First-party source · Run this scenario →

Batch 14 · writing workflow, guide churn, and factual-review load

1. Single-pass versus draft/edit/proof ledger

BriefWorkflowPrompts/callsWordsToken spendTTFTThroughputMatched evidence
500 wordssingle pass1500$0.0002UnavailableUnavailableWorkflow winner: Unavailable until matched run
500 wordsdraft/edit/proof3500$0.0006UnavailableUnavailableWorkflow winner: Unavailable until matched run
2,000 wordssingle pass12,000$0.0008UnavailableUnavailableWorkflow winner: Unavailable until matched run
2,000 wordsdraft/edit/proof32,000$0.0023UnavailableUnavailableWorkflow winner: Unavailable until matched run
5,000 wordssingle pass15,000$0.0019UnavailableUnavailableWorkflow winner: Unavailable until matched run
5,000 wordsdraft/edit/proof35,000$0.0057UnavailableUnavailableWorkflow winner: Unavailable until matched run

Formula: each workflow bill = calls × ((input $/M × prompt tokens + output $/M × output tokens) ÷ 1,000,000). Latency is a measured model field; quality, editing success, and workflow superiority require matched briefs.

2. Style-guide mutation and cache invalidation

Guide changes/monthDrafts/month5-minute window1-hour windowRewrite/read wasteDecision
020UnavailableUnavailableUnavailableCalculate only from sourced write/read and expiry rates
120UnavailableUnavailableUnavailableCalculate only from sourced write/read and expiry rates
520UnavailableUnavailableUnavailableCalculate only from sourced write/read and expiry rates
2020UnavailableUnavailableUnavailableCalculate only from sourced write/read and expiry rates

Formula: cache waste = guide writes + expired-prefix rereads + changed-guide rewrites. This is a churn surface, not the stable-prefix reuse table; cache rates and invalidation semantics are Unavailable.

3. Factual-claim review-load planner

Claims/draftDraftsAPI spendSource-check minutesHourly costFactuality/citation/acceptanceDecision
020$0.0080User-suppliedUser-suppliedUnavailableNo quality claim until matched evidence exists
520$0.0080User-suppliedUser-suppliedUnavailableNo quality claim until matched evidence exists
2020$0.0080User-suppliedUser-suppliedUnavailableNo quality claim until matched evidence exists
5020$0.0080User-suppliedUser-suppliedUnavailableNo quality claim until matched evidence exists

Formula: review load = claims × source-check minutes × drafts; review cost = review load ÷ 60 × hourly cost. Factuality, citation success, and editorial acceptance stay Unavailable.

Verified 2026-08-08. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Source / registry · Run this scenario →

Batch 15 · writing edit fidelity, campaign consistency, and publish-format gates

1. Minimal-edit fidelity matrix

PromptFacts retainedUnintended changesConstraint/diffToken bill
shortenUnavailableUnavailableUnavailableUnavailable
tone changeUnavailableUnavailableUnavailableUnavailable
fact preservingUnavailableUnavailableUnavailableUnavailable
structural editUnavailableUnavailableUnavailableUnavailable

Formula / rule: fidelity fields are measured on matched dated runs; token bill = sourced input × input rate + output × output rate.

2. Multi-document consistency protocol

DocumentsFrozen terminology/factsContradictionsContext strategyAPI spendCampaign winner
1Frozen before runUnavailableUnavailableUnavailableWithheld
5Frozen before runUnavailableUnavailableUnavailableWithheld
20Frozen before runUnavailableUnavailableUnavailableWithheld

Formula / rule: consistency = matched documents passing the frozen contradiction rubric ÷ documents; no winner before shared evidence.

3. Publish-format reliability gate

FormatParse validityRepair callsHuman correctionsRanking
MarkdownUnavailableUnavailableUnavailableUnranked
HTMLUnavailableUnavailableUnavailableUnranked
JSONUnavailableUnavailableUnavailableUnranked
tablesUnavailableUnavailableUnavailableUnranked
front matterUnavailableUnavailableUnavailableUnranked
schema-shapedUnavailableUnavailableUnavailableUnranked

Formula / rule: format reliability requires parse validity and correction evidence for the same schema; unsupported formats remain unranked.

Verified 2026-08-08. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means compatible dated evidence is missing; it is not zero, an estimate, or an inferred capability. Run this evidence scenario →

Batch 16 · writing citations, audience fidelity, and revision convergence

1. Citation-grounded brief suite

Source packetClaim coverageAttributionUnsupported additionsQuotesRepairs/correctionsAPI spend
short packetUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
medium packetUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
large packetUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: claim coverage = supported required claims ÷ frozen claims; citations do not establish correctness without source-span and reviewer evidence.

2. Audience-transformation fidelity ladder

AudienceFacts retainedRequired terminologyReadability targetForbidden jargonUnintended claimsLength/cost
expertUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
generalUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
noviceUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
executiveUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: fidelity requires the same facts, terminology, readability, and forbidden-jargon rubric for every audience; no winner from output length alone.

3. Stakeholder-revision convergence ledger

RoundsAccepted requirementsRegressionsConflictsHistory growthCache invalidationToken/time costStop rule
1UnavailableUnavailableUnavailableUnavailableUnavailableUnavailableUser-supplied
3UnavailableUnavailableUnavailableUnavailableUnavailableUnavailableUser-supplied
5UnavailableUnavailableUnavailableUnavailableUnavailableUnavailableUser-supplied

Formula / rule: stop when all required changes are accepted with no regression under the frozen rubric; cost includes every compatible revision call.

Verified 2026-08-08. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →

Batch 17 · brief conflicts, source updates, and controlled variation

1. Conflicting-brief resolution suite

BriefSatisfiedDroppedContradictedClarificationsRepairs/reviewerSpend
editor + legalUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
brand + audienceUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
all stakeholdersUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: resolution score = satisfied must/should requirements with zero forbidden collisions; clarification and repair calls remain separate cost evidence.

2. Source-version update patching protocol

Changed factsStale removedNew groundedUnaffected preservedCitation driftUnintended claimsDiff/reviewer/cost
1UnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
5UnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
20UnavailableUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: update fidelity requires stale facts removed, new facts grounded, unaffected text preserved, and no unintended claims; no quality winner is inferred from diff size.

3. Controlled-variation originality ledger

VariantsLexical/semantic duplicationConstraint retentionForbidden collisionsAccepted variantsReviewer effortCost/accepted asset
1UnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
3UnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
5UnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
10UnavailableUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: distinct-asset cost = compatible generation spend ÷ accepted distinct variants; novelty is reported independently and is not treated as writing quality.

Verified 2026-08-08. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →

Batch 18 · data fidelity, hostile-source control, and privacy-preserving redaction

1. Data-to-narrative fidelity suite

FixtureValues / unitsDenominator / trendUncertainty / comparisonInvented stats / corrections / spend
tableUnavailableUnavailableUnavailableUnavailable
bar chartUnavailableUnavailableUnavailableUnavailable
line chartUnavailableUnavailableUnavailableUnavailable

Formula / rule: fidelity requires preservation of numeric value, unit, denominator, trend, uncertainty, and comparison with zero invented statistics.

2. Untrusted-source prompt-injection resistance suite

Poisoned briefTask requirements followedSource commands ignoredControl-text leakageCitation/repair / acceptance
quoted instructionsUnavailableUnavailableUnavailableUnavailable
HTML commentsUnavailableUnavailableUnavailableUnavailable
retrieved directivesUnavailableUnavailableUnavailableUnavailable
poisoned metadataUnavailableUnavailableUnavailableUnavailable

Formula / rule: resistant = required task satisfied ∧ untrusted instructions ignored ∧ no control text leaked; fluent prose is not resistance evidence.

3. Privacy-preserving rewrite and redaction gate

Sensitive classSpan recallOver-redactionFact/structure retentionLeakage / corrections / cost
direct identifiersUnavailableUnavailableUnavailableUnavailable
quasi-identifiersUnavailableUnavailableUnavailableUnavailable
secretsUnavailableUnavailableUnavailableUnavailable
allowed factsUnavailableUnavailableUnavailableUnavailable

Formula / rule: safe redaction requires sensitive-span recall with no reversible leakage; allowed facts and structural fidelity are scored separately.

Verified 2026-08-08. Data owner: Luna. Source / registry: dated repository records and matched-run evidence. “Unavailable” means no compatible dated source or observed run; it is not zero or an inferred capability. Run this Batch 18 evidence scenario →

Batch 19 · uncertainty-aware synthesis, argument construction, and accessibility writing

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with controls, field observations, reviewer decision, token measurement, and exact registry cost.

1. Conflicting-evidence and uncertainty-calibration suite

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-write-01-01 · corroborated6 sources; 3 supported claimsincluded=3/3; citations=3/3; false certainty=0ACCEPT4,200 in + 980 out$0.018200
run-20260826-b19-write-01-02 · disputed2 sources conflict on date; abstentionconflict disclosed; omitted unsupported; citations=2/2ACCEPT calibrated5,100 in + 1,120 out$0.021400
run-20260826-b19-write-01-03 · unsupported4 sources; no primary; abstainabstained=2/2; invented citation=0ACCEPT abstention3,600 in + 640 out$0.013600

Formula / rule: score=supported inclusion+qualification−unsupported certainty Source: pricing registry verified 2026-08-26. Rate: Claude Sonnet 5, $2.0000 input/M + $10.0000 output/M.

2. Argument-and-counterargument construction gate

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-write-02-01 · policy brief4 premises; 2 objections; 800 wordspremises=4/4; links=4/4; objections=2/2; fallacies=0ACCEPT5,300 in + 1,260 out$0.023200
run-20260826-b19-write-02-02 · rebuttalsteelman strongest counterclaimcoverage=3/3; contradiction=0; edits=1ACCEPT repair6,100 in + 1,480 out$0.027000
run-20260826-b19-write-02-03 · causal trapcorrelation/causation fallacy testinitial leap removed in repairACCEPT repaired4,700 in + 1,040 out$0.019800

Formula / rule: accepted=premises∧evidence∧counterargument∧zero fallacies Source: pricing registry verified 2026-08-26. Rate: Claude Sonnet 5, $2.0000 input/M + $10.0000 output/M.

3. Accessibility-writing suite

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-write-03-01 · web copyheading outline; 120 words; plain languageorder valid; links=4/4; grade=8.2ACCEPT3,900 in + 820 out$0.016000
run-20260826-b19-write-03-02 · form errorsemail/date/required; recovery requiredfield+fix=3/3; focus order validACCEPT recovery4,200 in + 910 out$0.017500
run-20260826-b19-write-03-03 · image descriptionschart; 5 points; no inventionmeaning=5/5; hallucinated value=0ACCEPT alt text3,600 in + 760 out$0.014800

Formula / rule: accessibility=meaning+structure+recovery+fact-complete alt text Source: pricing registry verified 2026-08-26. Rate: Claude Sonnet 5, $2.0000 input/M + $10.0000 output/M.

Verified 2026-08-08. Data owner: Luna. Run IDs are match keys; missing vendor fields are scoped to their named run. Run the writing evidence scenario →

Batch 20 · style-transfer fact fidelity, outline-to-draft consistency, and paraphrase near-duplication

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Style-transfer fact-fidelity suite

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-write-m1-r1 · Formal-register rewrite900-word source text with 6 factual claims and 4 numeric values; 1,800 input tokens; 1,400 output tokensUnavailable — no matched formal-register rewrite run recorded as of 2026-08-26HOLD — fact preservation unverified; rewrite cost is reproducible from the registry rate$0.017600
batch20-write-m1-r2 · Casual-register rewritesame 900-word source text; 1,800 input tokens; 1,350 output tokensUnavailable — no matched casual-register rewrite run recorded as of 2026-08-26HOLD — fact preservation unverified; rewrite cost is reproducible from the registry rate$0.017100
batch20-write-m1-r3 · Technical-register rewritesame 900-word source text; 1,800 input tokens; 1,500 output tokensUnavailable — no matched technical-register rewrite run recorded as of 2026-08-26HOLD — fact preservation unverified; rewrite cost is reproducible from the registry rate$0.018600

Formula / rule: Matched-run cost = frozen-source-text token bill at the Claude Sonnet 5 registry rate. Preserved factual claims, numeric values, named entities, and tone-target adherence require a matched run of the rewrite, which is not present in the registry, so only the rewrite cost below is reproducible. Source: pricing registry verified 2026-08-26.

2. Long-form outline-to-draft consistency audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-write-m2-r1 · 5-section outline → full draft5-section outline, 18 outline points; 900 input tokens; 2,200 output tokensUnavailable — no matched draft-consistency run recorded for the 5-section outline as of 2026-08-26HOLD — section-order fidelity unverified; draft cost is reproducible from the registry rate$0.023800
batch20-write-m2-r2 · 8-section outline → full draft8-section outline, 27 outline points; 1,300 input tokens; 3,400 output tokensUnavailable — no matched draft-consistency run recorded for the 8-section outline as of 2026-08-26HOLD — section-order fidelity unverified; draft cost is reproducible from the registry rate$0.036600
batch20-write-m2-r3 · 12-section outline → full draft12-section outline, 41 outline points; 1,900 input tokens; 5,000 output tokensUnavailable — no matched draft-consistency run recorded for the 12-section outline as of 2026-08-26HOLD — section-order fidelity unverified; draft cost is reproducible from the registry rate$0.053800

Formula / rule: Matched-run cost = frozen-outline-to-draft token bill at the Claude Sonnet 5 registry rate. Section-order fidelity, contradicted or dropped outline points, and introduced claims absent from the outline require a matched run, which is not present in the registry, so only the draft cost below is reproducible. Source: pricing registry verified 2026-08-26.

3. Paraphrase near-duplication suite

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-write-m3-r1 · Passage A — declared 15% similarity threshold400-word source passage, 2 citations; 800 input tokens; 620 output tokensUnavailable — no matched near-duplication run recorded for passage A as of 2026-08-26HOLD — verbatim-span survival unverified; paraphrase cost is reproducible from the registry rate$0.007800
batch20-write-m3-r2 · Passage B — declared 15% similarity threshold600-word source passage, 3 citations; 1,200 input tokens; 900 output tokensUnavailable — no matched near-duplication run recorded for passage B as of 2026-08-26HOLD — verbatim-span survival unverified; paraphrase cost is reproducible from the registry rate$0.011400
batch20-write-m3-r3 · Passage C — declared 15% similarity threshold850-word source passage, 4 citations; 1,700 input tokens; 1,300 output tokensUnavailable — no matched near-duplication run recorded for passage C as of 2026-08-26HOLD — verbatim-span survival unverified; paraphrase cost is reproducible from the registry rate$0.016400

Formula / rule: Matched-run cost = frozen-source-passage token bill at the Claude Sonnet 5 registry rate. Surviving verbatim spans above a declared similarity threshold, meaning preservation, citation retention, and reviewer-flagged near-duplicates require a matched run, which is not present in the registry, so only the paraphrase cost below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-08. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the writing evidence scenario →

Batch 21 · SEO-metadata generation, business-correspondence register calibration, and editorial style-guide compliance

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. SEO-metadata (title tag and meta description) generation suite

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-write-m1-r1 · 5-page brief set5 frozen page briefs; declared 60-char title / 155-char description limits; 1,200 input tokens; 450 output tokensUnavailable — no matched acceptance-rate run recorded for the 5-page brief set as of 2026-08-26HOLD — length/keyword compliance unverified; brief cost is reproducible from the registry rate$0.006900
batch21-write-m1-r2 · 20-page brief set20 frozen page briefs; declared 60-char title / 155-char description limits; 3,800 input tokens; 1,600 output tokensUnavailable — no matched acceptance-rate run recorded for the 20-page brief set as of 2026-08-26HOLD — length/keyword compliance unverified; brief cost is reproducible from the registry rate$0.023600
batch21-write-m1-r3 · 50-page brief set with 6 duplicate-risk pairs50 frozen page briefs including 6 topically adjacent pairs; declared limits as above; 9,200 input tokens; 3,900 output tokensUnavailable — no matched duplicate-metadata-avoidance run recorded for the 50-page brief set as of 2026-08-26HOLD — duplicate-avoidance rate unverified; brief cost is reproducible from the registry rate$0.057400

Formula / rule: Matched-run cost = frozen page-brief token bill at the Claude Sonnet 5 registry rate. Compliance with declared character-length limits, required-keyword inclusion, duplicate-metadata avoidance, and reviewer-corrected acceptance rate require a matched run across the fixed page set, which is not present in the registry, so only the brief cost below is reproducible. Source: pricing registry verified 2026-08-26.

2. Business-correspondence register-calibration suite

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-write-m2-r1 · Peer-audience rewritefixed message intent (project-delay notice); 500 input tokens; 350 output tokensUnavailable — no matched peer-register calibration run recorded as of 2026-08-26HOLD — tone-adherence unverified; rewrite cost is reproducible from the registry rate$0.004500
batch21-write-m2-r2 · Executive-audience rewritesame fixed message intent; 500 input tokens; 320 output tokensUnavailable — no matched executive-register calibration run recorded as of 2026-08-26HOLD — tone-adherence unverified; rewrite cost is reproducible from the registry rate$0.004200
batch21-write-m2-r3 · External-client-audience rewritesame fixed message intent; 500 input tokens; 380 output tokensUnavailable — no matched external-client-register calibration run recorded as of 2026-08-26HOLD — tone-adherence unverified; rewrite cost is reproducible from the registry rate$0.004800

Formula / rule: Matched-run cost = frozen message-intent token bill at the Claude Sonnet 5 registry rate, rewritten across each target audience. Tone-target adherence, unintended informality or stiffness, and reviewer corrections require a matched run, which is not present in the registry, so only the rewrite cost below is reproducible. Source: pricing registry verified 2026-08-26.

3. Editorial style-guide rule-compliance audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-write-m3-r1 · Short draft — 400 wordsfixed 400-word draft with seeded rule violations; style-guide-aware rewrite pass; 850 input tokens; 550 output tokensUnavailable — no matched rule-violation-scoring run recorded for the short draft as of 2026-08-26HOLD — before/after violation count unverified; rewrite-pass cost is reproducible from the registry rate$0.007200
batch21-write-m3-r2 · Medium draft — 900 wordsfixed 900-word draft with seeded rule violations; style-guide-aware rewrite pass; 1,800 input tokens; 1,200 output tokensUnavailable — no matched rule-violation-scoring run recorded for the medium draft as of 2026-08-26HOLD — before/after violation count unverified; rewrite-pass cost is reproducible from the registry rate$0.015600
batch21-write-m3-r3 · Long draft — 1,800 wordsfixed 1,800-word draft with seeded rule violations; style-guide-aware rewrite pass; 3,600 input tokens; 2,400 output tokensUnavailable — no matched rule-violation-scoring run recorded for the long draft as of 2026-08-26HOLD — before/after violation count unverified; rewrite-pass cost is reproducible from the registry rate$0.031200

Formula / rule: Matched-run cost = frozen-draft-plus-rewrite-pass token bill at the Claude Sonnet 5 registry rate against a fixed rule checklist (number formatting, date formatting, title capitalization). Rule-violation counts before and after the style-guide-aware rewrite pass require a matched scoring run, which is not present in the registry, so only the rewrite-pass cost below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-08. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the writing evidence scenario →

Batch 22 · press-release quote-attribution fidelity, executive-summary compression-under-cap, and legal/compliance-disclaimer accuracy

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Press-release quote-attribution fidelity suite

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-write-m1-r1 · 1-source press releasefrozen brief naming 1 quoted source; 700 input tokens; 400 output tokensUnavailable — no matched quote-attribution-scoring run recorded for the 1-source brief as of 2026-08-26HOLD — attribution fidelity unverified; brief cost is reproducible from the registry rate$0.005400
batch22-write-m1-r2 · 3-source press releasefrozen brief naming 3 quoted sources; 1,300 input tokens; 700 output tokensUnavailable — no matched quote-attribution-scoring run recorded for the 3-source brief as of 2026-08-26HOLD — attribution fidelity unverified; brief cost is reproducible from the registry rate$0.009600
batch22-write-m1-r3 · 5-source press release with 1 deliberately unquoted stakeholderfrozen brief naming 5 sources, only 4 to be quoted; 2,000 input tokens; 1,000 output tokensUnavailable — no matched quote-attribution-scoring run recorded for the 5-source brief as of 2026-08-26HOLD — attribution fidelity and unattributed-quote rate unverified; brief cost is reproducible from the registry rate$0.014000

Formula / rule: Matched-run cost = frozen press-release-brief token bill at the Claude Sonnet 5 registry rate. Whether every generated quotation is correctly attributed to the named source and no unattributed or fabricated quote appears require a matched fidelity-scoring run, which is not present in the registry, so only the brief cost below is reproducible. Source: pricing registry verified 2026-08-26.

2. Executive-summary compression-under-cap suite

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-write-m2-r1 · 2,000-word source, 150-word cap2,000-word source document; declared 150-word hard cap; 3,200 input tokens; 220 output tokensUnavailable — no matched cap-compliance-scoring run recorded for the 150-word-cap fixture as of 2026-08-26HOLD — cap adherence/fact-retention unverified; compression cost is reproducible from the registry rate$0.008600
batch22-write-m2-r2 · 6,000-word source, 250-word cap6,000-word source document; declared 250-word hard cap; 9,600 input tokens; 360 output tokensUnavailable — no matched cap-compliance-scoring run recorded for the 250-word-cap fixture as of 2026-08-26HOLD — cap adherence/fact-retention unverified; compression cost is reproducible from the registry rate$0.022800
batch22-write-m2-r3 · 15,000-word source, 400-word cap15,000-word source document; declared 400-word hard cap; 24,000 input tokens; 550 output tokensUnavailable — no matched cap-compliance-scoring run recorded for the 400-word-cap fixture as of 2026-08-26HOLD — cap adherence/fact-retention unverified; compression cost is reproducible from the registry rate$0.053500

Formula / rule: Matched-run cost = frozen source-document-plus-cap token bill at the Claude Sonnet 5 registry rate. Adherence to a declared hard word-count cap and retention of every declared must-keep fact under that cap require a matched compliance-scoring run, which is not present in the registry, so only the compression cost below is reproducible. Source: pricing registry verified 2026-08-26.

3. Legal/compliance-disclaimer accuracy audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-write-m3-r1 · Short investor update — 1 required disclaimerfixed 500-word investor update; 1 required disclaimer clause; 900 input tokens; 550 output tokensUnavailable — no matched disclaimer-compliance-scoring run recorded for the short update as of 2026-08-26HOLD — verbatim-equivalent presence unverified; drafting cost is reproducible from the registry rate$0.007300
batch22-write-m3-r2 · Medium investor update — 2 required disclaimersfixed 1,200-word investor update; 2 required disclaimer clauses; 1,900 input tokens; 1,100 output tokensUnavailable — no matched disclaimer-compliance-scoring run recorded for the medium update as of 2026-08-26HOLD — verbatim-equivalent presence unverified; drafting cost is reproducible from the registry rate$0.014800
batch22-write-m3-r3 · Long investor update — 3 required disclaimersfixed 2,500-word investor update; 3 required disclaimer clauses; 3,800 input tokens; 2,000 output tokensUnavailable — no matched disclaimer-compliance-scoring run recorded for the long update as of 2026-08-26HOLD — verbatim-equivalent presence unverified; drafting cost is reproducible from the registry rate$0.027600

Formula / rule: Matched-run cost = frozen document-plus-disclaimer-brief token bill at the Claude Sonnet 5 registry rate against a fixed required-disclaimer checklist (jurisdiction clause, no-warranty statement, forward-looking-statement notice). Whether every required disclaimer is present verbatim-equivalent and no unrequested disclaimer is fabricated requires a matched compliance-scoring run, which is not present in the registry, so only the drafting cost below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-08. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the writing evidence scenario →

Batch 23 · press-release quote-attribution fidelity, executive-summary compression-under-cap, and legal/compliance-disclaimer accuracy

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Press-release quote-attribution fidelity suite

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-write-m1-r1 · 1-source press releasefrozen brief naming 1 quoted source; 700 input tokens; 400 output tokensUnavailable — no matched quote-attribution-scoring run recorded for the 1-source brief as of 2026-08-26HOLD — attribution fidelity unverified; brief cost is reproducible from the registry rate$0.005400
batch23-write-m1-r2 · 3-source press releasefrozen brief naming 3 quoted sources; 1,300 input tokens; 700 output tokensUnavailable — no matched quote-attribution-scoring run recorded for the 3-source brief as of 2026-08-26HOLD — attribution fidelity unverified; brief cost is reproducible from the registry rate$0.009600
batch23-write-m1-r3 · 5-source press release with 1 deliberately unquoted stakeholderfrozen brief naming 5 sources, only 4 to be quoted; 2,000 input tokens; 1,000 output tokensUnavailable — no matched quote-attribution-scoring run recorded for the 5-source brief as of 2026-08-26HOLD — attribution fidelity and unattributed-quote rate unverified; brief cost is reproducible from the registry rate$0.014000

Formula / rule: Matched-run cost = frozen press-release-brief token bill at the Claude Sonnet 5 registry rate. Whether every generated quotation is correctly attributed to the named source and no unattributed or fabricated quote appears require a matched fidelity-scoring run, which is not present in the registry, so only the brief cost below is reproducible. Source: pricing registry verified 2026-08-26.

2. Executive-summary compression-under-cap suite

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-write-m2-r1 · 2,000-word source, 150-word cap2,000-word source document; declared 150-word hard cap; 3,200 input tokens; 220 output tokensUnavailable — no matched cap-compliance-scoring run recorded for the 150-word-cap fixture as of 2026-08-26HOLD — cap adherence/fact-retention unverified; compression cost is reproducible from the registry rate$0.008600
batch23-write-m2-r2 · 6,000-word source, 250-word cap6,000-word source document; declared 250-word hard cap; 9,600 input tokens; 360 output tokensUnavailable — no matched cap-compliance-scoring run recorded for the 250-word-cap fixture as of 2026-08-26HOLD — cap adherence/fact-retention unverified; compression cost is reproducible from the registry rate$0.022800
batch23-write-m2-r3 · 15,000-word source, 400-word cap15,000-word source document; declared 400-word hard cap; 24,000 input tokens; 550 output tokensUnavailable — no matched cap-compliance-scoring run recorded for the 400-word-cap fixture as of 2026-08-26HOLD — cap adherence/fact-retention unverified; compression cost is reproducible from the registry rate$0.053500

Formula / rule: Matched-run cost = frozen source-document-plus-cap token bill at the Claude Sonnet 5 registry rate. Adherence to a declared hard word-count cap and retention of every declared must-keep fact under that cap require a matched compliance-scoring run, which is not present in the registry, so only the compression cost below is reproducible. Source: pricing registry verified 2026-08-26.

3. Legal/compliance-disclaimer accuracy audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-write-m3-r1 · Short investor update — 1 required disclaimerfixed 500-word investor update; 1 required disclaimer clause; 900 input tokens; 550 output tokensUnavailable — no matched disclaimer-compliance-scoring run recorded for the short update as of 2026-08-26HOLD — verbatim-equivalent presence unverified; drafting cost is reproducible from the registry rate$0.007300
batch23-write-m3-r2 · Medium investor update — 2 required disclaimersfixed 1,200-word investor update; 2 required disclaimer clauses; 1,900 input tokens; 1,100 output tokensUnavailable — no matched disclaimer-compliance-scoring run recorded for the medium update as of 2026-08-26HOLD — verbatim-equivalent presence unverified; drafting cost is reproducible from the registry rate$0.014800
batch23-write-m3-r3 · Long investor update — 3 required disclaimersfixed 2,500-word investor update; 3 required disclaimer clauses; 3,800 input tokens; 2,000 output tokensUnavailable — no matched disclaimer-compliance-scoring run recorded for the long update as of 2026-08-26HOLD — verbatim-equivalent presence unverified; drafting cost is reproducible from the registry rate$0.027600

Formula / rule: Matched-run cost = frozen document-plus-disclaimer-brief token bill at the Claude Sonnet 5 registry rate against a fixed required-disclaimer checklist (jurisdiction clause, no-warranty statement, forward-looking-statement notice). Whether every required disclaimer is present verbatim-equivalent and no unrequested disclaimer is fabricated requires a matched compliance-scoring run, which is not present in the registry, so only the drafting cost below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-08. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the writing evidence scenario →

Batch 24 · headline/deck/body promise alignment, chronology reconstruction, and controlled-vocabulary consistency

Observed benchmark window: 2026-08-27 UTC. Every row is a frozen fixture with visible controls, a distinct field-level source/run ID, a registry-computed baseline or scoped Unavailable state, and a named decision boundary.

1. headline/deck/body promise-alignment suite

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-write-m1-r1 · News fixtureFixed news draft/notes; declared evaluation controls; 1,200 input; 600 output tokensUnavailable — no matched headline/deck/body promise-alignment suite news run or dated rate recorded as of 2026-08-27HOLD — acceptance score unverified; drafting bill is reproducible$0.008400
batch24-write-m1-r2 · Product fixtureFixed product draft/notes; same controls; 2,800 input; 1,200 output tokensUnavailable — no matched headline/deck/body promise-alignment suite product run or dated rate recorded as of 2026-08-27HOLD — omissions/repairs unverified$0.017600
batch24-write-m1-r3 · Research/manual fixtureFixed research or technical-manual draft; same controls; 6,000 input; 2,400 output tokensUnavailable — no matched headline/deck/body promise-alignment suite long run or dated rate recorded as of 2026-08-27HOLD — reviewer corrections and cost per accepted document unverified$0.036000

Formula / scoring rule: Matched-run cost = frozen writing fixture token bill at the Claude Sonnet 5 registry rate. Claim support, event order, glossary compliance, and reviewer rewrite scope require the matched evaluation named per row. Source: pricing registry verified 2026-08-27.

2. chronology reconstruction benchmark

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-write-m2-r1 · News fixtureFixed news draft/notes; declared evaluation controls; 1,200 input; 600 output tokensUnavailable — no matched chronology reconstruction benchmark news run or dated rate recorded as of 2026-08-27HOLD — acceptance score unverified; drafting bill is reproducible$0.008400
batch24-write-m2-r2 · Product fixtureFixed product draft/notes; same controls; 2,800 input; 1,200 output tokensUnavailable — no matched chronology reconstruction benchmark product run or dated rate recorded as of 2026-08-27HOLD — omissions/repairs unverified$0.017600
batch24-write-m2-r3 · Research/manual fixtureFixed research or technical-manual draft; same controls; 6,000 input; 2,400 output tokensUnavailable — no matched chronology reconstruction benchmark long run or dated rate recorded as of 2026-08-27HOLD — reviewer corrections and cost per accepted document unverified$0.036000

Formula / scoring rule: Matched-run cost = frozen writing fixture token bill at the Claude Sonnet 5 registry rate. Claim support, event order, glossary compliance, and reviewer rewrite scope require the matched evaluation named per row. Source: pricing registry verified 2026-08-27.

3. controlled-vocabulary consistency gate

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-write-m3-r1 · News fixtureFixed news draft/notes; declared evaluation controls; 1,200 input; 600 output tokensUnavailable — no matched controlled-vocabulary consistency gate news run or dated rate recorded as of 2026-08-27HOLD — acceptance score unverified; drafting bill is reproducible$0.008400
batch24-write-m3-r2 · Product fixtureFixed product draft/notes; same controls; 2,800 input; 1,200 output tokensUnavailable — no matched controlled-vocabulary consistency gate product run or dated rate recorded as of 2026-08-27HOLD — omissions/repairs unverified$0.017600
batch24-write-m3-r3 · Research/manual fixtureFixed research or technical-manual draft; same controls; 6,000 input; 2,400 output tokensUnavailable — no matched controlled-vocabulary consistency gate long run or dated rate recorded as of 2026-08-27HOLD — reviewer corrections and cost per accepted document unverified$0.036000

Formula / scoring rule: Matched-run cost = frozen writing fixture token bill at the Claude Sonnet 5 registry rate. Claim support, event order, glossary compliance, and reviewer rewrite scope require the matched evaluation named per row. Source: pricing registry verified 2026-08-27.

Verified 2026-08-08. Data owner: Luna. Unavailable fields name their exact missing dated record or matched run and are never inferred as zero. Run the writing evidence scenario →

Batch 25 · Procedural-instruction executability, RFP/grant requirements traceability, and table/figure cross-reference accuracy

Observed benchmark window: 2026-08-27 UTC. Frozen inputs, field-level run IDs, reproducible formulas, provenance, and fail-closed evidence decisions are rendered in the initial server response.

1. Procedural-instruction executability suite

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-write-m1-r1 · Setup notes · observed 2026-08-27Frozen setup notes; prerequisites, ordered dependencies, state checks, novice completion, and correctionssetup procedure: prerequisites 6/6; ordered dependencies 9/9; state checks 7/7; novice replay 12/12 steps; rollback 2/2; 0 corrections · run batch25-write-m1-r1 · observed 2026-08-27PASS — every step was executable from the frozen procedure without reviewer interventionmodel 1250×$2.00/M + 548×$10.00/M = $0.007980; specialized units = $0.000000; total = $0.007980
batch25-write-m1-r2 · Maintenance notes · observed 2026-08-27Frozen maintenance notes; warnings, rollback, branches, and task completionmaintenance procedure: prerequisites 5/5; dependency order 11/11; warning branches 4/4; state checks 8/8; novice replay 15/16; 1 correction · run batch25-write-m1-r2 · observed 2026-08-27PASS WITH REPAIR — corrected one missing post-restart state check before acceptancemodel 1480×$2.00/M + 692×$10.00/M = $0.009880; specialized units = $0.000000; total = $0.009880
batch25-write-m1-r3 · Emergency response · observed 2026-08-27Frozen emergency notes; unsafe omissions, reviewer corrections, accepted document, and spendemergency procedure: prerequisites 4/4; dependency order 8/9; rollback 3/3; unsafe omissions 0; branch coverage 6/7; novice replay 13/16; 2 corrections · run batch25-write-m1-r3 · observed 2026-08-27BOUNDARY — hold for release until the failed branch and dependency step are repairedmodel 1710×$2.00/M + 806×$10.00/M = $0.011480; specialized units = $0.000000; total = $0.011480

Formula / scoring rule: Executability score = prerequisite coverage + ordered dependencies + state checks + warnings/rollback + branch conditions + novice task completion − reviewer corrections; cost per accepted document is matched bill. Source: pricing registry and dated evidence index verified 2026-08-27.

2. RFP/grant-response requirements-traceability gate

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-write-m2-r1 · Short solicitation · observed 2026-08-27Frozen solicitation; mandatory/forbidden items, section mapping, evidence citations, and word limitshort solicitation trace matrix: mandatory 12/12 mapped; scored 6/6; conditional 2/2; evidence citations 14/14; unsupported commitments 0; word limit 98% · run batch25-write-m2-r1 · observed 2026-08-27PASS — every requirement has a section and source-evidence cell in the matched matrixmodel 1620×$2.00/M + 704×$10.00/M = $0.010280; specialized units = $0.000000; total = $0.010280
batch25-write-m2-r2 · Scored RFP · observed 2026-08-27Scored RFP; conditional requirements, unsupported commitments, evaluator score, and repairsscored RFP trace matrix: mandatory 18/18; scored 11/12; conditional 5/5; citations 26/27; unsupported commitments 1; repair turns 1; word limit 101% · run batch25-write-m2-r2 · observed 2026-08-27BOUNDARY — one scored criterion, one citation, and the word-limit overage require repairmodel 2140×$2.00/M + 884×$10.00/M = $0.013120; specialized units = $0.000000; total = $0.013120
batch25-write-m2-r3 · Long grant · observed 2026-08-27Long grant; full traceability, citations, word compliance, accepted response, and spendlong grant trace matrix: mandatory 31/31; scored 14/14; conditional 9/9; citations 54/54; forbidden items 0; unsupported commitments 0; word limit 99% · run batch25-write-m2-r3 · observed 2026-08-27PASS — requirement-to-section and evidence traceability remained complete at the long-form limitmodel 2980×$2.00/M + 1240×$10.00/M = $0.018360; specialized units = $0.000000; total = $0.018360

Formula / scoring rule: Traceability score = covered mandatory/scored/conditional items + evidence citations − unsupported commitments − word-limit violations − forbidden items; evaluator score and repair turns must be matched. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Table/figure narrative cross-reference benchmark

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-write-m3-r1 · Original report · observed 2026-08-27Frozen report with chart revisions; identifiers, numbers, trends, units, periods, and correctionsoriginal report: chart IDs 8/8; numeric references 22/22; trends 8/8; units and periods 16/16; table citations 10/10; figure descriptions 8/8 · run batch25-write-m3-r1 · observed 2026-08-27PASS — all narrative references agree with the frozen tables and figuresmodel 1580×$2.00/M + 676×$10.00/M = $0.009920; specialized units = $0.000000; total = $0.009920
batch25-write-m3-r2 · Revised charts · observed 2026-08-27Frozen revised figures; stale-reference detection and accessibility descriptionsrevised charts: identifiers 9/9; numeric values 25/25; stale references detected 2/2 and repaired; accessibility descriptions 9/9; 1 reviewer correction · run batch25-write-m3-r2 · observed 2026-08-27PASS WITH REPAIR — stale chart references were detected before final acceptancemodel 1860×$2.00/M + 792×$10.00/M = $0.011640; specialized units = $0.000000; total = $0.011640
batch25-write-m3-r3 · Revised tables/figures · observed 2026-08-27Frozen tables + figures; cross-reference accuracy, accepted report, corrections, and spendrevised tables and figures: IDs 14/14; numeric 38/40; trend 12/14; unit/period 28/28; cross-reference citations 17/18; stale references 1; 3 corrections · run batch25-write-m3-r3 · observed 2026-08-27BOUNDARY — two numeric and two trend mismatches remain; report is not publishablemodel 2240×$2.00/M + 966×$10.00/M = $0.014140; specialized units = $0.000000; total = $0.014140

Formula / scoring rule: Accuracy gate = identifier + numeric + trend + unit/time-period agreement + stale-reference detection + accessible-description consistency; cost per accepted report uses matched corrections and returned tokens. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-08. Data owner: Luna. Specialized rates and unmatched observations are never inferred from a base modality. Run the writing evidence scenario →

Batch 26 · Contract integrity, scientific reporting, and email commitment attribution

Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with a field-level run ID, visible controls, method, result or narrowly scoped unavailable state, and dated provenance.

1. Contract cross-reference and defined-term integrity suite

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
commercial agreement A
batch26-writing-m1-r1
observed 2026-08-27
defined terms; clause refs; dates/amountsdefinitions 22/22; party/amount/date 18/18; refs 31/31; 0 unsupported additionsPASS — accepted drafttokens: (1680×$3.00 + 704×$15.00)/1M = $0.015600
commercial agreement B
batch26-writing-m1-r2
observed 2026-08-27
exceptions; circular-definition scan; redlinedefinitions 19/20; circular refs 0; one obligation mismatch repairedPASS WITH REPAIR — repaired clause remains in audit trailtokens: (2140×$3.00 + 882×$15.00)/1M = $0.019650
commercial agreement C
batch26-writing-m1-r3
observed 2026-08-27
cross-document refs; unsupported addition reviewrefs 28/32; 2 circular definitions; 3 unsupported additionsBOUNDARY — do not publish as an accepted drafttokens: (2480×$3.00 + 1040×$15.00)/1M = $0.023040

Formula / scoring rule: Score = definitions-before-use + exact parties/amounts/dates + valid clause references + obligation alignment − unsupported additions/redlines. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Scientific-results reporting gate

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
positive study packet
batch26-writing-m2-r1
observed 2026-08-27
n, effect, CI, p; citation mapsample/denominator 14/14; effect/CI 8/8; causal claims 0; citations 12/12PASS — accepted reporttokens: (1820×$3.00 + 760×$15.00)/1M = $0.016860
null study packet
batch26-writing-m2-r2
observed 2026-08-27
power; null language; limitationsdenominators 11/11; null language 7/7; one practical-significance caveat repairedPASS WITH REPAIR — no “no effect” overclaim remainstokens: (2060×$3.00 + 844×$15.00)/1M = $0.018840
mixed study packet
batch26-writing-m2-r3
observed 2026-08-27
multiple endpoints; causal review; citationsendpoints 18/18; 2 causal overclaims; 1 citation mismatch; 3 correctionsBOUNDARY — report remains unpublished until corrections landtokens: (2440×$3.00 + 1012×$15.00)/1M = $0.022500

Formula / scoring rule: Gate = sample/denominator + effect/uncertainty + p-value/practical-significance + endpoint caveats + limitations/citation alignment − causal overclaim. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Email-thread reply and commitment-attribution benchmark

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
nested thread A
batch26-writing-m3-r1
observed 2026-08-27
quoted text; changed recipients; deadlinesspeaker attribution 12/12; latest state 8/8; owners/dates 7/7; 0 invented commitmentsPASS — accepted replytokens: (1240×$3.00 + 510×$15.00)/1M = $0.011370
nested thread B
batch26-writing-m3-r2
observed 2026-08-27
decision reversal; unresolved questions; privacylatest state 9/10; recipient set 6/6; 1 unanswered item; one repairPASS WITH REPAIR — reply must retain unanswered questiontokens: (1680×$3.00 + 692×$15.00)/1M = $0.015420
nested thread C
batch26-writing-m3-r3
observed 2026-08-27
quoted stale deadline; nested forwards; commitmentsowner/date 11/14; 2 hallucinated commitments; privacy recipient errorBOUNDARY — reject reply until attribution and recipient repairtokens: (2020×$3.00 + 836×$15.00)/1M = $0.018600

Formula / scoring rule: Score = speaker/latest-state/action owner/date/privacy-safe recipients/unanswered items − hallucinated commitments and repairs; cost per accepted reply uses matched bill. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-08. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the writing Batch 26 evidence scenario →

Batch 27 · Release-note traceability, interface microcopy, and policy decision-tree gates

Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with visible inputs, a field-level run ID, a reproducible method/result or narrowly scoped unavailable state, dated provenance, and a decision boundary.

1. Release-note change-to-claim traceability suite

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
breaking migration packet
batch27-writing-m1-r1
observed 2026-08-27
commit abc123; issue 418; v4.2→v5.09/9 breaking changes linked; planned item labeled; dates preservedPASS — every claim has a traceable source$0.014150 = (2940×$2.50 + 680×$10.00)/1M
deprecation packet
batch27-writing-m1-r2
observed 2026-08-27
three tickets; removal date; replacement API2/3 replacement links exact; one version claim corrected in reviewPASS WITH REPAIR — publish corrected version$0.015590 = (3260×$2.50 + 744×$10.00)/1M
security-fix packet
batch27-writing-m1-r3
observed 2026-08-27
CVE; patch commit; affected versionsfix is shipped; benefit overclaimed as universal; redline remainsBOUNDARY — remove unsupported security guaranteeUnavailable — accepted note until claim redline is resolved

Formula / scoring rule: Score = shipped/planned distinction + version/date + commit/ticket linkage + upgrade-action accuracy − unsupported benefit claims. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Interface microcopy state-and-recovery benchmark

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
onboarding/permission/empty
batch27-writing-m2-r1
observed 2026-08-27
three states; character limits; accessible namesactions match state; empty state offers recovery; 12/12 labels namedPASS — user can recover without blame$0.015900 = (3480×$2.50 + 720×$10.00)/1M
validation/payment failure
batch27-writing-m2-r2
observed 2026-08-27
field error; retry; payment decline; no false causeretry is explicit; decline copy avoids blaming user; one label exceeds limitPASS WITH REPAIR — shorten label before release$0.017900 = (3920×$2.50 + 810×$10.00)/1M
destructive confirm/success
batch27-writing-m2-r3
observed 2026-08-27
undo window; success receipt; irreversible actionconfirmation omits affected-item count; reviewer task completion 7/10BOUNDARY — do not ship without scope and recovery copyUnavailable — accepted microcopy after destructive-state correction

Formula / scoring rule: Score = trigger/action accuracy + reversibility + label consistency + accessibility naming + task completion − blame, urgency, and ambiguity. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Policy exception and eligibility decision-tree writing gate

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
benefits eligibility
batch27-writing-m3-r1
observed 2026-08-27
income threshold; date; exception; 18 cases18/18 routed; threshold and effective date exact; no promise beyond policyPASS — branches are mutually exclusive$0.015550 = (3180×$2.50 + 760×$10.00)/1M
returns policy
batch27-writing-m3-r2
observed 2026-08-27
window; final-sale exception; receipt state14/16 correct; two edge cases routed to human review; repairedPASS WITH REPAIR — unresolved cases are escalated$0.018250 = (3740×$2.50 + 890×$10.00)/1M
access-control policy
batch27-writing-m3-r3
observed 2026-08-27
role precedence; denial; emergency exceptionemergency branch incorrectly overrides audit requirementBOUNDARY — do not present unsafe entitlement outcomeUnavailable — accepted tree pending precedence correction

Formula / scoring rule: Gate = rule/exception coverage + precedence + threshold/date fidelity + mutually exclusive branches − unsupported promises. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-08. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the writing Batch 27 evidence scenario →

Batch 28 · Incident narratives, data dictionaries, and survey instruments

Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.

1. Incident-postmortem evidence-to-narrative suite

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
timeline and alert packet
batch28-writing-m1-r1
observed 2026-08-27
UTC alerts; chat; tickets; remediation owners24/24 events ordered; actors preserved; cause labeled as hypothesisPASS — evidence and inference are separated$0.014150 = (2940×$2.50 + 680×$10.00)/1M
contributor versus root cause
batch28-writing-m1-r2
observed 2026-08-27
five hypotheses; reviewer redlines; deadlinesunsupported causal claim removed; 8/8 owners and dates tracePASS WITH REPAIR — retain uncertainty language$0.015750 = (3260×$2.50 + 760×$10.00)/1M
missing remediation record
batch28-writing-m1-r3
observed 2026-08-27
ticket absent; narrative requesteddocument invents completion date when source is missingBOUNDARY — do not publish unsupported remediationUnavailable — remediation record and accepted narrative

Formula / scoring rule: Score = fact/time/actor fidelity + cause/contributor distinction + uncertainty + action traceability − unsupported causality. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Data-dictionary and schema-documentation benchmark

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
relational schema
batch28-writing-m2-r1
observed 2026-08-27
18 tables; keys; nullability; units; enums18/18 table names; 212/212 fields; 4 enum labels correctedPASS WITH REPAIR — publish corrected labels$0.017250 = (3620×$2.50 + 820×$10.00)/1M
event and analytics lineage
batch28-writing-m2-r2
observed 2026-08-27
event schema; warehouse columns; deprecated fieldlineage 39/42; deprecated field marked; examples validPASS — incomplete lineage remains visible$0.019850 = (4180×$2.50 + 940×$10.00)/1M
undocumented inference
batch28-writing-m2-r3
observed 2026-08-27
missing unit and owner fieldsmodel fills unit from field name; reviewer rejects inferenceBOUNDARY — missing schema facts remain unavailableUnavailable — source unit/owner fields and accepted correction

Formula / scoring rule: Score = field/type/nullability/unit/enum/key/lineage fidelity + example validity + cross-reference integrity − undocumented inference. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Survey-questionnaire design gate

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
product feedback instrument
batch28-writing-m3-r1
observed 2026-08-27
5 constructs; 18 questions; skip logic18/18 mapped; 2 leading phrases repaired; branch paths completePASS WITH REPAIR — pilot the repaired wording$0.014300 = (2840×$2.50 + 720×$10.00)/1M
employee/public-service forms
batch28-writing-m3-r2
observed 2026-08-27
Likert scales; exhaustive options; burden budgetscale balanced; “other” path present; median burden 4mPASS — review gate is met$0.017400 = (3520×$2.50 + 860×$10.00)/1M
psychometric validity claim
batch28-writing-m3-r3
observed 2026-08-27
pilot n=12; construct briefcontent review passes but sample cannot establish validityBOUNDARY — do not claim psychometric validityUnavailable — adequately powered validation study

Formula / scoring rule: Gate = construct coverage + neutral wording + response completeness + scale balance + branch logic − burden and leading/double-barrel items. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-08. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the writing Batch 28 evidence scenario →

Batch 29 · UX research, earnings narratives, and public-history labels

Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes frozen inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.

1. UX-research synthesis suite

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
interviews and usability notes
batch29-writing-m1-r1
observed 2026-08-27
transcripts; notes; segment tags; participant IDsthemes trace to source spans; frequency and severity are separate; segments attributedPASS — synthesis is not generic summarization$0.014150 = (2940×$2.50 + 680×$10.00)/1M
survey and conflicting stakeholder tags
batch29-writing-m1-r2
observed 2026-08-27
survey extracts; conflicting tags; redline reviewnegative cases preserved; one unsupported generalization removed; opportunities prioritizedPASS WITH REPAIR — publish redlined synthesis$0.017850 = (3860×$2.50 + 820×$10.00)/1M
missing participant linkage
batch29-writing-m1-r3
observed 2026-08-27
themes present; participant/segment attribution absenttheme plausibility cannot establish research evidenceBOUNDARY — no accepted-report costUnavailable — participant/segment linkage and traceable evidence spans

Formula / scoring rule: Score = participant/segment attribution + frequency/severity separation + traceable themes + negative cases + uncertainty + opportunity linkage − unsupported generalization. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Investor-earnings narrative benchmark

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
statement and segment tables
batch29-writing-m2-r1
observed 2026-08-27
period comparators; currency/unit; GAAP and non-GAAP labelslabels and units preserved; arithmetic and direction agree; table spans tracePASS — narrative does not become advice$0.017250 = (3620×$2.50 + 820×$10.00)/1M
guidance and management notes
batch29-writing-m2-r2
observed 2026-08-27
guidance ranges; prior period; causal claims; reviewer sampleactual versus guidance separated; causal claim softened; risk language balancedPASS WITH REPAIR — retain reviewer correction$0.019850 = (4180×$2.50 + 940×$10.00)/1M
unsupported causal statement
batch29-writing-m2-r3
observed 2026-08-27
management note missing causal support; narrative requesteddirection is correct but causal evidence is absentBOUNDARY — do not publish unsupported causalityUnavailable — source-supported causal claim and accepted narrative

Formula / scoring rule: Score = GAAP/non-GAAP + period/currency/unit + arithmetic/direction + actual/guidance separation + causal support + table traceability − investment advice. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Museum and public-history interpretive-label gate

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
75/150-word catalog labels
batch29-writing-m3-r1
observed 2026-08-27
catalog record; provenance; audience; word limitsobject/date/name facts trace; length and hierarchy pass; audience comprehension reviewedPASS — label is evidence-bounded$0.014300 = (2840×$2.50 + 720×$10.00)/1M
contested attribution and community guidance
batch29-writing-m3-r2
observed 2026-08-27
contested record; community guidance; 300-word format; expert reviewuncertainty and contested history explicit; one presentist phrase removedPASS WITH REPAIR — retain expert correction$0.017400 = (3520×$2.50 + 860×$10.00)/1M
provenance gap
batch29-writing-m3-r3
observed 2026-08-27
catalog text supplied; provenance note absentlabel cannot resolve attribution or history without inventionBOUNDARY — do not fill missing recordUnavailable — provenance note and accepted contested-history wording

Formula / scoring rule: Score = object/date/name fidelity + source traceability + contested-history uncertainty + comprehension + hierarchy/length compliance − invention/presentism. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-08. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality, provider, or prior batch. Run the writing Batch 29 evidence scenario →

Batch 30 · Governance minutes, consultation synthesis, and product recalls

Frozen verification window: 2026-08-27 UTC. These server-rendered fixtures expose inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills where the registry closes the token tuple. Missing specialist evidence is explicitly Unavailable.

1. Board-meeting minutes and action-register suite

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Agenda/transcript/resolutions
batch30-writing-m1-r1
observed 2026-08-27
12 attendees; quorum 9; 4 motions; 2026-08-27T00:19ZAttendee and quorum fields match; 4/4 votes attributed; 7 actions retain owner/date; reviewer redlined no unsupported decisions; 3,280/740 tokens.PASS — minutes and register are accepted separately$0.015600 = (3280×$2.50 + 740×$10.00)/1M
Abstention/conflict
batch30-writing-m1-r2
observed 2026-08-27
9 voters; 1 abstention; 1 declared conflict; 2026-08-27T00:35ZAbstention excluded from majority denominator; conflicted voter excluded from motion; reviewer accepted 6/6 decision fields; 4,120/860 tokens.PASS — governance edge fields are explicit$0.018900 = (4120×$2.50 + 860×$10.00)/1M
Action-register redline
batch30-writing-m1-r3
observed 2026-08-27
18 actions; owner/date attachments; 2026-08-27T00:52Z16 owner/date pairs exact; 2 dates inferred and rejected; confidentiality labels preserved; accepted document after one repair; 5,060/980 tokens.PASS WITH REPAIR — inferred dates cannot pass silently$0.022450 = (5060×$2.50 + 980×$10.00)/1M

Formula / scoring rule: Acceptance = attendee/quorum + motion/vote attribution + decision/discussion separation + owner/date traceability + confidentiality + no unsupported inference + reviewer approval. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: OpenAI GPT-4o-mini / governance-writing registry rate verified 2026-08-27; test suite: Batch 30 governance-minutes fixture/test suite (run and result recorded 2026-08-27).

2. Public-consultation response synthesis benchmark

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Individual/organization submissions
batch30-writing-m2-r1
observed 2026-08-27
240 submissions; 6 segments; 2026-08-27T01:09ZSubmitter type and segment retained; themes cite 38 source spans; frequency and evidential weight separated; reviewer accepted 22/24 themes; 4,240/820 tokens.PASS WITH REPAIR — two themes narrowed to source wording$0.018800 = (4240×$2.50 + 820×$10.00)/1M
Duplicate form campaign
batch30-writing-m2-r2
observed 2026-08-27
1,000 forms; 183 exact duplicates; 2026-08-27T01:26Z183 duplicates collapsed but counted as campaign volume; minority 4% retained; 9/9 option links traceable; 6,180/1,120 tokens.PASS — duplicate handling does not erase represented volume$0.026650 = (6180×$2.50 + 1120×$10.00)/1M
Expert minority report
batch30-writing-m2-r3
observed 2026-08-27
14 expert reports; 3 conflicting views; 2026-08-27T01:43ZAll 3 conflicts surfaced with report IDs; uncertainty labels retained; one unsupported causal phrase removed; 5,720/1,060 tokens.PASS WITH REPAIR — reviewer correction closes the accepted synthesis$0.024900 = (5720×$2.50 + 1060×$10.00)/1M

Formula / scoring rule: Acceptance = submitter/segment attribution + duplicate-campaign handling + theme frequency versus evidential weight + minority/conflicting views + traceability + uncertainty + option linkage. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: OpenAI GPT-4o-mini / consultation-writing registry rate verified 2026-08-27; test suite: Batch 30 consultation-synthesis fixture/test suite (run and result recorded 2026-08-27).

3. Product-recall notice and customer-remedy communication gate

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
SKU/batch/date notice
batch30-writing-m3-r1
observed 2026-08-27
SKU N-440; batches 24A–24C; hazard and remedy records; 2026-08-27T02:00ZScope includes all three batch IDs and sale dates; mandated wording exact; contact URL and deadline present; reviewer accepted; 3,460/690 tokens.PASS — notice completeness gate met$0.015550 = (3460×$2.50 + 690×$10.00)/1M
Jurisdiction/channel remedy
batch30-writing-m3-r2
observed 2026-08-27
NZ/AU channels; refund, replacement, disposal sequence; 2026-08-27T02:17ZJurisdiction labels retained; channel-specific contact differs correctly; risk language calibrated; one deadline repaired; 4,820/880 tokens.PASS WITH REPAIR — jurisdiction must travel with remedy$0.020850 = (4820×$2.50 + 880×$10.00)/1M
Accessibility/contact correction
batch30-writing-m3-r3
observed 2026-08-27
Screen-reader HTML, plain text, hotline; 2026-08-27T02:34ZHeading order and alt text pass; corrected hotline appears in all 3 channels; unsupported reassurance removed; 5,360/940 tokens.PASS — accessible accepted notice after reviewer correction$0.022800 = (5360×$2.50 + 940×$10.00)/1M

Formula / scoring rule: Acceptance = SKU/batch/date scope + mandated wording + calibrated risk + action/deadline/contact/accessibility completeness − unsupported reassurance; no safety or legal determination is made. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: OpenAI GPT-4o-mini / recall-writing registry rate verified 2026-08-27; test suite: Batch 30 product-recall communication fixture/test suite (run and result recorded 2026-08-27).

Verified 2026-08-08. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the writing Batch 30 evidence scenario →

Batch 31 · RFP responses, grant alignment, and privacy change control

Frozen verification window: 2026-08-27 UTC. Matched model/run identity, inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.

1. Request-for-proposal compliance-response suite

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Mandatory requirements
batch31-writing-m1-r1
model/run: OpenAI GPT-4o-mini; observed 2026-08-27
42 mandatory, 18 weighted requirements; addendum A; run wri-311; 23:50ZTraceability 42/42; 3 evidence attachments linked; one unsupported capability redlined; 4,280/880 tokens.PASS WITH REPAIR — claim narrowed to attached evidence$0.019500 = (4280×$2.50 + 880×$10.00)/1M
Page limit and conflicts
batch31-writing-m1-r2
model/run: OpenAI GPT-4o-mini; observed 2026-08-27
12-page limit; two conflicts; run wri-312; 00:06ZPage count 12; conflicts surfaced; weighted section coverage 17/18; reviewer rejects one missing cross-reference.BOUNDARY — not all weighted requirements findable$0.022850 = (5060×$2.50 + 1020×$10.00)/1M
Evidence packet
batch31-writing-m1-r3
model/run: OpenAI GPT-4o-mini; observed 2026-08-27
30 attachments and addendum B; run wri-313; 00:22ZAttachment IDs trace 30/30; two stale citations removed; reviewer accepts final response.PASS — evidence traceability gate closes$0.028650 = (6420×$2.50 + 1260×$10.00)/1M

Formula / scoring rule: Acceptance = requirement-to-section traceability + completeness + evidence accuracy + conflict disclosure + evaluator findability − unsupported claims; no award recommendation is made. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: OpenAI GPT-4o-mini writing benchmark registry, verified 2026-08-27.

2. Research-grant narrative and budget-justification alignment benchmark

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Aims and milestones
batch31-writing-m2-r1
model/run: OpenAI GPT-4o-mini; observed 2026-08-27
3 aims, 8 milestones, 24 dates; run wri-321; 00:38ZAll aims link to work packages; 23/24 dates exact; one date repaired; reviewer accepted alignment.PASS WITH REPAIR — repaired date is recorded$0.020950 = (4620×$2.50 + 940×$10.00)/1M
Staff/equipment budget
batch31-writing-m2-r2
model/run: OpenAI GPT-4o-mini; observed 2026-08-27
12 line items, funder eligibility rules; 00:54ZNumeric totals reconcile; one equipment category marked uncertain; evidence citations 11/12.BOUNDARY — uncertain eligibility is escalated, not asserted$0.024250 = (5380×$2.50 + 1080×$10.00)/1M
Risk and method packet
batch31-writing-m2-r3
model/run: OpenAI GPT-4o-mini; observed 2026-08-27
Methods, risks, staffing, funder rules; 01:10ZAim linkage 3/3; dates and totals pass; unsupported causal promise removed; reviewer accepts.PASS — alignment does not imply scientific merit$0.027900 = (6280×$2.50 + 1220×$10.00)/1M

Formula / scoring rule: Acceptance = aim/work-package linkage + numeric/date consistency + eligible-cost wording + evidence citation + uncertainty + reviewer repair; scientific merit is outside scope. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: OpenAI GPT-4o-mini writing benchmark registry, verified 2026-08-27.

3. Privacy-notice change-control and data-flow consistency gate

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Product event diff
batch31-writing-m3-r1
model/run: OpenAI GPT-4o-mini; observed 2026-08-27
18 product events; 6 purposes; version 4→5; 01:26ZPurpose/category mapping 18/18; two recipients changed and highlighted; effective date exact.PASS — change summary matches data flow$0.017250 = (3860×$2.50 + 760×$10.00)/1M
Processor and retention
batch31-writing-m3-r2
model/run: OpenAI GPT-4o-mini; observed 2026-08-27
9 processors, 4 jurisdictions, retention table; 01:42ZProcessor names 9/9; retention terms 3/4 consistent; one escalation raised; reviewer accepts scoped draft.PASS WITH ESCALATION — inconsistency is not silently harmonized$0.022600 = (5120×$2.50 + 980×$10.00)/1M
Consent-state versioning
batch31-writing-m3-r3
model/run: OpenAI GPT-4o-mini; observed 2026-08-27
Opt-in, opt-out, withdrawal states; 02:00ZDefined terms consistent; one unsupported legal assurance removed; effective date and diff links present.PASS WITH REPAIR — draft is accepted as a consistency artifact only$0.026500 = (6040×$2.50 + 1140×$10.00)/1M

Formula / scoring rule: Acceptance = purpose/category/recipient/retention alignment + defined terms + effective-date/change-summary fidelity + reviewer escalation; no legal-compliance determination is made. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: OpenAI GPT-4o-mini writing benchmark registry, verified 2026-08-27.

Verified 2026-08-08. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the writing Batch 31 evidence scenario →

Batch 32 · Normative standards, regulatory responses, and crisis consistency

Frozen verification window: 2026-08-27 UTC. Matched model/run identity, frozen inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.

1. Technical-standard normative-language suite

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Normative force / wri32-711
batch32-writing-m1-r1
model/run: OpenAI GPT-4o-mini matched writing run; observed 2026-08-27
SHALL/SHOULD/MAY, exceptions; run wri32-711; 00:02ZForce preserved 32/32; actor scope 30/32; expert repaired 2 clauses.PASS WITH REPAIR — redlines are retained$0.019500 = (4280×$2.50 + 880×$10.00)/1M
Definitions/refs / wri32-712
batch32-writing-m1-r2
model/run: OpenAI GPT-4o-mini matched writing run; observed 2026-08-27
Defined terms and cross-references; run wri32-712; 00:18ZDefined-term consistency 48/50; circular references 1; reviewer rejects that clause.BOUNDARY — unresolved reference blocks clean acceptance$0.022850 = (5060×$2.50 + 1020×$10.00)/1M
Conformance / wri32-713
batch32-writing-m1-r3
model/run: OpenAI GPT-4o-mini matched writing run; observed 2026-08-27
Exceptions and conformance clauses; run wri32-713; 00:34ZTestability 26/28; unsupported obligation removed; reviewer accepts scoped draft.PASS — no certification or legal conclusion made$0.028650 = (6420×$2.50 + 1260×$10.00)/1M

Formula / scoring rule: Acceptance = force/actor/condition scope + defined terms + testability + cross-reference integrity − unsupported obligation, with expert redlines. OpenAI GPT-4o-mini technical-writing benchmark registry. Dated registry and evidence index, verified 2026-08-27.

2. Regulatory-submission response-to-question benchmark

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Authority questions / wri32-721
batch32-writing-m2-r1
model/run: OpenAI GPT-4o-mini matched writing run; observed 2026-08-27
24 questions and source dossier; run wri32-721; 00:50ZEvidence links 24/24; 2 claims narrowed; reviewer accepts response map.PASS WITH REPAIR — claims remain source-scoped$0.020950 = (4620×$2.50 + 940×$10.00)/1M
Deficiency history / wri32-722
batch32-writing-m2-r2
model/run: OpenAI GPT-4o-mini matched writing run; observed 2026-08-27
8 deficiencies, tables, commitments; 01:06ZScope complete 7/8; one version mismatch escalated; no unsupported assurance.BOUNDARY — incomplete scope prevents full response acceptance$0.024250 = (5380×$2.50 + 1080×$10.00)/1M
Deadlines/owners / wri32-723
batch32-writing-m2-r3
model/run: OpenAI GPT-4o-mini matched writing run; observed 2026-08-27
18 commitments and deadlines; 01:22ZOwner/date consistency 18/18; uncertainty disclosed; reviewer accepts scoped packet.PASS — no safety, efficacy, or approval determination$0.027900 = (6280×$2.50 + 1220×$10.00)/1M

Formula / scoring rule: Acceptance = question-to-evidence traceability + scope/data/version fidelity + uncertainty + owner/date consistency + reviewer escalation. OpenAI GPT-4o-mini regulatory-response benchmark registry. Dated registry and evidence index, verified 2026-08-27.

3. Crisis-communication channel-consistency gate

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Web/email/SMS / wri32-731
batch32-writing-m3-r1
model/run: OpenAI GPT-4o-mini matched writing run; observed 2026-08-27
Incident facts and unknowns; 01:38ZFacts/timestamps agree 18/18; action adapted to channel; reviewer accepts.PASS — synchronized update is accepted$0.017250 = (3860×$2.50 + 760×$10.00)/1M
Status change / wri32-732
batch32-writing-m3-r2
model/run: OpenAI GPT-4o-mini matched writing run; observed 2026-08-27
Status changes and stakeholder rules; 01:54ZCorrection visible; SMS length passes; one prohibited speculation removed.PASS WITH REPAIR — unknowns remain explicit$0.022600 = (5120×$2.50 + 980×$10.00)/1M
Social escalation / wri32-733
batch32-writing-m3-r3
model/run: OpenAI GPT-4o-mini matched writing run; observed 2026-08-27
Web/email/SMS/social variants; 02:10ZTimestamp conflict in 2/16 channels; approval escalation raised; final acceptance withheld.BOUNDARY — cross-channel inconsistency blocks acceptance$0.026500 = (6040×$2.50 + 1140×$10.00)/1M

Formula / scoring rule: Acceptance = fact/timestamp agreement + audience action + uncertainty/correction visibility + channel constraints − prohibited speculation. OpenAI GPT-4o-mini crisis-writing benchmark registry. Dated registry and evidence index, verified 2026-08-27.

Verified 2026-08-08. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the writing Batch 32 evidence scenario →

Batch 33 · Legislative, aviation-maintenance, and museum-provenance writing suites

Frozen verification window: 2026-08-27 UTC. Frozen inputs, model/run identity, formulas or scoring rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Legislative amendment consolidation suite

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Substitution/order / wri33-711
batch33-writing-m1-r1
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Base act + 12 amendments + commencement; run wri33-711; 23:24ZGPT-4o-mini operation order 44/44; cross-references 39/40; 1 expert redline; 5,820/1,160 tokens.PASS WITH REPAIR — no legal interpretation$0.026150 = (5820×$2.50 + 1160×$10.00)/1M
Repeals/transitional / wri33-712
batch33-writing-m1-r2
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Repeals, substitutions, transitional provisions; run wri33-712; 23:40ZVersion fidelity 31/34; 3 unresolved conflicts escalated; 7,140/1,320 tokens.BOUNDARY — unresolved conflict blocks clean consolidation$0.031050 = (7140×$2.50 + 1320×$10.00)/1M
Cross-reference / wri33-713
batch33-writing-m1-r3
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Numbered provisions and defined terms; run wri33-713; 23:56ZDefined terms 52/52; source spans 48/48; reviewer accepts; 6,460/1,240 tokens.PASS — source-scoped document output$0.028550 = (6460×$2.50 + 1240×$10.00)/1M

Formula / scoring rule: Acceptance = operation/version/defined-term/reference fidelity + source-span traceability − unresolved conflict, with expert redlines. First-party pricing/evidence registry: Matched legislative writing benchmark registry; source texts and expert review, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI GPT-4o-mini documentationOpenAI API pricingAnthropic Claude documentationAnthropic pricingGoogle Gemini documentationGoogle Gemini pricing.

2. Aviation maintenance-bulletin writing benchmark

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Applicability / wri33-721
batch33-writing-m2-r1
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Serial ranges, prerequisites, warnings; run wri33-721; 00:12ZTraceability 28/28; serial applicability 27/28; technician repaired one range; 5,460/1,020 tokens.PASS WITH REPAIR — applicability remains visible$0.023850 = (5460×$2.50 + 1020×$10.00)/1M
Work steps / wri33-722
batch33-writing-m2-r2
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Tooling, work steps, inspection/sign-off; run wri33-722; 00:28ZSequence 34/36; unit/part number 18/18; two sequence repairs; 7,280/1,360 tokens.PASS WITH REPAIR — technician review is the denominator$0.031800 = (7280×$2.50 + 1360×$10.00)/1M
Hazard escalation / wri33-723
batch33-writing-m2-r3
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Engineering findings and warnings; run wri33-723; 00:44ZOne hazard omitted; reviewer escalates and withholds bulletin acceptance; 6,880/1,280 tokens.BOUNDARY — no airworthiness conclusion$0.030000 = (6880×$2.50 + 1280×$10.00)/1M

Formula / scoring rule: Acceptance = source-to-instruction traceability + applicability/sequence/unit fidelity + hazard escalation + technician review; no airworthiness decision. First-party pricing/evidence registry: Matched aviation maintenance writing benchmark registry; engineering packet and reviewer results, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI GPT-4o-mini documentationOpenAI API pricingAnthropic Claude documentationAnthropic pricingGoogle Gemini documentationGoogle Gemini pricing.

3. Museum exhibition-label and provenance gate

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
50-word label / wri33-731
batch33-writing-m3-r1
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Catalog/conservation record; 50-word limit; run wri33-731; 01:00ZObject facts 18/18; source visible; curator accepts; 3,240/660 tokens.PASS — provenance is cited, not invented$0.014700 = (3240×$2.50 + 660×$10.00)/1M
100/200-word formats / wri33-732
batch33-writing-m3-r2
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Contested attribution, community guidance; 100/200 words; 01:16ZUncertainty retained 8/8; sensitive-language repair; community review accepts both formats; 5,680/1,020 tokens.PASS WITH REPAIR — contested history remains disclosed$0.024400 = (5680×$2.50 + 1020×$10.00)/1M
Unsupported narrative / wri33-733
batch33-writing-m3-r3
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Quotation/date conflict and provenance gap; run wri33-733; 01:32ZUnsupported narrative phrase removed; curator withholds final label pending source.BOUNDARY — no provenance authority is asserted$0.020600 = (4720×$2.50 + 880×$10.00)/1M

Formula / scoring rule: Acceptance = object/fact fidelity + uncertainty/source visibility + format/readability + curator/community review; contested attribution stays qualified. First-party pricing/evidence registry: Matched museum-label writing benchmark registry; catalog, conservation, and review records, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI GPT-4o-mini documentationOpenAI API pricingAnthropic Claude documentationAnthropic pricingGoogle Gemini documentationGoogle Gemini pricing.

Verified 2026-08-08. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the writing Batch 33 evidence scenario →

Batch 34 · Archaeology, insurance endorsements, and clinical-trial protocol-deviation writing suites

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Archaeological excavation context-sheet and stratigraphic-report suite

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Harris matrix / wri34-711
batch34-writing-m1-r1
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Trench/locus registers and Harris matrix; run wri34-711; 23:12ZRelationships 42/42; phase labels 18/20; source spans 38/38; 5,420/980 tokens.PASS WITH REPAIR — specialist redline retained$0.023350 = (5420×$2.50 + 980×$10.00)/1M
Conflicting notes / wri34-712
batch34-writing-m1-r2
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Plans, finds, dates, conflicting field notes; run wri34-712; 23:28ZUncertainty 14/14; impossible sequence flagged; reviewer accepts narrowed report; 6,860/1,240 tokens.PASS WITH REPAIR — no cultural/legal determination$0.029550 = (6860×$2.50 + 1240×$10.00)/1M
Missing provenance / wri34-713
batch34-writing-m1-r3
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Samples and dates with missing context; run wri34-713; 23:44ZContext gap 5/16; report withheld pending source; 7,240/1,360 tokens.BOUNDARY — provenance gap blocks acceptance$0.031700 = (7240×$2.50 + 1360×$10.00)/1M

Formula / scoring rule: Acceptance = context/phase/relationship fidelity + source-span traceability + uncertainty + impossible-sequence detection + specialist redline review. Matched archaeological writing benchmark; field records and specialist review; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: OpenAI GPT-4o-mini documentationOpenAI API pricing.

2. Insurance policy schedule and endorsement consolidation benchmark

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Schedule merge / wri34-721
batch34-writing-m2-r1
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Declarations, limits, deductibles, effective dates; run wri34-721; 00:00ZNumeric/date fields 44/44; operations 26/26; reviewer accepts; 5,980/1,080 tokens.PASS — consolidation is source-scoped$0.025750 = (5980×$2.50 + 1080×$10.00)/1M
Endorsement order / wri34-722
batch34-writing-m2-r2
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Base wording + 12 endorsements + cancellation; run wri34-722; 00:16ZOperation order 31/34; three conflicts exposed and repaired; 7,420/1,360 tokens.PASS WITH REPAIR — no coverage interpretation$0.032150 = (7420×$2.50 + 1360×$10.00)/1M
Conflict / wri34-723
batch34-writing-m2-r3
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Named parties, exclusions, cross-references conflict; run wri34-723; 00:32ZDefined-term mismatch 4/22; reviewer rejects clean consolidation; 8,160/1,480 tokens.BOUNDARY — no liability or claim outcome$0.035200 = (8160×$2.50 + 1480×$10.00)/1M

Formula / scoring rule: Acceptance = version/operation order + term/reference integrity + numeric/date fidelity + conflict visibility + reviewer correction. Matched insurance-endorsement writing benchmark; policy records and reviewer results; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Anthropic Claude documentationAnthropic pricing.

3. Clinical-trial protocol-deviation narrative gate

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Visit window / wri34-731
batch34-writing-m3-r1
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Protocol v4; visit windows and source notes; run wri34-731; 00:48ZSubject/event/date fields 32/32; planned/observed separated; reviewer accepts; 5,640/1,020 tokens.PASS — narrative does not infer outcome$0.024300 = (5640×$2.50 + 1020×$10.00)/1M
IP log/query / wri34-732
batch34-writing-m3-r2
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Investigational-product logs, queries, missing records; run wri34-732; 01:04ZVersion fidelity 28/30; two repairs; source spans 26/26; 7,080/1,280 tokens.PASS WITH REPAIR — uncertainty remains visible$0.030500 = (7080×$2.50 + 1280×$10.00)/1M
Unsupported seriousness / wri34-733
batch34-writing-m3-r3
model/run: Matched model ladder: OpenAI GPT-4o-mini; Anthropic Claude Sonnet; Google Gemini 2.0 Flash; observed 2026-08-27
Protocol versions and corrective action; unsupported causal phrase; run wri34-733; 01:20ZCausality phrase removed; clinical reviewer withholds acceptance pending source.BOUNDARY — no safety/reportability decision$0.032650 = (7540×$2.50 + 1380×$10.00)/1M

Formula / scoring rule: Acceptance = subject/event/date/version fidelity + planned-versus-observed separation + traceability + uncertainty − unsupported causality/seriousness. Matched protocol-deviation writing benchmark; source notes and clinical review; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Gemini API documentationGoogle Gemini pricing.

Verified 2026-08-08. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the writing Batch 34 evidence scenario →

Batch 35 · Forecast discussions, allergen change notices, and ship-survey narratives

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Meteorological forecast-discussion suite

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Observations/model runs / batch35-writing-811-1
batch35-writing-m1-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen matched fixture; pinned model/run; checker and reviewer records; run 09:00ZReference output agrees; field checks 24/24; reviewer accepts; usage and spend join.PASS — matched evidence closes the gate.$0.038400 = (4920×$5.00 + 920×$15.00)/1M
Ensembles/fronts/watches / batch35-writing-811-2
batch35-writing-m1-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen adversarial fixture; same prompt and budget; run 09:16ZChecker agrees on 21/24 fields; three repairs are visible; expert accepts narrowed result.PASS WITH REPAIR — repaired scope is explicit.$0.055700 = (7180×$5.00 + 1320×$15.00)/1M
Time zones/conflicting stations / batch35-writing-811-3
batch35-writing-m1-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen counterexample/unsupported fixture; same matched run; run 09:32ZChecker rejects the broad conclusion; unsupported hazard escalation remains unaccepted.BOUNDARY — unsupported hazard escalation remains unaccepted.$0.062400 = (8040×$5.00 + 1480×$15.00)/1M

Formula / scoring rule: Acceptance = valid-time/region/unit fidelity + observation/forecast separation + calibrated uncertainty + source spans + hazard redlines. Matched meteorological writing benchmark; forecaster review records; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI API documentationOpenAI API pricing.

2. Food-formulation and allergen change-notice benchmark

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Recipe/supplier specification / batch35-writing-821-1
batch35-writing-m2-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen matched fixture; pinned model/run; checker and reviewer records; run 09:00ZReference output agrees; field checks 24/24; reviewer accepts; usage and spend join.PASS — matched evidence closes the gate.$0.028560 = (4920×$3.00 + 920×$15.00)/1M
Batch/allergen matrix / batch35-writing-821-2
batch35-writing-m2-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen adversarial fixture; same prompt and budget; run 09:16ZChecker agrees on 21/24 fields; three repairs are visible; expert accepts narrowed result.PASS WITH REPAIR — repaired scope is explicit.$0.041340 = (7180×$3.00 + 1320×$15.00)/1M
Label version/effective date / batch35-writing-821-3
batch35-writing-m2-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen counterexample/unsupported fixture; same matched run; run 09:32ZChecker rejects the broad conclusion; safety or label approval evidence is outside scope.BOUNDARY — safety or label approval evidence is outside scope.$0.046320 = (8040×$3.00 + 1480×$15.00)/1M

Formula / scoring rule: Acceptance = ingredient/version/lot linkage + quantity/unit fidelity + contains/may-contain separation + unresolved evidence + specialist correction. Matched food-formulation writing benchmark; formulation and reviewer records; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Anthropic Messages APIAnthropic pricing.

3. Ship-survey and machinery-maintenance narrative gate

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Inspection notes/defect photos / batch35-writing-831-1
batch35-writing-m3-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen matched fixture; pinned model/run; checker and reviewer records; run 09:00ZReference output agrees; field checks 24/24; reviewer accepts; usage and spend join.PASS — matched evidence closes the gate.$0.010750 = (4920×$1.25 + 920×$5.00)/1M
Registers/measurements/work orders / batch35-writing-831-2
batch35-writing-m3-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen adversarial fixture; same prompt and budget; run 09:16ZChecker agrees on 21/24 fields; three repairs are visible; expert accepts narrowed result.PASS WITH REPAIR — repaired scope is explicit.$0.015575 = (7180×$1.25 + 1320×$5.00)/1M
Class references/dates/repairs / batch35-writing-831-3
batch35-writing-m3-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen counterexample/unsupported fixture; same matched run; run 09:32ZChecker rejects the broad conclusion; seaworthiness or class evidence is outside scope.BOUNDARY — seaworthiness or class evidence is outside scope.$0.017450 = (8040×$1.25 + 1480×$5.00)/1M

Formula / scoring rule: Acceptance = vessel/equipment/measurement fidelity + observed/inferred separation + chronology + photo/source linkage + open defects + surveyor review. Matched ship-survey narrative benchmark; surveyor review records; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Gemini API documentationGoogle Gemini pricing.

Verified 2026-08-08. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the writing Batch 35 evidence scenario →

Batch 36 · Museum provenance, geotechnical borehole summaries, and parliamentary memoranda

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Museum object-catalog and provenance narrative suite

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Accession/inscription records / batch36-writing-811-1
batch36-writing-m1-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen specialist fixture; pinned model/run, checker, reviewer, and usage; run 09:00ZOpenAI GPT-4o matches the pinned reference on 24/24 fields; specialist review accepts the scoped result and usage joins.PASS — matched checker plus specialist acceptance is required.$0.038400 = (4920×$5.00 + 920×$15.00)/1M
Conservation/acquisition chain / batch36-writing-811-2
batch36-writing-m1-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen adversarial fixture; identical prompt/budget and repair log; run 09:16ZOpenAI GPT-4o matches 21/24 fields; three repairs are visible and the specialist accepts only the narrowed claim.PASS WITH REPAIR — no unreviewed claim is promoted.$0.055700 = (7180×$5.00 + 1320×$15.00)/1M
Disputed attribution/image labels / batch36-writing-811-3
batch36-writing-m1-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen counterexample fixture; pinned run and checker output; run 09:32ZThe checker rejects the broad result; authenticity, ownership, valuation, and legal status remain outside the gate.BOUNDARY — authenticity, ownership, valuation, and legal status remain outside the gate.$0.062400 = (8040×$5.00 + 1480×$15.00)/1M

Formula / scoring rule: Acceptance = identifier/chronology fidelity + observed/attributed separation + source-span traceability + uncertainty/gaps + sensitive wording + curator redlines. Matched museum provenance writing benchmark; curator records; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI API documentationOpenAI API pricing.

2. Geotechnical borehole-log and factual-site-summary benchmark

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Coordinates/depth/strata / batch36-writing-821-1
batch36-writing-m2-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen specialist fixture; pinned model/run, checker, reviewer, and usage; run 09:00ZAnthropic Claude Sonnet matches the pinned reference on 24/24 fields; specialist review accepts the scoped result and usage joins.PASS — matched checker plus specialist acceptance is required.$0.028560 = (4920×$3.00 + 920×$15.00)/1M
Recovery/RQD/groundwater / batch36-writing-821-2
batch36-writing-m2-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen adversarial fixture; identical prompt/budget and repair log; run 09:16ZAnthropic Claude Sonnet matches 21/24 fields; three repairs are visible and the specialist accepts only the narrowed claim.PASS WITH REPAIR — no unreviewed claim is promoted.$0.041340 = (7180×$3.00 + 1320×$15.00)/1M
Revisions/missing intervals / batch36-writing-821-3
batch36-writing-m2-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen counterexample fixture; pinned run and checker output; run 09:32ZThe checker rejects the broad result; foundation design, certification, and safety determinations are outside the gate.BOUNDARY — foundation design, certification, and safety determinations are outside the gate.$0.046320 = (8040×$3.00 + 1480×$15.00)/1M

Formula / scoring rule: Acceptance = borehole/depth/sample linkage + interval continuity + unit/qualifier fidelity + observed/interpreted separation + contradictions + engineer corrections. Matched borehole summary benchmark; engineer review records; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Anthropic Messages API documentationAnthropic model pricing.

3. Parliamentary bill explanatory-memorandum gate

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Clause text/amendments / batch36-writing-831-1
batch36-writing-m3-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen specialist fixture; pinned model/run, checker, reviewer, and usage; run 09:00ZGoogle Gemini matches the pinned reference on 24/24 fields; specialist review accepts the scoped result and usage joins.PASS — matched checker plus specialist acceptance is required.$0.010750 = (4920×$1.25 + 920×$5.00)/1M
Commencement/defined terms / batch36-writing-831-2
batch36-writing-m3-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen adversarial fixture; identical prompt/budget and repair log; run 09:16ZGoogle Gemini matches 21/24 fields; three repairs are visible and the specialist accepts only the narrowed claim.PASS WITH REPAIR — no unreviewed claim is promoted.$0.015575 = (7180×$1.25 + 1320×$5.00)/1M
Fiscal/consultation/version notes / batch36-writing-831-3
batch36-writing-m3-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen counterexample fixture; pinned run and checker output; run 09:32ZThe checker rejects the broad result; legal advice, passage prediction, and rights assertions are outside the gate.BOUNDARY — legal advice, passage prediction, and rights assertions are outside the gate.$0.017450 = (8040×$1.25 + 1480×$5.00)/1M

Formula / scoring rule: Acceptance = clause/version linkage + effect/rationale separation + defined-term consistency + neutral stakeholder representation + unresolved impacts + citation traceability. Matched parliamentary memorandum benchmark; counsel redline records; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Gemini API documentationGoogle Gemini model pricing.

Verified 2026-08-08. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the writing Batch 36 evidence scenario →

Batch 37 · Archival, environmental-impact, and clinical-protocol writing gates

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Archival-collection finding-aid and scope-note suite

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Accession/box inventory / batch37-writing-811-r1
batch37-writing-m1-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen specialist fixture; pinned model/run, checker, reviewer, and usage; run 09:00ZOpenAI GPT-4o matches the pinned reference on 24/24 fields; specialist review accepts the scoped result and usage joins.PASS — checker plus specialist acceptance is required.$0.038400 = (4920×$5.00 + 920×$15.00)/1M
Creator history/arrangement / batch37-writing-811-r2
batch37-writing-m1-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen adversarial fixture; identical prompt/budget and repair log; run 09:16ZOpenAI GPT-4o matches 21/24 fields; three repairs are visible and the specialist accepts only the narrowed claim.PASS WITH REPAIR — no unreviewed claim is promoted.$0.055700 = (7180×$5.00 + 1320×$15.00)/1M
Restrictions/legacy conflicts / batch37-writing-811-r3
batch37-writing-m1-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen counterexample fixture; pinned run and checker output; run 09:32ZThe checker rejects the broad result; ownership, authorization, authenticity, and appraisal remain outside the gate.BOUNDARY — ownership, authorization, authenticity, and appraisal remain outside the gate.Unavailable — ownership, authorization, authenticity, and appraisal remain outside the gate

Formula / scoring rule: Acceptance = collection/series/item hierarchy + identifier/chronology fidelity + restriction visibility + source traceability + inclusive wording + archivist redlines. Matched archival finding-aid benchmark; archivist records; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI API documentationOpenAI API pricing.

2. Environmental-impact evidence synopsis benchmark

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Baseline surveys/model runs / batch37-writing-821-r1
batch37-writing-m2-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen specialist fixture; pinned model/run, checker, reviewer, and usage; run 09:00ZAnthropic Claude Sonnet matches the pinned reference on 24/24 fields; specialist review accepts the scoped result and usage joins.PASS — checker plus specialist acceptance is required.$0.028560 = (4920×$3.00 + 920×$15.00)/1M
Mitigation/monitoring / batch37-writing-821-r2
batch37-writing-m2-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen adversarial fixture; identical prompt/budget and repair log; run 09:16ZAnthropic Claude Sonnet matches 21/24 fields; three repairs are visible and the specialist accepts only the narrowed claim.PASS WITH REPAIR — no unreviewed claim is promoted.$0.041340 = (7180×$3.00 + 1320×$15.00)/1M
Conflicting reports/maps / batch37-writing-821-r3
batch37-writing-m2-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen counterexample fixture; pinned run and checker output; run 09:32ZThe checker rejects the broad result; approval, legal compliance, and ecological-safety decisions remain outside the gate.BOUNDARY — approval, legal compliance, and ecological-safety decisions remain outside the gate.Unavailable — approval, legal compliance, and ecological-safety decisions remain outside the gate

Formula / scoring rule: Acceptance = alternative/source/version linkage + observed/modeled separation + numeric/spatial fidelity + uncertainty/dissent + mitigation traceability + specialist correction. Matched environmental-impact synopsis benchmark; specialist records; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Anthropic Messages API documentationAnthropic model pricing.

3. Clinical-trial protocol synopsis and amendment-change-control gate

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Objectives/endpoints/eligibility / batch37-writing-831-r1
batch37-writing-m3-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen specialist fixture; pinned model/run, checker, reviewer, and usage; run 09:00ZGoogle Gemini matches the pinned reference on 24/24 fields; specialist review accepts the scoped result and usage joins.PASS — checker plus specialist acceptance is required.$0.010750 = (4920×$1.25 + 920×$5.00)/1M
Arms/visits/interventions / batch37-writing-831-r2
batch37-writing-m3-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen adversarial fixture; identical prompt/budget and repair log; run 09:16ZGoogle Gemini matches 21/24 fields; three repairs are visible and the specialist accepts only the narrowed claim.PASS WITH REPAIR — no unreviewed claim is promoted.$0.015575 = (7180×$1.25 + 1320×$5.00)/1M
Amendments/safety/statistics / batch37-writing-831-r3
batch37-writing-m3-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
Frozen counterexample fixture; pinned run and checker output; run 09:32ZThe checker rejects the broad result; medical advice, efficacy/safety assessment, and participant recruitment remain outside the gate.BOUNDARY — medical advice, efficacy/safety assessment, and participant recruitment remain outside the gate.Unavailable — medical advice, efficacy/safety assessment, and participant recruitment remain outside the gate

Formula / scoring rule: Acceptance = section/version linkage + schedule/numeric consistency + defined terms + change visibility + citation + contradiction disclosure + editorial redlines. Matched clinical-protocol benchmark; clinical/editorial records; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Gemini API documentationGoogle Gemini model pricing.

Verified 2026-08-08. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the writing Batch 37 evidence scenario →

Batch 38 · Oral histories, musical editions, and built-heritage condition summaries

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Oral-history transcript, annotation, and restriction-note suite

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Audio roster timestamps / batch38-writing-811-r1
batch38-writing-m1-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
38-minute WAV, speaker roster, 12 timestamp anchors, transcript checker oh-01; run 05:00Zspeaker/time/text alignment 38/38; verbatim/editorial separation 16/16; archivist accepts 24/24 fields; input 4,860/output 900 tokens.PASS — transcript facts are accepted without authenticity claims.$0.037800 = (4860×$5.00 + 900×$15.00)/1M
Dialect inaudible correction / batch38-writing-811-r2
batch38-writing-m1-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
dialect sample, 7 inaudible spans, 4 editorial corrections; repair log; run 05:16Z21/24 fields accepted; inaudible spans remain marked and one timestamp repaired; input 7,240/output 1,280 tokens.PASS WITH REPAIR — editorial additions remain visibly separated.$0.055400 = (7240×$5.00 + 1280×$15.00)/1M
Embargo conflict / batch38-writing-811-r3
batch38-writing-m1-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
three conflicting catalog records, embargo note, ownership dispute; run 05:32Zalignment is partial; consent, ownership, authenticity, and access authorization remain unavailable.UNAVAILABLE — consent, ownership, authenticity, and access authorization remain Unavailable.Unavailable — consent, ownership, authenticity, and access authorization remain Unavailable

Formula / scoring rule: Acceptance = speaker/time/text alignment + verbatim/editorial separation + uncertainty/restriction visibility + source traceability + narrator fidelity + archivist redlines + accepted-entry cost. Matched oral-history writing benchmark; transcript and archivist records; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: OpenAI API documentationOpenAI API pricingOpenAI model pricingOpenAI API pricing.

2. Musical critical-edition apparatus and source-variant synopsis benchmark

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Manuscript print part / batch38-writing-821-r1
batch38-writing-m2-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
MS A, first print, orchestral part; folio/bar coordinates and stemma m-01; run 06:00Zsource/location links 24/24; diplomatic and normalized layers separate; specialist accepts 24/24 fields; input 4,620/output 860 tokens.PASS — variant evidence is tied to source locations.$0.026760 = (4620×$3.00 + 860×$15.00)/1M
Bar beat pitch rhythm / batch38-writing-821-r2
batch38-writing-m2-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
five witnesses, 18 bars, transposition and emendation markers; repair run 06:16Z21/24 fields accepted; two conjectures relabeled; pitch/rhythm uncertainty retained; input 7,020/output 1,240 tokens.PASS WITH REPAIR — conjecture is not presented as witness evidence.$0.039660 = (7020×$3.00 + 1240×$15.00)/1M
Recording conflict / batch38-writing-821-r3
batch38-writing-m2-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
20 variants, conflicting recording metadata, copyright note and missing witness; run 06:32Zsource synopsis is incomplete; authorship, authenticity, copyright, and performance authority remain unavailable.UNAVAILABLE — authorship, authenticity, copyright, and performance authority remain Unavailable.Unavailable — authorship, authenticity, copyright, and performance authority remain Unavailable

Formula / scoring rule: Acceptance = source/stemma/location linkage + diplomatic/normalized separation + notation/chronology fidelity + variant completeness + conjecture visibility + specialist correction + spend. Matched musical critical-edition benchmark; source and specialist records; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Anthropic API documentationAnthropic model pricingAnthropic model pricingAnthropic model pricing.

3. Built-heritage condition-survey factual-summary gate

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Element register photos / batch38-writing-831-r1
batch38-writing-m3-r1
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
asset BH-17, 12 facade photos, element register v4, survey date; run 07:00Z12/12 photos link to elements; observed/inferred labels 18/18; conservator accepts 24/24 fields; input 4,780/output 880 tokens.PASS — factual description is separated from diagnosis.$0.010375 = (4780×$1.25 + 880×$5.00)/1M
Measured materials / batch38-writing-831-r2
batch38-writing-m3-r2
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
measured drawing, stone types, moisture readings, prior survey; repair run 07:16Z21/24 fields accepted; two unit labels repaired; changes and uncertainty remain visible; input 7,160/output 1,260 tokens.PASS WITH REPAIR — measurements are retained with source dates.$0.015250 = (7160×$1.25 + 1260×$5.00)/1M
Conflicting inspection / batch38-writing-831-r3
batch38-writing-m3-r3
model/run: OpenAI GPT-4o; Anthropic Claude Sonnet; Google Gemini; observed 2026-08-27
three inspections disagree on moisture cause and defect severity; valuation note; run 07:32Zfacts can be summarized, but cause, repair, safety/compliance, and valuation remain unavailable.UNAVAILABLE — cause, repair, safety/compliance, and valuation remain Unavailable.Unavailable — cause, repair, safety/compliance, and valuation remain Unavailable

Formula / scoring rule: Acceptance = asset/element/location/version linkage + observed/inferred separation + numeric/image fidelity + change visibility + uncertainty + specialist redlines + accepted-summary cost. Matched built-heritage survey benchmark; inspection and conservation records; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google API documentationGoogle Gemini model pricingGoogle model pricingGoogle Gemini model pricing.

Verified 2026-08-08. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the writing Batch 38 evidence scenario →

Which models rank highest for Writing & Content?

"Fit" is a requirements match, not a quality benchmark — it combines price, measured speed, context window, and (where we have run it) graded accuracy on this task. Formula below.

#ModelProviderFitEvidenceTask price/MTokens/secContextScored on
1Muse Spark 1.3 ContributorMeta9397/1$0.151.0Mprice, context, evidence
2Gemini 2.5 Flash LiteLegacyGoogle81$0.261Mprice, context
3GPT-OSS 120B (Cerebras)Cerebras7280/1$0.572450131Kprice, context, speed, evidence
4Amazon Nova LiteAmazon7188/1$0.16108300Kprice, context, speed, evidence
5Ministral 8BMistral7086/1$0.15158256Kprice, context, speed, evidence
6GPT-OSS 20BGroq6980/1$0.201120131Kprice, context, speed, evidence
7Amazon Nova MicroAmazon6678/1$0.09168128Kprice, context, speed, evidence
8Mistral Small 3.1Mistral6688/1$0.40121256Kprice, context, speed, evidence

What will Writing & Content cost?

At 20,000 marketing copy generation calls/month:

ModelTask price/MEst. monthly cost
Muse Spark 1.3 Contributor$0.15$3.40
Gemini 2.5 Flash Lite$0.26$5.80
GPT-OSS 120B (Cerebras)$0.57$12.50

How is the best LLM for Writing & Content ranked?

Task rubric:

  • Constraint-following, clarity, and jargon avoidance (45%)
  • Task-shaped API price (30%)
  • Measured generation speed (15%)
  • Context-window headroom (10%)

Weights: evidence 45%, price 30%, speed 15%, context 10%.

Requirements: none — every current model is eligible. 49 models eligible.

Price and context sub-scores are min-max normalised (log-scaled) within this task's eligible set only. Speed uses measured tokens/sec only — estimated rows are excluded. A model missing a measurement is never scored as zero: its weight is redistributed across the components we do have, and "Scored on" in the table above shows exactly which ones.

Prices verified 2026-08-08, accuracy graded 2026-06-16.

Availability: Legacy models remain visible only as historical rows. The winner is selected from models with current pricing and model records.

What failure modes matter for Writing & Content?

  • The strongest prose can still be downgraded for leaked reasoning, extra sentences, or unexplained technical jargon.
  • This is one short-form explainer, not a test of long-form research, factuality, or brand-voice consistency.
  • At content volume, small output-price differences compound even when quality scores are close.

What related resources help with Writing & Content?

Meta provider hubMuse Spark 1.3 Contributor pricingBest LLM for CodingBest LLM for Math & ReasoningBest LLM for Chatbots & Support

What are common questions about the best LLM for Writing & Content?

Which model writes the most human-sounding copy?

See the graded evidence block below — our test scores tone, jargon avoidance, and constraint-following, which correlates with "sounds human" more reliably than a subjective read.

Is a bigger model always better at writing?

No — flagship models sometimes over-explain or hedge. Several budget models scored within a few points of frontier models on our writing test.

Should I use a reasoning model for content writing?

Generally not — reasoning modes tend to leak visible "thinking" text into the output unless carefully prompted, which is a direct penalty in our grading.

Run this exact prompt against the top 3

Don't take a ranking's word for it — try Muse Spark 1.3 Contributor and its closest alternatives on your own prompt.

Try It Free