18 of the 335 figures in these field notes come from a sentence that names GPT-5.6. Each one is quoted verbatim from the write-up it appeared in, with the day that write-up went out.
Some of these are about GPT-5.6 and some only measure against it — a competitor’s price quoted next to GPT-5.6’s belongs here too, because that is the sentence someone searching for the comparison is looking for. The sentence tells you which is which.
This is not a GPT-5.6 price sheet. Nothing here is read off a vendor page. These are figures a working developer wrote down, and the date is the day the piece was published — not the day the price was in force. Model prices move fast: read every row as of its own date, and treat an old one as a lead rather than a quote.
- 54% — “Sol operates in an ultra-mode with internal multi-agents, which improves task completion accuracy and token efficiency; in agentic coding tasks it is 54% more token-efficient than comparable models.” (2026-08-21) →
- 43.1% — “On ReactBench, GPT 5.6 Sol and Fable 5 posted Pass@1 scores of 43.1% and 41.2%.” (2026-08-21) →
- 95% — “Luna answers fastest, hits 95% accuracy on basic question answering, and costs the least per batch.” (2026-08-19) →
- 54% — “Sol is the heavy one, with a claimed 54% better token efficiency than models at the same level and an Ultra mode that runs 4 sub-agents in parallel.” (2026-08-19) →
- 2 hours 41 minutes 35 seconds — “Sol’s Ultra mode took 2 hours 41 minutes 35 seconds to build one detailed 3D scene, running its sub-agents in parallel the whole time, and a run of that length is not something a fallback can politely interrupt halfway through.” (2026-08-19) →
- $1.43 per run — “One front-end benchmark put GPT-5.6 Sol at $1.43 per run against $9.05 for Fable 5.” (2026-08-19) →
- 95% — “Luna’s 95% accuracy rate at 1/3 the cost of Terra shows how much you can save by making the right choice.” (2026-08-12) →
- 95% — “Terra’s document processing capabilities might seem cost-effective initially, but Luna’s 95% accuracy on basic QA tasks and faster response times mean fewer errors and rework, whereas Terra’s higher failure rate on complex tasks can lead to time wasted fixing mistakes, and Luna’s superior accuracy and reliability make it a better long-term choice, even though Terra often requires more tokens for similar tasks.” (2026-08-12) →
- 95% — “For example, Luna’s 95% accuracy rate for basic questions drops when faced with more complex queries, and Terra’s document analysis accuracy can vary depending on document structure and content.” (2026-08-12) →
- 95% — “For example, Luna’s 95% accuracy in basic Q&A is great for customer service, but the error correction time can be a drawback.” (2026-08-12) →
- 95% — “Luna’s 95% accuracy rate for basic questions makes it ideal for simple tasks, while Terra’s strength shows up in document analysis and more complex scenarios.” (2026-08-12) →
- 95% — “Luna’s 95% accuracy rate reduces rework costs, and Terra’s higher token costs add up over time.” (2026-08-12) →
- 3x — “Terra requires 3x more tokens for equivalent tasks, and its 800ms response time makes it far less efficient for time-sensitive applications.” (2026-08-12) →
- 2x — “Terra’s ability to extract all data is ideal for document processing, but its 2x token usage may be a concern for large-scale projects.” (2026-08-12) →
- 2x — “Luna takes less developer time, which shows how much more efficient it is overall, and Terra’s 2x more tokens for equivalent tasks adds up quickly in larger projects.” (2026-08-12) →
- 43.1% — “ReactBench tests showed that GPT 5.6 Sol and Fable 5 had Pass@1 scores of only 43.1% and 41.2% respectively, indicating problems in real-world React projects.” (2026-08-12) →
- $1.43 per run — “In contrast, GPT-5.6 Sol, at $1.43 per run, achieves 43.1% accuracy in the same tests, suggesting that while cheaper models may save money upfront, they often result in longer, more costly development processes.” (2026-08-12) →
- $1.43 per run — “For instance, GPT-5.6 Sol, while more expensive at $1.43 per run, shows superior performance with a 43.1% accuracy rate in the same ReactBench tests, which shows that cheaper models may save money upfront but can lead to longer development cycles due to frequent errors and rework.” (2026-08-12) →
All 335 figures, every kind · JSON · CSV
Where these 18 came from
A GPT-5.6 figure that is not here yet? Say which metric, which unit, and where you read it — in one line. The form already knows it is about GPT-5.6.
Or is one of the 18 above already out of date? Say which one — the form already knows it is about GPT-5.6; you only have to say what the number is now.
All write-ups