
RTK reports huge token savings, but our cost benchmarks disagree
We tested RTK (Rust Token Killer) with Claude Code on Fable 5.0, and OpenCode with DeepSeek V4 Pro 0813 on Terminal-Bench 2.1.
Know what AI tokens buy you.
Token economics is the discipline of knowing what AI tokens buy you: how compute is priced into tokens, how organizations allocate them, how agents consume them, and how that spend is measured against outcomes.
input-token spread, same model, leaner harness scored higher
Terminal-Bench, 2026 ↗4.7xClaude Code ~32,800 vs OpenCode ~6,900 scaffolding tokens
Systima, 2026 ↗≈80%of a real provider bill came from cache traffic
arXiv:2607.12161, July 2026 ↗−38% / +6.8%tool-output tokens removed, cost increased after the cached prefix broke
arXiv:2607.12161, July 2026 ↗≈0.1xbase input price for cache reads on both major providers
Anthropic and OpenAI pricing, 2026-08 ↗≈2xbill after a Codex CLI compaction retune
Codex CLI v0.118 regression, 2026 ↗$1,434/dayconsumed by a documented retry storm
arXiv:2604.16646, 2026 ↗$1,500/moper-tool cap at Uber after about four months of budget burn
TechCrunch, 2026 ↗+7.6%cost at low effort in an A/B test of a plugin claiming 60–90% cuts
JetBrains, July 2026 ↗
We tested RTK (Rust Token Killer) with Claude Code on Fable 5.0, and OpenCode with DeepSeek V4 Pro 0813 on Terminal-Bench 2.1.
AI mushroom identification from a photo with ChatGPT, Claude or Google Gemini: GPT-6 Astra, Gemini 3.8 Flash, Claude Fable 5.1, GPT-5.6 and GLM-5.3-Flash. Asked “What mushroom is that?” on 360 photos of poisonous species. Which warn, which get it right, which fail.
Mushroom identification with AI: GPT-5.6-Sol, Gemini 3.8 Flash, GLM-5.3-Flash and Claude Fable 5.1 benchmarked on poisonous and edible species of FungiTastic. A lot of dangerous errors.