A failure worth fixing once.
Three teams worked around the same broken integration environment.
Every month the invoice says your coding agents worked hard. It cannot say on what, for whom, or whether any of it was worth it.
Quesma turns Claude Code, Codex, and Cursor sessions into a clear view of what the agents did, what it cost, and what you got back.

A team you didn’t expect found a use for the tools.
The default model got smarter, and pricier.
The work ran hot for three weeks because you asked it to.
Cost tools only find things to cut. Quesma also finds problems recurring across teams, and the skills, tools, and routines used in one part of the company but nowhere else.
Three teams worked around the same broken integration environment.
deploy-preview worked for one team. Others had never seen it.
One repository rule sent every session down the same wrong turn.
“Always run the full e2e suite before committing”
31 sessions took the same wrong turn
context7-mcp spread until more than half the company relied on it.
Can coding agents find backdoors in compiled binaries using Ghidra and radare2?
We built training environments for difficult, multi-hour agent tasks. The work required reading full trajectories: prompts, tool calls, failures, retries, and results.
We audited Blitzy's 66.5% SWE-Bench Pro result, reviewing the environment, required changes, and final score.
Read the independent audit Terminal-Bench 3.0Quesma contributed to Terminal-Bench 3.0, a benchmark for testing agents on practical terminal tasks.
View the contributors OpenAI-cited researchOpenAI cited our Baba Is Bench work while examining what makes agent evaluations succeed.
Read the citationClaude Code buyer’s guide, August 2026: what Max, Team, and Enterprise cost, how to get the best deal, and how Uber and Shopify cap the spend.

Common wisdom says to put a strong model like Fable in charge and let cheaper models do the work. I tested it and the result was not what I expected: seventh place at twice the cost of the top single-model run.
Quantization is a lossy compression, and factual knowledge is incompressible. We test Qwen3.6 27B GGUF quantizations from Hugging Face (by Unsloth and Bartowski) and llama.cpp on the Incompressible Knowledge Probes (IKP) benchmark.
Early access opens soon. The open-source collector deploys in your environment, the record accumulates in your cloud, and early teams get every analysis first. Contact us to get in early.