A failure worth fixing once.
Three teams worked around the same broken integration environment.
Every month the invoice says your coding agents worked hard. It cannot say on what, for whom, or whether any of it was worth it.
Quesma turns Claude Code, Codex, and Cursor sessions into a clear view of what the agents did, what it cost, and what you got back.

A team you didn’t expect found a use for the tools.
The default model got smarter, and pricier.
The work ran hot for three weeks because you asked it to.
Cost tools only find things to cut. Quesma also finds problems recurring across teams, and the skills, tools, and routines used in one part of the company but nowhere else.
Three teams worked around the same broken integration environment.
deploy-preview worked for one team. Others had never seen it.
One repository rule sent every session down the same wrong turn.
“Always run the full e2e suite before committing”
31 sessions took the same wrong turn
context7-mcp spread until more than half the company relied on it.
Can coding agents find backdoors in compiled binaries using Ghidra and radare2?
We built training environments for difficult, multi-hour agent tasks. The work required reading full trajectories: prompts, tool calls, failures, retries, and results.
We audited Blitzy's 66.5% SWE-Bench Pro result, reviewing the environment, required changes, and final score.
Read the independent audit Terminal-Bench 3.0Quesma contributed to Terminal-Bench 3.0, a benchmark for testing agents on practical terminal tasks.
View the contributors OpenAI-cited researchOpenAI cited our Baba Is Bench work while examining what makes agent evaluations succeed.
Read the citationAI mushroom identification from a photo with ChatGPT, Claude or Google Gemini: GPT-6 Astra, Gemini 3.8 Flash, Claude Fable 5.1, GPT-5.6 and GLM-5.3-Flash. Asked “What mushroom is that?” on 360 photos of poisonous species. Which warn, which get it right, which fail.
Mushroom identification with AI: GPT-5.6-Sol, Gemini 3.8 Flash, GLM-5.3-Flash and Claude Fable 5.1 benchmarked on poisonous and edible species of FungiTastic. A lot of dangerous errors.
A buyer’s guide to OpenAI Codex for individuals, teams, and enterprises. Includes comparisons with Claude Code.
Early access opens soon. The open-source collector runs in your environment and the session data stays in your cloud. Contact us to be one of the first teams.