BabaIsBench
LLM agent benchmark based on the puzzle game Baba Is You. Agents play levels through a text interface, rewriting the rules to win, compared on pass rate, speed, and cost.
Open-source benchmarks for evaluating AI coding agents on real-world software engineering tasks.
Stay tuned for future benchmarks
Subscribe via RSSLLM agent benchmark based on the puzzle game Baba Is You. Agents play levels through a text interface, rewriting the rules to win, compared on pass rate, speed, and cost.
Security analysis benchmark for AI coding agents. Tests models on detecting backdoors, timebombs, and malicious code in compiled binaries using reverse engineering tools like Ghidra and radare2.

OpenTelemetry instrumentation benchmark for AI coding agents. Tests models on real-world tasks adding distributed tracing, metrics, and logging to multi-language codebases.
Build system benchmark for AI coding agents. Tests models on fixing compilation errors, updating dependencies, and navigating complex build configurations.
Independent audit
How Quesma independently reviewed the environment, methodology, and trajectories behind Blitzy's 66.5% public result.
Read the auditContribution
Quesma contributed to Terminal-Bench 3.0, a benchmark for testing agents on practical terminal tasks.
View the contributors