Skip to main content

Quesma Benchmarks

Open-source benchmarks for evaluating AI coding agents on real-world software engineering tasks.

The score is only thestart of the record.

Blitzy SWE-Bench Pro audit

How Quesma independently reviewed the environment, methodology, and trajectories behind Blitzy's 66.5% public result.

Read the audit

Terminal-Bench 3.0 contribution

Quesma contributed to Terminal-Bench 3.0, a benchmark for testing agents on practical terminal tasks.

View the contributors