Glossary
Benchmark
Also: benchmark · SWE-bench
A public, standardized task set used to compare models and agents. Unlike your own evals, it reflects typical tasks, not yours.
Example
SWE-bench checks whether an agent can fix real issues from GitHub repositories so that the tests pass.