Benchmarks and Leaderboards

Public benchmarks and leaderboards for comparing AI models. Benchmarks such as SWE-bench, Terminal-Bench, and LiveBench are a useful first filter for capability in a specific domain. Cross-provider aggregators are useful for comparing capability, speed, and price across labs.