COLLECTION / BENCHMARK

Leaderboards

Scraped public quality scores from GitHub Pages leaderboards such as Terminal-Bench 2.1. These are not local speed measurements.

48 records
AIME 2024aime-2024mathcategory199scored models91.8top scoremath
AIME 2025aime-2025mathcategory276scored models99.2top scoremath
ARC-AGIarc-agireasoningcategory5scored models9top scorereasoning
BFCL (function calling)bfclagenticcategory86scored models72.7top scoreagentic
CERcerspeech-asrcategory317scored models100top scorespeech-asr
ChartQAchartqamultimodalcategory200scored models89.7top scoremultimodal
Chatbot Arenachatbot-arenachatcategory166scored models92.4top scorechat
CodeContestscodecontestscodingcategory8scored models42.1top scorecoding
Codeforcescodeforcescodingcategory5scored models95.3top scorecoding
DeepSWEdeepsweswecategory17scored models46.2top scoreswe
DocVQAdocvqamultimodalcategory218scored models96.4top scoremultimodal
GPQAgpqareasoningcategory421scored models89top scorereasoning
GPQA Diamondgpqa-diamondreasoningcategory324scored models91.9top scorereasoning
HellaSwaghellaswagknowledgecategory392scored models91.8top scoreknowledge
HumanEvalhumanevalcodingcategory429scored models94.1top scorecoding
Humanity's Last Examhleknowledgecategory330scored models54.7top scoreknowledge
LiveCodeBenchlivecodebenchcodingcategory476scored models89top scorecoding
MATHmathmathcategory1015scored models99.6top scoremath
Math Olympiadmath-olympiadmathcategory16scored models77.6top scoremath
MATH-500math-500mathcategory141scored models97.8top scoremath
MMBenchmmbenchmultimodalcategory85scored models88.2top scoremultimodal
MMLUmmluknowledgecategory1152scored models94.4top scoreknowledge
MMLU-Prommlu-proknowledgecategory465scored models88top scoreknowledge
MMMUmmmumultimodalcategory324scored models82top scoremultimodal
MOSmosspeech-ttscategory23scored models7.7top scorespeech-tts
MT-Benchmt-benchchatcategory178scored models89.2top scorechat
OSWorldosworldagenticcategory17scored models76.1top scoreagentic
QuALITYqualityknowledgecategory34scored models100top scoreknowledge
RTFrtfspeech-asrcategory51scored models80top scorespeech-asr
RULERrulerlong-contextcategory109scored models96.8top scorelong-context