Rankings
Leaderboards and benchmarks for comparing AI models across tasks, languages, and modalities.
MTEB Leaderboard
Massive Text Embedding Benchmark. Compares embedding models across 8 tasks: classification, clustering, pair classification, reranking, retrieval, STS, summarization, and bitext mining.
https://huggingface.co/spaces/mteb/leaderboardChatbot Arena
Crowdsourced LLM benchmark by LMSYS. Users vote on blind head-to-head model comparisons, producing Elo-style rankings across coding, reasoning, math, and instruction following.
https://lmarena.ai/?leaderboardOpen LLM Leaderboard
Hugging Face open LLM benchmark. Evaluates open-weight models on ARC, HellaSwag, MMLU, TruthfulQA, Winogrande, and GSM8K with reproducible evaluation pipelines.
https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboardArtificial Analysis
Independent AI model quality, speed, and pricing comparison. Real-time leaderboards for LLMs, image, and speech models across multiple providers.
https://artificialanalysis.aiLiveBench
Contamination-free LLM benchmark with regularly updated questions. Tests math, coding, reasoning, language, instruction following, and data analysis with verifiable ground-truth answers.
https://livebench.aiEvalPlus Leaderboard
Rigorous code generation benchmark. Evaluates LLMs on HumanEval+ and MBPP+ with extended test cases to catch false positives in code synthesis.
https://evalplus.github.io/leaderboard.htmlBigCodeBench
Benchmark for code generation with complex function calls. Tests LLMs on practical programming tasks requiring library usage, data processing, and multi-step reasoning.
https://huggingface.co/spaces/bigcode/bigcodebench-leaderboardVision Arena
Crowdsourced vision-language model benchmark. Blind pairwise comparisons for image understanding, visual QA, and multimodal reasoning with Elo ratings.
https://lmarena.ai/?visionSWE-bench
Software Engineering benchmark testing AI on real GitHub issues. Evaluates ability to understand codebases, locate bugs, and generate working patches for real-world repositories.
https://www.swebench.comAider Polyglot Leaderboard
Benchmark for AI code editing across multiple programming languages. Measures how well LLMs can modify existing code, fix bugs, and implement features via aider coding assistant.
https://aider.chat/docs/leaderboards/SEAL Leaderboards
Scale AI expert-driven evaluations. Tests models on hard prompts across coding, instruction following, and math with professional human graders.
https://scale.com/leaderboardOpen ASR Leaderboard
Automatic Speech Recognition benchmark on Hugging Face. Compares transcription models on word error rate across multiple datasets and languages.
https://huggingface.co/spaces/hf-audio/open_asr_leaderboard