Rankings

Leaderboards and benchmarks for comparing AI models across tasks, languages, and modalities.

databaseEmbeddings

MTEB Leaderboard

Massive Text Embedding Benchmark. Compares embedding models across 8 tasks: classification, clustering, pair classification, reranking, retrieval, STS, summarization, and bitext mining.

embeddingsretrievalstsclusteringhuggingface
https://huggingface.co/spaces/mteb/leaderboard open_in_new
forumLLM

Chatbot Arena

Crowdsourced LLM benchmark by LMSYS. Users vote on blind head-to-head model comparisons, producing Elo-style rankings across coding, reasoning, math, and instruction following.

llmelochatbothuman-evallmsys
https://lmarena.ai/?leaderboard open_in_new
emoji_eventsLLM

Open LLM Leaderboard

Hugging Face open LLM benchmark. Evaluates open-weight models on ARC, HellaSwag, MMLU, TruthfulQA, Winogrande, and GSM8K with reproducible evaluation pipelines.

llmopen-sourcemmluhuggingfacebenchmark
https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard open_in_new
speedMulti-modal

Artificial Analysis

Independent AI model quality, speed, and pricing comparison. Real-time leaderboards for LLMs, image, and speech models across multiple providers.

benchmarkpricingspeedqualitycomparison
https://artificialanalysis.ai open_in_new
updateLLM

LiveBench

Contamination-free LLM benchmark with regularly updated questions. Tests math, coding, reasoning, language, instruction following, and data analysis with verifiable ground-truth answers.

llmbenchmarkcontamination-freemathcoding
https://livebench.ai open_in_new
codeCode

EvalPlus Leaderboard

Rigorous code generation benchmark. Evaluates LLMs on HumanEval+ and MBPP+ with extended test cases to catch false positives in code synthesis.

codehumanevalmbppgenerationbenchmark
https://evalplus.github.io/leaderboard.html open_in_new
terminalCode

BigCodeBench

Benchmark for code generation with complex function calls. Tests LLMs on practical programming tasks requiring library usage, data processing, and multi-step reasoning.

codebenchmarkfunction-callinghuggingface
https://huggingface.co/spaces/bigcode/bigcodebench-leaderboard open_in_new
visibilityVision

Vision Arena

Crowdsourced vision-language model benchmark. Blind pairwise comparisons for image understanding, visual QA, and multimodal reasoning with Elo ratings.

visionmultimodalimageelolmsys
https://lmarena.ai/?vision open_in_new
bug_reportCode

SWE-bench

Software Engineering benchmark testing AI on real GitHub issues. Evaluates ability to understand codebases, locate bugs, and generate working patches for real-world repositories.

codesoftware-engineeringgithubbugfixbenchmark
https://www.swebench.com open_in_new
edit_noteCode

Aider Polyglot Leaderboard

Benchmark for AI code editing across multiple programming languages. Measures how well LLMs can modify existing code, fix bugs, and implement features via aider coding assistant.

codeeditingpolyglotaiderbenchmark
https://aider.chat/docs/leaderboards/ open_in_new
workspace_premiumLLM

SEAL Leaderboards

Scale AI expert-driven evaluations. Tests models on hard prompts across coding, instruction following, and math with professional human graders.

llmexpert-evalcodingmathscale
https://scale.com/leaderboard open_in_new
micAudio

Open ASR Leaderboard

Automatic Speech Recognition benchmark on Hugging Face. Compares transcription models on word error rate across multiple datasets and languages.

audiospeechasrtranscriptionhuggingface
https://huggingface.co/spaces/hf-audio/open_asr_leaderboard open_in_new