Model Leaderboard

Compare AI models by capability and cost-effectiveness

Popular Comparisons

Logical Reasoning

71/252 models

HLE: Complex reasoning and problem-solving

Use Cases: Complex decision-making, multi-step analysis, logical reasoning

Knowledge Q&A

62/252 models

AA-Omniscience: Knowledge Reliability & Hallucination

Use Cases: Expert Q&A, fact-checking, educational tutoring

Scientific Research

74/252 models

GPQA: Graduate-level science questions

Use Cases: Academic research, scientific writing, experiment design

AI Agent

41/252 models

τ³-Banking: Knowledge Retrieval & Multi-step Tool Use

Use Cases: Automated workflows, multi-tool invocation, complex task decomposition

SciCode

55/252 models

SciCode: Scientific coding challenges

Use Cases: Scientific computing, research code, data analysis scripts

Programming & Development

66/252 models

Terminal-Bench 2.1: Agent Coding & Terminal Operations

Use Cases: Agentic Coding, Shell Scripting, DevOps Automation

Instruction

53/252 models

IFEval: Instruction following accuracy

Use Cases: Precise task execution, format compliance, constraint adherence

Disclaimer: Rankings are for reference only and do not represent precise test results or constitute any purchase or usage advice. We do not guarantee the accuracy, completeness, or timeliness of the data.

Data Sources: Rankings are based on official technical reports and public evaluations from model providers.