# Compare AI Models Side-by-Side Across 14 Benchmarks

Select 2-5 models and compare them side-by-side across every benchmark.

Models: [GPT-5.2](/content/benchmark/models/gpt-5-2/index.html) [Claude Opus 4.6](/content/benchmark/models/claude-opus-4-6/index.html) [Gemini 3.1 Pro](/content/benchmark/models/gemini-3-1-pro/index.html) [DeepSeek R1 0528](/content/benchmark/models/deepseek-r1-0528/index.html)

---

## Benchmarks

### Chatbot Arena ELO (ELO)
Human preference ELO from blind head-to-head votes  
**Higher is better**  
1475 1496 1501 1375

### MMLU-Pro (%)
Massive Multitask Language Understanding professional benchmark  
**Higher is better**  
83.0% 82.0% 92.6% 84.0%

### HumanEval+ (%)
Code generation correctness with extended tests  
**Higher is better**  
95.0% 93.5% 93.0% 88.0%

### IFEval (%)
Strict instruction following accuracy on verifiable constraints  
**Higher is better**  
95.0% 91.0% 92.0% 88.0%

### MATH (%)
Competition mathematics problem solving  
**Higher is better**  
97.0% 92.0% 93.0% 96.0%

### SWE-bench Verified (%)
Real-world software engineering task resolution  
**Higher is better**  
80.0% 80.8% 80.6% 55.0%

### GPQA Diamond (%)
Graduate-level science Q&A by domain experts  
**Higher is better**  
92.4% 91.3% 94.1% 81.0%

### Output Speed (tok/s)
Tokens generated per second  
**Higher is better**  
90 68 65 40

### Time to First Token (ms)
Latency before first token arrives  
**Lower is better**  
380 1680 4200 900

### LiveCodeBench (%)
Live competitive programming benchmark  
**Higher is better**  
80.0% 72.0% 75.0% 58.0%

### Input Cost ($/1M)
Cost per 1M input tokens  
**Lower is better**  
$1.8 $5.0 $2.0 $0.55

### Aider Polyglot (%)
Multi-language code editing accuracy with real git repos  
**Higher is better**  
75.0%

### Output Cost ($/1M)
Cost per 1M output tokens  
**Lower is better**  
$14.0 $25.0 $12.0 $2.2

### BFCL (%)
Berkeley Function Calling Leaderboard — tool use accuracy  
**Higher is better**  
59.2% 70.4%

### GSM8K (%)
Grade school math word problems  
**Higher is better**  
99.0% 97.0% 98.0% 97.5%

### AIME 2025 (%)
American Invitational Mathematics Examination — competition math  
**Higher is better**  
100.0% 100.0% 95.0% 87.5%

### ARC-AGI (%)
Abstraction and Reasoning Corpus for general intelligence  
**Higher is better**  
54.2% 60.0% 77.1% 15.0%

### Context Length (tokens)
Maximum context window size  
**Higher is better**  
400K 200K 1.0M 131K

## AI Comparison Tool FAQ

### How many AI models can I compare at once?
You can compare 2 to 5 AI models side-by-side. Select models from the dropdown to add them to the comparison. This range lets you see meaningful differences without overwhelming the charts.

### What benchmarks are included in the comparison?
The comparison includes all 14+ benchmarks tracked by the AI Value Index: Chatbot Arena ELO, SWE-bench Verified, MMLU-Pro, HumanEval, MATH, GPQA Diamond, output speed, time to first token, input and output pricing, and more across General, Coding, Math, Reasoning, Speed, and Cost categories.

### Can I share my AI model comparison?
Yes. The URL updates as you select models, so you can copy and share the link with anyone. They will see the exact same comparison you created. You can also bookmark comparisons for later reference.

### What is the difference between radar and bar chart views?
The radar chart shows all metrics at once on a spider/polygon chart, making it easy to see overall strengths and weaknesses. Bar charts compare models on individual metrics with exact values. Use radar for a quick overview and bar charts for precise comparisons.

### Which AI models should I compare?
It depends on your use case. For best quality, compare flagship models like GPT-5.2, Claude Opus 4.6, and Gemini 2.5 Pro. For value, compare mid-range options like GPT-5, Claude Sonnet 4.6, and DeepSeek V3.2. For budget apps, compare GPT-5 Nano, Gemini Flash, and Qwen models.
