← Benchmarks
Best AI Models for Reasoning Benchmark
Ranking of AI models by Reasoning benchmark score.
| Rank | Model | Provider | Score | Dataset | Version | Date |
|---|---|---|---|---|---|---|
| 1 | DeepSeek R1 | DeepSeek | 71.5 | gpqa | paper-2025 | 2025-01-20 |
| 2 | Claude 3.5 Sonnet | Anthropic | 59.4 | gpqa | official-2024 | 2024-06-20 |
| 3 | GPT-4o | OpenAI | 53.6 | gpqa | official-2024 | 2024-05-13 |
| 4 | Llama 3.1 405B | Meta | 51.1 | gpqa | official-2024 | 2024-07-23 |
| 5 | Claude 3 Opus | Anthropic | 50.4 | gpqa | official-2024 | 2024-03-04 |
| 6 | Gemini 1.5 Pro | 46.5 | gpqa | techreport-2024 | 2024-05-14 | |
| 7 | GPT-4o mini | OpenAI | 40.2 | gpqa | official-2024 | 2024-07-18 |
| 8 | Claude 3 Haiku | Anthropic | 33.3 | gpqa | official-2024 | 2024-03-04 |
Source
- Dataset
- gpqa
- Version
- paper-2025
- Date
- 2025-01-20
Benchmark
Related Resources
Use Cases