← ベンチマーク
推論ベンチマークで最優秀のAIモデル
推論ベンチマークスコアによるAIモデルのランキング。
| 順位 | モデル | プロバイダー | スコア | データセット | バージョン | 日付 |
|---|---|---|---|---|---|---|
| 1 | DeepSeek R1 | DeepSeek | 71.5 | gpqa | paper-2025 | 2025-01-20 |
| 2 | Claude 3.5 Sonnet | Anthropic | 59.4 | gpqa | official-2024 | 2024-06-20 |
| 3 | GPT-4o | OpenAI | 53.6 | gpqa | official-2024 | 2024-05-13 |
| 4 | Llama 3.1 405B | Meta | 51.1 | gpqa | official-2024 | 2024-07-23 |
| 5 | Claude 3 Opus | Anthropic | 50.4 | gpqa | official-2024 | 2024-03-04 |
| 6 | Gemini 1.5 Pro | 46.5 | gpqa | techreport-2024 | 2024-05-14 | |
| 7 | GPT-4o mini | OpenAI | 40.2 | gpqa | official-2024 | 2024-07-18 |
| 8 | Claude 3 Haiku | Anthropic | 33.3 | gpqa | official-2024 | 2024-03-04 |
Source
- データセット
- gpqa
- バージョン
- paper-2025
- 日付
- 2025-01-20
Benchmark
Related Resources
Use Cases