← 벤치마크
추론 벤치마크 최고 AI 모델
추론 벤치마크 점수에 따른 AI 모델 순위.
| 순위 | 모델 | 공급자 | 점수 | 데이터셋 | 버전 | 날짜 |
|---|---|---|---|---|---|---|
| 1 | DeepSeek R1 | DeepSeek | 71.5 | gpqa | paper-2025 | 2025-01-20 |
| 2 | Claude 3.5 Sonnet | Anthropic | 59.4 | gpqa | official-2024 | 2024-06-20 |
| 3 | GPT-4o | OpenAI | 53.6 | gpqa | official-2024 | 2024-05-13 |
| 4 | Llama 3.1 405B | Meta | 51.1 | gpqa | official-2024 | 2024-07-23 |
| 5 | Claude 3 Opus | Anthropic | 50.4 | gpqa | official-2024 | 2024-03-04 |
| 6 | Gemini 1.5 Pro | 46.5 | gpqa | techreport-2024 | 2024-05-14 | |
| 7 | GPT-4o mini | OpenAI | 40.2 | gpqa | official-2024 | 2024-07-18 |
| 8 | Claude 3 Haiku | Anthropic | 33.3 | gpqa | official-2024 | 2024-03-04 |
Source
- 데이터셋
- gpqa
- 버전
- paper-2025
- 날짜
- 2025-01-20
Benchmark
Related Resources
Use Cases