← Benchmarks
Beste KI-Modelle im Schlussfolgerung-Benchmark
Ranking von KI-Modellen nach Schlussfolgerung-Benchmark-Score.
| Rang | Modell | Anbieter | Punktzahl | Datensatz | Version | Datum |
|---|---|---|---|---|---|---|
| 1 | DeepSeek R1 | DeepSeek | 71.5 | gpqa | paper-2025 | 2025-01-20 |
| 2 | Claude 3.5 Sonnet | Anthropic | 59.4 | gpqa | official-2024 | 2024-06-20 |
| 3 | GPT-4o | OpenAI | 53.6 | gpqa | official-2024 | 2024-05-13 |
| 4 | Llama 3.1 405B | Meta | 51.1 | gpqa | official-2024 | 2024-07-23 |
| 5 | Claude 3 Opus | Anthropic | 50.4 | gpqa | official-2024 | 2024-03-04 |
| 6 | Gemini 1.5 Pro | 46.5 | gpqa | techreport-2024 | 2024-05-14 | |
| 7 | GPT-4o mini | OpenAI | 40.2 | gpqa | official-2024 | 2024-07-18 |
| 8 | Claude 3 Haiku | Anthropic | 33.3 | gpqa | official-2024 | 2024-03-04 |
Source
- Datensatz
- gpqa
- Version
- paper-2025
- Datum
- 2025-01-20
Benchmark
Related Resources
Use Cases