← Benchmarks
Best AI Models for Coding Benchmark
Ranking of AI models by Coding benchmark score.
| Rank | Model | Provider | Score | Dataset | Version | Date |
|---|---|---|---|---|---|---|
| 1 | Claude 3.5 Sonnet | Anthropic | 92 | humaneval | official-2024 | 2024-06-20 |
| 2 | Mistral Large 2 | Mistral | 92 | humaneval | official-2024 | 2024-07-24 |
| 3 | GPT-4o | OpenAI | 90.2 | humaneval | official-2024 | 2024-05-13 |
| 4 | Llama 3.1 405B | Meta | 89 | humaneval | official-2024 | 2024-07-23 |
| 5 | Llama 3.3 70B | Meta | 88.4 | humaneval | official-2024 | 2024-12-06 |
| 6 | GPT-4o mini | OpenAI | 87 | humaneval | official-2024 | 2024-07-18 |
| 7 | Qwen2.5 72B | Alibaba | 85.9 | humaneval | official-2024 | 2024-09-19 |
| 8 | Claude 3 Opus | Anthropic | 84.9 | humaneval | official-2024 | 2024-03-04 |
| 9 | Gemini 1.5 Pro | 84.1 | humaneval | techreport-2024 | 2024-05-14 | |
| 10 | DeepSeek V3 | DeepSeek | 82.6 | humaneval | paper-2024 | 2024-12-26 |
| 11 | GLM-4 | Zhipu | 80.1 | humaneval | official-2024 | 2024-01-16 |
| 12 | Claude 3 Haiku | Anthropic | 75.9 | humaneval | official-2024 | 2024-03-04 |
| 13 | Gemini 1.5 Flash | 71.7 | humaneval | techreport-2024 | 2024-05-14 |
Source
- Dataset
- humaneval
- Version
- official-2024
- Date
- 2024-06-20
Benchmark
Related Resources
Use Cases