Vision Models
Models with vision (image understanding) capabilities, ranked by overall score.
Based on 10 AI models
| Model | Provider | Score | Price (in) | Context |
|---|---|---|---|---|
| Claude 3.5 Sonnet | Anthropic | 78.4 | $3.00 | 200 000 |
| Gemini 1.5 Pro | 76.9 | $1.25 | 2 000 000 | |
| GPT-4o | OpenAI | 73.1 | $2.50 | 128 000 |
| Gemini 2.0 Flash | 47.5 | $0.10 | 1 048 576 | |
| Gemini 2.5 Flash | 47.5 | $0.30 | 1 048 576 | |
| Gemini 2.5 Pro | 47.5 | $1.25 | 1 048 576 | |
| Gemini 3 Flash | 47.5 | $0.50 | 1 048 576 | |
| Gemini 3.1 Pro | 47.5 | $2.00 | 2 000 000 | |
| Claude 3.7 Sonnet | Anthropic | 44.8 | $3.00 | 200 000 |
| Claude Haiku 4.5 | Anthropic | 44.8 | $1.00 | 200 000 |
Frequently Asked Questions
What is the best model for Vision Models?
Based on our data, Claude 3.5 Sonnet ranks first for Vision Models with a score of 78.
How were these models selected?
Models are ranked by scenario-specific data: benchmark scores (coding/reasoning/math/vision), capability support, context window, or price. All data comes from our transparent model catalog, updated daily.
How fresh is this data?
Model data, prices, and rankings are refreshed on a regular schedule; ranking scores are snapshotted daily. Each model page shows its last verified date.