Key Takeaways
- ResNet-50 achieves 76.1% top-1 accuracy on ImageNet
- EfficientNet-B7 scores 84.3% top-1 on ImageNet
- ViT-Huge/14 reaches 88.55% top-1 on ImageNet-21k
- H100 SXM5 GPU delivers 1979 TFLOPS FP16 performance
- A100 80GB achieves 624 TFLOPS FP16 tensor
- Grok-1 314B model inference at 1.5x faster on custom stack
- GPT-4V achieves 85.5% accuracy on RealWorldQA
- LLaVA-1.5 13B scores 78.5% on ScienceQA
- Kosmos-2 scores 68.8% on OK-VQA
- GPT-4 achieves 86.4% accuracy on the MMLU benchmark
- Llama 2 70B scores 68.9% on MMLU
- Claude 2 scores 75.0% on MMLU
- Claude 3.5 Sonnet reaches 84.9% on HumanEval
- GPT-4o scores 90.2% on HumanEval pass@1
- o1-preview achieves 74.4% on AIME 2024
Across benchmarks, state of the art models deliver up to 88.6% ImageNet top one and strong multimodal question answering.
Related reading
01 · Category
Computer Vision20 stats
Computer Vision Interpretation
02 · Category
Efficiency and Inference21 stats
Efficiency and Inference Interpretation
03 · Category
Multimodal Models19 stats
Multimodal Models Interpretation
More related reading
04 · Category
Natural Language Processing24 stats
Natural Language Processing Interpretation
05 · Category
Reasoning and Mathematics20 stats
Reasoning and Mathematics Interpretation
Cite This Report
This report is designed to be cited. We maintain stable URLs and versioned verification dates. Copy the format appropriate for your publication below.
Elif Demirci. (2026, February 24). AI Benchmark Statistics. Gitnux. https://gitnux.org/ai-benchmark-statistics
Elif Demirci. "AI Benchmark Statistics." Gitnux, 24 Feb 2026, https://gitnux.org/ai-benchmark-statistics.
Elif Demirci. 2026. "AI Benchmark Statistics." Gitnux. https://gitnux.org/ai-benchmark-statistics.
Sources & references
20 datasets cited across this report · attribution is report-level

