Key Takeaways
- Llama 3.1 MMLU score 88.6% with 128k context.
- GPT-4o achieves 88.7% on MMLU benchmark.
- Claude 3.5 Sonnet GPQA score 59.4%.
- GPT-4o supports a context window of 128,000 tokens for input.
- Claude 3.5 Sonnet has a 200,000 token context window.
- Gemini 1.5 Pro offers up to 1 million tokens in context window.
- GPT-4 Turbo input speed 4000 tokens/sec.
- Llama 3.1 405B requires 810 GB VRAM for 128k context.
- Mixtral 8x22B uses 140 GB RAM at FP16 for full context.
- RAG systems with LlamaIndex reduce context by 70% via retrieval.
- LangChain RAG pipelines achieve 25% accuracy boost on HotpotQA.
- FAISS index retrieval latency averages 5ms for 1M docs.
- GPT-3.5 Turbo has 16,385 token context window.
- Llama 3.1 8B processes 50 tokens/second on A100 GPU.
- Mistral 7B Instruct achieves 70 tokens/sec inference speed.
Newer models top strong benchmarks and long contexts, while RAG techniques cut tokens and reduce hallucinations.
Related reading
01 · Category
Benchmark Performance Scores21 stats
Benchmark Performance Scores Interpretation
02 · Category
Context Window Capacities25 stats
Context Window Capacities Interpretation
03 · Category
Memory Consumption Stats20 stats
Memory Consumption Stats Interpretation
More related reading
04 · Category
Retrieval Augmentation Metrics20 stats
Retrieval Augmentation Metrics Interpretation
05 · Category
Token Processing Speeds25 stats
Token Processing Speeds Interpretation
Cite This Report
This report is designed to be cited. We maintain stable URLs and versioned verification dates. Copy the format appropriate for your publication below.
Marie Larsen. (2026, February 24). Model Context Protocol Statistics. Gitnux. https://gitnux.org/model-context-protocol-statistics
Marie Larsen. "Model Context Protocol Statistics." Gitnux, 24 Feb 2026, https://gitnux.org/model-context-protocol-statistics.
Marie Larsen. 2026. "Model Context Protocol Statistics." Gitnux. https://gitnux.org/model-context-protocol-statistics.
Sources & references
26 datasets cited across this report · attribution is report-level

