Key Takeaways
- GPT-3 pre-training compute: 3.14 × 10^23 FLOP.
- PaLM 540B pre-training compute: 2.5 × 10^25 FLOP.
- LLaMA 65B pre-training compute: 1.2 × 10^24 FLOP.
- GPT-3 dataset size: approximately 300 billion tokens.
- PaLM 540B dataset size: 780 billion tokens.
- LLaMA 65B dataset size: 1.4 trillion tokens.
- GPT-3 training energy: 1,287 MWh.
- PaLM 540B training energy: ~10,000 MWh estimate.
- LLaMA 65B training energy: 784 MWh.
- GPT-3 parameter count: 175 billion.
- PaLM parameter count: 540 billion.
- LLaMA parameter count: 65 billion.
- GPT-3 training cost estimate: $4.6 million.
- PaLM 540B training cost: approximately $8 million.
- LLaMA 65B training cost: under $100k on public clouds.
Model scale keeps rising, from GPT-3 to PaLM and LLaMA, but energy and cost vary drastically.
Related reading
01 · Category
Compute Resources24 stats
Compute Resources Interpretation
02 · Category
Dataset Sizes24 stats
Dataset Sizes Interpretation
03 · Category
Energy Consumption25 stats
Energy Consumption Interpretation
More related reading
04 · Category
Parameter Counts24 stats
Parameter Counts Interpretation
05 · Category
Training Costs20 stats
Training Costs Interpretation
Cite This Report
This report is designed to be cited. We maintain stable URLs and versioned verification dates. Copy the format appropriate for your publication below.
Elena Vasquez. (2026, February 24). AI Training Statistics. Gitnux. https://gitnux.org/ai-training-statistics
Elena Vasquez. "AI Training Statistics." Gitnux, 24 Feb 2026, https://gitnux.org/ai-training-statistics.
Elena Vasquez. 2026. "AI Training Statistics." Gitnux. https://gitnux.org/ai-training-statistics.
Sources & references
12 datasets cited across this report · attribution is report-level

