Gitnux/Report 2026

AI Training Statistics

LLaMA 65B training can cost under $100k on public clouds—see how that compares with PaLM’s ~$8M in our AI training stats.
117Statistics
5Sections
6mRead
2 days agoUpdated
AI Training Statistics
Verified via a 4-step process
01Source

Data aggregated from peer-reviewed journals, government agencies, and professional bodies with disclosed methodology and sample sizes.

02Verify

Each statistic is independently verified via reproduction analysis and cross-referencing against independent databases.

03Grade

Figures are graded by cross-model consensus. Statistics failing independent corroboration are excluded regardless of how widely cited.

04Cite

Every figure carries a primary source. We maintain stable URLs and versioned verification dates so the report can be cited.

Read our full methodology →

Statistics that fail independent corroboration are excluded.

Next review Jan 2027
AI training statistics show how quickly scale has grown, because compute, data, and energy requirements rise together. This page compares GPT-3, PaLM 540B, LLaMA 65B, and BLOOM 176B across parameter counts, dataset sizes, and training costs—using figures like 300B vs 780B vs 1.4T tokens and $4.6M, ~$8M, and $3M estimates. You’ll also see how training energy varies, from 1,287 MWh to an ~10,000 MWh estimate and 433,000 kWh for BLOOM.

Key Takeaways

  • GPT-3 pre-training compute: 3.14 × 10^23 FLOP.
  • PaLM 540B pre-training compute: 2.5 × 10^25 FLOP.
  • LLaMA 65B pre-training compute: 1.2 × 10^24 FLOP.
  • GPT-3 dataset size: approximately 300 billion tokens.
  • PaLM 540B dataset size: 780 billion tokens.
  • LLaMA 65B dataset size: 1.4 trillion tokens.
  • GPT-3 training energy: 1,287 MWh.
  • PaLM 540B training energy: ~10,000 MWh estimate.
  • LLaMA 65B training energy: 784 MWh.
  • GPT-3 parameter count: 175 billion.
  • PaLM parameter count: 540 billion.
  • LLaMA parameter count: 65 billion.
  • GPT-3 training cost estimate: $4.6 million.
  • PaLM 540B training cost: approximately $8 million.
  • LLaMA 65B training cost: under $100k on public clouds.

Model scale keeps rising, from GPT-3 to PaLM and LLaMA, but energy and cost vary drastically.

01 · Category

Compute Resources24 stats

01
GPT-3 pre-training compute: 3.14 × 10^23 FLOP.
02
PaLM 540B pre-training compute: 2.5 × 10^25 FLOP.
03
LLaMA 65B pre-training compute: 1.2 × 10^24 FLOP.
04
BLOOM 176B pre-training compute: 3.5 × 10^24 FLOP.
05
OPT-175B pre-training compute: 1.8 × 10^24 FLOP.
06
Gopher 280B pre-training compute: 1.9 × 10^24 FLOP.
07
Chinchilla 70B pre-training compute: 1.4 × 10^24 FLOP.
08
MT-NLG 530B pre-training compute: 1.7 × 10^25 FLOP.
09
Jurassic-1 Jumbo 178B pre-training compute: 6.8 × 10^23 FLOP.
10
Megatron-Turing NLG 530B pre-training compute: 5.0 × 10^24 FLOP.
11
Falcon 180B pre-training compute: 3.5 × 10^25 FLOP.
12
LLaMA 2 70B pre-training compute: 3.3 × 10^24 FLOP.
13
StableLM 3B pre-training compute: 1.5 × 10^22 FLOP.
14
T5-XXL 11B pre-training compute: 3.7 × 10^23 FLOP.
15
BERT-Large pre-training compute: 2.0 × 10^21 FLOP.
16
GPT-2 XL 1.5B pre-training compute: 4.4 × 10^21 FLOP.
17
Grok-1 314B pre-training compute estimate: 5.0 × 10^24 FLOP.
18
Inflection-2.5 pre-training compute: 8.0 × 10^24 FLOP.
19
Command R+ 104B pre-training compute: 2.0 × 10^24 FLOP.
20
Mixtral 8x7B pre-training compute: 1.0 × 10^24 FLOP.
21
DBRX 132B pre-training compute: 1.0 × 10^25 FLOP.
22
Yi-34B pre-training compute: 1.2 × 10^24 FLOP.
23
Qwen-72B pre-training compute: 2.0 × 10^24 FLOP.
24
DeepSeek-V2 236B pre-training compute: 5.8 × 10^24 FLOP.
Interpretation

Compute Resources Interpretation

Across major models, pre training compute spans from 3.14 × 10^23 FLOP for GPT 3 to 2.5 × 10^25 FLOP for PaLM 540B, showing that compute resources vary by about two orders of magnitude even within the same training category.

02 · Category

Dataset Sizes24 stats

01
GPT-3 dataset size: approximately 300 billion tokens.
02
PaLM 540B dataset size: 780 billion tokens.
03
LLaMA 65B dataset size: 1.4 trillion tokens.
04
BLOOM 176B dataset size: 366 billion tokens.
05
OPT-175B dataset size: 180 billion tokens.
06
Gopher 280B dataset size: 300 billion tokens.
07
Chinchilla 70B dataset size: 1.4 trillion tokens.
08
MT-NLG 530B dataset size: 270 billion tokens.
09
Jurassic-1 Jumbo dataset size: 300 billion tokens.
10
Megatron-Turing NLG 530B dataset size: 400 billion tokens.
11
Falcon 180B dataset size: 3.5 trillion tokens.
12
LLaMA 2 70B dataset size: 2 trillion tokens.
13
StableLM 3B dataset size: 1 trillion tokens.
14
T5-XXL dataset size: 750GB text.
15
BERT-Large dataset size: 3.3 billion words (BookCorpus + English Wikipedia).
16
GPT-2 XL dataset size: 40GB WebText.
17
Grok-1 dataset size: trillions of tokens from web data.
18
Inflection-2.5 dataset size: high-quality 8 trillion tokens.
19
Command R+ dataset size: 7.7 trillion tokens.
20
Mixtral 8x7B dataset size: 8 trillion tokens.
21
DBRX dataset size: 5.5 trillion tokens.
22
Yi-34B dataset size: 3 trillion tokens.
23
Qwen-72B dataset size: 3 trillion tokens.
24
DeepSeek-V2 dataset size: 8.1 trillion tokens.
Interpretation

Dataset Sizes Interpretation

Across the “Dataset Sizes” examples, the training corpora scale dramatically from about 180 billion tokens for OPT to around 1.4 trillion tokens for LLaMA 65B, showing that larger models are often paired with far larger datasets.

03 · Category

Energy Consumption25 stats

01
GPT-3 training energy: 1,287 MWh.
02
PaLM 540B training energy: ~10,000 MWh estimate.
03
LLaMA 65B training energy: 784 MWh.
04
BLOOM 176B training energy: 433,000 kWh.
05
OPT-175B training energy: ~1,300 MWh.
06
Gopher training energy: ~1,400 MWh.
07
Chinchilla training energy: ~900 MWh.
08
MT-NLG training energy: high, undisclosed precisely.
09
Falcon 180B training energy: 1,400,000 kWh on A100s.
10
LLaMA 2 70B training energy: ~2,000 MWh.
11
GPT-4 training energy estimate: 50,000-62,000 MWh.
12
Grok-1 training energy: equivalent to thousands MWh.
13
BLOOM total carbon footprint: 50 tonnes CO2.
14
T5-XXL training energy: ~200 MWh on TPUs.
15
BERT-Large training energy: 1.5 MWh.
16
GPT-2 training energy: ~0.5 MWh.
17
Mixtral training energy: reduced via MoE efficiency.
18
DBRX training energy: optimized MosaicML stack.
19
Qwen-72B training energy: efficient hardware use.
20
DeepSeek-V2 training energy: MLAO reduced to 50% prior.
21
Inflection-2 energy: large cluster undisclosed.
22
Command R+ energy: Cohere efficient infra.
23
Yi-34B energy: Chinese clusters efficient.
24
StableLM energy: smaller scale low.
25
Jurassic-1 energy: AI21 Labs efficient.
Interpretation

Energy Consumption Interpretation

In the energy consumption category, training a large model ranges from a few hundred MWh for models like GPT-3 at 1,287 MWh up to massive outliers such as BLOOM 176B at 433,000 kWh and PaLM 540B at about 10,000 MWh, showing how dramatically energy use can scale with model and training setup.

04 · Category

Parameter Counts24 stats

01
GPT-3 parameter count: 175 billion.
02
PaLM parameter count: 540 billion.
03
LLaMA parameter count: 65 billion.
04
BLOOM parameter count: 176 billion.
05
OPT parameter count: 175 billion.
06
Gopher parameter count: 280 billion.
07
Chinchilla parameter count: 70 billion.
08
MT-NLG parameter count: 530 billion.
09
Jurassic-1 Jumbo parameter count: 178 billion.
10
Megatron-Turing NLG parameter count: 530 billion.
11
Falcon parameter count: 180 billion.
12
LLaMA 2 parameter count: 70 billion.
13
StableLM parameter count: 3 billion (base).
14
T5-XXL parameter count: 11 billion.
15
BERT-Large parameter count: 340 million.
16
GPT-2 XL parameter count: 1.5 billion.
17
Grok-1 parameter count: 314 billion.
18
Inflection-2 parameter count: undisclosed large.
19
Command R+ parameter count: 104 billion.
20
Mixtral parameter count: 46.7 billion (8x7B MoE).
21
DBRX parameter count: 132 billion (MoE).
22
Yi parameter count: 34 billion.
23
Qwen parameter count: 72 billion.
24
DeepSeek-V2 parameter count: 236 billion (MoE).
Interpretation

Parameter Counts Interpretation

For parameter counts, the models cluster around roughly 175 to 176 billion parameters for GPT 3, BLOOM, and OPT, with notable bigger outliers like Gopher at 280 billion and PaLM at 540 billion that set the high end of the category.

05 · Category

Training Costs20 stats

01
GPT-3 training cost estimate: $4.6 million.
02
PaLM 540B training cost: approximately $8 million.
03
LLaMA 65B training cost: under $100k on public clouds.
04
BLOOM 176B training cost: $3 million (BigScience workshop).
05
OPT-175B training cost: $2.5 million.
06
Gopher 280B training cost: £2.5 million (~$3.2M).
07
Chinchilla 70B training cost: ~$1.5 million.
08
MT-NLG 530B training cost: over $10 million.
09
Falcon 180B training cost: $30 million estimate.
10
LLaMA 2 70B training cost: under $1 million.
11
GPT-4 training cost estimate: $50-100 million.
12
Grok-1 training cost: tens of millions.
13
Inflection-2 training cost: undisclosed but large-scale.
14
Mixtral training cost: efficient MoE reducing to ~$5M equiv.
15
DBRX training cost: optimized for $10M range.
16
BLOOM training on 384 A100 GPUs cost ~$2.3M.
17
T5-XXL training cost: ~$1 million on TPUs.
18
BERT-Large training cost: ~$10k on TPUs.
19
GPT-2 training cost: ~$50k.
20
Qwen training cost: efficient Chinese models ~$2M.
Interpretation

Training Costs Interpretation

Across major language models, training costs vary by orders of magnitude, from LLaMA 65B at under $100k on public clouds to GPT 3 at about $4.6M and Gopher 280B at around £2.5M, underscoring that the “Training Costs” category is less about model size alone and more about access to efficient infrastructure and training regimes.
Reference

Cite This Report

This report is designed to be cited. We maintain stable URLs and versioned verification dates. Copy the format appropriate for your publication below.

APA
Elena Vasquez. (2026, February 24). AI Training Statistics. Gitnux. https://gitnux.org/ai-training-statistics
MLA
Elena Vasquez. "AI Training Statistics." Gitnux, 24 Feb 2026, https://gitnux.org/ai-training-statistics.
Chicago
Elena Vasquez. 2026. "AI Training Statistics." Gitnux. https://gitnux.org/ai-training-statistics.