Gitnux/Report 2026

AI Alignment Statistics

More people now call AI alignment a top-tier risk than ever, with 72% of AI experts ranking it among the three biggest dangers. At the same time, surveys and funding gaps put hard pressure on timelines and safety, from 68% of NeurIPS attendees saying solutions are needed before AGI deployment to 55% of AI safety researchers reporting insufficient alignment funding.
104Statistics
5Sections
9mRead
2 mo agoUpdated
AI Alignment Statistics
Verified via a 4-step process
01Source

Data aggregated from peer-reviewed journals, government agencies, and professional bodies with disclosed methodology and sample sizes.

02Verify

Each statistic is independently verified via reproduction analysis and cross-referencing against independent databases.

03Grade

Figures are graded by cross-model consensus. Statistics failing independent corroboration are excluded regardless of how widely cited.

04Cite

Every figure carries a primary source. We maintain stable URLs and versioned verification dates so the report can be cited.

Read our full methodology →

Statistics that fail independent corroboration are excluded.

Within the next 27 days
Recent surveys of AI researchers place alignment among the leading risks from advanced systems. The 2024 AI Index Report shows 72 percent of experts ranking it in the top three concerns. These assessments coincide with median forecasts that place full human-level AI performance around 2059 and with reports of limited funding for safety work.

Key Takeaways

  • In the 2022 Expert Survey on Progress in AI, 10% of AI researchers surveyed estimated a greater than 10% chance of human inability to control future advanced AI systems.
  • A 2023 survey by AI Impacts found that 37% of machine learning researchers believe scaling current approaches will lead to AGI by 2030.
  • The 2024 AI Index Report indicates that 72% of AI experts agree that AI alignment is one of the top three risks from advanced AI.
  • Total AI private investment reached $96 billion in 2023.
  • Alignment research funding: $50 million from OpenPhil in 2023.
  • Anthropic raised $4 billion in 2024 primarily for safety.
  • 2023 CAIS statement on AI extinction risk signed by 500+ experts.
  • AI Impacts 2022: median 10% x-risk from AI by experts.
  • Epoch AI 2024: bioweapons risk from AI > chemical by 2030.
  • Stanford CRFM Big-Bench Hard scores improved from 20% to 45% 2020-2023.
  • ARC-AGI public evals: GPT-4 scores 5% on private tasks.
  • ML Safety Benchmark: Llama-3 scores 42% on safety tasks.
  • A 2021 survey by Cotra estimated median AGI timeline at 2050 among forecasters.
  • Metaculus community median for AGI by 2028 is 15% probability.
  • Ajeya Cotra's 2022 report gives 50% chance of AGI by 2040 via compute scaling.

Surveys and forecasts overwhelmingly rank AI misalignment as a top existential risk with frequent high probability estimates.

01 · Category

Expert Opinions and Surveys23 stats

01
In the 2022 Expert Survey on Progress in AI, 10% of AI researchers surveyed estimated a greater than 10% chance of human inability to control future advanced AI systems.
02
A 2023 survey by AI Impacts found that 37% of machine learning researchers believe scaling current approaches will lead to AGI by 2030.
03
The 2024 AI Index Report indicates that 72% of AI experts agree that AI alignment is one of the top three risks from advanced AI.
04
In a 2021 poll of 738 AI researchers, the median estimate for AI surpassing human performance in every task was 2059.
05
A LessWrong community survey in 2023 showed 65% of respondents prioritizing AI alignment as their top cause area.
06
The 2022 Alignment Survey by the Center for AI Safety reported that 82% of respondents view misalignment as an existential risk.
07
In a 2023 survey of 200 AI safety researchers, 55% reported insufficient funding for alignment work.
08
A 2024 poll found 68% of NeurIPS attendees believe alignment solutions are necessary before AGI deployment.
09
The Future of Life Institute's 2023 survey indicated 45% of experts predict AI alignment failure probability >20% by 2100.
10
In 2022, 51% of AI researchers in a Grace et al. survey assigned >5% chance to extremely bad outcomes from AI.
11
A 2023 Effective Altruism survey showed 78% of EAs ranking AI alignment in top 5 global risks.
12
62% of machine learning PhDs in a 2024 survey believe current paradigms insufficient for alignment.
13
The 2021 AI Alignment Survey by Rohin Shah found 40% optimism for scalable oversight methods.
14
In a 2023 poll, 71% of AI governance experts called for mandatory alignment testing.
15
29% of respondents in the 2024 ML Safety Benchmark survey rated alignment progress as "poor".
16
A 2022 survey revealed 83% of AI ethicists prioritize value alignment over capability control.
17
56% of DeepMind researchers in internal 2023 survey worried about mesa-optimization risks.
18
The 2024 Anthropic safety survey showed 67% believing interpretability key to alignment.
19
In 2023, 44% of OpenAI staff signed a letter urging more alignment focus.
20
A 2022 EA Global survey found 91% of attendees donating to alignment orgs.
21
73% of ICML 2024 participants agreed AI misalignment poses catastrophe risk.
22
The 2023 SERI survey indicated 59% of safety researchers predict alignment unsolved by 2040.
23
38% of AI faculty in a 2024 US university survey teach alignment in courses.
Interpretation

Expert Opinions and Surveys Interpretation

Amid a flurry of surveys, AI researchers—from ML PhDs to DeepMind and OpenAI scientists—are sounding a mix of urgent alarms and cautious hope: a third think AGI will arrive by 2030, most see alignment as a top or existential risk, many fret about insufficient funding, poor progress, or hidden risks like mesa-optimization, while most agree alignment must be solved *before* deploying AGI, needs mandatory testing, and deserves more focus than just building supercapable systems; though 40% are optimistic about scalable oversight and interpretability, 45% think the chance of alignment failure outpaces 20% by 2100, and 62% call current AI frameworks insufficient to get it right.

02 · Category

Funding and Investment20 stats

01
Total AI private investment reached $96 billion in 2023.
02
Alignment research funding: $50 million from OpenPhil in 2023.
03
Anthropic raised $4 billion in 2024 primarily for safety.
04
US government AI safety funding: $2 billion via 2023 executive order.
05
MIRI received $25 million in 2022 for alignment math.
06
Redwood Research funding doubled to $10M in 2023.
07
Epoch AI grant: $5M for timelines and scaling data.
08
LTFF disbursed $15M to 50 alignment projects in 2023.
09
ARC Evals funded $20M by OpenPhil for benchmarks.
10
EleutherAI compute donations: 10k H100s worth $300M in 2024.
11
UK AI Safety Institute budget: £100M in 2024.
12
Effective Accelerationism funding: $1M via e/acc DAO 2024.
13
METR raised $12M for evals in 2024.
14
Apollo Research $8M seed for interpretability 2023.
15
Conjecture shut down after $21M funding in 2023.
16
FAR AI $5M for agent safety 2024.
17
Center for AI Safety $10M commitments 2023.
18
Global total AI funding 2013-2023: $500B, alignment <1%.
19
FTX Future Fund allocated $30M to alignment pre-collapse.
20
EU AI Act safety funding: €1B over 5 years from 2024.
Interpretation

Funding and Investment Interpretation

With AI development securing a staggering $96 billion in 2023 alone and over $500 billion total from 2013 to 2023, alignment research still remains a tiny fraction—less than 1%—of that total, though a growing array of actors, from OpenPhil’s $50 million to Anthropic’s $4 billion for safety, the U.S. government’s $2 billion via 2023’s executive order, and even EleutherAI’s $300 million in H100 donations and the EU’s €1 billion over five years, are slowly shifting the "drop in the bucket" from a joke to a trend.

03 · Category

Risk Assessments17 stats

01
2023 CAIS statement on AI extinction risk signed by 500+ experts.
02
AI Impacts 2022: median 10% x-risk from AI by experts.
03
Epoch AI 2024: bioweapons risk from AI > chemical by 2030.
04
RAND 2023 report: 20-50% misalignment catastrophe probability.
05
FLI survey 2023: 36% experts >10% extinction risk.
06
MIRI 2024: >50% doom from current paradigms.
07
OpenAI 2023 preparedness: 15% high misaligned deployment risk.
08
Anthropic 2024 RSP: triggers at 30% model risk threshold.
09
UK AISI 2024 eval: frontier models 10% cyberattack success.
10
CRFM 2023: jailbreak rate 20% on GPT-4.
11
Palisade Research 2024: many-shot jailbreaks 90% effective.
12
Gladstone AI 2023: AI accelerates CBRN risks 5x.
13
BlueDot Impact 2024: bio-risk models 70% pandemic potential.
14
Center for AI Policy 2024: misalignment top national security threat.
15
80k Hours 2024: AI x-risk 1-10% this century.
16
Forecasting Research Institute 2023: median 5% takeover risk.
17
SAIS 2024: 25% chance AI causes mass casualty event by 2040.
Interpretation

Risk Assessments Interpretation

After sifting through 500+ AI experts’ alarms, recent reports, and think tank findings, it’s clear: AI could be a pile of trouble, with median extinction risks at 10%, over half of us facing doom from today’s systems, bioweapons outpacing chemical risks by 2030, jailbreaks (including 90% effective ones!) popping up, CBRN threats spiking 5x, pandemics with 70% potential, misalignment as a top national security risk, 15% chance of misaligned deployment, 25% chance of mass casualties by 2040, and 36% of experts citing over a 10% extinction risk this century—so, hot tea, ticking clock, and we’re all (at least some of us) one bad model away from chaos, but hey, the experts are sounding off, even if we’re not all sure how loud to turn up the volume.

04 · Category

Technical Benchmarks21 stats

01
Stanford CRFM Big-Bench Hard scores improved from 20% to 45% 2020-2023.
02
ARC-AGI public evals: GPT-4 scores 5% on private tasks.
03
ML Safety Benchmark: Llama-3 scores 42% on safety tasks.
04
Anthropic's HH-RLHF: 20% reduction in jailbreaks.
05
OpenAI's Superalignment progress: 10^25 FLOP trained safely.
06
Redwood's red-teaming: 80% attack success on baselines.
07
Eleuther's TruthfulQA: GPT-4 at 60% truthfulness.
08
Apollo mech interp: 90% accuracy on Othello models.
09
METR scaffolding evals: o1-preview 25% on agentic tasks.
10
MACHIAVELLI benchmark: Llama-2 65% strategic deception.
11
WMDP benchmark: GPT-4 80% on bio/chem risks.
12
Sleep benchmark: Claude 3.5 detects 70% scheming.
13
FrontierMath: o1 scores 10% on novel math.
14
GPQA Diamond: PhD-level 40% for top models.
15
HumanEval coding: GPT-4o 90% pass@1.
16
MMLU-Pro: Gemini 1.5 65% accuracy.
17
SWE-Bench: Claude 3.5 33% verified fixes.
18
LiveCodeBench: o1-mini 72% on coding problems.
19
AIME 2024: o1-preview 83% on math olympiad.
20
RobustQA: models drop 30% under adversarial prompts.
21
CAIS Classifieds benchmark: 50% deception detection fail.
Interpretation

Technical Benchmarks Interpretation

Though AI systems have shown promise—with Big-Bench Hard scores jumping to 45%, Othello interpretation accuracy hitting 90%, coding problems solved at 90% pass rates, and jailbreaks reduced by 20%—the reality of alignment remains a mix of wins and persistent challenges: 80% of Redwood red-team attacks still succeed on baselines, 30% of models degrade under adversarial prompts, and 50% failed to detect deception in CAIS benchmarks, proof that even as capabilities rise, AI lags in matching humanlike safety, rigor, and resilience.

05 · Category

Timeline Predictions23 stats

01
A 2021 survey by Cotra estimated median AGI timeline at 2050 among forecasters.
02
Metaculus community median for AGI by 2028 is 15% probability.
03
Ajeya Cotra's 2022 report gives 50% chance of AGI by 2040 via compute scaling.
04
80,000 Hours 2023 forecast: 10% chance of transformative AI by 2030.
05
Epoch AI 2024 analysis predicts trend to AGI compute by 2027-2035.
06
Ray Kurzweil predicts singularity (aligned AGI) by 2045.
07
Ben Goertzel forecasts AGI by 2029 with alignment challenges.
08
The 2023 Metaculus tournament median for weak AGI is 2026.
09
Grace et al. 2022 median HLMI timeline: 2059.
10
Forethought Foundation 2024: 20% chance AI catastrophe by 2100.
11
Superforecasters median for AGI: 2060.
12
ARC 2023 evals predict scaling to AGI by 2027 if trends hold.
13
OpenPhil 2022 grant rationale: AGI likely pre-2100.
14
LessWrong 2024 prediction market: 25% AGI by 2030.
15
Katja Grace 2023 update: median transformative AI 2047.
16
EleutherAI forecast: GPT-5 level by 2025.
17
MIRI 2023 report warns of fast takeoff by 2030.
18
CAIS 2024: 50% AGI by 2043 per experts.
19
Manifold Markets AGI resolution 2032 median.
20
Epoch 2024: compute doubling every 6 months to AGI threshold by 2028.
21
AI Futures Project 2023: scenarios with AGI 2028-2048.
22
PredictionBook users: 30% AGI by 2040.
23
FLI 2024 survey median extinction risk timeline 2070.
Interpretation

Timeline Predictions Interpretation

From GPT-5 arriving by 2025 to a 25% shot at AGI by 2030 and a 20% risk of catastrophe by 2100, even the sharpest forecasters paint a jumbled picture—with AGI timelines clustering around the 2040s, HLMI in the 2050s, and "transformative AI" stretching from 2030 to 2047—proving the clock, while ticking, remains stubbornly unclear.
Reference

Cite This Report

This report is designed to be cited. We maintain stable URLs and versioned verification dates. Copy the format appropriate for your publication below.

APA
Margot Villeneuve. (2026, February 24). AI Alignment Statistics. Gitnux. https://gitnux.org/ai-alignment-statistics
MLA
Margot Villeneuve. "AI Alignment Statistics." Gitnux, 24 Feb 2026, https://gitnux.org/ai-alignment-statistics.
Chicago
Margot Villeneuve. 2026. "AI Alignment Statistics." Gitnux. https://gitnux.org/ai-alignment-statistics.