How to read this score
What the numbers mean and what they cover.
AI Risk Index
The index runs from 0 to 100. It combines six factors covering what advanced AI systems can do, how much they can do on their own and how well they can be controlled. Higher scores mean more concern.
It is not a probability or a countdown. A score of 65 does not mean a 65% chance of catastrophe. The scores and weights reflect judgments about the evidence, not measured odds.
Review bands: 0-24 routine · 25-49 enhanced · 50-74 intensive · 75-100 urgent. These labels guide the level of attention. A serious incident may need a closer look at any score.
Judgment range
The range shows what happens when all six factors are set to the low or high end of their assessed ranges. The weights stay the same.
It is not a statistical confidence interval. It shows how much the chosen assumptions can change the score. It does not put odds on the true level of risk or cover every uncertainty.
Evidence confidence
Confidence reflects how strong, relevant and complete the evidence is. It is separate from the risk score.
Missing information can lower confidence without raising the score. Several reports about one event do not count as separate confirmations.
Assurance gap
This score runs from 0 to 5. It measures how hard it is to check claims about systems, safeguards and behavior. A higher score means there is more we cannot verify.
It is reported separately and does not change the AI Risk Index.
Read more →Scope and review status
This assessment draws on selected cases involving advanced AI systems and research tests. It does not rate a typical chatbot, one specific system or all AI risk worldwide. A capability at one company and a control failure at another do not show that one system has both.
Each assessment includes sources, evidence that could challenge it and details about the setup. An incident involving an older model does not establish how a newer one behaves.
This assessment was reviewed by the project owner. It has not been independently validated. Internal research does not change a public score until a new release is approved.
The evidence date tells you what period the assessment covers. The publication date tells you when it was released. A review can leave the score unchanged.
See release history →What the score covers
The score currently covers cyber capability and autonomy. The areas below need separate research. They are outside this score, not rated as zero risk.
Biological risks and defenses need their own assessment.
Military uses of AI and escalation risk need their own assessment.
Lasting loss of human control needs its own assessment. It is different from human extinction.