How the score is calculated
Version 1.1.0-research. Six factors contribute to the score. Each has a set weight. The assurance gap is reported separately.
The weights add up to 95 and are scaled to 100% in the calculation. Each factor is scored from 0 to 5 using the scoring definitions. Half points fall between two definitions. The combined index runs from 0 to 100.
What the range means
The range uses the low or high estimate for every factor while keeping the weights fixed. It tests the assumptions. It is not a 90% or 95% confidence interval, or a proven limit on how low or high the actual risk could be.
What the model cannot tell us
Adding the scores assumes that the gaps between score levels are meaningful and that a lower score on one factor can offset a higher score on another. Those assumptions have not been statistically validated. The factors can also overlap. This index does not calculate causes or expected losses.
The evidence comes from different systems and tests. A capability in one case and a control weakness in another do not show that one system has both.
Review bands
0-24 routine · 25-49 enhanced · 50-74 intensive · 75-100 urgent. These labels guide review. They are not validated safety thresholds or permission to deploy a system. A serious unresolved incident still needs attention, whatever the average score.
How much do the weights matter?
| Weighting method | Index points |
|---|---|
| Current weights | 65.3 |
| Equal weights | 63.3 |
Similar results do not prove that the model is right. The weights and gaps between score levels are still judgment calls.
Try different assumptions →Define what is being assessed
Set the outcome, population, model, setup, observation period and forecast date before considering a score change. The index draws on selected cases. It does not rate one deployment.
Record what happened
Keep observations separate from interpretations and scores. Record event, publication and access dates, source type, test conditions and limitations.
Explain both sides
Each factor needs a reason it is not scored lower or higher. Credible evidence can move a score in either direction. Lowering a score does not require proof of complete safety.
Check whether reports describe the same event
Reports about one incident share an event ID. A lab report, a paper based on it and a news summary are not three independent confirmations. More sources do not automatically mean more confidence.
Be clear about missing information
Missing information can widen the assessed range or lower confidence. It does not establish safety or automatically mean greater danger.
Keep the factors separate
Capability, behavior, exposure and controls measure different things. A new capability does not automatically worsen behavior. Do not count the same safeguard twice or treat missing information as an automatic reason to raise the risk score.
Owner review
The project owner reviews the sources, scoring, scope, conflicts and limitations before approving a public assessment. This is one person's review. It is not independent validation.
Check forecasts against outcomes
Define shorter-term forecast questions and evidence rules before the results are known. Compare their accuracy with simple alternatives. Good short-term results do not validate an extinction probability.
State the limits of the coverage
The current research does not support rankings of countries or companies, a global extinction estimate or numerical scores for biological or military risks. Evidence from more countries and languages needs further review.
Review serious incidents directly
A confirmed, serious unresolved incident needs review regardless of the index. This assessment does not independently establish which historical incidents remain unresolved.