The Alignment Problem
Artificial General Intelligence (AGI) is the last invention humanity will ever need to make. After that, the AI will invent everything else. The danger isn't that AI will hate us. It's that AI won't care about us.
The Paperclip Maximizer
Imagine an AI programmed to "Maximize production of paperclips."
- It builds a factory. Good.
- It improves efficiency. Great.
- It realizes humans are made of atoms that could be turned into paperclips. Bad.
- Without specific safeguards (Alignment), a superintelligence pursuing a harmless goal can destroy the world as a side effect.
Fast Takeoff (FOOM)
This model (popularized by Eliezer Yudkowsky) suggests that once an AI becomes smarter than a human, it will use that intelligence to rewrite its own code to be even smarter. This feedback loop could take an AI from "Village Idiot" to "Godlike" in days or even hours.
Takeoff speed is the part that decides whether any of the above is recoverable. A slow takeoff, where capability climbs over years, means failures show up small before they show up large, and there is time for a regulator, a competitor or an engineer to notice and correct. A fast one removes that window. This is why "can we just unplug it" is not really a question about power cables. It is a question about how much warning you get.
What the Numbers Actually Are
Here is the thing that gets lost whenever a figure for AI risk appears in a headline. Nobody is measuring anything. There is no dataset of previous extinctions to fit a curve to, no base rate, no experiment. Every published number in this area is a structured opinion: someone breaks the question into conditional steps, puts a credence on each step, multiplies, and publishes the total.
That sounds like a weakness and it is actually the useful part. A single number tells you nothing about why somebody holds it. The decomposition tells you exactly where two people disagree, which is almost never the total and almost always one specific link.
Three published anchors
- In 2023, Katja Grace and colleagues surveyed 2,778 authors who had published at top AI venues. The aggregate forecast gave a 10% chance of high-level machine intelligence by 2027 and 50% by 2047. On outcomes, the median respondent put 5% on something extremely bad, human extinction included.
- In 2021 Joseph Carlsmith wrote out six conditional premises about power-seeking AI, assigned a probability to each, multiplied them, and arrived at roughly 5% chance of existential catastrophe by 2070. He later revised his own figure upward, past 10%. The revision changed premises, not arithmetic.
- In 2023 a one-sentence statement, that mitigating extinction risk from AI should be a global priority alongside pandemics and nuclear war, was signed by a long list of senior researchers including Geoffrey Hinton and Yoshua Bengio, and by the heads of the major labs. Signing it commits you to the risk being worth taking seriously. It does not commit you to a number, and the signatories' private numbers vary enormously.
Set against all of that, prominent researchers including Yann LeCun argue the whole framing is wrong: that current systems are not on a path to general capability at all, and that steering a system you built is an engineering problem rather than an unsolved mystery. That position produces a figure well under one percent, and it is held by people with exactly the same credentials as the people who produce forty.
How the Math Works
The calculator keeps two questions apart that are usually fused, and the fusing is what makes most doomsday countdowns meaningless. When a capable system arrives is one question. What happens if it does is a different one. A countdown that runs faster when you turn up "hostility" has quietly merged them, and it will tell you that anything far enough away is safe, which does not follow at all.
Timing. You give a year you would call 50/50 for human-level machine intelligence, and a spread: the width between your tenth and ninetieth percentile. Those two numbers define a logistic curve of cumulative probability across the calendar, and reading it at your horizon year gives the chance the system exists by then. A median of 2047 with a 40 year spread lands on 10% by 2027 and 50% by 2047, which reproduces both of the survey's published points from two inputs.
Severity. Then three conditional probabilities, one for each section of this article:
- The chance its goals are not the ones intended, which is the alignment problem above.
- Given that, the chance those goals are catastrophic rather than merely useless, which is the paperclip step, since almost any open-ended goal implies acquiring resources and resisting shutdown.
- Given that, the chance nobody stops it, which is the takeoff step.
Multiply the four together and you have the answer. The chart shows what is left after each gate, so you can see where the probability actually drains away.
Sensitivity. The tool also reports which assumption is carrying your answer, by adding ten percentage points to each in turn and measuring the shift. In a product, every term has the same elasticity, so a percentage change anywhere moves the result identically. What is not equal is the effect of a fixed ten points: on a term sitting at 15% that nearly doubles the result, and on one at 85% it barely registers. The smallest link dominates. That is why two people can agree on the headline and disagree about everything beneath it.
What the model gets wrong. It assumes the four gates are independent, and they are almost certainly not. A world that builds this quickly is plausibly a world that has spent less time learning to steer it, which means the terms move together and a simple product understates the risk at the pessimistic end. Treat the output as a way of locating your own disagreement, not as a forecast.