Doom as a bad method not a utopia trade-off
Advanced AI is generally expected to have some very high variance outcomes—it might herald everything good, it might destroy humanity. For instance, here are 800 random AI researchers’ expectations about how good the future is, lined up:
https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/pDyLMRoi2BDq34rFe/erypdqcjn9n6nstm4mwz
From https://arxiv.org/pdf/2401.02843v1
As you can see, most AI researchers put a serious chunk of probability on very different overall outcomes: maybe doom, maybe utopia. This is common. Most people I know who think there is a serious chance of the destruction of humanity from AI also believe that if humanity isn’t destroyed, things might be insanely good.
I often hear people talk as if this means we are in a trade-off where the question is whether the good outweighs the bad. For instance, they look at the people above who think there’s a 10% chance of extinction and a 30% chance of utopia and round this off to ‘net positive on AI’.
That seems like a kind of wild error. Like considering yourself optimistic regarding driving at 200mph to your new job if you think there’s only a 10% chance you’ll die in a fiery crash on the way there, and a 30% chance this job will radically improve your life.
The things you should be comparing are driving at 200mph and driving at a normal speed! The things you should be comparing are attempting to attain advanced AI by the current route, and by other routes!
We can debate whether all the other routes are bad or impossible somehow, for instance if constraining projects that risk loss of human control risks sending humanity into an irrecoverable ruin. But I don’t think having ruled out such things is why people are usually thinking in trade-off terms.
Rather I think this error comes from a few things:
- It being simpler to think of ‘pros vs. cons’ and the topic being too abstract for people to intuitively notice that they are comparing pros of a long term outcome vs. cons of the https://worldspiritsockpuppet.substack.com/p/the-first-future-and-the-best-future?utm_source=publication-search there we have noticed
- Sloppiness about talking about P(doom). Saying ‘P(doom)’ encourages thinking as if ‘AI’ implies a particular chance of ‘doom’. We should more accurately think about ‘p(doom|’such and such route’), e.g. P(doom|advanced AI from scaling up LLMs). People usually mean ‘P(doom|current trajectory)’ with some ambiguity about whether the current trajectory includes our own actions.
- https://worldspiritsockpuppet.substack.com/p/ai-is-an-abundance-of-choice-not vs. a bunch of different things we could build
If you are bullish on some kind of advanced AI utopia, you should generally be lesskeen to try to achieve it via a careless route that leaves you at high risk of dying and losing it on the way there.
https://www.lesswrong.com/posts/pDyLMRoi2BDq34rFe/doom-as-a-bad-method-not-a-utopia-trade-off#comments
https://www.lesswrong.com/posts/pDyLMRoi2BDq34rFe/doom-as-a-bad-method-not-a-utopia-trade-off