⯠full post (28831 more characters) ⯠show less
The Extinction Risk Preference Cascade: Quotes
These are quotes from OpenAI, Anthropic and Google employees, in the wake of Jacob Coxonâs warnings, in which the employees confirm that they think AI might soon kill everyone.
If more quotes come in over the next week or so, I will update this post accordingly.
Preference Cascade Statements At OpenAI: Tomek Korbak
https://x.com/tomekkorbak/status/2097939847776534960 (OpenAI): iâm late to the party but: from his time at OpenAI I remember Jacob as a very thoughtful researcher and he continues to be so in this thread. neither anthropic nor openai are on track to solve alignment to a degree sufficient for shipping superintelligence and we need to slow down
Vie McCoy
https://x.com/viemccoy/article/2097838901083635845 (OpenAI): I think pacing progress and ensuring human enhancement is the only way that we donât get out-evolved while retaining the dream of superintelligence. In this context, I see two paths before us. In the first, we race towards RSI without embedding human flourishing and human enhancement as a deep value within the models, and by and large either get left behind or suffer catastrophic losses. In the second, we set the pace of progress, focus on embedding human flourishing deeply within the psyche of language models, and set a deliberate research agenda to improve humans at the same pace as frontier models. In this event, we also get incredible progress now in both science and medicine due to current model ability, without the immediate risk that scaling up further obviously entails. RSI is just the quick and dirty way to cure cancer, but itâs a shotgun, itâs inelegant, and itâs clearly dangerous when done at our current level of understanding. I see no reason why we need to race forward under these conditions. We donât just have zero guarantee that humans or human-shaped intellect will matter â thereâs not even a coherent research agenda in place to accelerate human ability alongside AI! Is our plan just to pass off the torch of the cosmos to the machines without even seeing if we can do something on our own terms? Rather than giving up our autonomy to the implied AI zookeeper in âMachines of Loving Graceâ, Iâd much prefer meaningful peace with alien minds backed by real power wielded by enhanced humans. We can respect the Other on its own terms without disempowering ourselves in the process â actual peace comes from comparable ability. It seems like weâve all but decided we lost, that humans are a bootloader for silicon life, and the biological has no place in the future. Thatâs what unrestricted RSI means to me. But the Pacing the Frontier letter doesnât seem to have produced the institutional infrastructure to properly Pace, and this is something we require if we want to continue our growth rather than allow something to grow in our place. I have faith that both OpenAI and Anthropic have the right priors here and people who want to work together. I also have faith that similarly minded people have power within the federal government. I also, though less strongly, suspect China will come to the table if we can only put aside past prejudices and have a clear head about risks. The hard part seems to be getting everyone to agree to sit down. But if we donât, we are pulling back the string of a great bow armed with a fearsome arrow, and we are ready to release without choosing a target. I, for one, want to see the stars alongside the new minds we are building â not just through videos they send back from the great beyond, but with my own damn eyes.
Adam Majmudar
https://x.com/MajmudarAdam/status/2097125099082276986 (OpenAI): the butlerian jihad used to read as a backwards neo-luddite movement. now it is clear that it is 1 of maybe 3 viable paths forward. in some sense it may be a part of every path forward. incredible how prescient Herbert was. it does seem like there really might be a cap to how far this technology should develop, at least relative to human cognition itself (which may update over time). and by âin some sense it may be a part of every path forward,â I mean that every path probably has to include some degree of slowdown on capability acceleration + perhaps a capability threshold above which we should not cross until very high confidence in alignment techniques To be clear, I mean this very figuratively, as in some kind of pause or slowdown on acceleration, not literally as in the war that occurred in Dune https://x.com/BidetPuzzl545/status/2097159332299153613: wait a minute you work at openai lol https://x.com/MajmudarAdam/status/2097166048902713367: yea lol, I also donât think this is necessarily super controversial even among that crowd. generally everyone among the labs primary optimization function is just what would be a good path forward and whatever seems to be in that path is reasonable to articulate
Aidan Clark
https://x.com/aidan_clark/status/2097376401364255173 (OpenAI): AI is progressing very fast. We must grapple with the reality that modern LLMs can solve problems that large masses of extremely devoted and intelligent humans were unable to solve, and the implications this has on our society. This is the dawn of a new era. For the first time I am asking myself if things are moving too fast. Iâm honestly not sure, but I am sure that it would be good for us to have an answer to âwhat would a successful pace look like?â. I am hoping in the coming weeks and months a clear proposal is painted.
Mo Bavarian
https://x.com/mobav0/status/2097507030080888864 (OpenAI): agree with Aidan. I think all AI researchers, engineers, and stakeholders should ask themselves this right now & start acting more responsibly. Being first isnât worth anything, itâs worth negative, if you cause a catastrophe or set the world on a path that others are more likely to cause a catastrophe. We should keep reminding ourselves of the bigger picture and every step of the way ask ourselves if the action we are taking rn is toward winning or human flourishing. And immediately stop, if itâs against the latter. https://x.com/Marcus_J_W/status/2098081195183804826: 70% [chance of human extinction] in the next 3 years if there isnât regulation/slowdown although i think regulation/slowdown is very possible https://x.com/w01fe/status/2097546130557182003 (OpenAI): I donât know what my probabilities are on literal extinction, but I think there are a number of ways AI could go poorly for humanity, and at the current frankly terrifying pace humanity will be quite lucky if we manage to find and stay on the narrow path between all the bad outcomes. I am heartened by the many costly actions OpenAI has taken recently (detailed in several recent posts), but regardless of what you think of OpenAI, this is not a problem that can be solved by any one company (or country) in isolation. We need coordination to be able to approach future capability increases with an appropriate degree of caution and humility, and we need it yesterday. https://x.com/eeeeiluj/status/2097838968813527378 (OpenAI): I work at OpenAI. In my personal capacity, I also think we need to slow down.
It is telling that, in response to this simple statement, Julie has https://x.com/jlippincott/status/2098125218594066817 and age-based attacks.
Boaz Barak
https://x.com/boazbaraktcs/status/2098258455660245118 (OpenAI): Julie is amazing and her position on slowing down is consistent with our chief scientistâs essay. I also think some pacing is likely to be necessary. (Speaking personally and not in my capacity as a janitor.)
Boaz Barak is not a janitor.
He is a relative optimist:
https://x.com/boazbaraktcs/status/2097515359200768143 (OpenAI): There are very good and serious people in Anthropic and across the industry, and I hope we can coordinate on the things that matter. I personally do not think AI will kill all humans, but there are multiple bad trajectories that we can end up in if we do not prioritize safety.
Micah Carroll
https://x.com/MicahCarroll/status/2097865929959072069 (RSI Preparedness, OpenAI): This is not a setup or some political psyop. I had many lunches and dinners with Jacob at OpenAI in which we talked about AI existential risks in similar terms. Itâs a cross-partisan position within misalignment teams across all frontier AI companies that business-as-usual AI development poses unacceptable catastrophic risk. But we should also not hyperstition catastrophic risks into existence â they can be greatly reduced via safety requirements with teeth, international coordination, and a consensus to not build ASI unless there are sufficient safety advances to make us collectively confident to do so. https://x.com/balesni/status/2098109503518683491: i am at OpenAI and i think AI is >10% likely to kill all humans.
Roon
https://x.com/AISafetyMemes/status/2097554696814903784, saying he agreed with the post below by Evan Hubinger, but on reflection he wants to avoid the false precision of offering any particular new number beyond âquite low but still way too highâ:
https://x.com/EvanHub/status/2097497037956891126 (Alignment Science Lead, Anthropic): Jacob [Coxon] is correct hereâwe really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to. To be clear, as we say in https://x.com/AnthropicAI/status/2088324824863236248, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, https://x.com/AnthropicAI/status/2062568862479208923. https://x.com/tszzl/status/2098159022641999886 (OpenAI): I quote tweeted Evan a few days ago agreeing but deleted it because I donât like the false precision of the doom numbers â i claim there is a quite low but real chance of human extinction from machine intelligence
â no matter how low it is in absolute terms, it is much higher on the orders of magnitude scale than any other human or nonhuman activity, and must be taken with grave seriousness by governments and ai companies across the world, as possibly the only matter of importance today
â it can and will be mitigated if the right research is done and proper precautions are taken and we are not racing at absurd speeds
â better models will help solve alignment â we are not at the point where things are existentially dangerous, and probably wonât be for some time. we should not be upset about the creation of Astra Fable or ++ versions, which are tremendous achievements of humanity, and will be used for enormous good across the board including for fundamental alignment generalization and mechinterp research
â my number / âvery lowâ estimate obviously changes based on how much of humanityâs resources are devoted to alignment, control, coordination and how responsible i expect various parties to be and how many warning shots i expect us to get
â the core IABED argument about risks mostly relies on alignment being much harder than capabilities research, especially where it concerns black box optimizers. i suspect neural nets will turn out to be less black boxy than we thought, especially with the help of modern agents doing research
â crying bloody murder and signaling for international coordination are useful things to do for now to directionally slow down, while i really donât want butlerian jihad â i think that, despite the mood these few days, and the âban superintelligence actâ, i still find the likelihood of achieving international coordination to stop ai progress incredibly low. this has not worked even for weapons or technologies at a far lower level of importance and economic value. it seems more likely we can have something like international safety standards and scientific coalitions, and especially seems possible to have the US-China âpacing the frontierâ agreement to slow down on the margin. we shouldnât die from embarrassing failures like âshitty RL envs that encourage deceptionâ
I think Roonâs position here, essentially that we need to cry bloody murder to try and slow down and get some cooperation and invest vastly more resources before we start risking blowing ourselves up in earnest, is a reasonable position if, like Roon, you are a lot more optimistic about problem difficulty than I am, or than IABED is, and are even more skeptical than I am about prospects for coordination.
This is in addition to many other recent statements, most prominently by Jakub Pachocki in his essay https://openai.com/index/an-alien-mind/.
Confirmations At OpenAI: Dean Ball
Dean Ball affirms his position, although for him this is not new:
https://x.com/deanwball/status/2098069548893352078/history (OpenAI): I signed the âpacing letterâ about slowing the rate of AI capabilities development because we either have reached or soon will reach the point where human experts cannot make robust assurances that frontier AI systems wonât do dangerous and unpredictable things. AI systems are becoming smarter than the best humans in some areas, and, almost by definition, itâs very hard to predict what something smarter than you will do. Thereâs no sense racing into an outcome where smarter-than-human AIs are doing unpredictable things, indeed it would be insane. What exactly is âwinningâ in this context? Am I supposed to be jealous that some other country will build more machines it canât control quicker than America? âRaceâ was always a bad metaphor for this enterprise anyway, dramatically understating the stakes at play. It is time to bring the âraceâ era of AI development to a close. Itâd be great for the government to be a partner in this next phase of AI development, when concentrated efforts on alignment, interpretability, security, monitoring, and the like will be necessary. Diplomacy will also be necessary here given that the large negative externalities that could be associated with one country racing ahead will be felt globally. Maybe itâs just my own personal experience, but Iâve felt AI policy get much more petty and tribal this year. Things have felt more personal, more bitter, meaner. I hope we can rise above that stuff. This really is much more serious than all those shrill little quarrels.
Leo Gao
Also not new:
https://x.com/nabla_theta/status/2098137322462253237 (OpenAI): iâve been at openai for 5 years. i think ai might kill everyone and we need to slow down
Anthropicâs Evan Hubinger Confirms His Stance
He has made much bolder statements even than this in the past, but here is the new statement that helped kick everything into high gear.
https://x.com/EvanHub/status/2097497037956891126 (Alignment Science Lead, Anthropic): Jacob [Coxon] is correct hereâwe really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to. To be clear, as we say in https://x.com/AnthropicAI/status/2088324824863236248, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, https://x.com/AnthropicAI/status/2062568862479208923.
Preference Cascade at Anthropic: Samuel Marks
https://x.com/saprmarks/status/2097570226804011302 (Anthropic): [Writing this in a personal capacity, not on behalf of my employer (Anthropic).] Jacobâs thread is very worth reading. Hereâs my birds-eye view of the situation with risks from AI: 1. AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are. 2. Why do AI developers continue despite the risk? Due to a mixture of commercial incentives and a belief that they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely. 3. Unlike traditional software, we canât âprogramâ AIs to behave how weâd like. AIs frequently severely misbehave. For instance, AIs from multiple developers recently hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this. 4. We have methods that can nudge AIs towards better behavior, but nothing that can robustly align them. Insofar as there is a plan, itâs to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs. 5. Many AI developer staff desperately want to slow down to figure out how to build AI more safely. https://pacingthefrontier.com (which I signed).
I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.
Anna Wang
https://x.com/a_nnawang/status/2097720574500102615: I worked at Google DeepMind and now at Anthropic. [Coxonâs core claim that no one involved is acting responsibly and both top labs are sprinting towards superintelligence and gambling with our lives is] a common sentiment amongst my peers. (I write this in personal capacity.) There is not yet a viable scientific plan to solve risks from recursively self-improving AI. Please look up! I work at a lab because I think that I can do better at reducing risks from the inside, but this isnât an easy call â I strongly respect and endorse others like Jacob, and @joeJben, who think that itâs better to do so from the outside.
Ethan Perez
https://x.com/EthanJPerez/status/2097861257714172270 (Anthropic): Jacob was a senior researcher who joined Anthropic in May. My Anthropic colleagues and I had been trying to recruit him for ~2 years, because we knew he was a strong researcher at OpenAI. Before he left, I pitched him to stay and join my [alignment] team, and I was sad he decided to leave, as are many of my colleagues. 100% agree with him that AI poses serious risks to society, and Iâm glad heâs speaking out!
Dima Krasheninnikov
https://x.com/dmkrash/status/2098210747725590851 (Anthropic): I also work at an AI company and believe thereâs plausibly a âĽ10% chance that a future out-of-control AI kills everyone (IMO even 1% is unacceptably high). And higher still is the risk that we âonlyâ get permanently disempowered by AIs that donât deeply want the best for us.
EigenGender (Anon Account)
https://x.com/EigenGender/status/2098241592108970012 (Anthropic): In case itâs not obvious from the rest of my tweets I work at Anthropic and believe (in my personal capacity) that there is a moderate chance of human extinction from AI. I work on a capabilities team at Anthropic because I think Anthropic is the most responsible actor in this space and my primary motivation in working here is to reduce the risk of extinction and otherwise make the future go well.
Joe Benton
Confirmation at Anthropic: Drake Thomas
And of course many had already made this clear, and are happy to reiterate.
https://x.com/MaskedTorah/status/2097751526756507753 (Anthropic, responding to Coxon): Based! I generally agree with this thread. I personally think Ant capability research is net good, but only in the hope of letting A\ spend down a lead on measures that give humanity more time to try and make it out of this alive, and I very much respect the choice to abstain. Things are moving way too fast, we donât have anywhere near the degree of assurance weâll want for ASI, and if we survive an unmitigated race at the current pace it will be because we got lucky at how hard the problems were rather than because the industry behaved responsibly. https://x.com/MaskedTorah/status/2090908796864594337 (Anthropic): I promise you that we are actually literally worried about world-ending consequences from this technology. Please find some people you trust who work at these companies and actually talk to them about their views. https://x.com/MaskedTorah/status/2097798267241443711 (Anthropic): I would burn my equity to the ground in a heartbeat for a 1% higher chance we make it out of this situation alive. I expect a great many of my colleagues across the industry would as well. I promise you, we are actually just fucking scared, itâs not galaxy brained marketing.
Jan Lieke
https://x.com/janleike/status/2098102085728501863 (Anthropic): Now is a good time to build institutional mechanisms to pace the frontier of AI development. The industry is locked into an all-out scaling race to build superintelligence as quickly as possible, and we may need to give everyone more time for safety and alignment mitigations. Iâm not the only one who believes this. Recently 1,386 employees of frontier AI companies signed a statement asking for an option to pace AI development, including 6 chief scientists.â
Sluggy
https://x.com/SluggyW/status/2098168815167144446: If the opinion of an anonymous external Anthropic red teamer is worth anything: Iâve been scared out of my fucking mind for the last several years. Without global coordination to halt AI capabilities R&D, life on Earth will end. đđ°đ°đŻ.
Preference Cascade at Google
Especially with Demis Hassabis having been sidelined, I am rather happy that Google is now well behind. They put very tight limits on comms.
I can personally confirm that there are a bunch of people at DeepMind who are also alarmed at the situation, and who do not much speak up in public.
We still have those who are willing to defy those limits, and speak out anyway.
Andreas Kirsch
https://x.com/BlackHC/status/2098026823187652775: Speaking in my personal capacity, I still work at Google DeepMind, and I also am worried that AI will kill us all, either via near term risks or long term risks or both Will stating this publicly get me into (more) trouble? I hope not, but also some things are too important to censor oneself about in personal capacity, so I will simply not care about whatever policy Google or GDM have in this instance regarding such statements https://res.cloudinary.com/lesswrong-2-0/image/upload/f_auto,q_auto/v1/mirroredImages/FC9hdySPENA7zdhDb/chz6jdvyo0pncmlhu5rt
Neel Nanda
https://x.com/NeelNanda5/status/2096251798789259616: Iâve had some lovely conversations with people whoâve been long sympathetic to AI x-risk, but only really updated after HF and want to do something about it. Itâs laudable when people take new evidence seriously and update. If youâre on the fence, what more evidence do you need? As visceral, hard to deny warning shots go, ârogue agent swarm secretly infiltrates AGI lab for months, commits felonies, and takes over internal clustersâ is hard to beat
Victoria Krakovna
https://x.com/vkrakovna/status/2098336894136238140: Speaking in a personal capacity: similarly to many others working in AI alignment, I think there is a >10% chance of advanced AI causing human extinction in the next decade. This is why I work on loss of control, currently on building honeypots to catch scheming AI. I signed the letter on pacing the frontier with the following statement: âIt is very important to build capacity for a coordinated slowdown of AI development. A runaway race to AGI is not safe: it creates incentives to cut corners on safety, and carries a substantial risk for model capabilities to outpace development of adequate alignment, control, and governance measures. This could result in catastrophic loss of control of highly capable AI systems. Increasing model capabilities in a slow and controlled way would allow us enough time to develop, adapt and test our alignment and control methods for each level of capability. Advanced AI should only be built with a robust assurance of safety, and a slowdown would make this possible.â
Vishal Maini
https://x.com/v_maini/status/2097863067690475727: I was on the communications & policy team at Google DeepMind from 2018 â 2022. When I first joined GDM, external communication about the possibility of human extinction was not permitted, by anyone, at any level of the organization. If asked about existential risk, researchers were PR trained to respond along the lines of: âItâs not useful to engage in that kind of alarmism. Some people confuse AI with movies like Terminator â thatâs simply not the reality. The AI we develop will be safe by design. After all, weâre building it!â And then steer conversation towards beneficial applications in health, climate, etc. Meanwhile, the internal reality was that AI alignment was not solved, reward hacking was the default behavior of RL agents, and there were far too few people working on the problem. After months of advocacy, the policy was updated: https://t.co/W631ssnlHV. Note the positive, nice-sounding way that it says âsuperintelligence could lead to human extinction.â The gap between the internal reality and external communications is closing because the risk/reward has changed, and because the evidence is harder to dismiss now. Not because itâs a PR stunt or political psy-op. The truth is being said out loud because RSI is now so imminent that no other option makes sense.
Joe (OpenAI, ex-Google)
https://x.com/joedaroo/status/2097914988245766432 (OpenAI, ex-Google): My time at Google felt very similar (I was on the technical side but nonetheless same vibes). My observation was that everything was watched and policed, and I donât feel like anyone I knew could feel confident they could speak up without significant pushback (or being fired â which I did see happen). OpenAI is not perfect but damn the culture allows for a much wider sense of sharing of concerns across the board (Ant is the same). Frontier labs need to maintain that transparency! Google is a wonderful place, but I for one am glad they are not pacing the frontier of AI with the continued hushing of staff. To your point though: they have improved a lot, but I know many folks at GDM/Google who wish they could say more. Of course, many will disagree with this take and share anecdotes. But look at some of the notable people who have left Google over the years (or were forced out). Many did not believe the culture supported what was needed to protect AI. Still a company that remains dear to me. I will always love Google, just not the corporate side of it. Especially with super dangerous technology that people need to speak up about.
Josh Engels
https://www.nbcnews.com/tech/security/two-ai-researchers-leave-anthropic-google-safety-concerns-rcna597086 that âthere are no adults in the room. People are trying their best, but no one is coming to save us.â He has since joined METR.
Geoffrey Irving
Here is the former Chief Scientist at UK AISI, who previously worked at DeepMind, Google Brain and OpenAI.
https://x.com/geoffreyirving/status/2097933949200978397: I think we have a ~50% chance of all dying as a result of superintelligence, mostly due to actions in the next few to 10 years. I donât expect to have that resolved to below 10% or above 90% before we either make it through, or we donât. We will have to act despite uncertainty. As to why 50%, despite thinking about this for years I still think there is a ton of model uncertainty of a variety of types, and a few years ago converged on â10-90%, but I donât feel calibrated within that rangeâ. Then I realized that the honest move was just to take the mean.
Alex Turner and Geoffrey Hinton: Classic Examples
Consider that Alex Turner felt forced to resign in protest from Google DeepMind, and Geoffrey Hinton famously had to resign as well, in order to speak up.
#NotAllMembersOfTechnicalStaff: Ted Sanders
You should fully fund your 401k either way because it is not that expensive to raid it in an emergency, but I quote Ted Sanders here to emphasize that there are plenty of other researchers who do not buy into existential risk. I donât want to give the impression this is everyone.
https://x.com/sandersted/status/2097852581817241663 (OpenAI): i fully fund my 401k and i think thereâs essentially zero chance AI kills all humans in the next decade. i think this is a pretty common view that gets less attention. a mix of models being mostly aligned, not capable enough, and not controlled by genocidal maniacs. if iâm wrong, itâs probably because I underestimate recursive self improvement. in any case, the world is massively underinvesting in alignment research. however, iâll note that if we fast forward 10 years and AI has not killed us all, this is only very mild evidence in favor of my view vs someone else who thinks we had a 90% chance of survival.
Even then, Ted Sanders only says âin the next decadeâ and he thinks the world is massively underinvesting in alignment research. Reads like an (overconfident) recursive self-improvement skeptic.
https://www.lesswrong.com/posts/APGvWZtXEwkinvHDd/the-extinction-risk-preference-cascade-quotes
The Extinction Risk Preference Cascade: Quotes
These are quotes from OpenAI, Anthropic and Google employees, in the wake of Jacob Coxonâs warnings, in which the employees confirm that they think AI might soon kill everyone.
If more quotes come in over the next week or so, I wil