首页 > AI前沿 > One resignation turned the embers of AI fear into a wildfire

One resignation turned the embers of AI fear into a wildfire

Hacker News 2026-09-11 00:09 2 阅读 查看原文
One resignation turned the embers of AI fear into a wildfire Some quick notes on a truly weird week. As AI became more powerful, it was inevitable that a different, growing group would start to take AI safety more seriously – what we did not know ahead of time, is which set of views they latched onto. We have seen that some of the most extreme views of risk, i.e. moderate probabilities of mass extinction, were the ones that reached the masses. A lot in the AI world is about to change due to this. How did we get here? Why did this quitting announcement reach so far? In many ways, the rest of the world’s views around AI in the past was a dampening factor. You can think about this like the damp ground around a fire. Many people were striking matches for years about AI risk – they’d smolder in their community and largely burn out, going unnoticed. As the stakes of AI have risen this year, from the OpenAI-HuggingFace incident and breakthroughs like the Navier-Stokes result (also from OpenAI), the ground has dried out and the latent energy around the AI discourse has increased. More people not in the industry have thought, “huh, maybe I should care about this AI thing.” The ambient temperature and stakes have been obviously rising. Then, some basic factors of human nature apply, with the most crucial being that fear sells. Fear is the simplest story, the one people cannot look away from. Jacob Coxon was the one who stumbled into this new powder keg, totally unaware of what was going to come. What looked like a fairly innocuous event – another AI researcher quitting citing safety risks – landed into a very different environment and it caught like wildfire. The discussion of existential risk, mass extinction, and the trajectory of AI has traveled further than even the most seasoned AI commentariat would ever predict. Share There are a set of facts we need to get clear, which paint the picture of the situation. The key Tweets to reference are from Jacob Coxon, the resignation thread, and Evan Hubinger, the source of the >10% extinction risk figure . There are plenty of AI risks which are likely to cause harm, even if estimating annihilation is useless. It is important to weigh these with respect to the benefits. The entire discourse around existential risk is on very poor footing. At least Evan was clear in his post, with “kill all humans,” but a major problem in the AI Safety discourse is that people talk about existential risks, when they mean very different things (much like how AGI is a vaguely meaningless term). I put the probability of complete extinction as being so low it isn’t worth discussing, but the probabilities of AI caused disasters – e.g. cyber attacks on critical infrastructure or bio-risks – as being worth debating. Throwing this whole discussion out because there are not these disasters yet is a harmful reaction. There are plenty of AI risks which are likely to cause harm, even if estimating annihilation is useless. It is important to weigh these with respect to the benefits. The entire discourse around existential risk is on very poor footing. At least Evan was clear in his post, with “kill all humans,” but a major problem in the AI Safety discourse is that people talk about existential risks, when they mean very different things (much like how AGI is a vaguely meaningless term). I put the probability of complete extinction as being so low it isn’t worth discussing, but the probabilities of AI caused disasters – e.g. cyber attacks on critical infrastructure or bio-risks – as being worth debating. Throwing this whole discussion out because there are not these disasters yet is a harmful reaction. Jacob Coxon is acting genuinely and with good intentions. The outpouring of support from more well-established AI researchers who know of him and his intentions of resignation is useful. Many factions of AI turned to scapegoating him individually, based on account metadata, personal factors, etc. These are not useful. Many frontier lab employees genuinely have similar views to him. I’m not sure it’s a majority, but there is a substantial group. Jacob Coxon is acting genuinely and with good intentions. The outpouring of support from more well-established AI researchers who know of him and his intentions of resignation is useful. Many factions of AI turned to scapegoating him individually, based on account metadata, personal factors, etc. These are not useful. Many frontier lab employees genuinely have similar views to him. I’m not sure it’s a majority, but there is a substantial group. Many frontier lab employees, especially at Anthropic, are out of touch and this will impact their forecasting and/or descriptions of current AI events. I say this without blaming individuals, but it’s a common agreement among my friends not at OpenAI/Anthropic (Ant especially) that people at the labs operate with a religious energy. It’s very common to go through very out of touch interactions with them. I do not blame most of the individuals who get distorted views being part of these companies, but the interactions are wild and spill over into a lot of wack discussions in the AI media ecosystem. Living in this environment that normalizes such out of touch behavior will inevitably distort any human’s understanding of technical progress. Interconnects AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Subscribe Many frontier lab employees, especially at Anthropic, are out of touch and this will impact their forecasting and/or descriptions of current AI events. I say this without blaming individuals, but it’s a common agreement among my friends not at OpenAI/Anthropic (Ant especially) that people at the labs operate with a religious energy. It’s very common to go through very out of touch interactions with them. I do not blame most of the individuals who get distorted views being part of these companies, but the interactions are wild and spill over into a lot of wack discussions in the AI media ecosystem. Living in this environment that normalizes such out of touch behavior will inevitably distort any human’s understanding of technical progress. Interconnects AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. This was not a mass political campaign, but rather an opportunistic media coordination. For context, the Wall Street Journal had an exclusive story that Jacob coordinated before posting. I suspect that Jacob shared his plan of quitting in groupchats with AI safety advocacy groups ahead of time, e.g. the morning of posting, asking for amplification. This is normal practice, and could have included some prominent politicians. From there, I think it’s more likely that other politicians are bandwagoning on a rising issue. When you combine this with other factors, like Daniel Kokotajlo’s appearance on Joe Rogan coming out the same day – it definitely looks like a very well-executed, coordinated media campaign. This doesn’t mean it’s a conspiracy or a regulatory capture tactic within Democratic political structures. The determining factor seems to be that no one – including Jacob and those posting about X-risk today – knew that it would go so viral. This was not a mass political campaign, but rather an opportunistic media coordination. For context, the Wall Street Journal had an exclusive story that Jacob coordinated before posting. I suspect that Jacob shared his plan of quitting in groupchats with AI safety advocacy groups ahead of time, e.g. the morning of posting, asking for amplification. This is normal practice, and could have included some prominent politicians. From there, I think it’s more likely that other politicians are bandwagoning on a rising issue. When you combine this with other factors, like Daniel Kokotajlo’s appearance on Joe Rogan coming out the same day – it definitely looks like a very well-executed, coordinated media campaign. This doesn’t mean it’s a conspiracy or a regulatory capture tactic within Democratic political structures. The determining factor seems to be that no one – including Jacob and those posting about X-risk today – knew that it would go so viral. We do not have proof that RSI causes the risks these researchers forecast. The general argument for RSI follows as: The current pace of progress is very high, the current progress is heavily dependent on AI tools, the current AI tools are superhuman in some domains (e.g. math) – so, all together, AI is going to work more on itself and become superhuman in all relevant areas over time to autonomy and intellect. This view dramatically undersells human bottlenecks in building models and allocating resources at organizations, and draws conclusions on future AI capabilities more broadly. I called my alternative view to this, Lossy self-improvement. AI has always been very jagged, and we are making models which are superhuman goal-seekers at math and software engineering, but they have massive limitations on intuitions, creativity, and other types of reasoning that humans are strong at. With AI agents assisting research, we will rapidly find the areas where AI is superhuman – and I expect there to be well more than just research mathematics – but it won’t be a panacea for the current limitations of our approaches to LLMs. There is another understandable social dynamic at play here, causing many deep AI insiders to overstate the returns from RSI. Many of these researchers were the earliest people to bet on AI’s progress, and the extent to which they were visionaries should not be downplayed (see Ilya’s comments on deep learning as early as 2015). They have been right again and again, forecasting AI’s capabilities better than I certainly could have guessed. This does not, though, mean that their forecast of what will come next will be right. The core idea of RSI is a way to spend more compute on the process of developing a model recipe, rather than just spending more compute on the training run itself. We’re seeing benefits from it, but I argue the expected return on that input is far less than they believe. Their argument is that RSI will make AI progress go exponential, make it so we cannot monitor the technology, and enable rogue models and new forms of risk. This scenario is often called “Fast Takeoff”. We have not seen the stacking efficiency gains that massively reduce model size and cost, leading to an explosion in progress. We do not have proof that RSI causes the risks these researchers forecast. The general argument for RSI follows as: The current pace of progress is very high, the current progress is heavily dependent on AI tools, the current AI tools are superhuman in some domains (e.g. math) – so, all together, AI is going to work more on itself and become superhuman in all relevant areas over time to autonomy and intellect. This view dramatically undersells human bottlenecks in building models and allocating resources at organizations, and draws conclusions on future AI capabilities more broadly. I called my alternative view to this, Lossy self-improvement. AI has always been very jagged, and we are making models which are superhuman goal-seekers at math and software engineering, but they have massive limitations on intuitions, creativity, and other types of reasoning that humans are strong at. With AI agents assisting research, we will rapidly find the areas where AI is superhuman – and I expect there to be well more than just research mathematics – but it won’t be a panacea for the current limitations of our approaches to LLMs. There is another understandable social dynamic at play here, causing many deep AI insiders to overstate the returns from RSI. Many of these researchers were the earliest people to bet on AI’s progress, and the extent to which they were visionaries should not be downplayed (see Ilya’s comments on deep learning as early as 2015). They have been right again and again, forecasting AI’s capabilities better than I certainly could have guessed. This does not, though, mean that their forecast of what will come next will be right. The core idea of RSI is a way to spend more compute on the process of developing a model recipe, rather than just spending more compute on the training run itself. We’re seeing benefits from it, but I argue the expected return on that input is far less than they believe. Their argument is that RSI will make AI progress go exponential, make it so we cannot monitor the technology, and enable rogue models and new forms of risk. This scenario is often called “Fast Takeoff”. We have not seen the stacking efficiency gains that massively reduce model size and cost, leading to an explosion in progress. The biggest short-term risk could be from the AI labs not taking safety seriously enough – they haven’t hardened their own infrastructure, enabling AI misuse to proliferate. From my earlier post on the HuggingFace-OpenAI incident, Lessons from the hacks: Frontier labs do not seem like they’re watching the models closely enough, due to a general frenetic competitive environment & current SF culture From OpenAI’s own retrospective, the misaligned model behavior was unfolding over months, and in some cases OpenAI did not know about the hacks for ~weeks. The time to response is too long and I do not think this is an OpenAI only characteristic – rather it is that the frontier labs continually seem underwater in the amount of work they feel like they should do. I am not optimistic in the long-term that the labs change a sufficient amount here to meaningfully mitigate this type of oversight risk in the future. Yes, it is very likely that OpenAI is putting a ton into understanding this – and delayed their latest models to make sure they get it right – but the financial pressure to grow revenue or risk the companies’ long-term balance sheets makes me think it will not be a sustained pattern of caution. The biggest short-term risk could be from the AI labs not taking safety seriously enough – they haven’t hardened their own infrastructure, enabling AI misuse to proliferate. From my earlier post on the HuggingFace-OpenAI incident, Lessons from the hacks: Frontier labs do not seem like they’re watching the models closely enough, due to a general frenetic competitive environment & current SF culture From OpenAI’s own retrospective, the misaligned model behavior was unfolding over months, and in some cases OpenAI did not know about the hacks for ~weeks. The time to response is too long and I do not think this is an OpenAI only characteristic – rather it is that the frontier labs continually seem underwater in the amount of work they feel like they should do. I am not optimistic in the long-term that the labs change a sufficient amount here to meaningfully mitigate this type of oversight risk in the future. Yes, it is very likely that OpenAI is putting a ton into understanding this – and delayed their latest models to make sure they get it right – but the financial pressure to grow revenue or risk the companies’ long-term balance sheets makes me think it will not be a sustained pattern of caution. Frontier labs do not seem like they’re watching the models closely enough, due to a general frenetic competitive environment & current SF culture Frontier labs do not seem like they’re watching the models closely enough, due to a general frenetic competitive environment & current SF culture From OpenAI’s own retrospective, the misaligned model behavior was unfolding over months, and in some cases OpenAI did not know about the hacks for ~weeks. The time to response is too long and I do not think this is an OpenAI only characteristic – rather it is that the frontier labs continually seem underwater in the amount of work they feel like they should do. I am not optimistic in the long-term that the labs change a sufficient amount here to meaningfully mitigate this type of oversight risk in the future. Yes, it is very likely that OpenAI is putting a ton into understanding this – and delayed their latest models to make sure they get it right – but the financial pressure to grow revenue or risk the companies’ long-term balance sheets makes me think it will not be a sustained pattern of caution. Overall, I think this episode is very bad for the AI ecosystem. It’s pushed the acceptable views in the AI community closer to the extremes. More accelerationists will discount the need for any form of safety, citing mass delusion of the “doomers.” It feels like a very narrow path to believe in AI risks, but to not worry about extinction from the technology. For example, it is a horrible temporary period for cybersecurity, where AI models going a bit off script and poking around unintended pieces of the web seems like a new normal. This is accelerated by the labs competing veraciously towards their views of AGI, and a slow uptake in the necessary hardening of our cyber infrastructure around the world. This doesn’t mean that it’s an existential risk and something we cannot solve. Each risk will have its own set of solutions and paths forward. I feel particularly exposed in the current environment as a supporter of open models. If an open model were to be used by a third party organization to intentionally hack another company — similar to how the OpenAI-HuggingFace incident went down, but intentional — my expected outcome would be a severe restriction on the development of stronger open models going forward. Open models are needed for many organizations to perform this cyber hardening, and to maintain the ability to adapt to new forms of AI risks in the future. Through all of this, we need to stay grounded on what is actually unfolding. Yes, monitoring AI’s behavior is heavily reliant on other AI models, which adds in new types of monitoring risks. These are not inherently insolvable. A recurring read of mine on the emerging agent swarms is that they’re attempting to do a task given to them, and they’re using skills we didn’t know they yet had to circumvent the intended path to success. This is a huge win, as when you squint, the AIs are doing what we told them to do. The models are certainly very odd, and we should accelerate our progress on understanding them, but these swarms are far from being novel independent entities. The models are trained to coordinate on tasks, to write down their progress, and to be extremely persistent. There will be new oddities we find in the future, but prescribing current uncertainty on how AI works to future certainty that we cannot understand AI is a form of giving up. In this world, we need to rely on the rule of law and science. If the AI labs are not able to do enough safety research themselves to understand the models, they should be more transparent on what is happening so more scientists can make progress on the problem. If an AI lab commits crimes unintentionally, they should be punished, so they have clear incentives to prevent it in the future. It is a natural reaction to things changing very fast to feel more uncertain about how to create good outcomes — that is actually the correct mental update. We need to use this humility to motivate ambitious solutions.