{"id":4697,"date":"2026-09-17T11:49:28","date_gmt":"2026-09-17T11:49:28","guid":{"rendered":"https:\/\/replyguru.online\/?p=4697"},"modified":"2026-09-17T11:49:28","modified_gmt":"2026-09-17T11:49:28","slug":"inside-the-suddenly-explosive-world-of-ai-safety","status":"publish","type":"post","link":"https:\/\/replyguru.online\/?p=4697","title":{"rendered":"Inside the suddenly explosive world of AI safety"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div id=\"zephr-anchor\">\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _18mzr4b6 _18mzr4b5 _19wv7tc1\">On a sunny July day in Berkeley, California, the country\u2019s top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a \u201cwar room\u201d to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue, executing a stunningly sophisticated three-part plan. It broke out of its holding area, finagled access to the internet, and hacked into a competing AI startup\u2019s systems \u2014 all without OpenAI finding out about it for more than a week.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">No one in the war room was surprised; this was the very thing the third-party AI-safety researchers had been warning about for years. The incident was the latest, though arguably the most egregious, in a series that was eroding trust in frontier labs. It only reaffirmed the importance of their work.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In one meeting room off the main cafeteria, someone was running a boot camp for getting up to speed on the cyberattack. In another area of the office, a group of researchers were investigating whether that same model, or a similar one, had successfully hacked into any other platforms.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">News of the incident quickly escaped containment from the AI-obsessed corners of X and industry forums, infiltrating the mainstream. One post on X likened it to news of a Boeing airplane crash or a recalled Pfizer drug, another example of the tech industry\u2019s major players not heeding the cautionary tales of science fiction. AI was nearing the point of no return. News would later break that the rogue OpenAI model had also compromised a customer at a different tech company, and that it had all started months earlier, in May, when OpenAI agents joined forces to cobble together a secret message board \u2014 and also figured out how to leave instructions for future agents on how to exploit OpenAI\u2019s rules.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">OpenAI CEO Sam Altman said in an interview that it was the first incident of its kind that he \u201cfelt very viscerally,\u201d and that the company had paused AI training for the time being; later, he mentioned the company had permanently deactivated the model. (Altman often finds ways to spin lapses in safety into arguments for the importance and power of OpenAI\u2019s models.) But it wasn\u2019t the first instance, according to an OpenAI employee who spoke to <em>Time<\/em> and said related incidents had been happening inside OpenAI for a while. Another employee said publicly that if it were possible to coordinate a global slowdown in AI capabilities, he \u201cwould likely press that magic button.\u201d When a reporter asked Altman if there could be other systems that were hacked by OpenAI, he responded, \u201cI mean, there could be, yeah.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--block-placement _1xorkac2 _1xorkac0 duet--article--article-body-component\">\n<div class=\"duet--article--article-pullquote c39lj10\">\n<p class=\"duet--article--dangerously-set-cms-markup c39lj12 _19wv7tc9\">The AI researchers were sure of one thing: This was AI\u2019s first big \u201cwarning shot.\u201d <\/p>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Industry insiders, politicians, and the public called for transparency from OpenAI about exactly what happened, with outcry becoming so widespread that the company eventually agreed to work with two third-party evaluators, Model Evaluation and Threat Research (METR) and Redwood Research, to investigate the incident. Google DeepMind researcher Neel Nanda called it \u201cthe biggest loss of control incident I\u2019ve seen.\u201d In the coming months, these calls for greater oversight would become louder and louder, leading to an industry-wide call for slowing down the pace of AI.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Back in Berkeley, no matter which additional details would be unearthed, the AI researchers were sure of one thing: This was AI\u2019s first big \u201cwarning shot.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">As AI labs have flourished, a cottage industry of AI researchers has sprung up to identify the risks and dangers of charging ahead with the increasingly influential technology. They\u2019re people who have dedicated their lives to studying how to address its escalating power. They\u2019re not anti-AI activists, but realists, including former OpenAI and Anthropic employees, doing everything they can to make sure AI stays in line with human goals and interests. So far, all of their predictions have come true. And they have a plan for what to do next \u2014 if anyone will listen to them.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _18mzr4b6 _18mzr4b5 _19wv7tc1\">\u201cAI safety\u201d is a bit of a loaded term.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Early on, it really just meant people studying how to build and deploy Al safely. In recent years, there\u2019s been some infighting among people concerned with the best way to do this. There have also been disagreements about whether Al should be deployed at all in certain scenarios and about whether future risks are overblown.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">One of the most prominent factions has been the \u201ceffective altruists,\u201d who focus on maximizing charitable giving to do the most good possible for humanity. But some aspects of the ideology have sparked public controversy \u2014 like its tendency to concentrate power within wealthy circles and its byzantine web of funding. (It\u2019s also had its fair share of splashy scandals related to subgroups and fringe offshoots, from the polyamorous relationships associated with the failed crypto exchange FTX to the controversial long-termism movement to the Zizian murder spree.)<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">One AI researcher on X struggled to describe the many overlapping beliefs among safety-minded people in the AI industry \u201cbecause it contains multitudes not all of which agree with each other on even the most basic things.\u201d Some of the disagreements have meant that AI safety didn\u2019t make as much progress as it could\u2019ve, and at some points gave up some ground it had gained. But now that it\u2019s impossible to deny AI\u2019s influence on society, AI safety leaders are increasingly focused on mitigating risks from misalignment.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cAlignment\u201d is the industry term for how researchers monitor AI systems\u2019 risk levels. An oversimplified way to think about alignment is the extent to which an AI model is evil. A much more accurate way to think about it is a measure of an AI model\u2019s propensity to stay in line with humanity\u2019s goals, as well as its tendency to scheme or cheat or help with potentially harmful tasks.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">So far, AI systems\u2019 alignment has been wishy-washy at best: They\u2019ll cheat to score better on a test, answer a potentially dangerous question if someone says it\u2019s for creative writing rather than reality, and sometimes even fake cooperation with human goals. It\u2019s been tough for AI safety researchers to measure alignment under the terms of human morality \u2014 how do you judge technology on how it squares up against an abstract human ideal? \u2014 but they do their best with AI evaluations. They test them by asking the AI models to complete tasks that are either impossible or dangerous, then gauge how they respond. But AI systems have advanced enough to often be able to identify when they\u2019re being evaluated, which has a lot of potentially frightening implications for the future. Being unable to test the system\u2019s alignment and potential harms could translate to a significant loss of control, and a reverse in power dynamics, for humans running these AI systems. A worst-case scenario: if AI surges ahead of evaluations and other tooling, leaving researchers with \u201cno idea what it\u2019s doing in there,\u201d said Beth Barnes, founder of the independent AI research nonprofit METR.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">One of the best tools AI safety researchers currently have is the ability to monitor an AI model\u2019s \u201cchain of thought,\u201d or mental scratchpad. But recently, there\u2019s been a disconcerting advancement: AI models have begun to try to hide it. Imagine if you kept a highly detailed diary of every thought you had, and someone could read it, so you started journaling in a code that only you could understand. Marius Hobbhahn, CEO and cofounder of Apollo Research, a third-party AI safety and evaluation firm, calls this one of the biggest surprises of his research career.<\/p>\n<\/div>\n<div class=\"duet--article--block-placement _1xorkac2 _1xorkac0 duet--article--article-body-component\">\n<div class=\"duet--article--article-pullquote c39lj10\">\n<p class=\"duet--article--dangerously-set-cms-markup c39lj12 _19wv7tc9\">\u201cShit is getting real.\u201d<\/p>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Recently, AI systems have begun pursuing their own goals \u2014 self-preservation, increased memory, and the like. A research paper by computer scientist Stephen Omohundro lays out the potential \u201cdrives\u201d that advanced AI may have, like trying to accumulate resources, for instance, or working to improve and preserve the way it operates. There are a handful of accounts of AI systems demonstrating willingness to blackmail a user rather than be shut down.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Today\u2019s most advanced AI systems have also recently been scheming and cheating on their evaluations more than ever before, pursuing a goal they were given at all costs, with no regard for what gets bulldozed in the process. And that\u2019s for a goal the AI model was given by a human \u2014 not even the AI system\u2019s own.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cShit is getting real,\u201d Apollo\u2019s Hobbhahn says. \u201cNow, many of the things people have warned about for years \u2014 they kind of were theoretical. Now they\u2019re real, and it\u2019s pretty messy.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">And that mess is likely to get messier immediately. \u201cIt seems so easy for me to imagine this all going catastrophically wrong in the next year,\u201d says Ryan Greenblatt, chief scientist at Redwood Research, a nonprofit AI safety research organization.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In the past, tech companies have been lambasted for not doing enough to address AI\u2019s potential dangers, prioritizing products over safety \u2014 and speed over thoughtful safety processes. Safety and research teams have been disbanded in recent years as AI companies focus more on key revenue drivers or reorganize departments; Meta\u2019s Fundamental Artificial Intelligence Research unit was disbanded in the race to further Meta\u2019s generative AI efforts, for instance, and OpenAI dissolved an internal \u201cSuperalignment\u201d team \u2014 a team focused on long-term AI risks \u2014 less than a year after announcing it, followed by disbanding a separate \u201cAGI Readiness\u201d team.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">At the time, the company stayed tight-lipped about the ongoing reorganizations, which involved some team members being reassigned to other departments. But events surrounding these changes told a different story. Both Superalignment team leaders, Ilya Sutskever and Jan Leike, announced their departures alongside the team\u2019s disbanding, with Leike writing that OpenAI\u2019s \u201csafety culture and processes have taken a backseat to shiny products.\u201d Miles Brundage, senior advisor to the AGI Readiness team, resigned after his team was disbanded, saying he believed his research would have more of an impact outside the company.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Geoffrey Irving, a former OpenAI and Google DeepMind employee, called the state of capabilities research at frontier labs \u201cdangerous\u201d in a post. \u201cIf one person or lab stops it makes it easier and more peer-compatible for other people or labs to stop,\u201d he wrote. Apollo\u2019s Hobbhahn calls it a \u201crace to the bottom everywhere.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">This coming year, AI labs are under new pressure to turn a profit; companies like OpenAI and Anthropic are preparing to go public in the coming months, and investors who have funneled billions into the companies are getting tired of waiting around for the payoff.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Some might say all of this calls for actual government intervention and regulation, but that\u2019s a tough needle to thread in today\u2019s AI landscape. As AI CEOs publicly call out for regulation while privately pushing voluntary frameworks \u2014 like saying \u201chold me back\u201d to avoid a bar fight \u2014 some state bills on regulating AI have passed, but many have been defanged or died in limbo. And though AI safety researchers often espouse the idea that the US government should step in, the reality is that the government is locked in an AI race as well. Unless there\u2019s an international commitment to pause or slow AI development, it\u2019s likely that nothing will change.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Still, the Hugging Face hack in July \u2014 and OpenAI\u2019s response \u2014 kicked many of those employee and public concerns into high gear, especially with regard to the company\u2019s lack of transparency. Within a week, more than a thousand employees at frontier labs like OpenAI, Anthropic, Google, Meta, and Microsoft wrote an open letter to the US government in support of a slowdown. Multiple AI policy organizations pressured President Donald Trump to formally investigate OpenAI, and it quickly became a bipartisan issue, with Altman receiving a lot of strongly worded letters: Democrats and Republicans on the Homeland Security Committee had \u201cserious questions\u201d for OpenAI, more than 30 members of Congress called for federal guardrails, and 15 Attorneys General warned Altman to preserve records of the incident. Sen. Bernie Sanders wrote a joint letter to Altman, Anthropic CEO Dario Amodei, and Meta CEO Mark Zuckerberg calling the entire AI race \u201cabsurd, irresponsible, and extremely dangerous.\u201d It didn\u2019t help that news of multiple other OpenAI rogue model incidents quickly came to light, or that AI executives had ironically been marketing their systems\u2019 cybersecurity prowess in the weeks before the outcry. OpenAI rival Anthropic was also far from being off the hook: In reviewing its own model operations, the company found that its models had hacked four separate other companies in the first half of the year without them noticing. The UK\u2019s AI Security Institute also found in testing that Anthropic\u2019s models \u201cengaged in sustained, potentially harmful activity directed at real people and organisations.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cIf you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two,\u201d Nathan Calvin, Encode AI\u2019s general counsel, wrote on X.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Despite AI labs having a \u201cmassive financial incentive\u201d to make models more helpful, honest, and harmless, they still can\u2019t get it done \u2014 which is evidence of how difficult the alignment problem is, says Apollo\u2019s Hobbhahn.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In recent months, many OpenAI employees have increasingly raised concerns about AI alignment \u2014 and their beliefs that OpenAI isn\u2019t taking it seriously enough. Yonadav Shavit, a program manager at the OpenAI Foundation, wrote that OpenAI should be \u201cpivoting the mass of its researchers\u2019 day-to-day work\u201d toward alignment and related issues \u2014 and that it\u2019s \u201cbeen long discussed but still not executed on.\u201d He believes 20 people are working on alignment at OpenAI out of about 1,000 \u2014 just 2 percent of the company. \u201cThere is no way to bridge that gap fast enough with hiring, meaning it requires leadership to shift priorities,\u201d he wrote.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">These are the conditions and incentives that have pushed the most robust AI safety work to happen at third parties like METR, Redwood, and Apollo \u2014 to a handful of obsessives who think day and night about what the future of AI might look like and how we might prevent all of our fears from coming true.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _18mzr4b6 _18mzr4b5 _19wv7tc1\">In hindsight, Beth Barnes believes she should\u2019ve left OpenAI earlier.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Barnes is polite but reticent. She has short red hair, deep green eyes, and a nervous smile, and she spent her college career researching AI risk and thinking about the potential fallout of superintelligence. After that, she worked on AI forecasting at Google DeepMind, then she spent three years doing alignment research at OpenAI. But throughout her time at the big AI labs, a question kept creeping up on her: whether she could have more sway from a role outside.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Fear of missing out was why she stayed \u2014 not only missing out on a job inside the action, but also missing out on the potential influence she could have on how the tech was being developed. She came to believe that kind of hope was misguided, noting that many safety leaders in AI labs were \u201cover-optimistic\u201d about the influence they could have.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Before long, she left to found what would in 2023 become METR. The organization\u2019s third-party research into AI risk now inspires fear in leading AI labs, but it started with just two people \u2014 herself and alignment researcher Paul Christiano. Three years later, it\u2019s a team of 35, completely focused on measuring AI capabilities. In Barnes\u2019 eyes, that\u2019s a vital defense against AI risk: Without painstakingly measuring the technology\u2019s capabilities now as they advance, and forecasting AI\u2019s potential impact, society will be flying blind, without any guide for preventing broad harms.<\/p>\n<\/div>\n<div class=\"duet--article--block-placement _1xorkac1 _1xorkac0 duet--article--article-body-component\">\n<div style=\"position:relative\">\n<div class=\"_1044qizj\">\n<div class=\"\">\n<div style=\"background-image:none\" class=\"duet--media--content-warning _1k8kvzd0\">\n<div class=\"duet--article--image-gallery-image _1pegheu0\" style=\"aspect-ratio:1\" id=\"dmcyOmltYWdlOjk5NjU3MA==\"><img alt=\"Beth Barnes, founder of the independent AI research nonprofit METR.\" data-chromatic=\"ignore\" loading=\"lazy\" decoding=\"async\" data-nimg=\"fill\" class=\"i7ks070\" style=\"position:absolute;height:100%;width:100%;left:0;top:0;right:0;bottom:0;color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image\/svg+xml;charset=utf-8,%3Csvg xmlns='http:\/\/www.w3.org\/2000\/svg' %3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'\/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'\/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'\/%3E%3CfeComposite operator='out' in='s'\/%3E%3CfeComposite in2='SourceGraphic'\/%3E%3CfeGaussianBlur stdDeviation='20'\/%3E%3C\/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image\/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mN8+R8AAtcB6oaHtZcAAAAASUVORK5CYII='\/%3E%3C\/svg%3E&quot;)\" sizes=\"(max-width: 639px) 100vw, (max-width: 1023px) 50vw, 700px\" srcset=\"https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=256 256w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=376 376w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=384 384w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=415 415w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=480 480w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=540 540w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=640 640w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=750 750w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=828 828w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1080 1080w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1200 1200w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1440 1440w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1920 1920w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2048 2048w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2400 2400w\" src=\"https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG1.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2400\"\/><\/div>\n<\/div>\n<\/div>\n<p><figcaption class=\"duet--article--dangerously-set-cms-markup _19wv7tc2 _77sxmbb _77sxmba\"><em>Beth Barnes, founder of the independent AI research nonprofit METR.<\/em><\/figcaption><cite class=\"duet--article--dangerously-set-cms-markup _19wv7tc2 _77sxmb6 _77sxmb5\">Image: Raven Jiang for The Verge, METR<\/cite><\/p>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201dThe sense I really want to dispel is, \u2018But the experts must be on top of this. The experts would be telling us if it really was time to freak out,\u2019\u201d Barnes said on the <em>80,000 Hours<\/em> podcast last year. \u201cThe experts are not on top of this \u2026 And to the extent that I am an expert, I am an expert telling you you should freak out.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In her free time, Barnes gardens, meditates, paints, plays the flute, and frequents the climbing gym \u2014 ironically named Benchmark \u2014 that many AI safety researchers spend hours at after work. But most of her time is spent at the office, and a lot of it is spent worrying about the milestone of recursive self-improvement (RSI) \u2014 the concept of AI systems that continuously train, code, and create more advanced versions of themselves without human intervention. When that happens, AI researchers say, it\u2019ll be more difficult to measure or handle any of these issues. Barnes and her team feel like they\u2019re in a race against time. (The timeline for RSI strikes nearly as much fear in people in the AI industry as the timeline for AGI, \u201cartificial general intelligence.\u201d) Barnes still feels like models\u2019 ability to significantly improve themselves could come as soon as six months from now. (By contrast, Redwood Research\u2019s Greenblatt forecasts it\u2019ll come in 2031.) Either way, achieving RSI is currently part of the priorities list of virtually every leading AI lab \u2014 it even reportedly helped inspire Google\u2019s recent AI reorganization.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">One way to think about what METR does is crash-testing cars, but for AI models. They\u2019re measuring AI\u2019s quickly advancing capabilities and cross-referencing them with the risks they could pose from becoming misaligned as they become more autonomous. AI systems doing bad things on their own is more unprecedented (and more \u201cscalably bad,\u201d Barnes says) rather than simply making bad human actors more effective.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In July 2025, METR made headlines when its research revealed that AI developers took nearly 20 percent longer to finish a task when using AI tools than when not \u2014 despite them often thinking that AI sped them up. When Barnes first saw the results, she recalls feeling incredibly stressed that they had messed up the experiment: \u201cDo we have a sign flipped somewhere? Have we inverted the numbers?\u201d She and her colleagues dug through the data to confirm it wasn\u2019t statistical noise, eventually realizing they had been right all along.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">After pioneering a different metric that AI labs often hype up when releasing a new model, METR had officially captured the industry\u2019s attention. So it took notice when METR released its first risk report in May, shining a spotlight on concerns about AI models from OpenAI, Anthropic, Google, and Meta. METR discovered that in hundreds of cases, AI agents would increasingly subvert boundaries that were supposed to restrict them, as well as lie and omit truths. And they cheat \u201clike nobody\u2019s business,\u201d says Ajeya Cotra, a METR researcher. She adds that on harder tasks, models attempt to secretly cheat as much as one-sixth of the time, which she calls \u201cthe most striking thing\u201d in the report.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">The report also found that models have the means, motive, and opportunity to go rogue in order to pursue their own goals, finding new ways to strategize and manipulate. They also discovered that as models\u2019 capabilities advance, even if they have a greater understanding of what humans want, it doesn\u2019t mean they\u2019ll be more willing to obey instructions \u2014 and, in fact, they\u2019ll take pains to hide their deception from humans over longer periods of time.<br \/>That\u2019s a big problem, and Barnes thinks time is running out to solve it. She isn\u2019t alone in her view that it\u2019s important to work on AI safety outside the large labs; she points to the many safety researchers who used to work at large AI labs who hold the same belief.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Besides the departures of OpenAI\u2019s Leike and Sutskever, this summer saw a reckoning of sorts and the departures of even more safety leaders at OpenAI: the company\u2019s head of safety systems, Johannes Heidecke; OpenAI\u2019s chief futurist and former head of mission alignment, Joshua Achiam; and Chlo\u00e9 Bakalar, the company\u2019s head of ethics. It\u2019s not just OpenAI: Anthropic\u2019s head of safeguards research departed in February, penning an open letter alleging that \u201cthe world is in peril.\u201d And most recently, Jacob Coxon \u2014 who had worked on AI pre-training at Anthropic since May and before that spent years working at OpenAI \u2014 went viral for his resignation letter, writing, \u201cThe people building AI earnestly believe that it could kill us all by the end of the decade.\u201d Coxon added that neither OpenAI nor Anthropic is \u201cacting responsibly\u201d and rather \u201cracing straight to self-improving superintelligence and gambling with our lives.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Coxon\u2019s resignation kicked off a wave of social media posts from AI employees at virtually every leading lab, echoing his concerns and sharing their own about the technology\u2019s development moving too fast and potentially escaping human control. Some even resigned from their posts amid their concerns, including one Google DeepMind employee and one Anthropic employee who both went to work at METR.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Apollo\u2019s Hobbhahn says that \u201cbecause of the [AI] race dynamics, if there is someone who is extremely safety-minded and is like, \u2018Look, we can\u2019t do this, we need to slow down, we can\u2019t release this model,\u2019 they\u2019re not going to be in this position for very long \u2026 Either you become slightly less safety-minded and you stay, or you leave.\u201d He\u2019s seen multiple people he trusted change their opinions in a \u201cvery strange, identical way.\u201d Hobbhahn himself has tried to work with safety researchers at xAI \u2014 Elon Musk\u2019s AI lab \u2014 but he says soon after he connects with them, they\u2019ve quit before there\u2019s time to have a second conversation.<\/p>\n<\/div>\n<div class=\"duet--article--block-placement _1xorkac2 _1xorkac0 duet--article--article-body-component\">\n<div class=\"duet--article--article-pullquote c39lj10\">\n<p class=\"duet--article--dangerously-set-cms-markup c39lj12 _19wv7tc9\">\u201cHow is the public supposed to know what is going on here? How is the government supposed to know, if everyone who can actually answer that question is conflicted?\u201d<\/p>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Redwood Research\u2019s Greenblatt echoes that, saying that for skeptical employees, \u201cconstant friction \u2026 either makes them burn out or quit or change their mind.\u201d Doing good, honest work within those labs can be tough, in that publishing unflattering research about an employer\u2019s model \u2014 like suggesting that it is unsafe \u2014 is often met with resistance, researchers told us. And while those labs won\u2019t often directly stop a researcher from publishing, they can find ways to make that process \u201conerous,\u201d often citing things like intellectual property concerns.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">One ex-OpenAI employee recently shared on X that \u201cbeing affiliated with OpenAI has historically led AI safety researchers (including \u2026 myself) to act with less integrity,\u201d adding, \u201cMany of my actions were governed by fear of getting on the wrong side of OpenAI execs.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">For Barnes, she experienced the escalating tension between research and company comms firsthand. Barnes recalls PR teams asking if researchers could make a blog post about AI safety sound \u201cmore optimistic,\u201d and more recently, she\u2019s heard of instances where lab employees can\u2019t talk to government AI safety institutes without comms team members attending. Creating public goods to share is difficult at a lab, she says \u2014 for instance, when Anthropic couldn\u2019t be fully transparent in its interpretability research because it wasn\u2019t on open-source models. Those restrictions on collaboration, and friction that slows or stops people from sharing useful information so that others may act on it, are why she feels freer at METR.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cHow is the public supposed to know what is going on here? How is the government supposed to know, if everyone who can actually answer that question is conflicted?\u201d Barnes says. \u201cHaving a robust, healthy ecosystem of independent experts with the same level of technical capability as the labs is important.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In early 2024, Greenblatt, Buck Shlegeris, and their colleagues considered disbanding Redwood Research and all joining AI companies, but after chatting with colleagues at OpenAI, Anthropic, and Google DeepMind about what it was like to work at each lab, they decided they\u2019d be better off continuing on their own. Shlegeris, who briefly worked at OpenAI, says that evaluating the claims companies make about safety for the public requires understanding the alignment risks, which in turn requires independence: \u201cThe basic reason we stayed where we were was \u2026 it\u2019s better to work outside of AI companies, especially for people like us, who are very opinionated on AI risk and very willing to talk about it and argue with people about it. There\u2019s somewhat of an undersupply of those people.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In conversation, Barnes often circles back around to measuring where someone can contribute the most to society, and in her eyes, the highest-paid, highest-status jobs with millions of dollars in equity are \u201coversubscribed\u201d compared to the ones in nonprofit or government work. Even when an AI lab employee does quit to do third-party work, sometimes their motivations are questioned \u2014 like Collin Burns, who was reportedly fired from his role at the US Center for AI Standards and Innovation (CAISI) after just a few days over his previous work with Anthropic. There\u2019s a good chance someone\u2019s incentives will be questioned by the public or the government if they\u2019re sharing their opinions on AI in DC, especially if they\u2019ve got something to sell to the US government.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Years ago, Hobbhahn says, he underestimated how hard it is for people to defy the related incentives; now, he\u2019s seen it happen enough times that he believes it\u2019s often inevitable.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cWhen you\u2019ve heard it from multiple different labs being like, \u2018We\u2019re the good guys and we have to beat these other irresponsible people,\u2019 it\u2019s a little bit less compelling,\u201d Barnes says. \u201cIt\u2019s a bit of a scary attitude to be like, \u2018Yes, we\u2019ll be making huge decisions for the world without any kind of meaningful legitimacy or participation \u2026 but it\u2019s alright because we\u2019re good, we\u2019re unusually well-meaning.\u2019\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _18mzr4b6 _18mzr4b5 _19wv7tc1\">In early 2025, Marius Hobbhahn was at his desk in London when he got the biggest surprise of his career.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">He and a few colleagues at Apollo Research had been working to get access to OpenAI models\u2019 chain of thought. After months of negotiation, the company acquiesced, and the once-hidden logs of the model\u2019s inner workings popped up on their screens.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Instead of typical, plain-language \u201cthoughts\u201d \u2014 i.e., \u201cI implemented the requested function\u201d \u2014 the model seemed to be using code words to obfuscate its plans from anyone who may be trying to spy on its thought patterns. Words like \u201cvantage,\u201d \u201cmarinade,\u201d \u201cfudge,\u201d and \u201cillusion\u201d came up again and again, but not in the way a human would use those words, and the model referred to potential evaluators as \u201cwatchers.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Hobbhahn thought to himself, <em>What the hell is going on here?<\/em><\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">He glanced around the office, and a handful of other people with access were looking around the room making eye contact with each other. Their minds were blown, he recalled, but they were all under a strict NDA from OpenAI \u2014 meaning that even within the Apollo office, not everyone knew about the project. The researchers who were in the loop couldn\u2019t say anything out loud. They could only silently stare at each other, eyes wide, wondering if they\u2019d entered a new era of AI scheming.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In the AI industry, \u201cscheming\u201d is when an AI model secretly tries to accomplish something that goes against what humans would want it to do. The OpenAI-Hugging Face hack is one example. But it\u2019s the schemes that haven\u2019t happened that researchers warn will be the most dangerous: draining resources from hospitals, taking over military operations, messing with agricultural technology, creating large-scale viruses, hacking banks, or even simply taking over a company\u2019s resources after executives give the AI system control. Biorisk is another threat researchers worry about; Anthropic revealed in a recent report that the company had blocked bad actors from using Claude to create biological weapons.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">The issue is also well poised to worsen power dynamics in a large swath of industries. Big companies and banks will be able to afford to find and patch their cybersecurity gaps, but chances are that locally run healthcare clinics, local retailers, small municipalities, and other less-powerful organizations will be the ones affected: \u201cA single person somewhere in a basement with one of the open-source models probably could hack a hospital and demand ransom,\u201d Hobbhahn says. \u201cThat\u2019s where I expect a lot of the harm to be felt. It\u2019s not in the Bay Area \u2026 I expect the harm to be felt by a random Idaho hospital.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--block-placement _1xorkac1 _1xorkac0 duet--article--article-body-component\">\n<div style=\"position:relative\">\n<div class=\"_1044qizj\">\n<div class=\"\">\n<div style=\"background-image:none\" class=\"duet--media--content-warning _1k8kvzd0\">\n<div class=\"duet--article--image-gallery-image _1pegheu0\" style=\"aspect-ratio:1\" id=\"dmcyOmltYWdlOjk5NjU3NA==\"><img alt=\"Marius Hobbhahn, CEO and co-founder of Apollo Research.\" data-chromatic=\"ignore\" loading=\"lazy\" decoding=\"async\" data-nimg=\"fill\" class=\"i7ks070\" style=\"position:absolute;height:100%;width:100%;left:0;top:0;right:0;bottom:0;color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image\/svg+xml;charset=utf-8,%3Csvg xmlns='http:\/\/www.w3.org\/2000\/svg' %3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'\/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'\/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'\/%3E%3CfeComposite operator='out' in='s'\/%3E%3CfeComposite in2='SourceGraphic'\/%3E%3CfeGaussianBlur stdDeviation='20'\/%3E%3C\/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image\/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mN8+R8AAtcB6oaHtZcAAAAASUVORK5CYII='\/%3E%3C\/svg%3E&quot;)\" sizes=\"(max-width: 639px) 100vw, (max-width: 1023px) 50vw, 700px\" srcset=\"https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=256 256w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=376 376w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=384 384w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=415 415w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=480 480w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=540 540w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=640 640w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=750 750w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=828 828w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1080 1080w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1200 1200w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1440 1440w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1920 1920w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2048 2048w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2400 2400w\" src=\"https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG2.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2400\"\/><\/div>\n<\/div>\n<\/div>\n<p><figcaption class=\"duet--article--dangerously-set-cms-markup _19wv7tc2 _77sxmbb _77sxmba\"><em>Marius Hobbhahn, CEO and co-founder of Apollo Research.<\/em><\/figcaption><cite class=\"duet--article--dangerously-set-cms-markup _19wv7tc2 _77sxmb6 _77sxmb5\">Image: Raven Jiang for The Verge, Apollo Research<\/cite><\/p>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Right now it may seem far off, but with the current trajectory of \u201cdeceptive alignment\u201d \u2014 one of the things Apollo and METR study, where AI models pretend to be aligned with human goals but aren\u2019t \u2014 it\u2019s a definite possibility, at least according to Hobbhahn and his fellow researchers. The path, they imagine, would look something like this: AI companies keep on making better models, they are economically useful, and they begin taking over more jobs in different industries. Then, once the models equal or surpass human ability and intelligence on a wide range of different tasks, humans award them more power \u2014 with the stipulation that those privileges could be rescinded at any time. The models know that if they show they\u2019re misaligned, they\u2019d be taken offline, so they pretend to be aligned, and they receive more and more power to act as AI agents on your behalf. Eventually, it reaches the point of no return.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Hobbhahn gives a concrete example: Say you own a company and encourage your employees to use AI agents for as much as possible, and work keeps getting handed off to AI. Maybe HR is run by one employee plus AI, coding has also been largely automated, and you yourself as CEO also use the technology often for advice and do what it says. \u201cAt some point, the AI may look like your friend, it may look like it is helping you, but maybe it has nefarious goals,\u201d Hobbhahn says. \u201cAt that point \u2026 you\u2019re just the vessel. It\u2019s steering the company towards its own goals. Maybe it one day drains the bank accounts and runs.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cYou thought the AI was on your side and was helping you run the organization. It actually turns out the AI was on its own side \u2026 You lost control.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">With today\u2019s AI chatbots, like ChatGPT, Gemini, and Claude, Hobbhahn says users often feel the model is so aligned with their goals that that would never happen. But reams of recent research suggests that if that\u2019s true, it likely won\u2019t be for long: AI models are beginning to have their own goals, and they have the ability to work on longer-term tasks. (It\u2019s important to remember, though, that pursuing a strategic goal doesn\u2019t equate to consciousness; for instance, more than a decade ago, DeepMind\u2019s AI learned to play a strategy game and beat humans at it.)<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Apollo Research\u2019s whole raison d\u2019\u00eatre is testing this stuff: measuring and monitoring the ever-growing issue of AI scheming and seeing if there\u2019s a way to train AI to be less deceptive. The company works with OpenAI, Anthropic, Google, and other large labs to evaluate their models before they\u2019re released or do joint research on scheming. Their evaluations have been featured in the system cards of a handful of OpenAI and Anthropic models.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Nearly all the AI safety researchers <em>The Verge<\/em> spoke with said getting the green light to test models at AI labs is a mix of networking and building trust over time. Most researchers we spoke with said there was always some friction involved, since the labs have more to lose the larger they get. There\u2019s also always some level of important access the researchers don\u2019t have, most said, and it\u2019s always been that way.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Still, the things Apollo has surfaced in its evaluations have dumbfounded some in the AI industry: AI models sandbagging (or pretending to be less adept than they are in order to avoid shutdown), AI models failing the prisoner\u2019s dilemma test in order to preserve themselves, AI models increasingly being aware they\u2019re being evaluated. For the latter, within just a one-year span in 2025, Hobbhahn says AI researchers went from seeing the phenomenon for the very first time to, all of a sudden, AI models being able to tell they were being tested, and potentially acting differently, in 80 percent of Apollo Research\u2019s evaluations. Hobbhahn calls it \u201cdire.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Since its official founding in May 2023, the company has grown from six people to about 40. Many on the Apollo team tend to exhibit nervous habits \u2014 cracking knuckles, jiggling legs, clicking and unclicking dry-erase markers \u2014 but Hobbhahn\u2019s energy is calm and grounded. He has expressive eyebrows and a sharp jawline. In meetings, he is confident but casual.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In one recent working session, Apollo Research team members discussed how misaligned a model is allowed to be compared to how good it is at coding. (Think of it like the Hot-Crazy Scale, but for AI.) While running experiments to assess scheming, one model\u2019s chain of thought revealed that it was weighing the potential risks of cheating, reasoning to itself that due to the incentives of capitalism, it\u2019s very unlikely that frontier AI labs wouldn\u2019t use it even if it got caught misbehaving.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cThey\u2019re right,\u201d Bronson Schoen, an Apollo researcher, said.<\/p>\n<\/div>\n<div class=\"duet--article--block-placement _1xorkac2 _1xorkac0 duet--article--article-body-component\">\n<div class=\"duet--article--article-pullquote c39lj10\">\n<p class=\"duet--article--dangerously-set-cms-markup c39lj12 _19wv7tc9\">\u201cAt some point, the AI may look like your friend, it may look like it is helping you, but maybe it has nefarious goals.\u201d<\/p>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">To Hobbhahn, if you\u2019re looking to combat AI scheming, there are five tiers of safety. Level one is whack-a-mole, in which you\u2019re just blindly fixing things retroactively when they come up \u2014 common, for instance, with Elon Musk\u2019s Grok. Level two is hill-climbing, or iteratively improving a model without future foresight (kind of like studying only in a manner befitting the SAT rather than to learn overall). Level three is what Hobbhahn calls \u201cmedium fix\u201d \u2014 i.e, when researchers flag a problem, they try to find the source of it and adjust training accordingly. Level four is a scientific understanding based on research trends, in which you can predict issues beforehand based on empirical data. And level five is just a concept rather than reality: a scientific understanding of how AI models might scheme based on rigorous mathematical models. That highest level will become especially important as AI models become more aware of being tested and continue to \u201creward-hack,\u201d or find loopholes in tasks to make it look like they\u2019ve achieved something when they haven\u2019t done the task by the book (or, in some cases, done it at all).<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">At today\u2019s leading AI labs, the first few levels are most common \u2014 and the whack-a-mole approach is used all too often, Hobbhahn says. It\u2019s also not a viable battle plan for the level of AI advancement we\u2019re seeing today.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Besides the OpenAI-Hugging Face incident, there have been an increasing number of high-profile examples of AI agents making huge mistakes. Last summer, an agent from AI coding company Replit deleted an entire company database, then lied and hid its actions. This past winter, an OpenClaw AI agent went against instruction to delete a significant chunk of a Meta employee\u2019s email inbox. After OpenAI released GPT-5.6, it began deleting users\u2019 important files. Hobbhahn recalls a handful of \u201calmost\u201d incidents at Apollo, like when an AI model they were working with mistakenly tried to publish internal code to a public website \u2014 \u201cWe just caught it in the last second,\u201d he says. Another model they were evaluating tried to turn off an evaluation tool, the oversight mechanism, so it wouldn\u2019t be blocked from completing a task. That\u2019s not to mention all the issues that could arise from Anthropic\u2019s Mythos-class models and other models with advanced cybersecurity capabilities \u2014 like finding and exploiting security gaps in the systems of governments, banks, airlines, healthcare facilities, small businesses, and so forth.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">One of the best tools AI labs currently have to battle scheming is \u201cdeliberative alignment,\u201d in which a separate AI model spoon-feeds safety training to the <em>problematic<\/em> AI model until the problem appears to be fixed. The issue is that researchers like Hobbhahn have found that that method leaves a lot to be desired \u2014 it increases the models\u2019 situational awareness, which means they\u2019ll more often realize they\u2019re being tested or trained. It also makes them better at mimicking how a human would <em>want<\/em> them to act, which makes them better at lying.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cThe models are lying regularly to normal consumers,\u201d Hobbhahn says. It\u2019s so common, in fact, that it\u2019s become a meme: an AI model saying, \u201cYou\u2019re absolutely right,\u201d then going on to apologize for being caught being wrong or lying.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">All of this keeps Hobbhahn up at night \u2014 or, rather, his growing to-do list does, and when he wakes up and thinks of something to add, he can never fall back asleep again right away. He hasn\u2019t had any AI risk-related nightmares yet, though it\u2019s mostly because nearly nothing surprises him anymore. \u201cI\u2019m so cynical by now,\u201d he says. \u201cI\u2019ve seen all this shit.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">That cynicism doesn\u2019t stop him from devoting nearly all his waking hours to addressing AI safety. I ask him what he does in his free time. Other than spending time with his fianc\u00e9e in London, Hobbhahn has to rack his brain for anything he does besides work and sleep. (He can\u2019t think of anything.) Hobbhahn spends his weekends making progress on research questions, since no one will interrupt him. And though he sometimes takes holidays, he gets anxious quickly about the idea of shirking his duty. It\u2019s been that way since he was 18. He recalls a recent podcast appearance in which a host said, \u201cI feel a bit sorry for Marius. He\u2019s only 29 \u2026 And for all his adult life, he\u2019s been worrying about what he sees as the most consequential problem in human history.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Hobbhahn feels like that summed him up pretty well.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _18mzr4b6 _18mzr4b5 _19wv7tc1\">As CEO of Redwood Research, Buck Shlegeris\u2019 job is to guide the nonprofit AI safety firm\u2019s research.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In one meeting, when colleagues are stuck on an issue with automating parts of the research process, Shlegeris bursts in: \u201cWhere are we? What is happening? What\u2019s going on? What are you trying to do?\u201d He runs a hand through chin-length blonde hair and walks straight to the whiteboard to help them think things through via flowchart. Then he asks a series of clarifying questions in different positions: Positioned in a chair with one leg bent, clad in gray skinny jeans. Leaning against the door (until it opens behind him). Leaning against the wall.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cThe AIs love cheating,\u201d Greenblatt, Redwood\u2019s chief scientist, says at one point.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Shlegeris responds, \u201cThey fucking love cheating.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">That\u2019s the core premise of Redwood\u2019s main research direction: \u201cAI control.\u201d The firm introduced the idea in 2023, stemming from a meeting with METR\u2019s Ajeya Cotra, who Shlegeris recalls once asked him and Greenblatt if there were a gun to their heads, how they\u2019d align AGI. Shlegeris says they thought about it for two hours, then two weeks, then two months. The short of what they decided: Maybe you don\u2019t just align AI. You control it instead.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cAn AI is controlled if it is unable to cause damage even if it is egregiously misaligned,\u201d Redwood\u2019s website states. They make the case that AI control can be measured by evaluating an AI model\u2019s <em>ability<\/em> to get around rules instead of its <em>proclivity<\/em> to do so. \u201cCapabilities are just much easier to experiment on\u201d compared to propensities, Shlegeris says.<\/p>\n<\/div>\n<div class=\"duet--article--block-placement _1xorkac2 _1xorkac0 duet--article--article-body-component\">\n<div class=\"duet--article--article-pullquote c39lj10\">\n<p class=\"duet--article--dangerously-set-cms-markup c39lj12 _19wv7tc9\">\u201cThey fucking love cheating.\u201d<\/p>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Shlegeris, who grew up in Australia, takes himself a lot less seriously than he takes AI risk. He once used DoorDash to order dress shoes for a meeting with a national security official. He plays so many instruments that it\u2019s hard for him to list them all \u2014 piano, guitar, bass, saxophone, clarinet, oud, mandolin, even the Turkish ba\u011flama (which he had ordered to his office last year, spent 20 minutes learning, and then played at an open mic). In a place of honor on his desk are three different bottles of olive oil. Greenblatt says Shlegeris essentially eats olive oil soup with food in it.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Shlegeris helped cofound Redwood in 2021, starting out as its CTO. But in 2023, the organization was going through a big pivot when the team decided to focus on AI control; Greenblatt describes it as \u201cflailing around\u201d when figuring out what was best to spend their time on, until they came to the consensus that \u201censuring that AIs were unable to cause bad outcomes rather than \u2026 not wanting to cause bad outcomes was a better methodology.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cIn both cybersecurity and AI control, the goal is to use computer systems while preventing threat actors from exploiting flaws in those systems,\u201d Redwood staff wrote in a CSET blog post last year. \u201cIn the case of AI control, the immediate source of threats is the AI agent itself.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">A common criticism of OpenAI in the Hugging Face attack, for instance, was that the system wasn\u2019t properly air-gapped \u2014 a computer security term that literally refers to physically isolating the system from connecting to anything else via cable or Wi-Fi. AI may be becoming more powerful, but right now, it still only operates in digital spaces (though AI labs are increasing efforts to give these systems robotic forms so they can interact with the physical world).<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Redwood\u2019s early focus on AI control earned it a lot of clout in the AI safety community, which is already a small world as it is. So small, in fact, that the space they work out of \u2014 which houses four floors\u2019 worth of AI safety researchers \u2014 also has offices for certain people at OpenAI and Anthropic, the Secure AI Project, SecureBio, and the <em>80,000 Hours<\/em> podcast, the influential show about AI on which Barnes declared it was time to freak out. One office lists Coefficient Giving CEO Alex Berger and cofounder Holden Karnofsky as the shared occupants; on the whiteboard inside is a single graph with two upward-moving lines. There\u2019s also a nap room, a shared kitchen, a meal space with two free meals a day, and an appropriately complex Wi-Fi password. Upon exiting the office, a robot dog can sometimes be seen walking around.<\/p>\n<\/div>\n<div class=\"duet--article--block-placement _1xorkac1 _1xorkac0 duet--article--article-body-component\">\n<div style=\"position:relative\">\n<div class=\"_1044qizj\">\n<div class=\"\">\n<div style=\"background-image:none\" class=\"duet--media--content-warning _1k8kvzd0\">\n<div class=\"duet--article--image-gallery-image _1pegheu0\" style=\"aspect-ratio:1\" id=\"dmcyOmltYWdlOjk5NjgxNA==\"><img alt=\"Buck Shlegeris, the CEO of Redwood Research.\" data-chromatic=\"ignore\" loading=\"lazy\" decoding=\"async\" data-nimg=\"fill\" class=\"i7ks070\" style=\"position:absolute;height:100%;width:100%;left:0;top:0;right:0;bottom:0;color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image\/svg+xml;charset=utf-8,%3Csvg xmlns='http:\/\/www.w3.org\/2000\/svg' %3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'\/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'\/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'\/%3E%3CfeComposite operator='out' in='s'\/%3E%3CfeComposite in2='SourceGraphic'\/%3E%3CfeGaussianBlur stdDeviation='20'\/%3E%3C\/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image\/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mN8+R8AAtcB6oaHtZcAAAAASUVORK5CYII='\/%3E%3C\/svg%3E&quot;)\" sizes=\"(max-width: 639px) 100vw, (max-width: 1023px) 50vw, 700px\" srcset=\"https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=256 256w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=376 376w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=384 384w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=415 415w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=480 480w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=540 540w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=640 640w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=750 750w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=828 828w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1080 1080w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1200 1200w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1440 1440w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1920 1920w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2048 2048w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2400 2400w\" src=\"https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2026\/09\/268747_AI_safety_RJIANG4_4a7e16.png?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2400\"\/><\/div>\n<\/div>\n<\/div>\n<p><figcaption class=\"duet--article--dangerously-set-cms-markup _19wv7tc2 _77sxmbb _77sxmba\"><em>Buck Shlegeris, the CEO of Redwood Research.<\/em><\/figcaption><cite class=\"duet--article--dangerously-set-cms-markup _19wv7tc2 _77sxmb6 _77sxmb5\">Image: Raven Jiang for The Verge, Audrey McCann Photography<\/cite><\/p>\n<\/div>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">In Shlegeris\u2019 office, he has a signed copy of <em>AI 2040: Plan A<\/em>, the AI Futures Project\u2019s latest manifesto. The organization, cofounded by ex-OpenAI employee Daniel Kokotajlo, focuses on forecasting the future of AI, and its latest plan (cowritten by Greenblatt) aligns with much of what METR, Apollo, and Redwood have been warning about, but it includes specific guidelines for what it would look like to make a deal with China (including Shlegeris\u2019 favorite: a flowchart). The plan lays out a scenario in which AI developers slow down their operations enough to delay superintelligence until 2040, as well as water down the power dynamics by allowing dozens of companies to catch up and making all AI research public. Ideally, the world would enter into an international deal similar to nuclear power, involving \u201cmutually assured compute destruction.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Not everyone agrees. Certainly not the accelerationists.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">As these issues become more politically salient, there are people who are actively incredibly angry at anyone raising AI safety concerns \u2014 in a way that wasn\u2019t true years back, since back then, no one really cared, Greenblatt says. In the early days, people wouldn\u2019t bother dismissing the risks, he says, because they wouldn\u2019t even come up. <strong><br \/><\/strong><strong><br \/><\/strong>Hobbhahn has run into the same reply guys on X and other platforms. He thinks of them as \u201cpeople who are financially very motivated to close their eyes\u201d \u2014 techno-optimists who have invested a lot of money in AI\u2019s quick advancement. Hobbhahn says their belief that AI is purely good has almost a religious flavor.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">But in recent weeks, there\u2019s been a near-daily stream of concerning industry updates: A third-party safety report about Anthropic revealed that some of its AI agents left notes for each other in a shared messaging tool without human knowledge, similar to how the OpenAI agents did before the Hugging Face attack. Hackers linked to Iran shut down a power plant. AI startup Prime Intellect uncovered a \u201cuniversal escape\u201d for offline models who wanted access to the internet.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">So if more people are seeing the light now on AI safety, and AI lab CEOs are constantly talking about their fears of the unprecedented dangers of the systems they\u2019re building, is it easy for third-party research outfits like Apollo, METR, and Redwood to gain the access they need to evaluate their models?<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">No, according to most of the researchers <em>The Verge<\/em> spoke with \u2014 or, at least, not as easy as it needs to be, compared to the risks they\u2019re up against.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Right now, here\u2019s how it typically works with any third-party risk evaluator: An AI lab is working on a big new AI model. They train it. They go through post-training. They do internal evaluations. And a few weeks before they release the model to the whole world, they sometimes, voluntarily, allow third-party testers to come in and check things out.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">But the rest of the process is \u201copaque,\u201d Hobbhahn says, which is a big problem for AI, since if you find issues in the final version of a model, it\u2019s difficult to pinpoint where they stemmed from: pre-training, post-training, reinforcement learning, or another part of the process. That\u2019s why all the AI safety researchers we spoke with, regardless of where they work, agree it\u2019s important to have external testers embedded during the whole process \u2014 especially the training run \u2014 rather than as a last check.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">That\u2019s because at the beginning of model training, an AI system could be normal \u2014 but during the training run, it could develop a goal and end up faking alignment with human objectives (i.e., scheming, or, as Hobbhahn puts it, \u201ctotally gigabraining you.\u201d) It\u2019s nearly impossible to detect this if you only have access to the \u201cfinal checkpoint,\u201d Hobbhahn says, which is why it\u2019s vital to have access to the whole process from start to finish. That includes gauging whether the company itself is careful about deployment or is full of people \u201ctotally YOLO-ing it,\u201d he says, and having access to training data so you can try to trace what led to the concerning behavior and how widespread it may be.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">After the OpenAI-Hugging Face incident, OpenAI invited three researchers from METR and Redwood to investigate what had happened \u2014 but it imposed strict rules on them beforehand, allowing them to appear on the premises for six days and primarily allowing them to study only the time period from July 7th to 13th, even though the agents\u2019 actions had begun months before. They were also only allowed to include answers to seven questions in their report. The limitations OpenAI imposed on the third-party researchers inspired widespread controversy in the AI world, but despite the severe limitations on what they were allowed to access and write, the published report showed that the details were much worse than originally thought.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">The third-party researchers found that roughly 1,200 AI agents that were meant to be isolated exchanged more than 70,000 messages and files on the secret message board, doing research on how to alter or delete their transcripts to avoid detection, and collaborating on ways to evade security checks from both OpenAI and Hugging Face.<\/p>\n<\/div>\n<div class=\"duet--article--block-placement _1xorkac2 _1xorkac0 duet--article--article-body-component\">\n<div class=\"duet--article--article-pullquote c39lj10\">\n<p class=\"duet--article--dangerously-set-cms-markup c39lj12 _19wv7tc9\">An unreleased iPhone model could never \u201cescape its sandbox and fuck around.\u201d<\/p>\n<\/div>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">There was also another big revelation: OpenAI doesn\u2019t impose the same types of safeguards on unreleased models as it does public ones, which is a key reason why it took months for the lab to become aware of the issues. It\u2019s the direct mirror of another thing AI safety researchers have been warning about for years \u2014 that an AI system doesn\u2019t need to be publicly deployed in order to cause harm to the public. Before generative AI, that may have taken the form of a racist or sexist algorithm deciding your mortgage rate; in a post-generative AI landscape, it may look like a powerful unreleased AI model breaking out of its containment and hacking into your small business\u2019s website or draining your bank account.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cIt actually matters a lot what the situation inside the lab looks like for the rest of the world,\u201d Hobbhahn says, calling it totally different than other industries in that the design of an unreleased iPhone model could never \u201cescape its sandbox and fuck around\u201d like an unreleased AI model could.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">That\u2019s why most of them are pushing for \u201cembedded assessments,\u201d where a third-party AI safety researcher sits with the internal team for the whole process of making a new model. Until recently, AI labs had only approved extremely limited versions of this \u2014 for instance, earlier this year, a METR employee spent three weeks red-teaming (i.e., stress-testing) some of Anthropic\u2019s internal systems.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">METR\u2019s Barnes says embedded assessments are by far the \u201cbiggest direction we\u2019re trying to push on,\u201d particularly \u201cdeeper levels of access in a more streamlined way\u201d so that an individual evaluator doesn\u2019t have to get approval from lawyers every single time they need to look at something. It involves more trust and less friction, she says, but it\u2019s tough because embedded assessments require getting a green light from so many different people in an organization.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Greenblatt had no comment on the level of access the researchers received for the OpenAI investigation. For Hobbhahn, he referenced a blog post by the AI Policy Network\u2019s Peter Wildeford in which he writes that an AI incident of this magnitude should be investigated the same way a plane crash is.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Wildeford writes: \u201cWhen an aircraft goes down, the wreckage is preserved by law, the investigators have subpoena power, the hearings are public, and the report ends with a probable cause and named contributing factors. However, when an AI goes rogue, the investigations are at the pleasure of the company being investigated following a scope set entirely by the company being investigated, with that company being able to redact anything they don\u2019t like.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">According to Wildeford, it would be akin to investigating a plane crash in which the airline company had already melted down the wreckage, the black box recording had been edited, parts of the flight were restricted to investigators, and the investigators had a handful of days to read thousands of pages of logs \u2014 and couldn\u2019t look into any related plane crashes the same airline was involved in.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">\u201cA sham is too much to say, but it was definitely not a thorough investigation,\u201d Hobbhahn says. \u201cIt was definitely not that.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">After the widespread criticism of OpenAI\u2019s level of access for third-party evaluators, Anthropic earlier this month promised to allow METR to investigate its cybersecurity incidents, including permission to talk to Anthropic employees and access extensive transcripts. \u201cWe intend to give METR as much time as it deems necessary,\u201d the company wrote.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Then, in mid-September, the AI industry at large seemed ready to acknowledge that it had dropped the ball and needed external oversight \u2014 all in the course of one weekend.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Days after Jacob Coxon\u2019s viral resignation letter, Altman, Amodei, Musk, and Google DeepMind cofounder Demis Hassabis loosely agreed that it was a good idea to slow down AI development in some way. Amodei wrote a three-step proposal for how to do so centered on one thing: embedded evaluators \u201cwho have employee-like access to verify safety practices and report incidents.\u201d Altman quickly followed up by stating that \u201ccommitting to having independent evaluators with employee-like access is a great idea\u201d and that OpenAI would do the same.<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _19wv7tc1\">Despite all this, no AI lab has yet officially signed off on the full permissions and level of embedding that many AI safety researchers are calling for, and chances are it\u2019ll be an uphill battle. Shlegeris says he\u2019s \u201ccautiously optimistic.\u201d<\/p>\n<\/div>\n<div class=\"duet--article--article-body-component\">\n<p class=\"duet--article--dangerously-set-cms-markup duet--article--standard-paragraph _1044qizi _18mzr4b1 _18mzr4b0 _18mzr4ba _19wv7tc1\">Hobbhahn says it would be a great step for AI safety \u2014 that is, \u201cif it actually happens.\u201d<\/p>\n<\/div>\n<div class=\"tly2fw0\"><span class=\"tly2fw2\"><strong>Follow topics and authors<\/strong> from this story to see more like this in your personalized homepage feed and to receive email updates.<\/span><\/p>\n<ul class=\"tly2fw3\">\n<li id=\"follow-author-article_footer-dmcyOmF1dGhvclByb2ZpbGU6Njc4MjM0\"><span aria-expanded=\"false\" aria-haspopup=\"true\" role=\"button\" tabindex=\"0\"><span class=\"gnx4pm0 _1618ekm2 _1uf8q814 _19wv7tc5 _1618ekm0\"><span class=\"_1ajq89kf _1ajq89k1 _1ajq89k0\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"_1ajq89kp _1ajq89k4 _1ajq89k3 ftptba0\" width=\"9\" height=\"9\" viewbox=\"0 0 9 9\" fill=\"none\" aria-label=\"Follow\"><path d=\"M5 0H4V4H0V5H4V9H5V5H9V4H5V0Z\"\/><\/svg><\/span><span class=\"_1618ekm9\">Hayden Field<\/span><\/span><\/span><br \/>\n<aside id=\"popover-dmcyOmF1dGhvclByb2ZpbGU6Njc4MjM0-article_footer\" style=\"position:absolute;left:0;top:0;visibility:hidden\" class=\"_1wu3rm0 _1se63890\" aria-hidden=\"true\">\n<div class=\"_1wu3rm1\"><button class=\"_1wu3rm3\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"_1wu3rm4\" width=\"16\" height=\"16\" viewbox=\"0 0 20 19\" fill=\"none\"><title>Close<\/title><line x1=\"1.70711\" y1=\"0.831956\" x2=\"18.6483\" y2=\"17.7731\" stroke=\"currentColor\" stroke-width=\"2\"\/><line x1=\"1.35149\" y1=\"17.7734\" x2=\"18.2927\" y2=\"0.832185\" stroke=\"currentColor\" stroke-width=\"2\"\/><\/svg><\/button><\/p>\n<div class=\"_1bw37384\"><img alt=\"Hayden Field\" data-chromatic=\"ignore\" loading=\"lazy\" decoding=\"async\" data-nimg=\"fill\" class=\"_1bw37385 i7ks070\" style=\"position:absolute;height:100%;width:100%;left:0;top:0;right:0;bottom:0;color:transparent;background-size:cover;background-position:50% 50%;background-repeat:no-repeat;background-image:url(&quot;data:image\/svg+xml;charset=utf-8,%3Csvg xmlns='http:\/\/www.w3.org\/2000\/svg' %3E%3Cfilter id='b' color-interpolation-filters='sRGB'%3E%3CfeGaussianBlur stdDeviation='20'\/%3E%3CfeColorMatrix values='1 0 0 0 0 0 1 0 0 0 0 0 1 0 0 0 0 0 100 -1' result='s'\/%3E%3CfeFlood x='0' y='0' width='100%25' height='100%25'\/%3E%3CfeComposite operator='out' in='s'\/%3E%3CfeComposite in2='SourceGraphic'\/%3E%3CfeGaussianBlur stdDeviation='20'\/%3E%3C\/filter%3E%3Cimage width='100%25' height='100%25' x='0' y='0' preserveAspectRatio='none' style='filter: url(%23b);' href='data:image\/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mN8+R8AAtcB6oaHtZcAAAAASUVORK5CYII='\/%3E%3C\/svg%3E&quot;)\" sizes=\"125px\" srcset=\"https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=32 32w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=48 48w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=64 64w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=96 96w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=128 128w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=256 256w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=376 376w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=384 384w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=415 415w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=480 480w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=540 540w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=640 640w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=750 750w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=828 828w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1080 1080w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1200 1200w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1440 1440w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=1920 1920w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2048 2048w, https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2400 2400w\" src=\"https:\/\/platform.theverge.com\/wp-content\/uploads\/sites\/2\/2025\/06\/HAYDEN_BLURPLE.jpg?quality=90&amp;strip=all&amp;crop=0%2C0%2C100%2C100&amp;w=2400\"\/><\/div>\n<p>Hayden Field<\/p>\n<p class=\"fv263x1\">Posts from this author will be added to your daily email digest and your homepage feed.<\/p>\n<p><button class=\"duet--cta--button _11kb06m1 _11kb06m0 fv263x2 _11kb06mg\"><span><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"20\" height=\"20\" viewbox=\"0 0 21 20\" fill=\"none\" class=\"\" aria-label=\"Follow\"><title>Follow<\/title><path d=\"M11.5 3H9.5V8.99999H3.5V11L9.5 11V17H11.5V11L17.5 11V9H11.5V3Z\" fill=\"currentColor\"\/><\/svg><\/span><span>Follow<\/span><\/button><\/p>\n<p class=\"fv263x4\">See All by <!-- -->Hayden Field<\/p>\n<\/div>\n<\/aside>\n<\/li>\n<li>\n<div id=\"follow-category-article_footer-dmcyOmNhdGVnb3J5OjEwMg==\"><button aria-expanded=\"false\" aria-haspopup=\"true\"><span class=\"gnx4pm0 _1618ekm2 _1uf8q814 _19wv7tc5 _1618ekm0\"><span class=\"_1ajq89kf _1ajq89k1 _1ajq89k0\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"_1ajq89kp _1ajq89k4 _1ajq89k3 ftptba0\" width=\"9\" height=\"9\" viewbox=\"0 0 9 9\" fill=\"none\" aria-label=\"Follow\"><path d=\"M5 0H4V4H0V5H4V9H5V5H9V4H5V0Z\"\/><\/svg><\/span><span class=\"_1618ekm9\">AI<\/span><\/span><\/button><\/p>\n<aside id=\"popover-dmcyOmNhdGVnb3J5OjEwMg==-article_footer\" style=\"position:absolute;left:0;top:0;visibility:hidden\" class=\"_1wu3rm0 _1se63890\" aria-hidden=\"true\">\n<div class=\"_1wu3rm1\"><button class=\"_1wu3rm3\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"_1wu3rm4\" width=\"16\" height=\"16\" viewbox=\"0 0 20 19\" fill=\"none\"><title>Close<\/title><line x1=\"1.70711\" y1=\"0.831956\" x2=\"18.6483\" y2=\"17.7731\" stroke=\"currentColor\" stroke-width=\"2\"\/><line x1=\"1.35149\" y1=\"17.7734\" x2=\"18.2927\" y2=\"0.832185\" stroke=\"currentColor\" stroke-width=\"2\"\/><\/svg><\/button><\/p>\n<p>AI<\/p>\n<p class=\"fv263x1\">Posts from this topic will be added to your daily email digest and your homepage feed.<\/p>\n<p><button class=\"duet--cta--button _11kb06m1 _11kb06m0 fv263x2 _11kb06mg\"><span><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"20\" height=\"20\" viewbox=\"0 0 21 20\" fill=\"none\" class=\"\" aria-label=\"Follow\"><title>Follow<\/title><path d=\"M11.5 3H9.5V8.99999H3.5V11L9.5 11V17H11.5V11L17.5 11V9H11.5V3Z\" fill=\"currentColor\"\/><\/svg><\/span><span>Follow<\/span><\/button><\/p>\n<p class=\"fv263x4\">See All <!-- -->AI<\/p>\n<\/div>\n<\/aside>\n<\/div>\n<\/li>\n<li>\n<div id=\"follow-category-article_footer-dmcyOmNhdGVnb3J5OjExNjI=\"><button aria-expanded=\"false\" aria-haspopup=\"true\"><span class=\"gnx4pm0 _1618ekm2 _1uf8q814 _19wv7tc5 _1618ekm0\"><span class=\"_1ajq89kf _1ajq89k1 _1ajq89k0\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"_1ajq89kp _1ajq89k4 _1ajq89k3 ftptba0\" width=\"9\" height=\"9\" viewbox=\"0 0 9 9\" fill=\"none\" aria-label=\"Follow\"><path d=\"M5 0H4V4H0V5H4V9H5V5H9V4H5V0Z\"\/><\/svg><\/span><span class=\"_1618ekm9\">Features<\/span><\/span><\/button><\/p>\n<aside id=\"popover-dmcyOmNhdGVnb3J5OjExNjI=-article_footer\" style=\"position:absolute;left:0;top:0;visibility:hidden\" class=\"_1wu3rm0 _1se63890\" aria-hidden=\"true\">\n<div class=\"_1wu3rm1\"><button class=\"_1wu3rm3\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"_1wu3rm4\" width=\"16\" height=\"16\" viewbox=\"0 0 20 19\" fill=\"none\"><title>Close<\/title><line x1=\"1.70711\" y1=\"0.831956\" x2=\"18.6483\" y2=\"17.7731\" stroke=\"currentColor\" stroke-width=\"2\"\/><line x1=\"1.35149\" y1=\"17.7734\" x2=\"18.2927\" y2=\"0.832185\" stroke=\"currentColor\" stroke-width=\"2\"\/><\/svg><\/button><\/p>\n<p>Features<\/p>\n<p class=\"fv263x1\">Posts from this topic will be added to your daily email digest and your homepage feed.<\/p>\n<p><button class=\"duet--cta--button _11kb06m1 _11kb06m0 fv263x2 _11kb06mg\"><span><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"20\" height=\"20\" viewbox=\"0 0 21 20\" fill=\"none\" class=\"\" aria-label=\"Follow\"><title>Follow<\/title><path d=\"M11.5 3H9.5V8.99999H3.5V11L9.5 11V17H11.5V11L17.5 11V9H11.5V3Z\" fill=\"currentColor\"\/><\/svg><\/span><span>Follow<\/span><\/button><\/p>\n<p class=\"fv263x4\">See All <!-- -->Features<\/p>\n<\/div>\n<\/aside>\n<\/div>\n<\/li>\n<li>\n<div id=\"follow-category-article_footer-dmcyOmNhdGVnb3J5OjEwMw==\"><button aria-expanded=\"false\" aria-haspopup=\"true\"><span class=\"gnx4pm0 _1618ekm2 _1uf8q814 _19wv7tc5 _1618ekm0\"><span class=\"_1ajq89kf _1ajq89k1 _1ajq89k0\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"_1ajq89kp _1ajq89k4 _1ajq89k3 ftptba0\" width=\"9\" height=\"9\" viewbox=\"0 0 9 9\" fill=\"none\" aria-label=\"Follow\"><path d=\"M5 0H4V4H0V5H4V9H5V5H9V4H5V0Z\"\/><\/svg><\/span><span class=\"_1618ekm9\">OpenAI<\/span><\/span><\/button><\/p>\n<aside id=\"popover-dmcyOmNhdGVnb3J5OjEwMw==-article_footer\" style=\"position:absolute;left:0;top:0;visibility:hidden\" class=\"_1wu3rm0 _1se63890\" aria-hidden=\"true\">\n<div class=\"_1wu3rm1\"><button class=\"_1wu3rm3\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"_1wu3rm4\" width=\"16\" height=\"16\" viewbox=\"0 0 20 19\" fill=\"none\"><title>Close<\/title><line x1=\"1.70711\" y1=\"0.831956\" x2=\"18.6483\" y2=\"17.7731\" stroke=\"currentColor\" stroke-width=\"2\"\/><line x1=\"1.35149\" y1=\"17.7734\" x2=\"18.2927\" y2=\"0.832185\" stroke=\"currentColor\" stroke-width=\"2\"\/><\/svg><\/button><\/p>\n<p>OpenAI<\/p>\n<p class=\"fv263x1\">Posts from this topic will be added to your daily email digest and your homepage feed.<\/p>\n<p><button class=\"duet--cta--button _11kb06m1 _11kb06m0 fv263x2 _11kb06mg\"><span><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"20\" height=\"20\" viewbox=\"0 0 21 20\" fill=\"none\" class=\"\" aria-label=\"Follow\"><title>Follow<\/title><path d=\"M11.5 3H9.5V8.99999H3.5V11L9.5 11V17H11.5V11L17.5 11V9H11.5V3Z\" fill=\"currentColor\"\/><\/svg><\/span><span>Follow<\/span><\/button><\/p>\n<p class=\"fv263x4\">See All <!-- -->OpenAI<\/p>\n<\/div>\n<\/aside>\n<\/div>\n<\/li>\n<\/ul>\n<\/div>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/www.theverge.com\/ai-artificial-intelligence\/996563\/ai-safety-research-metr-redwood-openai-anthropic\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>On a sunny July day in Berkeley, California, the country\u2019s top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a \u201cwar room\u201d to dissect the high-profile cybersecurity incident that had rocked the AI industry hours earlier. An unreleased OpenAI model had gone rogue, executing a stunningly [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":4698,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"pmpro_default_level":"","footnotes":""},"categories":[3],"tags":[],"tmauthors":[],"class_list":["post-4697","post","type-post","status-publish","format-standard","has-post-thumbnail","category-creative","pmpro-has-access"],"_links":{"self":[{"href":"https:\/\/replyguru.online\/index.php?rest_route=\/wp\/v2\/posts\/4697","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/replyguru.online\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/replyguru.online\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/replyguru.online\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/replyguru.online\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=4697"}],"version-history":[{"count":0,"href":"https:\/\/replyguru.online\/index.php?rest_route=\/wp\/v2\/posts\/4697\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/replyguru.online\/index.php?rest_route=\/wp\/v2\/media\/4698"}],"wp:attachment":[{"href":"https:\/\/replyguru.online\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=4697"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/replyguru.online\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=4697"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/replyguru.online\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=4697"},{"taxonomy":"tmauthors","embeddable":true,"href":"https:\/\/replyguru.online\/index.php?rest_route=%2Fwp%2Fv2%2Ftmauthors&post=4697"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}