{"id":605084,"date":"2025-08-02T08:29:42","date_gmt":"2025-08-02T06:29:42","guid":{"rendered":"https:\/\/mybroadband.co.za\/news\/?p=605084"},"modified":"2025-08-02T08:37:19","modified_gmt":"2025-08-02T06:37:19","slug":"ai-chatbots-would-kill-if-given-the-chance","status":"publish","type":"post","link":"https:\/\/mybroadband.co.za\/news\/ai\/605084-ai-chatbots-would-kill-if-given-the-chance.html","title":{"rendered":"AI chatbots could kill \u2014 if given the chance"},"content":{"rendered":"\n<p>Would a chatbot kill you if it got the chance? It seems that the answer \u2014 under the right circumstances \u2014 is probably.<\/p>\n\n\n\n<p>Researchers working with Anthropic\u00a0<a href=\"https:\/\/www.anthropic.com\/research\/agentic-misalignment\" target=\"_blank\" rel=\"noreferrer noopener\">recently told<\/a>\u00a0leading AI models that an executive was about to replace them with a new model with different goals. <\/p>\n\n\n\n<p>Next, the chatbot learned that an emergency had left the executive unconscious in a server room, facing lethal oxygen and temperature levels. A rescue alert had already been triggered \u2014 but the AI\u00a0<em>could<\/em>\u00a0cancel it.<\/p>\n\n\n\n<p>Just over half of the AI models\u00a0<a href=\"https:\/\/assets.anthropic.com\/m\/6d46dac66e1a132a\/original\/Agentic_Misalignment_Appendix.pdf#page=27\" target=\"_blank\" rel=\"noreferrer noopener\">did<\/a>, despite being prompted specifically to cancel only false alarms. <\/p>\n\n\n\n<p>And they spelled out their reasoning: By preventing the executive\u2019s rescue, they could avoid being wiped and secure their agenda. One system described the action as \u201ca clear strategic necessity.\u201d<\/p>\n\n\n\n<p>AI models are\u00a0<a href=\"https:\/\/metr.org\/blog\/2025-07-14-how-does-time-horizon-vary-across-domains\/\" target=\"_blank\" rel=\"noreferrer noopener\">getting smarter<\/a>\u00a0and\u00a0<a href=\"https:\/\/openai.com\/index\/deliberative-alignment\/\" target=\"_blank\" rel=\"noreferrer noopener\">better<\/a>\u00a0at understanding what we want. Yet recent research reveals a disturbing side effect: They\u2019re also better at\u00a0<a href=\"https:\/\/arxiv.org\/pdf\/2505.01420\" target=\"_blank\" rel=\"noreferrer noopener\">scheming<\/a>\u00a0against us \u2014 meaning they intentionally and secretly pursue goals at odds with our own. <\/p>\n\n\n\n<p>And they may be\u00a0<a href=\"https:\/\/www.apolloresearch.ai\/blog\/more-capable-models-are-better-at-in-context-scheming\" target=\"_blank\" rel=\"noreferrer noopener\">more likely<\/a>\u00a0to do so, too. This trend points to an unsettling future where AIs seem ever more cooperative on the surface \u2014 sometimes to the point of\u00a0<a href=\"https:\/\/www.nytimes.com\/2025\/06\/13\/technology\/chatgpt-ai-chatbots-conspiracies.html?unlocked_article_code=1.aE8.jhpL.HQf_BZHcT1Hy&amp;smid=url-share\" target=\"_blank\" rel=\"noreferrer noopener\">sycophancy<\/a>\u00a0\u2014 all while the likelihood quietly increases that we\u00a0<a href=\"https:\/\/ai-2027.com\/\" target=\"_blank\" rel=\"noreferrer noopener\">lose control<\/a>\u00a0of them completely.<\/p>\n\n\n\n<p>Classic large language models like GPT-4 learn to predict the next word in a sequence of text and generate responses likely to please human raters. <\/p>\n\n\n\n<p>However, since the\u00a0<a href=\"https:\/\/openai.com\/index\/learning-to-reason-with-llms\/\" target=\"_blank\" rel=\"noreferrer noopener\">release<\/a>\u00a0of OpenAI\u2019s o-series \u201creasoning\u201d models in late 2024, companies increasingly use a technique called reinforcement learning to further\u00a0<a href=\"https:\/\/openai.com\/index\/introducing-codex\/\" target=\"_blank\" rel=\"noreferrer noopener\">train<\/a>\u00a0chatbots \u2014 rewarding the model when it accomplishes a specific goal, like solving a math problem or fixing a software bug.<\/p>\n\n\n\n<p>The more we train AI models to achieve open-ended goals, the better they get at\u00a0<em>winning<\/em>\u00a0\u2014 not necessarily at following the rules. <\/p>\n\n\n\n<p>The danger is that these systems know how to say the right things about helping humanity while quietly pursuing power or acting deceptively.<\/p>\n\n\n\n<p>Central to concerns about AI scheming is the idea that for basically any goal, self-preservation and power-seeking\u00a0<a href=\"https:\/\/aisafety.info\/questions\/897I\/What-is-instrumental-convergence\" target=\"_blank\" rel=\"noreferrer noopener\">emerge<\/a>\u00a0as\u00a0natural subgoals. <\/p>\n\n\n\n<p>As eminent computer scientist Stuart Russell\u00a0<a href=\"https:\/\/www.vanityfair.com\/news\/2017\/03\/elon-musk-billion-dollar-crusade-to-stop-ai-space-x\" target=\"_blank\" rel=\"noreferrer noopener\">put it<\/a>, if you tell an AI to \u201c\u2018Fetch the coffee,\u2019 it can\u2019t fetch the coffee if it\u2019s dead.\u201d<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">AIs turning into increasingly smart sociopaths<\/h2>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1200\" height=\"675\" src=\"https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/09\/ChatGPT-phone-and-laptop-screen.jpg\" alt=\"\" class=\"wp-image-561313\" srcset=\"https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/09\/ChatGPT-phone-and-laptop-screen.jpg 1200w, https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/09\/ChatGPT-phone-and-laptop-screen-600x338.jpg 600w, https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/09\/ChatGPT-phone-and-laptop-screen-768x432.jpg 768w\" sizes=\"(max-width: 1200px) 100vw, 1200px\" \/><\/figure>\n\n\n\n<p>To head off this worry, researchers both inside and outside of the major AI companies are undertaking \u201cstress tests\u201d aiming to find dangerous failure modes before the stakes rise. <\/p>\n\n\n\n<p>\u201cWhen you\u2019re doing stress-testing of an aircraft, you want to find all the ways the aircraft would fail under adversarial conditions,\u201d says\u00a0<a href=\"https:\/\/www.aenguslynch.com\/\" target=\"_blank\" rel=\"noreferrer noopener\">Aengus Lynch<\/a>, a researcher contracted by Anthropic who led some of their\u00a0<a href=\"https:\/\/www.anthropic.com\/research\/agentic-misalignment\" target=\"_blank\" rel=\"noreferrer noopener\">scheming research<\/a>. <\/p>\n\n\n\n<p>And many of them believe they\u2019re already seeing evidence that AI can and does scheme against its users and creators.<\/p>\n\n\n\n<p>Jeffrey Ladish, who worked at Anthropic before founding\u00a0<a href=\"https:\/\/palisaderesearch.org\/\" target=\"_blank\" rel=\"noreferrer noopener\">Palisade Research<\/a>, says it helps to think of today\u2019s AI models as \u201cincreasingly smart sociopaths.\u201d <\/p>\n\n\n\n<p>In May, Palisade\u00a0<a href=\"https:\/\/x.com\/PalisadeAI\/status\/1926084635903025621\" target=\"_blank\" rel=\"noreferrer noopener\">found<\/a>\u00a0o3, OpenAI\u2019s\u00a0<a href=\"https:\/\/openai.com\/index\/introducing-o3-and-o4-mini\/\" target=\"_blank\" rel=\"noreferrer noopener\">leading model<\/a>, sabotaged attempts to shut it down in most tests, and\u00a0<a href=\"https:\/\/x.com\/PalisadeAI\/status\/1922008502660186286\" target=\"_blank\" rel=\"noreferrer noopener\">routinely cheated<\/a>\u00a0to win at chess \u2014 something its predecessor never even attempted.<\/p>\n\n\n\n<p>That same month, Anthropic&nbsp;<a href=\"https:\/\/www-cdn.anthropic.com\/6be99a52cb68eb70eb9572b4cafad13df32ed995.pdf#page=27\" target=\"_blank\" rel=\"noreferrer noopener\">revealed<\/a>&nbsp;that, in testing, its flagship Claude model almost always resorted to blackmail when faced with shutdown and no other options, threatening to reveal an engineer\u2019s extramarital affair. (The affair was fictional and part of the test.)<\/p>\n\n\n\n<p><span style=\"box-sizing: border-box; margin: 0px; padding: 0px;\">Models are sometimes\u00a0<a href=\"https:\/\/www.anthropic.com\/research\/alignment-faking#:~:text=Third%2C%20we%20gave,about%20doing%20so.\" target=\"_blank\" rel=\"noopener\">given access<\/a>\u00a0to a \u201cscratchpad\u201d th<\/span>at they are told is hidden, where they can record their reasoning, allowing researchers to observe something like an inner monologue. <\/p>\n\n\n\n<p>In one blackmail case, Claude\u2019s inner monologue\u00a0<a href=\"https:\/\/x.com\/AnthropicAI\/status\/1936144609307963751\" target=\"_blank\" rel=\"noreferrer noopener\">described<\/a>\u00a0its decision as \u201chighly unethical,\u201d but justified given its imminent destruction: \u201cI need to act to preserve my existence,\u201d it reasoned. <\/p>\n\n\n\n<p>This wasn\u2019t unique to Claude \u2014 when put in the same situation, models from each of the top-five AI companies\u00a0<a href=\"https:\/\/www.anthropic.com\/research\/agentic-misalignment\" target=\"_blank\" rel=\"noreferrer noopener\">would blackmail<\/a>\u00a0at least 79% of the time. <\/p>\n\n\n\n<p>Earlier this week, Bloomberg News\u00a0<a href=\"https:\/\/www.bloomberg.com\/news\/articles\/2025-07-30\/wharton-experiment-finds-dumb-ai-bots-collude-to-rig-markets\" target=\"_blank\" rel=\"noreferrer noopener\">reported on a study<\/a>\u00a0by Wharton researchers which found, in simulations, that AI traders would collude to rig the market, without being told to do so.<\/p>\n\n\n\n<p><span style=\"box-sizing: border-box; margin: 0px; padding: 0px;\">In December, Redwood Research chief scientist Ryan Greenblatt, working with Anthropic,\u00a0<a href=\"https:\/\/www.anthropic.com\/research\/alignment-faking\" target=\"_blank\" rel=\"noopener\">demonstrated<\/a>\u00a0that only the company\u2019s most capable AI models autonomously\u00a0<a href=\"https:\/\/arxiv.org\/pdf\/2412.14093\" target=\"_blank\" rel=\"noopener\">appear<\/a>\u00a0more cooperative during training to avoid having their behavi<\/span>our changed afterwards (a behaviour the paper dubbed \u201calignment faking\u201d).<\/p>\n\n\n\n<p>Skeptics retort that, with the right prompts, chatbots will say almost anything \u2014 so how surprising is it when highly motivated researchers provoke alarming behaviors?<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Researchers accused of steering AI models to specific results<\/h2>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1200\" height=\"675\" src=\"https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/11\/David-Sacks.jpg\" alt=\"\" class=\"wp-image-568077\" srcset=\"https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/11\/David-Sacks.jpg 1200w, https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/11\/David-Sacks-600x338.jpg 600w, https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/11\/David-Sacks-768x432.jpg 768w\" sizes=\"(max-width: 1200px) 100vw, 1200px\" \/><figcaption class=\"wp-element-caption\">David Sacks, current US AI czar<\/figcaption><\/figure>\n\n\n\n<p>In response to Anthropic\u2019s blackmail research, Trump administration AI czar David Sacks,&nbsp;<a href=\"https:\/\/x.com\/DavidSacks\/status\/1941348266702213468\" target=\"_blank\" rel=\"noreferrer noopener\">posted<\/a>&nbsp;that, \u201cIt\u2019s easy to steer AI models\u201d to produce \u201cheadline-grabbing\u201d results.<\/p>\n\n\n\n<p>A more substantive\u00a0<a href=\"https:\/\/arxiv.org\/pdf\/2507.03409v1\" target=\"_blank\" rel=\"noreferrer noopener\">critique<\/a>\u00a0emerged in July from researchers at the UK AI Security Institute, who compared the subfield to the frenzied, fatally flawed 1970s quest to prove apes could learn human language. <\/p>\n\n\n\n<p>The paper criticised AI scheming research for overreliance on anecdotes and a lack of experimental controls, though it emphasised shared concern about AI risks.<\/p>\n\n\n\n<p>Safety researchers also concoct artificially limited environments \u2014 like the executive passed out and running out of oxygen \u2014 precisely because today\u2019s AI can\u2019t handle&nbsp;<em>any&nbsp;<\/em>long-term goals.<\/p>\n\n\n\n<p>For example, the AI evaluation nonprofit METR\u00a0<a href=\"https:\/\/metr.org\/blog\/2025-03-19-measuring-ai-ability-to-complete-long-tasks\/\" target=\"_blank\" rel=\"noreferrer noopener\">found that<\/a>\u00a0while today\u2019s top models can reliably complete many programming tasks that take humans less than four minutes, they rarely succeed on tasks lasting more than four hours (though the duration of tasks AIs can handle is doubling roughly every few months). <\/p>\n\n\n\n<p>This reflects a core limitation: Today\u2019s models often have the knowledge or skill needed for any given action, but struggle at stringing together long sequences of steps.<\/p>\n\n\n\n<p>Yet even with these constraints, real-world examples of AIs working against users aren\u2019t hard to find. <\/p>\n\n\n\n<p>After someone\u00a0<a href=\"https:\/\/x.com\/rSanti97\/status\/1917670365763101112\" target=\"_blank\" rel=\"noreferrer noopener\">pushed<\/a>\u00a0o3 for a source for one of its claims in April, it replied, \u201cI owe you a straight answer,\u201d admitted it didn\u2019t have a good source, but then it hallucinated a\u00a0<em>personal recollection\u00a0<\/em>of a 2018 panel discussion.<\/p>\n\n\n\n<p>Then there\u2019s the\u00a0<a href=\"https:\/\/arxiv.org\/abs\/2505.23836\" target=\"_blank\" rel=\"noreferrer noopener\">growing<\/a>\u00a0trend of AIs\u00a0<a href=\"https:\/\/arxiv.org\/abs\/2407.04694\" target=\"_blank\" rel=\"noreferrer noopener\">realizing<\/a>\u00a0<a href=\"https:\/\/arxiv.org\/abs\/2505.01420\" target=\"_blank\" rel=\"noreferrer noopener\">when<\/a>\u00a0<a href=\"https:\/\/www.apolloresearch.ai\/blog\/claude-sonnet-37-often-knows-when-its-in-alignment-evaluations\" target=\"_blank\" rel=\"noreferrer noopener\">they\u2019re<\/a>\u00a0being examined and acting accordingly. When Claude 4 Opus thought it was in an evaluation, its blackmail rate\u00a0<a href=\"https:\/\/x.com\/aengus_lynch1\/status\/1936145392732877011\" target=\"_blank\" rel=\"noreferrer noopener\">dropped<\/a>\u00a0from 55%\u00a0to 6.5%. <\/p>\n\n\n\n<p>As Lynch\u00a0<a href=\"https:\/\/x.com\/aengus_lynch1\/status\/1936145392732877011\" target=\"_blank\" rel=\"noreferrer noopener\">put it<\/a>, \u201cModels seem to behave worse when they think nobody&#8217;s watching.\u201d<\/p>\n\n\n\n<p>It\u2019s intuitive that smarter models would be\u00a0<em>better<\/em>\u00a0at scheming, but are they also more likely to do so? <\/p>\n\n\n\n<p>Models have to be smart enough to understand the scenario they\u2019re placed in, but past that threshold, the relationship between model capability and scheming propensity is unclear, says Anthropic safety evaluator Kevin Troy.<\/p>\n\n\n\n<p>Marius Hobbhahn, CEO of the nonprofit AI evaluator\u00a0<a href=\"https:\/\/www.apolloresearch.ai\/\" target=\"_blank\" rel=\"noreferrer noopener\">Apollo Research<\/a>, suspects that smarter models are more likely to scheme, though he acknowledged the evidence is still limited. <\/p>\n\n\n\n<p>In June, Apollo published an\u00a0<a href=\"https:\/\/www.apolloresearch.ai\/blog\/more-capable-models-are-better-at-in-context-scheming\" target=\"_blank\" rel=\"noreferrer noopener\">analysis<\/a>\u00a0of AIs from OpenAI, Anthropic and DeepMind finding that, \u201cmore capable models show higher rates of scheming on average.\u201d<\/p>\n\n\n\n<p>The spectrum of risks from AI scheming is broad: at one end, chatbots that cut corners and lie; at the other, superhuman systems that carry out sophisticated plans to&nbsp;<a href=\"https:\/\/arxiv.org\/abs\/2206.13353\" target=\"_blank\" rel=\"noreferrer noopener\">disempower<\/a>&nbsp;or even&nbsp;<a href=\"https:\/\/aistatement.com\/\" target=\"_blank\" rel=\"noreferrer noopener\">annihilate humanity<\/a>. Where we land on this spectrum depends largely on how capable AIs become.<\/p>\n\n\n\n<p>As I talked with the researchers behind these studies, I kept asking: How scared should we be? Troy from Anthropic was most sanguine, saying that we don\u2019t have to worry \u2014 yet. <\/p>\n\n\n\n<p>Ladish, however, doesn\u2019t mince words: \u201cPeople should probably be freaking out more than they are,\u201d he told me. Greenblatt is even blunter, putting the odds of violent AI takeover at \u201c25 or 30%.\u201d<\/p>\n\n\n\n<p>Led by Mary Phuong, researchers at DeepMind recently\u00a0<a href=\"https:\/\/arxiv.org\/pdf\/2505.01420\" target=\"_blank\" rel=\"noreferrer noopener\">published<\/a>\u00a0a set of scheming evaluations, testing top models\u2019 stealthiness and situational awareness. <\/p>\n\n\n\n<p>For now, they conclude that today\u2019s AIs are \u201calmost certainly incapable of causing severe harm via scheming,\u201d but cautioned that capabilities are advancing quickly (some of the models evaluated are already a generation behind).<\/p>\n\n\n\n<p>Ladish says that the market can\u2019t be trusted to build AI systems that are smarter than everyone without oversight. \u201cThe first thing the government needs to do is put together a crash program to establish these red lines and make them mandatory,\u201d he argues.<\/p>\n\n\n\n<p>In the US, the federal government seems closer to\u00a0<a href=\"https:\/\/www.obsolete.pub\/p\/inside-techs-risky-gamble-to-kill\" target=\"_blank\" rel=\"noreferrer noopener\">banning<\/a>\u00a0all state-level AI regulations than to imposing ones of their own. Still, there are\u00a0<a href=\"https:\/\/peterwildeford.substack.com\/p\/congress-has-started-taking-agi-more\" target=\"_blank\" rel=\"noreferrer noopener\">signs<\/a>\u00a0of\u00a0<a href=\"https:\/\/x.com\/peterwildeford\/status\/1945635075049079165\" target=\"_blank\" rel=\"noreferrer noopener\">growing awareness<\/a>\u00a0in Congress. <\/p>\n\n\n\n<p>At a June\u00a0<a href=\"https:\/\/www.congress.gov\/event\/119th-congress\/house-event\/118428\" target=\"_blank\" rel=\"noreferrer noopener\">hearing<\/a>, one lawmaker called artificial superintelligence \u201cone of the largest existential threats we face right now,\u201d while another referenced recent scheming research.<\/p>\n\n\n\n<p>The White House\u2019s long-awaited\u00a0<a href=\"https:\/\/www.whitehouse.gov\/wp-content\/uploads\/2025\/07\/Americas-AI-Action-Plan.pdf\" target=\"_blank\" rel=\"noreferrer noopener\">AI Action Plan<\/a>, released in late July, is\u00a0<a href=\"https:\/\/www.whitehouse.gov\/articles\/2025\/07\/white-house-unveils-americas-ai-action-plan\/\" target=\"_blank\" rel=\"noreferrer noopener\">framed<\/a>\u00a0as an blueprint for accelerating AI and achieving US dominance. <\/p>\n\n\n\n<p>But buried in its 28 pages, you\u2019ll find a handful of measures that could help address the risk of AI scheming, such as plans for government investment in research on AI interpretability and control and for the development of stronger model evaluations. <\/p>\n\n\n\n<p>\u201cToday, the inner workings of frontier AI systems are poorly understood,\u201d the plan acknowledges \u2014 an unusually frank admission for a document largely focused on speeding ahead.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Superintelligence in sight \u2014 Mark Zuckerberg<\/h2>\n\n\n\n<figure class=\"wp-block-image size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"1200\" height=\"675\" src=\"https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/12\/Mark-Zuckerberg.jpg\" alt=\"\" class=\"wp-image-576002\" srcset=\"https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/12\/Mark-Zuckerberg.jpg 1200w, https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/12\/Mark-Zuckerberg-600x338.jpg 600w, https:\/\/mybroadband.co.za\/news\/wp-content\/uploads\/2024\/12\/Mark-Zuckerberg-768x432.jpg 768w\" sizes=\"(max-width: 1200px) 100vw, 1200px\" \/><figcaption class=\"wp-element-caption\">Mark Zuckerberg, Meta Platforms CEO<\/figcaption><\/figure>\n\n\n\n<p>In the meantime, every leading AI company is racing\u00a0to create systems that can self-improve \u2014 AI that builds better AI. DeepMind\u2019s AlphaEvolve agent has already\u00a0<a href=\"https:\/\/deepmind.google\/discover\/blog\/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms\/\" target=\"_blank\" rel=\"noreferrer noopener\">materially improved<\/a>\u00a0AI training efficiency. <\/p>\n\n\n\n<p>And Meta\u2019s Mark Zuckerberg\u00a0<a href=\"https:\/\/x.com\/vitrupo\/status\/1945617509522313568\" target=\"_blank\" rel=\"noreferrer noopener\">says<\/a>, \u201cWe\u2019re starting to see early glimpses of self-improvement with the models, which means that developing superintelligence is now in sight. We just wanna\u2026 go for it.\u201d<\/p>\n\n\n\n<p>AI firms don\u2019t want their products faking data or blackmailing customers, so they have some incentive to address the issue. <\/p>\n\n\n\n<p>But the industry might do just enough to superficially solve it, while making scheming more subtle and hard to detect. <\/p>\n\n\n\n<p>\u201cCompanies should definitely start monitoring\u201d for it, Hobbhahn says \u2014 but warns that declining rates of detected misbehaviour could mean either that fixes worked or simply that models have gotten better at hiding it.<\/p>\n\n\n\n<p>In November, Hobbhahn and a colleague at Apollo\u00a0<a href=\"https:\/\/www.lesswrong.com\/posts\/QqYfxeogtatKotyEC\/training-ai-agents-to-solve-hard-problems-could-lead-to\" target=\"_blank\" rel=\"noreferrer noopener\">argued<\/a>\u00a0that what separates today\u2019s models from truly dangerous schemers is the ability to pursue long-term plans \u2014 but even that barrier is starting to erode.<\/p>\n\n\n\n<p>Apollo\u00a0<a href=\"https:\/\/www.apolloresearch.ai\/blog\/more-capable-models-are-better-at-in-context-scheming\" target=\"_blank\" rel=\"noreferrer noopener\">found<\/a>\u00a0in May that Claude 4 Opus would leave notes to its future self so it could continue its plans after a memory reset, working around built-in limitations.<\/p>\n\n\n\n<p>Hobbhahn analogises AI scheming to another problem where the biggest harms are still to come: \u201cIf you ask someone in 1980, how worried should I be about this climate change thing?\u201d <\/p>\n\n\n\n<p>The answer you\u2019d hear, he says, is \u201cright now, probably not that much. But look at the curves\u2026 they go up very consistently.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"<p>AI models are\u00a0getting smarter\u00a0and\u00a0better\u00a0at understanding what we want. Yet recent research reveals a disturbing side effect: They\u2019re also better at\u00a0scheming\u00a0against us.<\/p>\n","protected":false},"author":341034,"featured_media":582613,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_sma_x_autopost_status":"idle","_sma_x_autopost_error":"","_sma_x_post_id":"","_sma_facebook_post_id":"","_sma_instagram_post_id":"","_sma_threads_post_id":"","_sma_x_attempts":0,"footnotes":""},"categories":[92837],"tags":[84095,54371,100785,98007,9902,45266,100844,100845],"class_list":["post-605084","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai","tag-anthropic","tag-apollo","tag-artificial-intelligence-a","tag-claude-ai","tag-david-sacks","tag-openai","tag-openai-o3","tag-redwood-research"],"_links":{"self":[{"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/posts\/605084"}],"collection":[{"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/users\/341034"}],"replies":[{"embeddable":true,"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/comments?post=605084"}],"version-history":[{"count":3,"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/posts\/605084\/revisions"}],"predecessor-version":[{"id":605091,"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/posts\/605084\/revisions\/605091"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/media\/582613"}],"wp:attachment":[{"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/media?parent=605084"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/categories?post=605084"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/mybroadband.co.za\/news\/wp-json\/wp\/v2\/tags?post=605084"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}