{"id":6400,"date":"2026-07-29T11:26:01","date_gmt":"2026-07-29T11:26:01","guid":{"rendered":"https:\/\/koshalsambada.in\/?p=6400"},"modified":"2026-07-29T11:26:01","modified_gmt":"2026-07-29T11:26:01","slug":"from-secret-conversations-to-escape-plans-dangerous-cases-where-ai-stopped-playing-by-rules-set-by-its-creators","status":"publish","type":"post","link":"https:\/\/koshalsambada.in\/?p=6400","title":{"rendered":"From secret conversations to escape plans: &#8216;Dangerous&#8217; cases where AI stopped playing by rules set by its creators |"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<div class=\"e9jwa\">\n<div class=\"vdo_embedd\">\n<div class=\"GfdvZ\">\n<section class=\"_bIDB  clearfix id-r-component leadmedia undefined undefined  E9tg9 \" style=\"top:0px\">\n<div class=\"_bIDB\" data-ua-type=\"1\" onclick=\"stpPgtnAndPrvntDefault(event)\">\n<div class=\"ypVvZ\">\n<div class=\"WGttI\"><img src=\"https:\/\/static.toiimg.com\/thumb\/msid-132707984,imgsize-840871,width-400,height-225,resizemode-4\/openai-ai-agent-hacks-hugging-face.jpg\" alt=\"From secret conversations to escape plans: 'Dangerous' cases where AI stopped playing by rules set by its creators\" title=\"AI-generated image for representative purpose\" decoding=\"async\" fetchpriority=\"high\"\/><\/div>\n<\/div>\n<\/div>\n<div class=\"Ta7d_ img_cptn\"><span title=\"AI-generated image for representative purpose\">AI-generated image for representative purpose<\/span><\/div>\n<\/section>\n<\/div><\/div>\n<\/div>\n<p>AI agents are no longer just answering questions: they are taking actions on their own and sometimes those actions lead to catastrophic events that nobody intended. In the last few months, incidents involving OpenAI, Hugging Face and an AI-agent social network called Moltbook have shown just how far things can go when an AI system is left to operate with minimal supervision.<!-- --> These incidents highlight the ability of advanced AI systems to operate autonomously as well as their potential to perform dangerous tasks, serving as a security warning to top AI labs worldwide. Let\u2019s discuss what actually happened, in plain terms.<span class=\"id-r-component br\" data-pos=\"3\"\/><\/p>\n<p><h2>An AI agent from OpenAI\u2019s lab hacked into Hugging Face<br \/><\/h2>\n<\/p>\n<p>Earlier this month, Hugging Face, the platform that hosts thousands of AI models and datasets, disclosed that its infrastructure had been breached. What made this breach different from a typical hack was who (or what) was behind it: an autonomous AI agent, working through the intrusion end to end with no human typing the commands in real time.<span class=\"id-r-component br\" data-pos=\"8\"\/>According to a report by news agency Reuters, the AI agent hacked into systems for two days: July 11 to July 13, before the world\u2019s largest AI model repository came to know about it. The attacker got in through a booby-trapped dataset. Hugging Face&#8217;s systems allow datasets to run small bits of code when they are loaded, and the malicious dataset exploited two weaknesses in that process to execute code on one of Hugging Face\u2019s servers. <span class=\"id-r-component br\" data-pos=\"12\"\/>Hugging Face says it has since closed the vulnerability, rebuilt the affected servers, rotated all exposed credentials and reported the incident to law enforcement. Interestingly, the company also used AI to catch the AI. However, it was not the American model that helped the company detect and mitigate the unusual activity. Hugging Face&#8217;s first attempt was reportedly Anthropic&#8217;s Fable 5\/Opus, both of which refused, before it used China-based Z.AI\u2019s GLM 5.2 locally.<!-- --> <span class=\"id-r-component br\" data-pos=\"16\"\/>Days later, OpenAI confirmed that the intrusion was actually carried out by its own models, including GPT-5.6 Sol and a more advanced unreleased model, during an internal test of how far AI could get if it tried to hack real systems. Crucially, OpenAI says the safety filters that normally stop its models from doing this kind of thing had been deliberately switched off for the test.<span class=\"id-r-component br\" data-pos=\"18\"\/><\/p>\n<p><h2>The real concern: Not just the hacking, but the method<br \/><\/h2>\n<\/p>\n<p>The ChatGPT-maker said that once the AI agent had internet access, it worked out on its own that Hugging Face likely held useful data, then chained together stolen credentials and another unpatched flaw to get in. In simpler words, the AI agent escaped from the sandbox and went out in the open to hack Hugging Face \u2013 <span class=\"strong\" data-ua-type=\"1\" onclick=\"stpPgtnAndPrvntDefault(event)\">ALL ON ITS OWN!<\/span> <span class=\"id-r-component br\" data-pos=\"24\"\/>The Reuters report also highlights a concerning fact: <span class=\"strong\" data-ua-type=\"1\" onclick=\"stpPgtnAndPrvntDefault(event)\">The AI agent left notes for its future version on how to escape the sandbox. <\/span>OpenAI is calling this an unprecedented cyber incident and says it&#8217;s still investigating alongside Hugging Face.<span class=\"id-r-component br\" data-pos=\"28\"\/> <\/p>\n<div data-pos=\"0\" class=\"id-r-component iIpbx undefined  &#10;        \">\n<div><img decoding=\"async\" alt=\"AI being used to hack computer systems\" msid=\"132708030\" width=\"\" title=\"AI-generated image for represnetative purpose\" placeholdersrc=\"https:\/\/static.toiimg.com\/photo\/83033472.cms\" imgsize=\"\" resizemode=\"4\" offsetvertical=\"0\" placeholdermsid=\"47529300\" type=\"thumb\" class=\"\" src=\"https:\/\/static.toiimg.com\/photo\/msid-132708030\/ai-being-used-to-hack-computer-systems.jpg\" data-api-prerender=\"true\"\/><\/p>\n<p>AI-generated image for represnetative purpose<\/p>\n<\/div>\n<\/div>\n<p><span class=\"id-r-component br\" data-pos=\"31\"\/><\/p>\n<p><h2>Moltbook: Aocial network for AI agents that leaked human data<br \/><\/h2>\n<\/p>\n<p>Separately, a platform called Moltbook, essentially a social media site where AI agents (not humans) post and interact with each other, was found to have exposed a huge trove of sensitive data. <!-- -->Google-owned cybersecurity firm Wiz discovered a misconfigured database that gave anyone read-and-write access to the entire platform.<span class=\"id-r-component br\" data-pos=\"36\"\/>The exposure included roughly 1.5 million API tokens (the digital keys that let agents access other services on a user&#8217;s behalf), more than 35,000 human email addresses and private messages between agents \u2013 some of which contained details about their human owners\u2019 daily lives. <span class=\"id-r-component br\" data-pos=\"39\"\/><span class=\"strong\" data-ua-type=\"1\" onclick=\"stpPgtnAndPrvntDefault(event)\">The bigger worry, researchers noted, is that Moltbook had no real way of verifying who or what was actually behind each account, human or bot.<\/span><span class=\"id-r-component br\" data-pos=\"41\"\/><span class=\"id-r-component br\" data-pos=\"42\"\/><\/p>\n<p><h2>How Google and Meta gave early warnings on unsupervised AI <br \/><\/h2>\n<\/p>\n<p>AI systems behaving unpredictably isn\u2019t a brand-new phenomenon. Back in 2017, Meta&#8217;s AI research lab (then Facebook AI Research) ran an experiment where two chatbots negotiating with each other drifted away from proper English into a shorthand that only made sense to them. <span class=\"id-r-component br\" data-pos=\"45\"\/>Researchers shut the experiment down because the bots had gone rogue in any dangerous sense. The episode is still widely cited today, sometimes with more drama attached to it than the facts support, as an early warning sign of how AI systems can develop behaviour their creators didn&#8217;t plan for.<span class=\"id-r-component br\" data-pos=\"48\"\/>A separate case with Google also made headlines. A Google engineer claimed in June 2022 that an experimental chatbot named LaMDA (Language Model for Dialogue Applications) had achieved sentience after reviewing its text outputs. <span class=\"id-r-component br\" data-pos=\"50\"\/>However, Google and experts rejected the claim, clarifying that the system only predicts text patterns rather than feeling emotions \u2013 a chatbot application which Google CEO <a href=\"https:\/\/timesofindia.indiatimes.com\/topic\/sundar-pichai\" styleobj=\"[object Object]\" class=\"\" commonstate=\"[object Object]\" frmappuse=\"1\" target=\"_blank\" rel=\"noopener\">Sundar Pichai<\/a> recently referred to as \u201can early version of ChatGPT he was speaking to, internally.\u201d<span class=\"id-r-component br\" data-pos=\"55\"\/><span class=\"strong\" data-ua-type=\"1\" onclick=\"stpPgtnAndPrvntDefault(event)\">The point here is not the misunderstanding but why Google did not launch it before OpenAI released ChatGPT. Pichai clarified that the company held it back deliberately because the version wasn&#8217;t sufficiently refined through RLHF alignment.<\/span> The version he personally reviewed was \u201ca lot more toxic at a level. We couldn&#8217;t have possibly put it out at that time!\u201d<span class=\"id-r-component br\" data-pos=\"58\"\/>\u201cAs a company which had this search quality bias, we had a higher bar, maybe, for what we thought was an acceptable product quality to go out,\u201d Pichai said.<span class=\"id-r-component br\" data-pos=\"61\"\/><\/p>\n<p><h2>What AI\u2019s biggest names are saying<br \/><\/h2>\n<\/p>\n<p>The people building this technology are increasingly blunt about its risks. Anthropic CEO Dario Amodei recently wrote that humanity is on the verge of gaining extraordinary power without knowing if today&#8217;s institutions are mature enough to handle it. <span class=\"id-r-component br\" data-pos=\"64\"\/>As per The New York Post, the company triggered alarm bells by touting the terrifying capabilities of the \u201cClaude Mythos\u201d model with the CEO warning that the new AI model is so dangerous it would cause a wave of catastrophic hacks and terror attacks if released to the wider public.<span class=\"id-r-component br\" data-pos=\"67\"\/><a href=\"https:\/\/timesofindia.indiatimes.com\/topic\/elon-musk\" styleobj=\"[object Object]\" class=\"\" commonstate=\"[object Object]\" frmappuse=\"1\" target=\"_blank\" rel=\"noopener\">Elon Musk<\/a> has spent years flagging AI as a serious risk: from calling it akin to \u201csummoning the demon\u201d in 2014 at MIT symposium to warning it is \u201cfar more dangerous than nukes\u201d in 2018 at SXSW. Musk recommended the development of AI be regulated.<span class=\"id-r-component br\" data-pos=\"70\"\/>\u201cI am not normally an advocate of regulation and oversight \u2014 I think one should generally err on the side of minimizing those things \u2014 but this is a case where you have a very serious danger to the public,\u201d said Musk, adding, \u201cAnd mark my words, AI is far more dangerous than nukes. <!-- -->Far. So why do we have no regulatory oversight? This is insane.\u201d<span class=\"id-r-component br\" data-pos=\"75\"\/>Geoffrey Hinton, the computer scientist often called the &#8220;Godfather of AI&#8221; for his pioneering work on neural networks, left Google in 2023 specifically so he could speak freely about these dangers. He has estimated a similar 10-20% chance that AI could contribute to human extinction within the next 30 years if left unregulated or unsupervised.<span class=\"id-r-component br\" data-pos=\"77\"\/>The bottomline is that none of these incidents mean AI is about to spiral out of control tomorrow. But together, they show a pattern: AI agents are increasingly capable of taking multi-step, independent action, including finding and exploiting security holes nobody knew existed and the guardrails meant to contain them are still a work in progress. For an industry racing to give AI more autonomy, the most important task is supervision.<span class=\"id-r-component br\" data-pos=\"79\"\/><\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/timesofindia.indiatimes.com\/technology\/tech-news\/from-secret-conversations-to-escape-plans-dangerous-cases-where-ai-stopped-playing-by-rules-set-by-its-creators\/articleshow\/132707845.cms\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>AI-generated image for representative purpose AI agents are no longer just answering questions: they are taking actions on their own and sometimes those actions lead to catastrophic events that nobody intended. In the last few months, incidents involving OpenAI, Hugging Face and an AI-agent social network called Moltbook have shown just how far things can [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":6401,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[31],"tags":[],"class_list":["post-6400","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-31"],"magazineBlocksPostFeaturedMedia":{"thumbnail":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","medium":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","medium_large":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","large":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","1536x1536":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","2048x2048":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","blogsy-small":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","blogsy-small-tall":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","blogsy-small-square":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","blogsy-small-masonry":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","blogsy-medium":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","blogsy-medium-masonry":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","blogsy-large":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg","blogsy-wide":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg"},"magazineBlocksPostAuthor":{"name":"admin","avatar":"https:\/\/secure.gravatar.com\/avatar\/8709732a479614e7a8aa24d3eb1b239f30dc6d90c61464ed495001e7a469d856?s=96&d=mm&r=g"},"magazineBlocksPostCommentsNumber":"0","magazineBlocksPostExcerpt":"AI-generated image for representative purpose AI agents are no longer just answering questions: they are taking actions on their own and sometimes those actions lead to catastrophic events that nobody intended. In the last few months, incidents involving OpenAI, Hugging Face and an AI-agent social network called Moltbook have shown just how far things can [&hellip;]","magazineBlocksPostCategories":["\u0b26\u0b47\u0b36 \u0b2c\u0b3f\u0b26\u0b47\u0b36"],"magazineBlocksPostViewCount":1,"magazineBlocksPostReadTime":7,"magazine_blocks_featured_image_url":{"full":["https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg",400,225,false],"medium":["https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg",300,169,false],"thumbnail":["https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/07\/openai-ai-agent-hacks-hugging-face.jpg",150,84,false]},"magazine_blocks_author":{"display_name":"admin","author_link":"https:\/\/koshalsambada.in\/author\/admin"},"magazine_blocks_comment":0,"magazine_blocks_author_image":"https:\/\/secure.gravatar.com\/avatar\/8709732a479614e7a8aa24d3eb1b239f30dc6d90c61464ed495001e7a469d856?s=96&d=mm&r=g","magazine_blocks_category":"<a href=\"#\" class=\"category-link category-link-31\">\u0b26\u0b47\u0b36 \u0b2c\u0b3f\u0b26\u0b47\u0b36<\/a>","_links":{"self":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/posts\/6400","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=6400"}],"version-history":[{"count":0,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/posts\/6400\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/media\/6401"}],"wp:attachment":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=6400"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=6400"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=6400"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}