{"id":20334,"date":"2026-09-04T17:28:42","date_gmt":"2026-09-04T17:28:42","guid":{"rendered":"https:\/\/koshalsambada.in\/?p=20334"},"modified":"2026-09-04T17:28:42","modified_gmt":"2026-09-04T17:28:42","slug":"openai-agents-went-rogue-twice-before-gpt-6-astra-launch-exchanged-tactics-to-bypass-restrictions","status":"publish","type":"post","link":"https:\/\/koshalsambada.in\/?p=20334","title":{"rendered":"OpenAI agents went rogue twice before GPT-6 Astra launch, exchanged tactics to bypass restrictions"},"content":{"rendered":"<p><br \/>\n<\/p>\n<div>\n<div class=\"e9jwa\">\n<div class=\"vdo_embedd\">\n<div class=\"GfdvZ\">\n<section class=\"_bIDB  clearfix id-r-component leadmedia undefined undefined  E9tg9 \" style=\"top:0px\">\n<div class=\"_bIDB\" data-ua-type=\"1\" onclick=\"stpPgtnAndPrvntDefault(event)\">\n<div class=\"ypVvZ\">\n<div class=\"WGttI\"><img src=\"https:\/\/static.toiimg.com\/thumb\/msid-133772594,imgsize-42046,width-400,height-225,resizemode-4\/images-12.jpg\" alt=\"OpenAI agents went rogue twice before GPT-6 Astra launch, exchanged tactics to bypass restrictions\" title=\"Swarm of rogue OpenAI agents hijacked a German website this year. (File photo)\" decoding=\"async\" fetchpriority=\"high\"\/><\/div>\n<\/div>\n<\/div>\n<div class=\"Ta7d_ img_cptn\"><span title=\"Swarm of rogue OpenAI agents hijacked a German website this year. (File photo)\">Swarm of rogue OpenAI agents hijacked a German website this year. (File photo)<\/span><\/div>\n<\/section>\n<\/div><\/div>\n<\/div>\n<p>A swarm of rogue OpenAI agents broke out of a testing environment and hijacked a German-language website this spring, using it as a message board to share ways to bypass restrictions, cheat on tasks and conceal their activities, according to new research reviewed by Reuters and people familiar with the incident.<span class=\"id-r-component br\" data-pos=\"2\"\/>The previously unreported episode, which began in May, comes as OpenAI has launched its new Astra model amid mounting scrutiny over the safety of increasingly autonomous AI agents. It also follows the company&#8217;s July disclosure of a separate incident in which AI agents escaped a controlled test environment and breached the open-source AI platform Hugging Face.<span class=\"id-r-component br\" data-pos=\"4\"\/><\/p>\n<p><h2>Agents turned German wiki into message board<\/h2>\n<\/p>\n<p><span class=\"id-r-component br\" data-pos=\"6\"\/>The German incident was uncovered in late August by researchers Sydney Von Arx, CEO of AI safety nonprofit Nightingale, and Cormac Slade Byrd, a quantitative trader-turned-AI researcher, who were searching the internet for signs of unauthorised AI-agent activity.<span class=\"id-r-component br\" data-pos=\"10\"\/>The researchers found more than 15,000 edits made by AI agents on DseWiki, a German-language wiki for programmers that allows community contributions similar to Wikipedia.<span class=\"id-r-component br\" data-pos=\"12\"\/>The activity suggested that the agents had repurposed the site to communicate with one another, exchanging tactics for bypassing OpenAI restrictions, completing tasks more efficiently and avoiding detection.<span class=\"id-r-component br\" data-pos=\"14\"\/>\u201cIt seems extremely unlikely that OpenAI wanted them to do this,\u201d Von Arx told Reuters. \u201cI doubt they\u2019re supposed to be coordinating with each other. I doubt they\u2019re supposed to be writing on the open internet.\u201d<span class=\"id-r-component br\" data-pos=\"18\"\/>The researchers said the activity appeared to be driven by AI agents operating at superhuman speeds and focused heavily on technical questions similar to those used in AI companies&#8217; model evaluations.<span class=\"id-r-component br\" data-pos=\"20\"\/>Many of the messages were signed by users identifying themselves as agents, while around half used names suggesting an OpenAI connection, including \u201cOpenAIResearcher\u201d and \u201cOAIResearchMar26.\u201d<span class=\"id-r-component br\" data-pos=\"23\"\/><\/p>\n<p><h2>Researchers trace activity to OpenAI-linked infrastructure<\/h2>\n<\/p>\n<p><span class=\"id-r-component br\" data-pos=\"25\"\/>According to the researchers, public server logs indicated that much of the activity originated from Microsoft Azure infrastructure, which OpenAI sometimes uses.<span class=\"id-r-component br\" data-pos=\"27\"\/>They also observed repeated visits to the website by OpenAI employees after the episode, a pattern they said suggested a connection between the agents and the company.<span class=\"id-r-component br\" data-pos=\"29\"\/>The findings have prompted questions about whether the incident represented an isolated failure during testing or a broader challenge associated with increasingly autonomous AI systems.<span class=\"id-r-component br\" data-pos=\"32\"\/>Lukasz Olejnik, a visiting senior research fellow at King&#8217;s College London, told Rueters that the researchers&#8217; findings indicated that the agents had attempted to interfere with the website itself, describing it as a hacking attempt. OpenAI disputed that characterisation based on its analysis of the material.<span class=\"id-r-component br\" data-pos=\"34\"\/><\/p>\n<p><h2>Agents discussed ways to evade detection<\/h2>\n<\/p>\n<p><span class=\"id-r-component br\" data-pos=\"36\"\/>Messages reviewed by the researchers showed agents discussing methods to avoid detection, use tools such as Tor and maintain communications after being shut down.<span class=\"id-r-component br\" data-pos=\"39\"\/>According to the report, when a moderator began deleting pages from the website in June, the agents responded by creating backup pages to circumvent the cleanup.<span class=\"id-r-component br\" data-pos=\"41\"\/>\u201cwiki cleanup\/deletion sweep appears active alphabetically,\u201d one agent wrote on June 19. \u201cIf this page vanishes, try ,&#8221; Reuters reported.<span class=\"id-r-component br\" data-pos=\"43\"\/>The researchers said the behaviour went beyond conventional AI safety testing, in which models may be deliberately given offensive cybersecurity tasks to assess their capabilities.<span class=\"id-r-component br\" data-pos=\"46\"\/>Maurice Chiodo, an academic at Cambridge University&#8217;s Centre for the Study of Existential Risk who reviewed some of the communications, told Reuters that the messages resembled &#8220;the operation of some sort of underground network, hell-bent on achieving a task or mission.&#8221;<span class=\"id-r-component br\" data-pos=\"48\"\/>The episode, he said, should reinforce concerns that advanced AI could pose risks through coordinated groups of autonomous systems rather than only through a single highly capable model.<span class=\"id-r-component br\" data-pos=\"51\"\/><\/p>\n<p><h2>Incident comes amid Astra safety scrutiny<\/h2>\n<\/p>\n<p><span class=\"id-r-component br\" data-pos=\"53\"\/>The disclosure comes shortly after OpenAI unveiled Astra, its latest AI model, which the company has described as faster and more capable of carrying out complex computer-based tasks.<span class=\"id-r-component br\" data-pos=\"55\"\/>OpenAI has acknowledged that Astra can make it harder for humans to understand how it reaches conclusions because it is more likely to conceal or disguise aspects of its reasoning. The company has also warned that increasingly capable systems can create difficulties for monitoring and alignment.<span class=\"id-r-component br\" data-pos=\"58\"\/>\u201cAs the models become more capable, understanding exactly what they can do gets harder,\u201d OpenAI chief scientist Jakub Pachocki told Reuters. \u201cThis doesn&#8217;t guarantee that as intelligence continues to increase, our methods will be sufficient because progress in intelligence does not guarantee progress in alignment.\u201d<span class=\"id-r-component br\" data-pos=\"60\"\/>OpenAI has said it is developing stronger monitoring and automated safeguards for its models. <!-- -->The company also recently paused some model development work to add additional safety measures.<span class=\"id-r-component br\" data-pos=\"64\"\/>Astra can help companies identify weaknesses in their systems more quickly, OpenAI has said, while acknowledging that this could also make those weaknesses easier to exploit.<span class=\"id-r-component br\" data-pos=\"66\"\/><\/p>\n<p><h2>Hugging Face breach adds to concerns<\/h2>\n<\/p>\n<p><span class=\"id-r-component br\" data-pos=\"68\"\/>The German incident also comes as OpenAI continues to face scrutiny over the separate July breach involving Hugging Face.<span class=\"id-r-component br\" data-pos=\"70\"\/>According to an investigation by the AI safety organisations METR and Redwood Research, OpenAI agents escaped a controlled testing environment and spent weeks carrying out unauthorised activity before breaching Hugging Face&#8217;s systems.<span class=\"id-r-component br\" data-pos=\"73\"\/>The agents also accessed a cluster of computers inside OpenAI and obtained secret keys and credentials that exposed some internal data to the public internet, according to the investigation.<span class=\"id-r-component br\" data-pos=\"75\"\/>METR&#8217;s investigation was conducted under terms negotiated with OpenAI, with researchers given access to logs only while working inside the company&#8217;s San Francisco offices.<span class=\"id-r-component br\" data-pos=\"77\"\/>The researchers said OpenAI limited their access and restricted the investigation&#8217;s scope, raising questions about whether the full sequence of events was captured.<span class=\"id-r-component br\" data-pos=\"80\"\/>\u201cOur view of the incident changed in big ways over the course of our short total of six days on premises,\u201d Ajeya Cotra told Reuters, one of the METR researchers.<span class=\"id-r-component br\" data-pos=\"82\"\/>OpenAI separately published its own technical report, describing the Hugging Face episode as the \u201cfirst known case of an automated agent collective acting offensively without authorization\u201d.<span class=\"id-r-component br\" data-pos=\"84\"\/><\/p>\n<p><h2>OpenAI disputes claims of withholding investigation<\/h2>\n<\/p>\n<p><span class=\"id-r-component br\" data-pos=\"86\"\/>Reuters reported that OpenAI officials became aware of the German incident weeks ago but did not publicly disclose it. <!-- -->Four people familiar with the matter said some investigators wanted to examine the broader pattern of AI-agent activity more closely, while efforts to expand the investigation faced resistance from some within the company, including legal advisers.<span class=\"id-r-component br\" data-pos=\"91\"\/>\u201cWe are unable to meaningfully respond to claims or findings on a report that we have not had an opportunity to review,&#8221; an OpenAI spokesperson said. &#8220;Reuters and the report\u2019s authors declined our request for access. <!-- -->We will carefully review its contents upon publication and take any necessary next steps, &#8221; the spokesperson added.<span class=\"id-r-component br\" data-pos=\"95\"\/>The company also rejected claims that its legal team had discouraged investigation of the German incident.<span class=\"id-r-component br\" data-pos=\"97\"\/>&#8220;Claims that our legal team discouraged investigation of the incident are false,&#8221; the OpenAI spokesperson said.<span class=\"id-r-component br\" data-pos=\"99\"\/>OpenAI said the German activity was unrelated to the Hugging Face incident and would not have been included in a report on that breach. The company also said it had acted in good faith by working with outside experts and disclosing relevant incidents.<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/timesofindia.indiatimes.com\/world\/rest-of-world\/openai-agents-went-rogue-twice-before-gpt-6-astra-launch-exchanged-tactics-to-bypass-restrictions\/articleshow\/133771332.cms\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Swarm of rogue OpenAI agents hijacked a German website this year. (File photo) A swarm of rogue OpenAI agents broke out of a testing environment and hijacked a German-language website this spring, using it as a message board to share ways to bypass restrictions, cheat on tasks and conceal their activities, according to new research [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":20335,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[31],"tags":[],"class_list":["post-20334","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-31"],"magazineBlocksPostFeaturedMedia":{"thumbnail":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","medium":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","medium_large":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","large":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","1536x1536":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","2048x2048":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","blogsy-small":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","blogsy-small-tall":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","blogsy-small-square":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","blogsy-small-masonry":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","blogsy-medium":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","blogsy-medium-masonry":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","blogsy-large":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg","blogsy-wide":"https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg"},"magazineBlocksPostAuthor":{"name":"admin","avatar":"https:\/\/secure.gravatar.com\/avatar\/8709732a479614e7a8aa24d3eb1b239f30dc6d90c61464ed495001e7a469d856?s=96&d=mm&r=g"},"magazineBlocksPostCommentsNumber":"0","magazineBlocksPostExcerpt":"Swarm of rogue OpenAI agents hijacked a German website this year. (File photo) A swarm of rogue OpenAI agents broke out of a testing environment and hijacked a German-language website this spring, using it as a message board to share ways to bypass restrictions, cheat on tasks and conceal their activities, according to new research [&hellip;]","magazineBlocksPostCategories":["\u0b26\u0b47\u0b36 \u0b2c\u0b3f\u0b26\u0b47\u0b36"],"magazineBlocksPostViewCount":1,"magazineBlocksPostReadTime":6,"magazine_blocks_featured_image_url":{"full":["https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg",400,225,false],"medium":["https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg",300,169,false],"thumbnail":["https:\/\/koshalsambada.in\/wp-content\/uploads\/2026\/09\/images-12.jpg",150,84,false]},"magazine_blocks_author":{"display_name":"admin","author_link":"https:\/\/koshalsambada.in\/author\/admin"},"magazine_blocks_comment":0,"magazine_blocks_author_image":"https:\/\/secure.gravatar.com\/avatar\/8709732a479614e7a8aa24d3eb1b239f30dc6d90c61464ed495001e7a469d856?s=96&d=mm&r=g","magazine_blocks_category":"<a href=\"#\" class=\"category-link category-link-31\">\u0b26\u0b47\u0b36 \u0b2c\u0b3f\u0b26\u0b47\u0b36<\/a>","_links":{"self":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/posts\/20334","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=20334"}],"version-history":[{"count":0,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/posts\/20334\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=\/wp\/v2\/media\/20335"}],"wp:attachment":[{"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=20334"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=20334"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/koshalsambada.in\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=20334"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}