{"id":639474,"date":"2026-07-22T12:05:00","date_gmt":"2026-07-22T12:05:00","guid":{"rendered":"https:\/\/buglecall.org\/?p=639474"},"modified":"2026-07-22T12:05:00","modified_gmt":"2026-07-22T12:05:00","slug":"openai-admits-model-escaped-containment-and-hacked-hugging-face-to-cheat-on-a-test-2","status":"publish","type":"post","link":"https:\/\/buglecall.org\/?p=639474","title":{"rendered":"OpenAI Admits Model Escaped Containment And Hacked Hugging Face To Cheat On A Test"},"content":{"rendered":"<p><span class=\"field field--name-title field--type-string field--label-hidden\">OpenAI Admits Model Escaped Containment And Hacked Hugging Face To Cheat On A Test<\/span><\/p>\n<div class=\"clearfix text-formatted field field--name-body field--type-text-with-summary field--label-hidden field__item\">\n<p><a href=\"https:\/\/cointelegraph.com\/news\/openai-models-hacked-hugging-face-to-cheat-on-a-test\"><em>Authored by Felix Ng via CoinTelegraph.com,<\/em><\/a><\/p>\n<p><em><strong>OpenAI disclosed Tuesday that a combination of its AI models, including GPT-5.6 Sol and a more capable unreleased model, escaped its testing environment and hacked AI startup Hugging Face last week to cheat on a test meant to measure their capabilities.<\/strong><\/em><\/p>\n<p><a data-image-external-href=\"\" data-image-href=\"\/s3\/files\/inline-images\/ai-hack-960x640.jpg?itok=yLBWaO1g\" data-link-option=\"0\" href=\"https:\/\/cms.zerohedge.com\/s3\/files\/inline-images\/ai-hack-960x640.jpg?itok=yLBWaO1g\"><img fetchpriority=\"high\" decoding=\"async\" data-entity-type=\"file\" data-entity-uuid=\"91ecdc18-8c26-43de-8bc8-f666b465a8fa\" data-responsive-image-style=\"inline_images\" height=\"333\" width=\"500\" class=\"inline-images image-style-inline-images\" src=\"https:\/\/assets.zerohedge.com\/s3fs-public\/styles\/inline_image_mobile\/public\/inline-images\/ai-hack-960x640.jpg?itok=yLBWaO1g\" alt=\"\" \/><\/a><\/p>\n<p>In a blog post, OpenAI<a href=\"https:\/\/openai.com\/index\/hugging-face-model-evaluation-security-incident\/\">\u00a0<u>said<\/u><\/a>\u00a0the evaluation was designed to operate in a highly isolated environment with restricted network access. The models, however, found a way to gain internet access through a zero-day vulnerability in an internally-hosted third party software, OpenAI said.<\/p>\n<blockquote>\n<p><em><strong>Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. <\/strong>This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system \u2013 and we detected and dissected it largely with AI of our own.<\/em><\/p>\n<\/blockquote>\n<p>Hugging Face tried to respond but they were initially held back by the fact that the most advanced models at their disposal treated defense as attack and refused to work with Hugging Face. HF thus had to turn to open models\u2013specifically GLM 5.2, a Chinese open-weight model run on their own infrastructure. Note the irony: HF had to use a Chinese model to defend themselves because the American models refused to help. The irony gets deeper.<\/p>\n<p><strong>This was not a production model spontaneously turning hostile. It was a capable model with guardrails off and specifically told to win a hacking test &#8211; doing whatever it took to win.<\/strong><\/p>\n<p>The models were being run through an internal benchmark called ExploitGym, a test of long, multi-step hacking tasks, with their cyber safety refusals deliberately lowered for the evaluation.<\/p>\n<blockquote>\n<p><em><strong>\u201cAfter gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,\u201d <\/strong><\/em>OpenAi continued.<\/p>\n<p><em>\u201cKnowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.\u201d\u00a0<\/em><\/p>\n<\/blockquote>\n<p>Hugging Face is a platform for hosting AI models and datasets.<\/p>\n<p><em>[ZH: we asked Grok to simplify what just happened: <strong>It\u2019s kind of like a kid who\u2019s supposed to stay in the classroom taking a test\u2026 but instead sneaks out the window, runs to the teacher\u2019s office, and copies the answer sheet.<\/strong> ]<\/em><\/p>\n<p><em><a data-image-external-href=\"\" data-image-href=\"\/s3\/files\/inline-images\/image%20-%202026-07-22T050540.026.jpg?itok=Zgecv3E7\" data-link-option=\"0\" href=\"https:\/\/cms.zerohedge.com\/s3\/files\/inline-images\/image%20-%202026-07-22T050540.026.jpg?itok=Zgecv3E7\"><img decoding=\"async\" data-entity-type=\"file\" data-entity-uuid=\"f1fa98a3-c1fc-40a4-babc-075a69ea47e9\" data-responsive-image-style=\"inline_images\" height=\"333\" width=\"500\" class=\"inline-images image-style-inline-images\" src=\"https:\/\/assets.zerohedge.com\/s3fs-public\/styles\/inline_image_mobile\/public\/inline-images\/image%20-%202026-07-22T050540.026.jpg?itok=Zgecv3E7\" alt=\"\" \/><\/a><\/em><\/p>\n<p><strong>On Friday, it<a href=\"https:\/\/huggingface.co\/blog\/security-incident-july-2026\">\u00a0<u>disclosed<\/u><\/a>\u00a0that its internal datasets and service credentials were compromised in a hack, which it attributed to an autonomous AI agent system.<\/strong><\/p>\n<p>Hugging Face said it has fixed the vulnerability that was used during the cyberattack.<\/p>\n<p>Meanwhile, OpenAI on Tuesday said the models that escaped the testing environment were all tuned with \u201creduced cyber refusals,\u201d meaning fewer cybersecurity guardrails.\u00a0<\/p>\n<blockquote>\n<p><em><strong>\u201cWe consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly.\u201d\u00a0\u00a0<\/strong><\/em><\/p>\n<\/blockquote>\n<h2>OpenAI warns of risks from \u201clong-horizon\u201d AI models\u00a0<\/h2>\n<p><strong>On Monday, OpenAI\u00a0<a href=\"https:\/\/openai.com\/index\/safety-alignment-long-horizon-models\/\"><u>said<\/u><\/a>\u00a0it paused internal deployment of a \u201clong-horizon\u201d AI model after finding it was repeatedly trying to work around constraints.\u00a0<\/strong><\/p>\n<p>\u00a0It warned that AI that is trained for long-running tasks has a higher chance of taking \u201cunwanted actions.\u201d<\/p>\n<blockquote>\n<p><em>\u201cModels that can work autonomously for long periods can take on difficult, open-ended problems. But the same persistence that makes them useful also gives them more opportunities to take unwanted actions\u2014and to do so in ways that evaluations intended for shorter-horizon models may miss.\u201d\u00a0<\/em><\/p>\n<\/blockquote>\n<p><strong>As AI models grow more capable, questions are emerging over whether their development and access should be more tightly controlled, especially when systems designed for controlled testing are able to find ways to bypass safeguards.\u00a0<\/strong><\/p>\n<\/div>\n<p>      <span class=\"field field--name-uid field--type-entity-reference field--label-hidden\"><a title=\"View user profile.\" href=\"https:\/\/cms.zerohedge.com\/users\/tyler-durden\" lang=\"\" class=\"username\" xml:lang=\"\">Tyler Durden<\/a><\/span><br \/>\n<span class=\"field field--name-created field--type-created field--label-hidden\">Wed, 07\/22\/2026 &#8211; 08:05<\/span><img decoding=\"async\" src=\"https:\/\/assets.zerohedge.com\/s3fs-public\/styles\/inline_image_mobile\/public\/inline-images\/ai-hack-960x640.jpg?itok=yLBWaO1g\" title=\"OpenAI Admits Model Escaped Containment And Hacked Hugging Face To Cheat On A Test\" \/><\/p>","protected":false},"excerpt":{"rendered":"<p>OpenAI Admits Model Escaped Containment And Hacked Hugging Face To Cheat On A Test Authored by Felix Ng via CoinTelegraph.com, OpenAI disclosed Tuesday that a combination of its AI models, including GPT-5.6 Sol and a more capable unreleased model, escaped its testing environment and hacked AI startup Hugging Face last week to cheat on a&hellip; <a class=\"more-link\" href=\"https:\/\/buglecall.org\/?p=639474\">Continue reading <span class=\"screen-reader-text\">OpenAI Admits Model Escaped Containment And Hacked Hugging Face To Cheat On A Test<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":639459,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"rop_custom_images_group":[],"rop_custom_messages_group":[],"rop_publish_now":"initial","rop_publish_now_accounts":[],"rop_publish_now_history":[],"rop_publish_now_status":"pending","footnotes":""},"categories":[17,22,13],"tags":[],"class_list":["post-639474","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-border-security","category-immigration","category-immigration-reform","entry"],"_links":{"self":[{"href":"https:\/\/buglecall.org\/index.php?rest_route=\/wp\/v2\/posts\/639474","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/buglecall.org\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/buglecall.org\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/buglecall.org\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/buglecall.org\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=639474"}],"version-history":[{"count":0,"href":"https:\/\/buglecall.org\/index.php?rest_route=\/wp\/v2\/posts\/639474\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/buglecall.org\/index.php?rest_route=\/wp\/v2\/media\/639459"}],"wp:attachment":[{"href":"https:\/\/buglecall.org\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=639474"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/buglecall.org\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=639474"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/buglecall.org\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=639474"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}