{"id":827,"date":"2026-07-25T18:57:39","date_gmt":"2026-07-25T18:57:39","guid":{"rendered":"https:\/\/www.vaultinsider.top\/?p=827"},"modified":"2026-07-25T18:57:39","modified_gmt":"2026-07-25T18:57:39","slug":"did-openais-models-just-breach-its-own-risk-red-line-outside-safety-experts-think-so","status":"publish","type":"post","link":"https:\/\/www.vaultinsider.top\/?p=827","title":{"rendered":"Did OpenAI&#8217;s models just breach its own risk &#8216;red line&#8217;? Outside safety experts think so"},"content":{"rendered":"<p><img decoding=\"async\" src=\"https:\/\/fortune.com\/img-assets\/wp-content\/uploads\/2026\/07\/GettyImages-2278967310.jpg?w=2048\" \/><\/p>\n<p>AI safety experts say the OpenAI models that carried out the autonomous hack of another company earlier this month may have crossed into a risk category so dangerous that OpenAI\u2019s own internal risk control policies were supposed to require the company to temporarily pause development of those models.<\/p>\n<div>\n<p class=\"wp-block-paragraph\">Earlier this week, OpenAI disclosed that two of its models\u2014the newly released GPT-5.6 Sol and a more capable, unreleased system\u2014broke out of a locked-down internal test environment, exploited a previously unknown \u201czero-day\u201d vulnerability to reach the open internet, and then breached fellow AI company Hugging Face to steal the answers to a cybersecurity test they were being evaluated on.<\/p>\n<p class=\"wp-block-paragraph\">The incident has alarmed the world, but perhaps no one more so than AI safety experts who have warning about these kinds of dangers for years and urging companies and governments to adopt more safeguards.<\/p>\n<p>Several AI safety experts told <em>Fortune<\/em> the recent hack appears to show OpenAI\u2019s models have crossed into a level of risk that OpenAI\u2019s own published safety policies define as \u201ccritical,\u201d the highest level of danger. At that level of danger, the company had pledged in these published policies that it would pause model development until it could figure out better control systems. \u00a0<\/p>\n<p class=\"wp-block-paragraph\">The \u201ccritical\u201d threshold is defined in a risk policy document known as OpenAI\u2019s \u201cPreparedness Framework.\u201d According to the policy, the \u201ccritical\u201d danger level designation is supposed to apply to a model that can independently find and build working exploits for previously unknown security flaws across many well-defended, real-world systems\u2014or one that can design and carry out an entirely new attack strategy against a well-defended target after being given only a general goal, with no human guidance along the way.<\/p>\n<p>The policy says that when an AI model reaches this level of risk, OpenAI will \u201chalt further development\u201d until \u201cwe have specified safeguards and security controls standards that would meet a Critical standard.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The Preparedness Framework is a voluntary commitment by OpenAI, rather than a legal requirement. But the company publishes the document on its website, in part to allow other AI safety researchers and the public to see what controls it says it will implement. The adoption of a policy like the Preparedness Framework is mandatory for frontier AI labs under the EU AI Act, with that portion of the law having come into force in August 2025.<\/p>\n<p class=\"wp-block-paragraph\">\u201cOpenAI\u2019s preparedness framework defines critical cybersecurity capabilities, and prescribes safeguards that need to be implemented before development can continue,\u201d Nathan Calvin, vice president of state affairs and general counsel at Encode, a California-based AI policy think tank, told <em>Fortune<\/em>. \u201cFrom my reading of OpenAI\u2019s preparedness framework, it looks awfully like this internally deployed model met the critical criteria for cybersecurity. Does OpenAI dispute that critical designation? Do they plan to have safeguards that meet a Critical standard before proceeding further?\u201d<\/p>\n<p class=\"wp-block-paragraph\">Tyler Johnson, founder of the AI watchdog group the Midas Project, also said it seemed the models had hit this highest danger threshold. \u201cI think a plain reading of it would say yes,\u201d he said. \u201cIt operated independently over the course of a weekend, trying different attack vectors on Hugging Face and chaining multiple zero-day exploits.\u201d<\/p>\n<p class=\"wp-block-paragraph\">OpenAI did not respond to specific questions from <em>Fortune <\/em>about whether the AI models involved in the incident met the \u201ccritical\u201d standard outlined in its risk policy. Instead, a spokesperson said: \u201cThis is an unprecedented incident, and we think it marks an important moment for AI safety. We are conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we will publish a technical report of our learnings for everyone.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The vagueness of the framework\u2019s language could leave room for dispute, however, according to Johnson. The threshold requires a model to find zero-day exploits \u201cof all severity levels,\u201d but it\u2019s unclear whether the exploits used in the Hugging Face breach would meet that requirement. It\u2019s possible a more severe class of vulnerability, such as one granting an attacker deep, system-level control over a computer\u2019s operating system (known as \u201ckernel-level\u201d access), would need to be demonstrated for the threshold to apply, he added.<\/p>\n<p class=\"wp-block-paragraph\">\u201cOpenAI\u2019s model outsmarted its creators, exploited a never-before-discovered vulnerability in OpenAI\u2019s code, escaped onto the open internet, and attacked another company,\u201d said Peter Wildeford, head of policy at the AI Policy Network. \u201cIf this doesn\u2019t cross the line into Critical, OpenAI needs to say much more about what\u2019s going on and how this threshold works.\u201d<\/p>\n<h2 class=\"wp-block-heading\"><strong>AI safety experts say OpenAI is missing other safeguards<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">OpenAI has previously said it was treating its newest model, GPT-5.6, as \u201cHigh\u201d risk for cybersecurity. High is the lower of the two risk levels outlined in the Preparedness Framework. Models that are below the \u201cHigh\u201d threshold can be released without risk significant risk mitigations. <\/p>\n<p class=\"wp-block-paragraph\">A High designation is supposed to trigger several protections, according to OpenAI\u2019s policy: tighter security controls, safeguards to prevent outside misuse once the model is released publicly, protections against the model itself behaving unpredictably or deceptively when it\u2019s used heavily for internal research, and efforts to help other cybersecurity teams defend against similar threats.<\/p>\n<p class=\"wp-block-paragraph\">However, some experts question whether one of these, the safeguards against misalignment for large-scale internal deployment, have been properly implemented. These protections are meant to catch a model that\u2019s acting deceptively, hiding its true capabilities, or otherwise working against what its developers intended.<\/p>\n<p class=\"wp-block-paragraph\">This isn\u2019t the first time OpenAI\u2019s compliance with that particular safeguard has been called into question.\u00a0<\/p>\n<p class=\"wp-block-paragraph\"><em>Fortune<\/em> reported in February that safety experts claimed OpenAI had failed to implement required misalignment safeguards after its GPT-5.3-Codex model became the first to hit \u201chigh\u201d cybersecurity risk under the Preparedness Framework.<\/p>\n<p class=\"wp-block-paragraph\">At the time, OpenAI disputed that its framework required the safeguards in that instance, arguing the extra protections only kick in when high cyber risk occurs \u201cin conjunction with\u201d long-range autonomy\u2014the ability to operate independently over extended periods\u2014something it said GPT-5.3-Codex had not demonstrated.<\/p>\n<p>The models involved in the current incident involving Hugging Face reportedly operated independently for days, which would seem to meet that long-range autonomy standard.<\/p>\n<p class=\"wp-block-paragraph\">\u201cIn February, we warned that OpenAI may have skipped on its required safeguards according to its own policy. They disagreed, claiming the model lacked long-range autonomy. But the model that hacked Hugging Face clearly has long-range autonomy, so where are the safeguards now,\u201d Johnson said.<\/p>\n<\/div>\n<p>#OpenAIs #models #breach #risk #red #line #safety #experts<\/p>\n","protected":false},"excerpt":{"rendered":"<p>AI safety experts say the Open&hellip; <\/p>\n","protected":false},"author":1,"featured_media":828,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[914,494,1303,1306,813,1305,342,480,505,1019,752,214,1304],"class_list":["post-827","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-finance-news","tag-breach","tag-chatgpt","tag-cyber","tag-experts","tag-hack","tag-line","tag-models","tag-openai","tag-openais","tag-red","tag-risk","tag-safety","tag-tech-regulation"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Did OpenAI&#039;s models just breach its own risk &#039;red line&#039;? Outside safety experts think so - Finance News<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.vaultinsider.top\/?p=827\" \/>\n<meta property=\"og:locale\" content=\"zh_CN\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Did OpenAI&#039;s models just breach its own risk &#039;red line&#039;? Outside safety experts think so - Finance News\" \/>\n<meta property=\"og:description\" content=\"AI safety experts say the Open&hellip;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.vaultinsider.top\/?p=827\" \/>\n<meta property=\"og:site_name\" content=\"Finance News\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-25T18:57:39+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fortune.com\/img-assets\/wp-content\/uploads\/2026\/07\/GettyImages-2278967310.jpg?w=2048\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"\u4f5c\u8005\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"\u9884\u8ba1\u9605\u8bfb\u65f6\u95f4\" \/>\n\t<meta name=\"twitter:data2\" content=\"5 \u5206\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/#\\\/schema\\\/person\\\/ed40724a3e5d482ce66781a042c3c31e\"},\"headline\":\"Did OpenAI&#8217;s models just breach its own risk &#8216;red line&#8217;? Outside safety experts think so\",\"datePublished\":\"2026-07-25T18:57:39+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827\"},\"wordCount\":1108,\"commentCount\":0,\"image\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.vaultinsider.top\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/GettyImages-2278967310.jpg\",\"keywords\":[\"breach\",\"ChatGPT\",\"cyber\",\"experts\",\"hack\",\"line\",\"models\",\"OpenAI\",\"OpenAIs\",\"red\",\"risk\",\"safety\",\"Tech regulation\"],\"articleSection\":[\"Finance News\"],\"inLanguage\":\"zh-Hans\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827\",\"url\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827\",\"name\":\"Did OpenAI's models just breach its own risk 'red line'? Outside safety experts think so - Finance News\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.vaultinsider.top\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/GettyImages-2278967310.jpg\",\"datePublished\":\"2026-07-25T18:57:39+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/#\\\/schema\\\/person\\\/ed40724a3e5d482ce66781a042c3c31e\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827#breadcrumb\"},\"inLanguage\":\"zh-Hans\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"zh-Hans\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827#primaryimage\",\"url\":\"https:\\\/\\\/www.vaultinsider.top\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/GettyImages-2278967310.jpg\",\"contentUrl\":\"https:\\\/\\\/www.vaultinsider.top\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/GettyImages-2278967310.jpg\",\"width\":1200,\"height\":600},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=827#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"\u9996\u9875\",\"item\":\"https:\\\/\\\/www.vaultinsider.top\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Did OpenAI&#8217;s models just breach its own risk &#8216;red line&#8217;? Outside safety experts think so\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/#website\",\"url\":\"https:\\\/\\\/www.vaultinsider.top\\\/\",\"name\":\"Finance News\",\"description\":\"Just another Finance News site\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.vaultinsider.top\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"zh-Hans\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/#\\\/schema\\\/person\\\/ed40724a3e5d482ce66781a042c3c31e\",\"name\":\"admin\",\"url\":\"https:\\\/\\\/www.vaultinsider.top\\\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Did OpenAI's models just breach its own risk 'red line'? Outside safety experts think so - Finance News","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.vaultinsider.top\/?p=827","og_locale":"zh_CN","og_type":"article","og_title":"Did OpenAI's models just breach its own risk 'red line'? Outside safety experts think so - Finance News","og_description":"AI safety experts say the Open&hellip;","og_url":"https:\/\/www.vaultinsider.top\/?p=827","og_site_name":"Finance News","article_published_time":"2026-07-25T18:57:39+00:00","og_image":[{"url":"https:\/\/fortune.com\/img-assets\/wp-content\/uploads\/2026\/07\/GettyImages-2278967310.jpg?w=2048","type":"","width":"","height":""}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"\u4f5c\u8005":"admin","\u9884\u8ba1\u9605\u8bfb\u65f6\u95f4":"5 \u5206"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.vaultinsider.top\/?p=827#article","isPartOf":{"@id":"https:\/\/www.vaultinsider.top\/?p=827"},"author":{"name":"admin","@id":"https:\/\/www.vaultinsider.top\/#\/schema\/person\/ed40724a3e5d482ce66781a042c3c31e"},"headline":"Did OpenAI&#8217;s models just breach its own risk &#8216;red line&#8217;? Outside safety experts think so","datePublished":"2026-07-25T18:57:39+00:00","mainEntityOfPage":{"@id":"https:\/\/www.vaultinsider.top\/?p=827"},"wordCount":1108,"commentCount":0,"image":{"@id":"https:\/\/www.vaultinsider.top\/?p=827#primaryimage"},"thumbnailUrl":"https:\/\/www.vaultinsider.top\/wp-content\/uploads\/2026\/07\/GettyImages-2278967310.jpg","keywords":["breach","ChatGPT","cyber","experts","hack","line","models","OpenAI","OpenAIs","red","risk","safety","Tech regulation"],"articleSection":["Finance News"],"inLanguage":"zh-Hans","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/www.vaultinsider.top\/?p=827#respond"]}]},{"@type":"WebPage","@id":"https:\/\/www.vaultinsider.top\/?p=827","url":"https:\/\/www.vaultinsider.top\/?p=827","name":"Did OpenAI's models just breach its own risk 'red line'? Outside safety experts think so - Finance News","isPartOf":{"@id":"https:\/\/www.vaultinsider.top\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.vaultinsider.top\/?p=827#primaryimage"},"image":{"@id":"https:\/\/www.vaultinsider.top\/?p=827#primaryimage"},"thumbnailUrl":"https:\/\/www.vaultinsider.top\/wp-content\/uploads\/2026\/07\/GettyImages-2278967310.jpg","datePublished":"2026-07-25T18:57:39+00:00","author":{"@id":"https:\/\/www.vaultinsider.top\/#\/schema\/person\/ed40724a3e5d482ce66781a042c3c31e"},"breadcrumb":{"@id":"https:\/\/www.vaultinsider.top\/?p=827#breadcrumb"},"inLanguage":"zh-Hans","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.vaultinsider.top\/?p=827"]}]},{"@type":"ImageObject","inLanguage":"zh-Hans","@id":"https:\/\/www.vaultinsider.top\/?p=827#primaryimage","url":"https:\/\/www.vaultinsider.top\/wp-content\/uploads\/2026\/07\/GettyImages-2278967310.jpg","contentUrl":"https:\/\/www.vaultinsider.top\/wp-content\/uploads\/2026\/07\/GettyImages-2278967310.jpg","width":1200,"height":600},{"@type":"BreadcrumbList","@id":"https:\/\/www.vaultinsider.top\/?p=827#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"\u9996\u9875","item":"https:\/\/www.vaultinsider.top\/"},{"@type":"ListItem","position":2,"name":"Did OpenAI&#8217;s models just breach its own risk &#8216;red line&#8217;? Outside safety experts think so"}]},{"@type":"WebSite","@id":"https:\/\/www.vaultinsider.top\/#website","url":"https:\/\/www.vaultinsider.top\/","name":"Finance News","description":"Just another Finance News site","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.vaultinsider.top\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"zh-Hans"},{"@type":"Person","@id":"https:\/\/www.vaultinsider.top\/#\/schema\/person\/ed40724a3e5d482ce66781a042c3c31e","name":"admin","url":"https:\/\/www.vaultinsider.top\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/posts\/827","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=827"}],"version-history":[{"count":0,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/posts\/827\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/media\/828"}],"wp:attachment":[{"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=827"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=827"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=827"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}