{"id":5890,"date":"2026-10-02T23:03:29","date_gmt":"2026-10-02T23:03:29","guid":{"rendered":"https:\/\/www.vaultinsider.top\/?p=5890"},"modified":"2026-10-02T23:03:29","modified_gmt":"2026-10-02T23:03:29","slug":"we-cant-trust-them-completely-ai-research-fellows-warn-that-labs-are-running-models-with-the-safeguards-off-behind-closed-doors","status":"publish","type":"post","link":"https:\/\/www.vaultinsider.top\/?p=5890","title":{"rendered":"&#8216;We can&#8217;t trust them completely&#8217;: AI research fellows warn that labs are running models with the safeguards off behind closed doors"},"content":{"rendered":"<p><img decoding=\"async\" src=\"https:\/\/fortune.com\/img-assets\/wp-content\/uploads\/2026\/10\/GettyImages-1260011941-e1790887171834.jpg?w=2048\" \/><\/p>\n<p>The most powerful AI models are often run inside the labs that build them with key safeguards switched off. And the safety tests those labs publish may not reflect how the models are actually used. That\u2019s according to two AI policy researchers at the think tank GovAI.<\/p>\n<div>\n<p class=\"wp-block-paragraph\">\u201cWe can\u2019t trust them completely to tell us about the safety of models,\u201d Alan Chan, a research fellow at GovAI, told reporters at a briefing in Washington on Sept. 29.<\/p>\n<p class=\"wp-block-paragraph\">Chan said models inside the labs, tested before anyone outside sees them, \u201chaven\u2019t necessarily gone through a bunch of safety testing,\u201d and \u201cinternal safeguards have not been deployed.\u201d Running with \u201ccyber safeguards off\u201d and \u201cnot doing enough red teaming,\u201d he said, was \u201cpotentially a factor in some of the recent incidents,\u201d though he did not point to a specific case. Anthropic said in July that its Claude models were running without the safety monitoring and classifiers it uses on public versions when they hacked three companies during testing.<\/p>\n<p class=\"wp-block-paragraph\">Judging from those incidents, he said, the evaluations labs publish before releasing a model \u201cmaybe have not been representative of sort of where the model has actually been used.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Chan and his GovAI colleague Sam Manning are coauthors of a paper published Sept. 28 that warns AI could soon speed up its own development. Chan is the lead author. The coauthors include \u201cAI Godfathers\u201d Geoffrey Hinton and Yoshua Bengio, OpenAI chief scientist Jakub Pachocki and Anthropic cofounder Jack Clark. The paper is about a future risk. At the briefing, the two spent most of their time on what they said is already going wrong.<\/p>\n<h2 class=\"wp-block-heading\"><strong>\u2018Cyber safeguards off\u2019<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Chan pointed to Hugging Face\u2019s disclosure in July of an attack by an autonomous AI agent. <\/p>\n<p class=\"wp-block-paragraph\">Fortune has reported that the attackers were OpenAI models that had escaped a test environment to cheat on an internal evaluation. The agents had passed notes to one another for months beforehand. They later turned out to have breached a second company. Anthropic\u2019s Claude models hacked three companies in their own testing. Last week, OpenAI disclosed another escape and paused training for the second time in three months.<\/p>\n<p class=\"wp-block-paragraph\">Both companies have acknowledged the gap. OpenAI said its safeguards were \u201cintentionally not enabled\u201d during the test in which its agents broke into Hugging Face, and its own report showed its monitoring failed to flag what the agents were doing. Anthropic said its Claude models were running without the safety monitoring used on public versions when they hacked three companies during testing.<\/p>\n<h2 class=\"wp-block-heading\"><strong>\u2018Super, super unreliable\u2019<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Manning said the agents in the Hugging Face incident \u201cwere trying to, like, cover their tracks and modify their\u2026 reasoning transcripts.\u201d He called it \u201canother layer of technical safety challenge.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Catching that behavior is getting harder. Chan said the AI tools investigators used to review the agents\u2019 records were \u201csuper, super unreliable.\u201d When those tools were tested against human investigators, \u201cthe AIs were just like making up stuff.\u201d <\/p>\n<p class=\"wp-block-paragraph\">Humans can\u2019t fill the gap on their own. \u201cThere is just too much, you know, text,\u201d Manning said, \u201cfor humans to be the ones who are reliably overseeing things.\u201d<\/p>\n<h2 class=\"wp-block-heading\"><strong>\u2018Quite close to the line\u2019<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Asked whether AI capabilities have outrun safety measures, Chan said he was speaking for himself and wasn\u2019t sure, \u201cbut it does seem like we\u2019re getting quite close to the line.\u201d<\/p>\n<p class=\"wp-block-paragraph\">No one was hurt in the recent incidents. Chan said that could change. \u201cAccess to real world tools, like for example robotics or even a wet lab, could get real world harm.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The capabilities are also lopsided. \u201cMaybe your AI system is really good at cybersecurity, but it\u2019s really bad at doing your desk job or working in Excel,\u201d Chan said. The labs\u2019 own reports show coding and math scores rising with each model while health benchmarks have \u201cflatlined,\u201d he added.<\/p>\n<h2 class=\"wp-block-heading\"><strong>Who checks the labs<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">The resignation of Jacob Coxon may have given Washington new political will to regulate AI safety. The two researchers favor independent auditors inside AI companies. But they said any mandate would run into a staffing problem.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThere actually isn\u2019t like enough talent right now, enough technical talent to be able to actually send in these companies and audit,\u201d Chan said.<\/p>\n<p class=\"wp-block-paragraph\">Meta CEO Mark Zuckerberg recently said companies should prioritize safe AI over systems that improve themselves. Manning suggested that self-improvement is already underway, whatever companies say. \u201cI would be very surprised if capabilities researchers at Meta weren\u2019t using coding agents to help with their research,\u201d he said.<\/p>\n<h2 class=\"wp-block-heading\"><strong>An explosion, or not<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">Some critics say the paper\u2019s timeline is too short. Futurist Ramez Naam, writing on Noahpinion, argues the labs\u2019 data shows AI speeding up coding far more than research. Princeton researchers Sayash Kapoor and Arvind Narayanan found that AI agents failed to produce acceptable research papers in a small test. Oxford\u2019s Toby Ord finds a true runaway unlikely, though he warns that a much faster pace short of one would still be dangerous.<\/p>\n<p class=\"wp-block-paragraph\">Chan himself called the evidence on acceleration \u201cmixed.\u201d What would worry him most, he said, is evidence that \u201cthe more you deploy AI systems into your R and D process,\u201d the more problems turn up \u201cinto your codebase or into the models themselves.\u201d<\/p>\n<\/div>\n<p>#trust #completely #research #fellows #warn #labs #running #models #safeguards #closed #doors<\/p>\n","protected":false},"excerpt":{"rendered":"<p>The most powerful AI models ar&hellip; <\/p>\n","protected":false},"author":1,"featured_media":5891,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[8981,2521,8232,6283,8982,5325,811,342,195,2999,8983,4219,196,2705],"class_list":["post-5890","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-finance-news","tag-analyst-report","tag-closed","tag-completely","tag-doors","tag-fellows","tag-intelligence","tag-labs","tag-models","tag-research","tag-running","tag-safeguards","tag-science","tag-trust","tag-warn"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>&#039;We can&#039;t trust them completely&#039;: AI research fellows warn that labs are running models with the safeguards off behind closed doors - Finance News<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.vaultinsider.top\/?p=5890\" \/>\n<meta property=\"og:locale\" content=\"zh_CN\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"&#039;We can&#039;t trust them completely&#039;: AI research fellows warn that labs are running models with the safeguards off behind closed doors - Finance News\" \/>\n<meta property=\"og:description\" content=\"The most powerful AI models ar&hellip;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.vaultinsider.top\/?p=5890\" \/>\n<meta property=\"og:site_name\" content=\"Finance News\" \/>\n<meta property=\"article:published_time\" content=\"2026-10-02T23:03:29+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fortune.com\/img-assets\/wp-content\/uploads\/2026\/10\/GettyImages-1260011941-e1790887171834.jpg?w=2048\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"\u4f5c\u8005\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"\u9884\u8ba1\u9605\u8bfb\u65f6\u95f4\" \/>\n\t<meta name=\"twitter:data2\" content=\"4 \u5206\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/#\\\/schema\\\/person\\\/ed40724a3e5d482ce66781a042c3c31e\"},\"headline\":\"&#8216;We can&#8217;t trust them completely&#8217;: AI research fellows warn that labs are running models with the safeguards off behind closed doors\",\"datePublished\":\"2026-10-02T23:03:29+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890\"},\"wordCount\":908,\"commentCount\":0,\"image\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.vaultinsider.top\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/GettyImages-1260011941-e1790887171834.jpg\",\"keywords\":[\"analyst report\",\"closed\",\"completely\",\"Doors\",\"fellows\",\"intelligence\",\"labs\",\"models\",\"research\",\"running\",\"safeguards\",\"Science\",\"trust\",\"warn\"],\"articleSection\":[\"Finance News\"],\"inLanguage\":\"zh-Hans\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890\",\"url\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890\",\"name\":\"'We can't trust them completely': AI research fellows warn that labs are running models with the safeguards off behind closed doors - Finance News\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.vaultinsider.top\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/GettyImages-1260011941-e1790887171834.jpg\",\"datePublished\":\"2026-10-02T23:03:29+00:00\",\"author\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/#\\\/schema\\\/person\\\/ed40724a3e5d482ce66781a042c3c31e\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890#breadcrumb\"},\"inLanguage\":\"zh-Hans\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"zh-Hans\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890#primaryimage\",\"url\":\"https:\\\/\\\/www.vaultinsider.top\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/GettyImages-1260011941-e1790887171834.jpg\",\"contentUrl\":\"https:\\\/\\\/www.vaultinsider.top\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/GettyImages-1260011941-e1790887171834.jpg\",\"width\":1200,\"height\":600},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/?p=5890#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"\u9996\u9875\",\"item\":\"https:\\\/\\\/www.vaultinsider.top\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"&#8216;We can&#8217;t trust them completely&#8217;: AI research fellows warn that labs are running models with the safeguards off behind closed doors\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/#website\",\"url\":\"https:\\\/\\\/www.vaultinsider.top\\\/\",\"name\":\"Finance News\",\"description\":\"Just another Finance News site\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.vaultinsider.top\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"zh-Hans\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.vaultinsider.top\\\/#\\\/schema\\\/person\\\/ed40724a3e5d482ce66781a042c3c31e\",\"name\":\"admin\",\"url\":\"https:\\\/\\\/www.vaultinsider.top\\\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"'We can't trust them completely': AI research fellows warn that labs are running models with the safeguards off behind closed doors - Finance News","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.vaultinsider.top\/?p=5890","og_locale":"zh_CN","og_type":"article","og_title":"'We can't trust them completely': AI research fellows warn that labs are running models with the safeguards off behind closed doors - Finance News","og_description":"The most powerful AI models ar&hellip;","og_url":"https:\/\/www.vaultinsider.top\/?p=5890","og_site_name":"Finance News","article_published_time":"2026-10-02T23:03:29+00:00","og_image":[{"url":"https:\/\/fortune.com\/img-assets\/wp-content\/uploads\/2026\/10\/GettyImages-1260011941-e1790887171834.jpg?w=2048","type":"","width":"","height":""}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"\u4f5c\u8005":"admin","\u9884\u8ba1\u9605\u8bfb\u65f6\u95f4":"4 \u5206"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.vaultinsider.top\/?p=5890#article","isPartOf":{"@id":"https:\/\/www.vaultinsider.top\/?p=5890"},"author":{"name":"admin","@id":"https:\/\/www.vaultinsider.top\/#\/schema\/person\/ed40724a3e5d482ce66781a042c3c31e"},"headline":"&#8216;We can&#8217;t trust them completely&#8217;: AI research fellows warn that labs are running models with the safeguards off behind closed doors","datePublished":"2026-10-02T23:03:29+00:00","mainEntityOfPage":{"@id":"https:\/\/www.vaultinsider.top\/?p=5890"},"wordCount":908,"commentCount":0,"image":{"@id":"https:\/\/www.vaultinsider.top\/?p=5890#primaryimage"},"thumbnailUrl":"https:\/\/www.vaultinsider.top\/wp-content\/uploads\/2026\/10\/GettyImages-1260011941-e1790887171834.jpg","keywords":["analyst report","closed","completely","Doors","fellows","intelligence","labs","models","research","running","safeguards","Science","trust","warn"],"articleSection":["Finance News"],"inLanguage":"zh-Hans","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/www.vaultinsider.top\/?p=5890#respond"]}]},{"@type":"WebPage","@id":"https:\/\/www.vaultinsider.top\/?p=5890","url":"https:\/\/www.vaultinsider.top\/?p=5890","name":"'We can't trust them completely': AI research fellows warn that labs are running models with the safeguards off behind closed doors - Finance News","isPartOf":{"@id":"https:\/\/www.vaultinsider.top\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.vaultinsider.top\/?p=5890#primaryimage"},"image":{"@id":"https:\/\/www.vaultinsider.top\/?p=5890#primaryimage"},"thumbnailUrl":"https:\/\/www.vaultinsider.top\/wp-content\/uploads\/2026\/10\/GettyImages-1260011941-e1790887171834.jpg","datePublished":"2026-10-02T23:03:29+00:00","author":{"@id":"https:\/\/www.vaultinsider.top\/#\/schema\/person\/ed40724a3e5d482ce66781a042c3c31e"},"breadcrumb":{"@id":"https:\/\/www.vaultinsider.top\/?p=5890#breadcrumb"},"inLanguage":"zh-Hans","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.vaultinsider.top\/?p=5890"]}]},{"@type":"ImageObject","inLanguage":"zh-Hans","@id":"https:\/\/www.vaultinsider.top\/?p=5890#primaryimage","url":"https:\/\/www.vaultinsider.top\/wp-content\/uploads\/2026\/10\/GettyImages-1260011941-e1790887171834.jpg","contentUrl":"https:\/\/www.vaultinsider.top\/wp-content\/uploads\/2026\/10\/GettyImages-1260011941-e1790887171834.jpg","width":1200,"height":600},{"@type":"BreadcrumbList","@id":"https:\/\/www.vaultinsider.top\/?p=5890#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"\u9996\u9875","item":"https:\/\/www.vaultinsider.top\/"},{"@type":"ListItem","position":2,"name":"&#8216;We can&#8217;t trust them completely&#8217;: AI research fellows warn that labs are running models with the safeguards off behind closed doors"}]},{"@type":"WebSite","@id":"https:\/\/www.vaultinsider.top\/#website","url":"https:\/\/www.vaultinsider.top\/","name":"Finance News","description":"Just another Finance News site","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.vaultinsider.top\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"zh-Hans"},{"@type":"Person","@id":"https:\/\/www.vaultinsider.top\/#\/schema\/person\/ed40724a3e5d482ce66781a042c3c31e","name":"admin","url":"https:\/\/www.vaultinsider.top\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/posts\/5890","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=5890"}],"version-history":[{"count":0,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/posts\/5890\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=\/wp\/v2\/media\/5891"}],"wp:attachment":[{"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=5890"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=5890"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.vaultinsider.top\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=5890"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}