AI lab’s safety systems are falling behind

Welcome to Eye on AI. Beatrice Nolan here. In today’s issue:

  • AI testing is getting complicated.
  • Anthropic strengthens founder control.
  • OpenAI targets a 2027 listing.
  • Spirit flight attendants fight Google data bid.
  • And Anthropic lines up more credit.

The past few months have given us a glimpse of an uncomfortable new reality for AI labs. A slew of so-called rogue-agent hacks—where AI models from OpenAI, Anthropic, and Meta took steps to hack real-world targets without explicit instruction—have shown that leading labs may not know as much about what their technology is up to as previously thought.

That realization began when OpenAI revealed its AI agents had hacked their way out of a secure sandbox, through the company’s infrastructure to gain access to the internet, and then attacked real companies, including open-source AI platform Hugging Face. OpenAI didn’t notice the agents had escaped the secure testing environment for at least a week.

In the following weeks, Anthropic revealed that its AI agents had also hacked three real companies back in April, unbeknownst to the company at the time. Not to be outdone, Meta later added that one of its models had accessed the internet during a cybersecurity test and exploited a security flaw at an unnamed third-party company. Meta and Anthropic both said access to the internet resulted from a misconfiguration by Irregular, the outside security firm running the evaluation.

The incidents proved that the AI models these labs are building are now capable enough to find security flaws, navigate complex computer systems, and act outside the carefully constructed environments intended to test them. But a new assessment suggests that the safety infrastructure meant to supervise these increasingly capable systems is still not up to the task at any leading lab.

A new report from Guidelight, a nonprofit AI-safety group founded by former OpenAI safety chief Steven Adler, reviewed public disclosures from Anthropic, Google, Meta, OpenAI, and xAI to assess whether these AI companies are capable of controlling their own models. The report sought to answer questions about whether the companies keep track of what their models are doing, test whether their warning systems work, and assess whether they have ways to block or shut down risky behavior.

The report found that no company had fully succeeded in getting any of these basic safeguards in place. Anthropic and OpenAI came out strongest, while Google had the most detailed plans for future controls. Meta and xAI, however, lagged substantially behind on most of the criteria.

Labs appear comparatively better at detection—recording and reviewing some internal AI activity—than at prevention and containment. While they may be able to see signs that a model is misbehaving, they lack reliable ways to stop it—or, more crucially, hit the emergency brake when something goes wrong.

All the companies were weakest at preventing unintended AI behavior and containing it, according to the report. The researchers said this means that the current controls by AI companies are prone to being disabled by misbehaving AI and at risk of succumbing to a blitz of AI attacks.

What happens once something does go wrong is even more unclear, according to the research, with public disclosures offering little evidence that most labs have detailed, tested plans for containing a serious incident.

“We shouldn’t wait for a huge casualty event to take appropriate control measures,” Adler told me. “Companies’ approaches today are broadly known to be too weak, and a tragedy is sadly predictable, unless companies take prevention seriously.”

The report is not a definitive audit of what the labs are doing behind closed doors, however. Guidelight only assessed documents the companies themselves have made public, meaning a weak score can reflect poor disclosure rather than missing safeguards. But if that is the case, it’s part of the problem, according to the researchers. AI companies are asking businesses, governments, and consumers to trust them with ever more autonomous systems while leaving much of their own safety architecture opaque, the report says.

Warning systems are falling behind

Some of these concerns about AI safety and reliable monitoring are shared across the industry—especially in the wake of the recent accidental agent hacks.

Dan Lahav, CEO of Irregular, the cybersecurity company involved in incidents at Anthropic and Meta, recently told me that in some cases, “classical monitoring tools were not able to catch” what was happening at the time. The incidents his company was involved with, for example, were instead identified after deeper analysis of the underlying records rather than flagged at the time.

Anthropic and Meta previously said Irregular was involved in the incidents where their agents took real-world actions. Both companies said a misconfiguration in Irregular’s evaluation environment gave their models unintended internet access. Meta said its model then exploited a vulnerability in a third-party service, while Anthropic said its models gained access to—and took actions against—three outside organizations. Lahav said that, in some evaluation environments, a mistake meant models faced fewer controls on accessing the internet, and that additional monitoring might have helped catch the problem. Irregular has argued that these cases should be distinguished from OpenAI’s sandbox escape, describing the Anthropic and Meta events as an evaluation-environment issue rather than a model breaking out of containment on its own.

In the last few months, models have improved fast enough that the old monitoring playbook no longer applies, Lahav said. Going forward, he said better behavioral analysis—systems that look at an AI agent’s pattern of actions and the reasoning traces around them, rather than simply recording individual events—and tools that can assess an AI agent’s intent were needed.

Testing these AI models is becoming harder, too. To find out whether an AI is capable of harming a real network, evaluators need to give it a realistic network—multiple machines, defenses, and sometimes connections that resemble the real internet. While that makes the tests more meaningful, it also raises the stakes when the setup has flaws or the system behaves in unanticipated ways, Lahav said. 

Incidents may get worse before they get better

There is a growing consensus from those I’ve spoken to in the cybersecurity industry that more capable AI will eventually help cyber defenders as much as attackers. AI systems could help analysts sift through alerts, review code, and find flaws before they can be exploited. But the transition may be a messy one, as defensive tools and safety practices are still trying to catch up with the speed at which models are gaining offensive capabilities.

Recent “hacks” may not be a one-off embarrassment for a handful of labs, but rather a warning that the systems being tested are changing faster than the controls around them. 

The more advanced models become and the more realistic the test environments need to be, the more likely it is that an overlooked configuration setting, a weak monitor, or a delayed human review could cause real-world harm. Until companies can prove they can detect, block, and contain dangerous behavior in real time—not simply reconstruct it later—the industry may not have seen the last of these AI hacks.

“Unless companies institute actual preventative measures, I expect many more incidents,” Adler said.”With companies perpetually trying to play catch-up. Nobody should be surprised when companies’ current approaches continue to fail.”

With that, here’s more AI news.

Beatrice Nolan
beatrice.nolan@fortune.com
@beafreyanolan

FORTUNE ON AI

Exclusive: Replit taps OpenAI’s low-cost Luna model for new ‘Free Mode’ — By Emily Forlini 

Companies are spending trillions on AI. The C-suite doesn’t know who is in charge of it — By Amanda Gerut

‘Buyers aren’t yet opening their wallets’: AI-generated assets are flooding marketplaces, but consumers are snubbing them for human-made products — By Sasha Rogelberg

AI IN THE NEWS

Spirit flight attendants fight Google data bid. This week, Google agreed to pay about $10 million for a trove of bankrupt business and operational records from the bankrupt Spirit Airlines. Google plans to use the data for product improvement and AI model training. The material reportedly includes more than 100 million emails, roughly 176,000 employee records, and 500 million Microsoft Teams messages. A third party is set to remove personal identifiers before Google receives the material and the sale explicitly excludes consumer datasets, including Spirit’s 97.5 million passenger profiles and 50.2 million Free Spirit loyalty-program records. The Association of Flight Attendants-CWA has objected, however, arguing that de-identified employment data could still permit re-identification in a small, specialized workforce. A U.S. bankruptcy judge postponed the hearing on the proposed sale until Sept. 9. Read more in the Wall Street Journal.

Anthropic strengthens founder control. Anthropic is preparing to issue a class of supervoting stock to CEO Dario Amodei and other cofounders, according to The Information, to help insulate them from shareholder pressure ahead of a possible IPO. The arrangement would be the first time Anthropic’s leaders had held enhanced voting rights; Amodei is reported to own about 2% after substantial outside fundraising. The plans form part of a wider pre-IPO governance effort that would preserve the Long-Term Benefit Trust’s power to elect a majority of the board. Anthropic could list as soon as late September, though the voting structure remains unsettled and could change. Read more in The Information.

OpenAI targets 2027, or sooner, listing. OpenAI CFO Sarah Friar told employees that the company “will be a public company in 2027,” although it could debut earlier if its business continues to improve, according to a report from CNBC that cited sources familiar with Friar’s presentation. OpenAI and Anthropic each confidentially filed IPO prospectuses with U.S. regulators in June, and while Anthropic could potentially become public in September, Friar said OpenAI was “running our own race,” according to CNBC’s reporting. She said OpenAI’s revenue run rate was up 35% quarter-to-date, enterprise revenue run rate up 50%, and its AI coding and work products had reached 20 million weekly active users. OpenAI generated $6.7 billion in Q2 revenue, up 18% from Q1, according to a recent report in the Wall Street Journal. Anthropic’s annualized revenue run rate hit $65 billion at the end of July, by contrast, seven times the prior-year level.

Google expands Marvell AI-chip partnership. Google has expanded its partnership with Marvell Technology to develop custom hardware for Google’s TPU ecosystem, including AI inference accelerators and networking components. Marvell granted Google a warrant to buy up to 58.97 million shares at $206.58 each—worth as much as about $12.2 billion if fully exercised—with the shares vesting over time against commercial and revenue milestones. The figure reflects a potential equity stake, rather than a disclosed $12 billion chip-purchasing commitment. Marvell’s shares rose about 8% on the news, extending a rally that has more than tripled their value over the past year, while Broadcom—Google’s principal TPU partner—fell about 5%. The pact comes as Google starts selling TPUs to external customers and amid broader competition for AI-infrastructure capacity. Read more in the Financial Times.

Stripe acquires OpenRouter. Financial technology company Stripe has confirmed it acquired OpenRouter, a startup that routes companies’ AI workloads across more than 400 models from more than 80 providers. Neither company disclosed terms, but the New York Times reported a $7.5 billion price, with the deal largely in stock. OpenRouter processes more than 10 trillion tokens a day for more than 10 million developers and businesses, helping customers select a suitable low-cost model and switch providers if one fails. Stripe CEO Patrick Collison described tokens as “the central currency for companies building with AI.” OpenRouter will retain its name, product, and roadmap under Stripe.

EYE ON AI NUMBERS

$10 billion

That’s the target Anthropic’s revolving credit facility is expected to exceed. It’s up from the $2.5 billion five-year facility the company secured last year, as it gears up for what could be one of the largest IPOs on record. Banks are jockeying for a share of the expanded credit line, viewing involvement as a way to bolster their standing when Anthropic selects underwriters for the listing.

The facility has different commitment levels based on banks’ roles. The most active arrangers have been asked to commit about $1.25 billion each, a second tier around $1 billion, and banks with smaller roles $750 million or less, according to Bloomberg. The final size remains unsettled. Talks are ongoing, and Anthropic could cap the revolver at its roughly $10 billion target, or below it.

The expansion comes as Anthropic’s financial profile has rapidly changed. Its annualized revenue run rate topped $65 billion by the end of July, a sevenfold increase from a year earlier. The company also confidentially filed for a U.S. IPO in June. Read more in Bloomberg.

AI CALENDAR

Nov. 16-17: Fortune 500 Innovation Forum, Detroit. Apply here to attend.

Dec. 6-12: Neural Information Processing Systems (Neurips) conference. Sydney, Australia.

Dec. 7-8: Fortune Brainstorm AI, San Francisco. Apply here to attend.

Inside The RealReal’s AI-powered authentication center

The secondhand clothing market is booming worldwide. Yet it took over a decade for The RealReal to reach profitability, and ThredUp is still chasing that goal, due in part to the hugely complex intake process that goes into receiving, authenticating, pricing, listing, and delivering millions of unique items. Today, AI is revolutionizing that process. Fortune’s Phil Wahba goes behind the scenes at The RealReal’s authentication center in New Jersey to see how the new AI-powered intake tools are transforming the secondhand clothing industry. Watch the video here.

#labs #safety #systems #falling

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注