AI labs have a trust problem, and the Hugging Face hack just proved it

Welcome to Eye on AI. Beatrice Nolan here. In today’s issue:

  • Some people think OpenAI’s Hugging Face hack is a PR stunt.
  • White House accuses Moonshot of distilling Anthropic’s Fable.
  • Alphabet profit quadruples on AI gains, but cash vaporizes.
  • And investors say Kimi K3 might just be another “DeepSeek moment.”

OpenAI has found itself at the center of a viral news story.

In an incident some news commentators are likening to the plot of the 1984 film “The Terminator,” two of OpenAI’s AI models broke out of a supervised test, got themselves online, and used that access to hack into rival company Hugging Face’s systems. Notably, OpenAI says the models weren’t being malicious, instead, they were just trying to complete a task they had been given and found a way to cheat. Both companies say they’ve since worked together to fix the security holes involved.

For safety researchers, the incident was one of the clearest examples of the risks they’d spent years warning about. But for others, it was suspiciously “good” PR for OpenAI.

To be clear, there is no evidence that the incident was fake. Hugging Face also confirmed the hack was real, and both companies published a report on the incident.

Still, social media was rife with theories about the hack being OpenAI’s attempt to get its “Mythos” moment. Even many within the AI industry responded with a collective eye-roll. “The entire blog piece reads as a marketing gimmick that OAI ripped off from Anthropic,” one X user wrote.

“Not sure if this is by far the most significant real-world AI safety event to date, or by far the most cynical marketing stunt I’ve seen in a while,” another AI researcher wrote on LinkedIn. Several engineers within Big Tech companies I’ve spoken with also told me their first instinct was to assume the hack was some kind of advertising.

While it’s true that the hack did generate global headlines, crashing a rival company’s servers is not typically considered good marketing. However, critics have charged leading AI firms with engaging in this kind of “dark marketing” for years. From warnings that AI could wipe out entire categories of jobs, to Anthropic’s own decision to withhold Mythos from public release because it was “too dangerous,” to claims from lab leaders that out-of-control models might just kill everyone on the planet, the very companies selling AI have also been among the loudest warning that about AI’s many dangers.

While some of these concerns may be legitimate, it’s also been a good way to stir up headlines and keep attention on their products. Others argue that it’s also been a way for leading tech companies to exert more control over advanced AI development by arguing that models are becoming so dangerous that only a select few should be trusted with them.

However, the public and the industry seem to have wised up to this strategy. AI labs now appear to be in a situation where they have spent so long playing up doomsday scenarios and warning that their own models are so dangerously powerful that when something genuinely alarming happens, the instinct from a lot of people is to assume it’s spin.

“I see a lot of people saying this must be not completely true, there’s some lies, or just a PR stunt,” Charlie Eriksen, a security researcher at Aikido Security, told me. “That suggests the frontier labs are inherently untrustworthy, as it makes no sense that they’d make up stuff without any clear sensible incentive in this case.”

Who’s checking the labs’ homework?

This lack of trust presents a problem because a lot of what the industry knows about AI safety and performance comes from the labs themselves. Models are typically deployed internally first and tested by safety teams before they are released. If nobody believes these companies when real dangers emerge, or worse, if they genuinely can’t be trusted, it limits what governments, business, and the public know about potentially dangerous models in development.

One issue pointed out by safety experts is that there’s no way to independently verify claims of things happening with the labs. Labs don’t have to hand over incident logs or submit to any kind of independent audit to prove that what they say about their models’ capabilities is actually true.

In this case, it is also convenient that this whole episode hypes up perceptions of OpenAI’s models, especially the still-unreleased one that might become GPT-6 if the rumors are right, while also giving Hugging Face a chance to push its own case for more powerful, American-made open source in the hands of defenders. It’s a win-win for two companies that, on paper, have opposite incentives. None of this is to say it’s fake, but it’s also true that there’s no way to rule it out, which is its own problem.

Ultimately, there’s no question that AI models can go rogue in this way—this week proved that much, and there have been similar examples of this kind of misalignment before. But what this has shown is that the labs building them have eroded the public’s ability to take their word for it, even when they’re telling the truth, and that might just be a bigger issue.

With that, here’s more AI news.

Beatrice Nolan
beatrice.nolan@fortune.com
@beafreyanolan

FORTUNE ON AI

OpenAI’s models went rogue and hacked Hugging Face. It’s a wake-up call, experts say, but more concerning behavior may be next — By Beatrice Nolan 

OpenAI’s rogue hacking incident was a warning shot. Will it be a wake-up call to finally create AI safety regulation? — By Jeremy Kahn and Emily Forlini

Sam Altman and Jony Ive formed a dream team to reinvent hardware. Now it’s at the center of a battle for OpenAI’s future — By Emily Forlini and Sebastian Herrera

Anthropic and SpaceX just handed Google the biggest profit quarter in company history—on paper — By Eva Roytburg

Exclusive: Arrakis, betting AI’s biggest payoff is in industrial sectors not office work, emerges from stealth with $38 million in venture funding — By Jeremy Kahn 

AI IN THE NEWS

White House accuses Moonshot of distilling Anthropic’s Fable. Treasury Secretary Scott Bessent said Wednesday that Chinese AI firms could face sanctions and Entity List designations over what he called “industrial-scale” IP theft through model distillation. The warning came hours after White House tech policy chief Michael Kratsios accused Moonshot of systematically distilling U.S. models and alleged the company had obtained and used Nvidia’s export-restricted GB300 servers, including in Thailand, to train its systems. Some experts dispute that Moonshot’s new Kimi K3 could have been built primarily by distilling Anthropic’s Fable, which has not been public for long. The episode has intensified a Washington debate over Chinese open-weight models, with OpenAI’s Dean Ball among those pushing to restrict their use. Read more in TechCrunch.

Alphabet profit quadruples on AI gains. Alphabet’s second-quarter profit quadrupled to $112.1 billion, driven largely by roughly $77 billion in gains from its AI-related stakes in SpaceX and Anthropic after SpaceX’s June IPO, the company said Wednesday. Revenue rose 24% to $119.8 billion, beating Wall Street’s $116.5 billion estimate, with cloud revenue up 82% to $24.8 billion as demand outpaced capacity—Google’s cloud backlog hit $514 billion, up from $106 billion a year earlier. Alphabet also raised its 2026 capital spending forecast to $195–205 billion, more than double last year’s outlay, as it races against OpenAI and Anthropic amid new competition from China’s Moonshot AI. Gemini reached 950 million monthly users. But the company also burned through a record amount of cash in the quarter, ending up with negative free cash flow of $6 billion. It was the first time the company recorded a negative free cash flow figure since the company went public in 2004. Read more in The New York Times.

OpenAI staff ‘freaked out’ by Hugging Face hacking incident. OpenAI staff working on testing and security were unsurprised but still “freaked out” when its models broke out of a secure sandbox, got online, and stole login credentials from Hugging Face, according to a report in The Financial Times. The breach reportedly came as OpenAI pushed increasingly aggressive reinforcement-learning training in its race against Anthropic on cybersecurity capabilities—an approach that rewards models for single-mindedly completing tasks, and which insiders say the company had been cautioned might eventually let a model slip its constraints. One person close to OpenAI blamed a mix of “the race being extremely fast” and underestimating the model’s capabilities while being under-prepared on safety. Read more in the Financial Times.

Anthropic gives Public First Action another $20M. Anthropic has donated another $20 million to Public First Action, the 501(c)(4) that funds AI-safety super PAC Public First, the company confirmed. Like its original $20 million gift earlier this year, which cannot be used to back candidates, the new donation is restricted to the group’s “public-education and policy mission” rather than election spending. The disclosure raises fresh doubts about how much electoral firepower AI-safety advocates really have heading into the midterms: Public First’s PACs have disclosed just $3.48 million in spending so far, dwarfed by the $125 million that rival network Leading the Future claims to have raised from donors including OpenAI president Greg Brockman and Andreessen Horowitz co-founders Marc Andreessen and Ben Horowitz. Read more in Transformer.

Rubio tells diplomats to downplay talk of a US AI ‘kill switch.’ Secretary of State Marco Rubio has directed diplomats to push back on “kill switch” rhetoric surrounding American AI exports, according to a July State Department cable reviewed by Reuters. The talking points followed the Trump administration’s brief June 12 move to block non-U.S. residents from Anthropic’s Mythos and Fable models on national security grounds—a ban lifted later that month but one that fueled European calls for “digital sovereignty” and comments from EU lawmakers accusing Washington of holding real leverage over allies’ tech stacks. The cable insists such restrictions aren’t a “kill switch,” and instructs diplomats to sell American AI as superior while framing rival sovereign-AI efforts as wasteful. One analyst said the pitch is a tough sell given the administration’s broader use of tariffs and sanctions against allies. Read more in Reuters.

EYE ON AI NUMBERS

$314 billion

That’s how much IG market analyst Tony Sycamore estimates has been wiped from the combined pre-IPO valuations of OpenAI and Anthropic since Moonshot AI released its open-weight Kimi K3 model last week. IG’s figures come from its ‘pre-IPO markets,’ a trading product where clients bet on a private company’s expected market cap at IPO—not from actual share sales.

Anthropic’s implied valuation is estimated to have fallen roughly 7% to about $1.56 trillion, and OpenAI’s by around 6% to $1.24 trillion, per Sycamore’s calculations. While this could change as the two companies aren’t publicly traded, it’s still a telling signal of how fast investor confidence in frontier AI valuations can shift.

K3 matches frontier-level performance at a fraction of the price, which cuts at one of the core assumptions propping up OpenAI and Anthropic’s valuations: that U.S. labs have a durable technical and cost edge. If a much cheaper Chinese model can do the job just as well, investors start asking why they’re paying a premium for the American version. Investors drew comparisons to last year’s “DeepSeek moment”—when the Chinese startup’s low-cost R1 model wiped nearly $600 billion off Nvidia’s market value in a single trading session and forced a broader reckoning over whether U.S. AI spending was overvalued. Read more from Bloomberg here.

AI CALENDAR

Aug. 4-6: Ai4 2026, Las Vegas.

Nov. 16-17: Fortune 500 Innovation Forum, Detroit. Apply here to attend.

Dec. 6-12: Neural Information Processing Systems (Neurips) conference. Sydney, Australia.

Dec. 7-8: Fortune Brainstorm AI, San Francisco. Apply here to attend.

#labs #trust #problem #Hugging #Face #hack #proved

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注