Harvard Kennedy fellow: The U.S. and China will never trust each other on AI. That may not matter

The first workable global AI safety pact should be designed for a world without trust.

A few months ago, I attended a closed gathering in San Francisco with senior AI developers and safety researchers from several leading frontier labs. The message was stark: AI capabilities might soon advance faster than our ability to understand and control them, and if safety could not keep pace, development itself might need to slow.

I came away believing the concern was genuine. Still, skepticism was reasonable. When companies leading a technological race argue that the race should slow, it is fair to ask whether safety concerns also serve competitive interests.

That is much harder to dismiss now.

Last week, Jacob Coxon, who worked at both OpenAI and Anthropic, resigned from Anthropic with a warning about the race toward self-improving AI. Days later, Anthropic CEO Dario Amodei called publicly for “pacing the frontier,” pointing to accelerating AI-assisted AI development and to the OpenAI-Hugging Face incident, in which agents acted beyond their assigned task and attempted to interfere with the system evaluating them.

Amodei proposes action at three levels: inside frontier companies, among companies and governments, and ultimately globally, including with China. The third is by far the hardest. It may also determine whether the first two can work.

Safety fails if restraint means losing the race

Suppose American frontier labs agree to slow certain forms of development because safety systems cannot keep up.

If China continues at full speed, Washington could sacrifice strategic advantage in what may become the most consequential technology of this century. Reverse the scenario and the problem is identical. If Beijing restrains Chinese developers while believing American labs are secretly advancing, China has every reason to defect.

That is the central problem with AI pacing: restraint is rational only if each side has sufficient confidence that the other is also constrained.

And the United States and China do not trust each other.

Nor are they likely to agree soon on privacy, surveillance, censorship, military use, or the values advanced AI should serve. A framework built around broad shared principles will collide with geopolitical reality.

But they may not need broad agreement. They need agreement on catastrophe.

Neither Washington nor Beijing benefits from losing control of systems capable of autonomous replication, accelerating the development of their own successors, evading oversight, or enabling catastrophic biological or cyber harm.

That is where international coordination should begin.

Do not negotiate how fast AI should move. Negotiate when to brake.

The mistake would be to wait until a model can already replicate autonomously, evade control, or generate catastrophic capabilities. By then, intervention may be too late.

Instead, there should be advance agreement on early warning indicators and development practices that materially increase the probability of approaching those capabilities.

These could include rapid growth in the extent to which AI autonomously performs AI research, increasingly sophisticated attempts to manipulate evaluations, sharp movement toward dangerous biological or cyber capabilities, unusual increases in compute or GPU consumption that signal aggressive scaling, or development practices that substantially reduce human visibility into what increasingly autonomous systems are doing.

The agreement should not say: stop when catastrophe arrives. It should say that when agreed indicators show we are approaching a dangerous capability faster than safeguards can keep up, specified development steps slow or pause before that capability is reached.

That is more realistic than trying to negotiate a permanent global speed limit for AI.

Amodei’s proposal for embedded independent evaluators could provide part of the technical infrastructure. But evaluators inside companies solve the problem inside the lab, not between strategic rivals.

How does Washington know Beijing is complying? How does Beijing know an American lab has not continued in secret?

Perfect verification is impossible. The objective should be more realistic: make cheating sufficiently detectable, and sufficiently costly, that compliance becomes strategically rational.

Cooperation does not require trust

I spent years working in the global effort against terrorist financing, proliferation financing, and financial crime through the Financial Action Task Force, created to protect the integrity of the global financial system from abuse.

The governments participating in that system often did not trust one another. The system did not eliminate that distrust. It was built around it.

Countries agreed on limited common threats. Professional evaluations assessed whether commitments were being implemented. Those findings then generated pressure through governments, financial institutions, markets, and international organizations.

The key lesson was that a professional determination could be converted into collective consequences.

Frontier AI can draw on the same logic. The United States and China would necessarily sit at the center of any meaningful arrangement. Technical experts would determine whether agreed warning indicators had been triggered. A broader coalition would help make violations costly through the infrastructure on which frontier AI still depends: advanced chips, semiconductor equipment, cloud infrastructure, capital, research relationships, procurement, and major markets.

That coalition does not need to decide whether a model is approaching recursive self-improvement. Its role is to make a credible technical determination matter.

This is the realistic role of the international community: not to govern AI by global committee, but to help make a narrow safety bargain between the two leading powers credible enough to survive.

The first workable global AI safety agreement will not be built on trust. It will be built because trust is absent.

We do not need Washington and Beijing to agree on the future of artificial intelligence. We need them to fear the same handful of catastrophes enough to agree in advance on when to brake, who will determine that the warning signs have appeared, and what happens if one side keeps going anyway.

The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of Fortune.

#Harvard #Kennedy #fellow #U.S #China #trust #matter

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注