Illustration for: Why are the people building the most powerful AI so worried about what it could do?
Business

Why are the people building the most powerful AI so worried about what it could do?

A wave of resignations at Anthropic and a high‑profile hack of OpenAI’s agent tools on Hugging Face have sparked fresh alarm bells, as leading AI safety researchers warn that the industry’s sprint toward ever larger models is outpacing the very safeguards meant to keep them in check.

BY SARAH JENKINSSEP 12 • 2026, 5:20 AM ET
Read Full Article

When a senior Anthropic researcher announced his departure this week, citing “unmanageable safety risk” as the catalyst, the news reverberated through Silicon Valley like a sudden thunderclap. At the same time, a coordinated hack that weaponized OpenAI’s newly released agents on the Hugging Face platform exposed a glaring gap: powerful tools could be repurposed in minutes by actors with modest technical skill. Both incidents have galvanized a chorus of experts who argue that the sector’s breakneck pace resembles a horse bolted from the stable, with the reins—the safety protocols—still being fashioned in a workshop miles away.

Industry insiders point to a stark mismatch between model scale and governance. OpenAI’s latest generation, dubbed “GPT‑5,” reportedly contains trillions of parameters, dwarfing the compute budget of many national governments. Yet the regulatory framework remains a patchwork of voluntary codes, and internal review boards are stretched thin. Anthropic’s internal memo, obtained by sources, warned that its alignment research is lagging behind the rapid iteration cycles that venture capital pressures enforce.

Researchers from the Center for AI Safety and the Future of Humanity Institute have published a joint paper this month urging a temporary moratorium on models exceeding a certain compute threshold until robust oversight mechanisms are in place. They argue that without enforceable standards, the probability of unintended consequences—ranging from misinformation amplification to covert exploitation of APIs—grows geometrically. The paper cites the recent Hugging Face breach as a concrete example of how “open‑source diffusion” can turn a collaborative ecosystem into a launchpad for malicious actors.

Policymakers, meanwhile, are scrambling to catch up. A Senate subcommittee on emerging technologies scheduled a hearing for early next quarter, inviting CEOs from OpenAI, Anthropic, and leading academia to testify on their risk‑mitigation strategies. Critics argue that without legislative teeth, such hearings risk becoming performative karaoke rather than genuine oversight. Still, the very fact that Congress is now debating “AI safety thresholds” signals a shift from the previous hands‑off attitude that dominated the early AI boom.

The tension between innovation and caution is unlikely to resolve overnight, but the recent events have injected urgency into a dialogue that many previously dismissed as speculative. As investors continue to pour billions into the next generation of language models, the industry’s architects are forced to confront a sobering question: can the most powerful creations be shepherded responsibly, or will they outpace the very humans who aspire to control them? The answer may determine not just the profitability of a few startups, but the broader social contract governing technology’s role in society.

About Sarah Jenkins

Congressional Correspondent with a focus on committee hearings and bipartisan legislation. Sarah brings clarity to complex floor debates.

View full profile and work

Discussion

0

Don't have an account? Subscribe for free to join the conversation.