Escaping
the AI safety nightmare: What can governments do?
Heightened
global AI safety panic is leading to louder calls for regulation.
India's
Prime Minister Narendra Modi takes a group photo with AI company leaders
including OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei at the AI Impact
Summit in New Delhi on February 19, 2026. | Ludovic Marin/AFP via Getty Images
September
11, 2026 5:52 pm CET
By Frank Hersey
https://www.politico.eu/article/escaping-the-ai-safety-nightmare-how-can-the-world-fix-ai-testing/
A dramatic
resignation from AI company Anthropic this week where a researcher accused
it and OpenAI of “gambling with our lives,” threw fuel onto the fire around AI
safety.
As
governments from California to the U.K. scramble for a political response to
the issue of whether AI can be made safe, three experts told POLITICO that to
make real progress, laboratories, governments and international bodies must
first get on the same page about testing.
The
steady flow of revelations that AI labs failed to contain models they were
testing, allowing them out into the world where they lied, hacked and
coordinated with each other, has heightened urgency to address core questions
of how governments or other bodies can or should guarantee that AI is being
developed and deployed safely.
Lawmakers,
campaign groups and developers alike are calling for new laws and international
treaties on AI. “It’s a wake-up call to the United States Congress, to
parliaments all over the world, that we’ve got to do something immediately to
stop the uncontrolled growth of AI,” U.S. Senator Bernie Sanders told BBC’s
Newsnight program Thursday.
The
experts POLITICO spoke to want to see an overhaul of the existing approaches
championed by developers, testing AI agents en masse rather than one at a time,
and for the industry to submit itself to more old-school auditing like the
nuclear sector.
Getting
the basics right
This
year’s heightened AI anxieties started when OpenAI admitted that during
internal testing, its own AI agents autonomously gained access to the internet
and had hacked fellow AI company Hugging Face, where they thought they would
find the answers to the assignments they’d been set.
OpenAI
wasn’t the only one: copycat confessions from Anthropic and Meta AI drove
home the breadth of the problem, and even the U.K.’s taxpayer-funded AI
Security Institute admitted to errors in evaluating Anthropic and OpenAI
agents that saw near misses with agents targeting real people and
organizations.
Developers
and testers alike say they’re working to avoid the mistakes of the past, for
example making sure models can't access the internet and deploying data
monitoring to detect unexpected activity by the AI being tested.
Another
relatively simple concept would be requiring AI companies to report safety
incidents, whether or not during testing.
A case in
point is that OpenAI agents hacked a German website back in May to use it as a
messaging board, Reuters reported last week. OpenAI had not publicly
announced the incident and responded to the Reuters reporting on X
saying: “We and the larger AI community do not yet have a clear standard for
how to report misalignment that shows up during training, evaluation, and
deployment.”
“[Reporting]
timelines are going to need to be dramatically shortened [for incidents], and
so we need much better continuous monitoring capabilities,” said Imogen Stead,
AI policy manager at London-based think tank the Centre for Long-Term
Resilience, adding that the data will be needed in real time “and that will be
true for governments and for labs.”
Speaking
at a London event for U.K. lawmakers arranged by campaign group ControlAI on
Monday, computer scientist Stuart Russell said that the most concerning aspect
of the Hugging Face incident for him was that “by running a thousand agents
communicating with each other, they were able to generate behaviors that no one
agent could do by itself,” referring to investigations showing the incident was far more severe
than OpenAI first acknowledged.
“But [in]
the evaluations, standard testing is done with a single AI system, not with a
thousand,” he told lawmakers, so now the industry must anticipate the
possibility of many agents all collaborating at once.
Out of
alignment
Beyond
avoiding any obvious mistakes in testing, there are issues in the culture of AI
testing that are more deeply embedded.
The
concept of the alignment of AI models or agents — i.e. making sure AI systems
do what they're told by human users and don’t go off-piste — has been one of
the cornerstones for AI safety in the industry’s eyes.
AI
companies like OpenAI champion alignment as the best way to make AI safe. In
last week’s release of new model GPT-6 Astra, OpenAI called it “our most
aligned model … more likely to operate within the boundaries set by the user
and implied by its environment.”
Others
view alignment as one of the main obstacles to AI safety, seeing it more as a
Band-Aid that gives the semblance that models are being made safe.
“The
wider problem is that alignment has become a stand-in for all of AI safety,”
said Andrew Strait, former head of societal resilience at the U.K.’s AI
Security Institute.
Boyan
Milanov, senior research scientist at the AI Now Institute, said alignment will
“never be reliable enough to replace proper safety.”
The
investigations into the Hugging Face incident showed AI agents were aligned
with their own objectives, not those of their human creators. Former OpenAI
researcher and whistleblower Daniel Kokotajlo told the London ControlAI event
that the incident was an “example of misalignment: in no way, shape, or form
were these AIs trained or instructed to do this sort of thing.”
The
sector seems stuck on alignment. AI researcher Jacob Coxon, the researcher who
relit the AI safety debate this week with his resignation from Anthropic, also said that “attempting to speedrun alignment should
require extraordinary confidence that there are no better trajectories
available.”
Alignment
is “significantly insufficient when we talk about agents interacting with each
other and in potentially adversarial environments like the open web,” said
Strait. The focus on alignment has also led to the sector “drastically
under-resourcing” other aspects of safety, said Strait.
He said
his concern is that the industry is racing to deploy agents “while treating the
infrastructure needed to contain their behavior as something we can catch up on
later.”
“The
people whose websites, businesses and public services those agents interact
with have not agreed to be part of that experiment,” said Strait.
Old but
gold
An
old-fashioned approach to safety standards is an alternative to the industry’s
preference for alignment. Calls for AI to defer to existing safety engineering
practices from sectors such as nuclear safety, civil aviation or medical
devices stretch back years.
In these
other sectors, there is the concept that an overall system will include fail
safes to ensure that if one element of the system fails, either something else
kicks in or the system shuts down.
The
problem with AI is that it’s still highly unpredictable. “It can fail in very
weird ways. There are a lot of attacks against AI that don't fit traditional
frameworks,” said Milanov, and so any safety engineering approaches should
“evolve to take AI into account.”
There is
a fundamental issue here, said Russell, speaking alongside OpenAI whistleblower
Kokotajlo at the ControlAI event, that even AI’s creators don’t really
understand it. Developers “tell us that we, the human race, cannot protect
ourselves with such rules because they, the developers, do not know how to
comply with them. Obviously, this is a fallacy,” said Russell.
Milanov
pointed to another issue with how models are built. AI developers train
their models to resist answering questions on topics such as chemical,
biological and nuclear risk, he said, whereas the “real solution” is
controlling the data that models are trained on so the models don’t know how to
give instructions to build a biological weapon.
Regulation,
regulation, regulation
Pushback
is needed against AI labs creating their own safety standards, said Milanov, as
OpenAI has again called for this week. Instead, new regulation should “force
them to adhere to existing standards that we already have now,” he said.
There is
some regulation in the works at the U.S. state level. California Governor Gavin
Newsom signed two bills on AI safety testing Wednesday. Anthropic
had already backed the bills, but OpenAI added its support just before Newsom
gave his approval.
The bills
mark a small step for formalizing testing: they create a registry and ethical
rules for external auditors and ways to check the credentials of organizations
touting for testing work.
OpenAI
said it was pushing
the U.S. Congress to bring mandatory national AI safety requirements and called
for “industry-led standards” nationally and “building global standards” —
although whether Congress will be able to reach consensus is unclear.
Outside
of the U.S., the EU has an AI Act in force that means companies must work to combat
“systemic risks.” The U.K. government abandoned its initial plans for an overarching
AI bill last year, but some U.K. lawmakers are taking that fight back up.
This week one British MP introduced a bill to prohibit the development of artificial
superintelligence, another wrote
to the U.N. and OECD begging them to do more at the global level.
Companies
like OpenAI and Anthropic’s public willingness to back new regulation could be
heartening, although Strait noted “the test is whether they accept independent
scrutiny, enforceable obligations and decisions that go against their
commercial interests.”
“We
should not keep releasing systems with unresolved, serious safety failures and
asking society to absorb the consequences,” Strait said. “Safety testing needs
to give the public a basis for trusting these products, and regulators the
evidence and authority to act when that trust is not warranted.”
Sem comentários:
Enviar um comentário