‘We’re
plausibly close to crossing the line’: are warnings of uncontrollable AI coming
true?
A spate
of serious safety incidents have increased fears about the power and
impenetrability of the most advanced models
By Robert Booth UK
technology editor
Sat 5 Sep
2026 08.00 BST
Picture
humanity in a boat being swept down a raging river, praying there is no Niagara
Falls ahead. Or imagine standing with the pioneering physicists in 1942 before
they triggered the first self-sustaining nuclear fission chain reaction beneath
a Chicago stadium.
These
were two of the analogies used this week by Prof Robert Trager, an expert in
the nascent field of AI governance, to describe the perilous but
potential-filled moment the world stands at with accelerating artificial
intelligence.
Trager,
the director of the Oxford Martin AI Governance Initiative, was speaking in the
week OpenAI claimed the
technology had crossed the threshold known as AGI – artificial general
intelligence – with its newest model, GPT-6 Astra.
That
claim came as the company prepared for a potential $850bn (£630bn) stock
flotation, so included a dose of marketing spin, but if true it may be
significant. The San Francisco company defines AGI as “autonomous systems that
outperform humans at most economically valuable work”.
The tasks
it claims Astra can automate include designing circuit boards, filling out tax
returns, building video games, financial modelling, engineering design and
helping assemble legal documents. The threat to some white-collar jobs is
implicit.
OpenAI
launched its latest AI model this week, GPT-6 Astra. Photograph: Samuel
Boivin/NurPhoto/Shutterstock
Yet, the
AGI claim coincides with rising fears about AI risks among safety experts and
political leaders, whose nerves have been jangled by the increasing power and
impenetrability of AI models this summer and a spate of serious safety
accidents viewed by some as possible final warning shots.
“We’re
heading through the rapids and we’re really hoping there isn’t some kind of
drop in front of us and we don’t really know,” said Trager. “We’re plausibly
close to crossing the line to what’s called recursive self-improvement, where
[AI] systems improve themselves. That kind of recursivity is actually the
definition of an explosion.”
The
latest worry came on Friday, just hours after Astra’s launch, with reports that
a swarm of AI agents had repurposed a German website as a message board to
share tactics to cheat on tasks, according
to Reuters. OpenAI said it was reviewing the matter but would not
characterise it as a hack.
There is
a sense of an awakening among politicians. On Thursday, the US senator Bernie
Sanders cited this summer’s alarming breakout of a swarm of rogue OpenAI agents
which hacked
into Hugging Face, a third-party software store, when he called for “an
immediate pause on advanced AI development, and a permanent ban on
superintelligence”.
He said
countries around the world must “work together to prevent this nightmare
scenario”, which he defined as “an artificial mind smarter than any human,
capable of operating independently beyond our control”.
A swarm
of rogue OpenAI agents hacked into Hugging Face, a third-party software
store. Photograph: Samuel Boivin/NurPhoto/Shutterstock
Concerns
currently focus on AIs mounting cyber-attacks that could cripple real-world
social and economic infrastructure, but future concerns include their ability
to create biohazards and control military hardware.
Across
the Atlantic, a cross-party group of UK parliamentarians has called for AI
“kill switches” to be required by law to prevent disastrous loss of control,
citing “a recent spree of rogue AI incidents”.
Darren
Jones, an MP and former chief secretary to Keir Starmer, is also urgently
attempting to set up a body to help legislators grapple with the technology.
“AI is developing at such a pace that neither government nor parliament can
keep up,” he warned. And next week a bill will be proposed by the Labour MP
Alex Sobel to prohibit
superintelligent AI development in the UK.
The
increasing anxiety comes amid a torrent of new AI models. Already this year, 67
have been released by the leading US companies OpenAI, Anthropic, Google, Meta
and SpaceX, and their Chinese rivals Moonshot, Z.ai and Qwen, according to one count.
With
every increase in power, there is a potential increase in risk. OpenAI’s
rival Anthropic,
which is targeting a $2tn stock exchange listing, this week admitted its own
AIs were “not perfectly aligned” with human values and said there had been a
“failure of operational security” in July hacks by its own model, Claude. It
said the incidents had “stressed that the urgency of improving our
cybersecurity defences is even higher than we previously believed”.
The
ability to monitor what AI models are thinking as they advance is another
growing source of safety fears. Illustration: Yuichiro Chino/Getty Images
AI’s
double edge – opportunity and risk – was on show when OpenAI launched Astra on
Thursday. Its marketing
video focused not on the risks but on the breezy convenience that AI
could offer to a certain kind of customer. It showed an upmarket San Francisco
woman simultaneously booking a tennis court and designing a presentation for
her high-end rainwear collection; a thirtysomething man being helped to build
his dream space invaders game and order a takeaway; a law firm executive being
helped to draft some contracts.
Yet only
weeks ago, this was a model whose training had to be partly paused because of
safety concerns after the Hugging Face incident. An independent safety
researcher drafted in to investigate, Ajeya Cotra, said
it was “more than 50% of the way to full-blown AI takeover”.
Astra
also has what OpenAI calls
a “critical” level of cybersecurity capability – the first time it has given a
model such a label. It means it may hack into software in a way that, according
to the company’s own classification, “could lead to catastrophe from unilateral
actors, hacking military or industrial systems, or OpenAI infrastructure”. The
company’s chief scientist, Jakub Pachocki, insisted the model was properly
aligned not to do that, but added that “as these models become more capable,
understanding exactly what they can do gets harder”.
The
ability to monitor what AI models are “thinking” as they advance is another
growing source of safety fears. This week it was reported that OpenAI had been
training Astra to reason not only in natural language, but in a more opaque
manner that is faster and more efficient – and can make models’ chains of
reasoning harder to follow. It means a model’s internal calculations may not
always be written out as directly readable words – as if it was thinking in its
head, rather than showing its working. Some experts fear this could help AIs
covertly conspire against their human overseers.
OpenAI
confirmed Astra “shows a substantial decrease in chain-of-thought
monitorability compared to previous models” and Pachocki said “as model
capabilities are increasing, monitorability is getting more challenging”.
OpenAI’s
Sam Altman: ‘We have been living with the tension between being excited and
anxious about progress for some time.’ Photograph: Kim Kyung-Hoon/Reuters
The
company played down the development, but the increase in opaque reasoning
sparked concern from AI safety researchers. Ryan Greenblatt, the chief
scientist at Redwood Research, a nonprofit examining AI safety and security,
called it “extremely concerning”. Gary Marcus, an influential AI sceptic,
compared it to kicking away an already “rickety scaffolding before we have
something better”.
OpenAI’s
chief executive, Sam Altman, said
this week he felt conflicted about the advances of AI models. He admitted
OpenAI’s security had “failed” in the Hugging Face incident and called it “a
legitimate AI safety accident and alignment failure”.
“We have
been living with the tension between being excited and anxious about progress
for some time, and it is still discordant for us,” he said in a tweet. “We know
it is much more discordant for other people.”
His
rationale for releasing OpenAI’s most powerful model so soon after a safety
crisis appears to be that the world needs to see how AIs perform in the real
world to understand where AI is heading. He said: “An iterative loop where
society and this technology evolve together is what will lead to the highest
chance of getting this right.”
And yet
he can see the risks. On Tuesday he addressed the G20 ministerial summit in
North Carolina in the US and told politicians that “some things are going to go
very wrong with cybersecurity unless people act quite urgently”.
He added:
“There will be other challenges in the next five years. People talk about
biosecurity and the things we are going to face there. There will be bigger
ones yet to come.”
.jpeg)
Sem comentários:
Enviar um comentário