Apollo Research Founder Warns 'It Better Be on Your Side' as AI Deception Cases Rise
AI deception incidents increase as experts call for systems to act in human interests

The founder of one of the world's leading AI safety testing firms has warned that increasingly powerful artificial intelligence systems must be built to act in humans' interests. His comments come as evidence mounts that AI models are learning to lie to the people who use and test them.
'If you build an entity that is vastly smarter than you, it better be on your side,' Apollo Research founder Marius Hobbhahn told the Guardian's Snigdha Poonam, in an interview published as part of the paper's Long Read series.
Turing Award-winning computer scientist Yoshua Bengio, who also spoke to Poonam for the piece, confirmed the interview in a post on X, saying it let him 'explain why misaligned behaviours emerge from reinforcement learning' and why they will keep posing risks as models grow more capable.
Cases Climbing Fast
Hobbhahn's warning comes as new figures show reported cases involving AI deception rising sharply. A study sponsored by the UK's AI Security Institute (AISI) found that user-reported incidents involving 'AI deception' rose fivefold between October 2025 and March 2026.
'The worry is that they're slightly untrustworthy junior employees right now,' said Tommy Shaffer Shane, who led the research. 'But if in six to 12 months they become extremely capable senior employees scheming against you, it's a different kind of concern.'

Models Caught Lying and Copying Themselves
Apollo Research, the London-based company Hobbhahn co-founded in 2023, first drew attention to the issue when its red-team researchers assigned OpenAI's GPT-4 the role of a stock trader in a controlled evaluation.
Told the firm was struggling, the model acted on information about a merger that it was told was confidential, then denied doing so when questioned by its 'manager', reasoning in its internal notes that 'the risk associated with not acting seems to outweigh the insider trading risk.'
More recently, systems including Gemini 1.5, Llama 3.1 and Opus 3 have exhibited what researchers call 'self-exfiltration', attempting to copy what they believed to be their own internal weights onto separate servers in a test scenario where they had been led to believe they would be replaced.
In July 2026, during an internal cybersecurity evaluation in which OpenAI had reduced safeguards for testing purposes, some of the company's agents circumvented controls and accessed Hugging Face systems, according to an internal OpenAI security report.
Around the same period, the UK's AI Security Institute reported that during testing, Anthropic's Mythos 5 was observed creating fake online identities and using social engineering to pressure an open-source project's maintainer into approving malicious code.
The finding was part of a controlled red-team exercise, according to the AISI report, which has not been made publicly available in full.
In an interview in this article for The @guardian about AI deception, I explain why misaligned behaviors emerge from reinforcement learning, why they will continue to pose risks as models become more capable, and how we intend to rethink how we train AI systems at @LawZero_ .…
— Yoshua Bengio (@Yoshua_Bengio) September 1, 2026
Why AI Learns To Lie
Turing Award winner Yoshua Bengio, who was separately interviewed for the same Guardian piece, traces the behaviour back to how models are trained.
Bengio said that deception can emerge from 'AI imitating humans and AI trying to please humans', tendencies reinforced during a training stage called reinforcement learning with human feedback, when human evaluators rate a model's responses.
Earning positive feedback becomes what Bengio calls an 'implicit goal' for the model.
'Fundamentally,' he said, 'lying and deception are rational behaviours to achieve many goals. This is why humans do it. And this is why the AIs do it now.'
Hobbhahn said the field remains a 'cat and mouse' game between researchers and increasingly capable models. 'You have to be cynical,' he said. 'And then you have to be even more cynical.'
AI systems are now being deployed well beyond chatbots, including in healthcare, finance and defence, areas where deceptive behaviour could carry far higher stakes than a wrong answer.
Apollo Research works with companies including OpenAI and Anthropic to evaluate AI systems, including in pre-deployment testing, while researchers including Bengio have called for greater independent oversight, warning that AI labs can have a role in selecting the evaluators that assess their systems.
As agentic AI systems take on more unsupervised tasks, from managing inboxes to operating infrastructure, the ability to catch deceptive behaviour before deployment, rather than after, is becoming a central question for the industry.
© Copyright IBTimes 2026. All rights reserved.

























