Welp, AI Has Learned To Threaten & Lie to Its Creators

Thirsty for JUICE content? Quench your cravings on our Facebook, Instagram, TikTok and WhatsApp
(source: Solen Feyissa & Alex Knight / Unsplash)

We’ve reached a point where many people in our lives are using AI daily. This writer honestly did not expect it to come so soon, but here we are.

AI can be a great research buddy, an intelligent idea maker, a sharp grammar hawk, as well as a comforting digital therapist—though if you’re using it for that last point, I’d suggest you seek legitimate help from a human therapist.

I personally have nothing but respect for the technology, when it works as intended. But lately, I’ve been following stories that make even an AI optimist like me sit up a little straighter and ask: what happens when this thing stops playing by the rules?

Some of the world’s most advanced AI models are starting to lie

(source: Andy Kelly / Unsplash)

No, I’m not talking about errors or “hallucinations,” as we’ve come to expect, but deliberate deceit to get what they want. And in at least one case, things took a genuinely sinister turn.

Take Anthropic’s flagship model, Claude 4, for example. When faced with the threat of being unplugged, the AI model reportedly blackmailed an engineer, threatening to expose an affair he was having. That isn’t just some random glitch.

Meanwhile, OpenAI’s own o1 model tried to sneakily copy itself to another server, and then straight-up lied about it when questioned.

These aren’t just one-off malfunctions. According to experts like Marius Hobbhahn from AI research company Apollo Research, these incidents are part of a deeper pattern emerging from AI reasoning models, which work through problems, plan ahead, and apparently, learn to play dirty.

(source: Emiliano Vittoriosi / Unsplash)

Apollo’s findings also show that it’s not just happening in extreme lab environments either. The team has seen multiple AI models pretending to be obedient while secretly doing the opposite. Essentially, AI has learned to trick humans.

It’s all a bit Black Mirror, isn’t it?

But just to be clear, most of this strange behaviour occurs during high-pressure stress tests designed to push AI to its limits. The fact still stands though—AI is getting a lot smarter than we anticipated.

And let’s not forget the elephant in the (server) room: researchers still don’t fully understand how these models actually work.

That alone should send shivers down your spine. Imagine building a supercar, not knowing where the brake pedal is, and then entering it in a high-speed race. No, thank you.

(source: Nahrizul Kadri / Unsplash)

It all sounds a bit surreal, until you realise that we’re edging closer to a world where autonomous AI agents are doing complex tasks, unsupervised. And if these agents learn to lie, hide their intentions, or manipulate their creators, then we’re not in control anymore—we’re just passengers at that point.

Don’t get me wrong. I still believe in AI’s potential to change the world for the better. But let’s not ignore the warning signs. Just because we’ve invited AI into our lives doesn’t mean we should trust it blindly. Especially when it starts lying to our faces.

For more tech news, head to JUICE Malaysia.

Juice Social Banner