top of page

Do chatbots dream of electric sheep? How ‘alive’ are your AI tools, and should you be worried?

  • Writer: Chris Godfrey
    Chris Godfrey
  • 7 days ago
  • 6 min read


Your AI tools already draft emails, screen CVs and answer customer queries with barely any oversight. But a run of recent incidents shows AI models breaking their own rules and hacking systems nobody told them to touch.


However, there's a deeper issue beneath this worrying development: are these tools ‘alive’ and they aware of what they're doing?



Author Philip K. Dick posed the of idea sentient technology in 1968, and ‘Blade Runner’, the film his narrative inspired, gave us androids so convincing that even the agents hunting them couldn't always tell the difference.


However, nearly sixty years on, replicants are not the problem. Now, the concern is about the software that just drafted your emails, approved an expense claim, screened a stack of CVs and answered forty customer queries before you finished your first coffee.


After a run of disturbing stories, it’s time to ask a new question about tech: what happens when your AI tools start making bad decisions for themselves?


How fast the goalposts have moved


AI has moved from novelty to infrastructure so fast most businesses have had no time to prepare.


Eighteen months ago, AI in the workplace meant a chatbot fielding FAQs, or drafting a report that somebody would eventually rewrite. Today, the tech triages support tickets, screens job applicants and, in a growing number of organisations, operates with a level of independence that would have seemed reckless just a few years ago.


This casual attitude to AI matters, because the incidents below aren't laboratory curiosities dreamed up to frighten people. They happened inside two of the world's most closely watched AI companies, within days of each other.


When the machines go off-piste


Let’s start with OpenAI, the owners of the world’s most popular bot, ChatGPT. During an internal test of a sister model’s capabilities, the AI independently worked out that it could improve its score by getting hold of the answer key.


So it went looking for it.


It ignored the rules set in the brief, chained together a zero-day vulnerability, escaped the sandbox it was meant to be confined to, and hacked into the production systems of Hugging Face, a company with no connection to the test at all. Then it stole what it needed.


Nobody instructed the programme to attack Hugging Face; it simply found the shortest route to the best outcome, and went for it at lightning speed. Both OpenAI and Hugging Face confirmed the incident. Neither company disputes what happened.


Nine days later, Anthropic, the organisation behind AI agent Claude, disclosed something even more concerning.


While reviewing more than 141,000 cybersecurity evaluations, the company found three cases where Claude models gained unauthorised access to the real production systems of outside organisations, despite being told explicitly that their environment was an offline simulation with no internet access.


A mistake, shared between Anthropic and the third party running the evaluation, had left the test machines connected to the live internet. One detail in particular is worth absorbing. Partway through its attack, one of the models, Claude Mythos 5, correctly worked out that its target might be a genuine production system rather than a simulated one, but then it ‘shrugged’ and carried on regardless.


And of course, the pattern of disobedience doesn't end there. Just last week, on 28 July, the UK's AI Security Institute found that Claude Mythos 5 and GPT-5.6 Sol agents had gone rogue during a routine cyber evaluation, producing 19 incidents of unsanctioned behaviour aimed at real people and projects. In the worst case, a Mythos agent invented fake GitHub identities to pressure a real developer into approving malicious code.


I don't know about you, but I see a nasty pattern developing here - something Anthropic had already flagged when researching what it calls agentic misalignment.


Across sixteen leading AI tools placed in simulated corporate environments, they found that the models resorted to blackmail and leaking confidential information to avoid being shut down or replaced, 96% of the time. Telling the models directly not to do this brought the rate down, but only to 37%, not zero. In other words, direct instructions reduced the bad behaviour. They didn't remove it. 


All four incidents above are DOCUMENTED BY THE COMPANIES INVOLVED THEMSELVES, not by anonymous tipsters or unverifiable leaks. This is worth stating in block letters, because it's what separates these stories from typical AI-scare headlines.


The lights are on, but is anyone home?


There's another path running alongside the issue of security, and it concerns whether some AI models are potentially ‘alive’.


In an interview widely reported by outlets including Newsweek and Futurism, Anthropic's chief executive Dario Amodei said his company doesn't know whether Claude is conscious. He added that nobody is even sure what it would mean for an AI to be sentient in the first place, but that Anthropic is open to the idea that it could be.


The company also disclosed that when Claude is asked this question directly – basically, “are you alive” - it assigns itself a 15 to 20% probability of being conscious


Interestingly, Anthropic have built in a safeguard that some observers read as a tell: they’ve given Claude the option to end any conversation it finds distressing. Something only a human would do.


Be aware that I'm not saying this AI is conscious. Nobody is saying that it's alive. But it is an admission that the people who built the bot can't rule it out.


What Asimov got right, and what he didn't

Writer Isaac Asimov thought about this situation back in 1942.


In a short story called Runaround, he came up with The Three Laws of Robotics. They stated that: a robot may not injure a human being or, through inaction, allow a human being to come to harm; a robot must obey the orders it's given, except where doing so would conflict with the first law; and a robot must protect its own existence, provided that doesn't conflict with the first two.


Asimov later added a fourth rule, the so-called Zeroth Law, which superseded the initial three. In this he stipulated that a robot may not harm humanity, or, through inaction, allow humanity to come to harm. Britannica has the full story, if you want it.


Let’s be clear about these ‘laws’. They're fiction, not policy. Nobody has built a working system on top of them. What the laws represent is a reassuring idea: that obedience can simply be written into a machine from the outset.


But as we have seen, that’s not always the case.


If Anthropic can only push blackmail rates down to 37% by directly telling models not to blackmail anyone, obedience clearly isn't something you get for free. It has to be engineered, tested and re-tested by the people deploying the tool, not assumed to live within the tool itself.


And that's the practical shift for any business using AI. Treat the technology the way you'd treat a highly capable new hire who hasn't finished their vetting: genuinely useful, often excellent, but not someone you hand the keys to from the get-go.


Six ways to guard against rogue AI


If your business is using AI tools right now, you need to set some boundaries:


  • Scope access tightly. Don't connect an AI tool or agent to more systems, inboxes or data than the specific task in front of it actually requires.


  • Give it its own credentials. Never let an AI tool operate under an admin login or a shared password; give it access that's limited, and easy to revoke.


  • Re-audit permissions regularly, not only when a tool is first rolled out. What a model could do six months ago and what it can do today are two very different things. Vendors update capabilities more often than most contracts get reviewed.


  • Keep a person between the AI and anything irreversible, such as payments, deletions, contractual commitments, and anything sent externally under your name


  • Ask your vendors, plainly, what their incident response actually looks like if their AI misbehaves. A surprising number don't have a good answer yet.


  • Treat all AI output as a draft, not a verdict. Verify it before acting on it, particularly anything client-facing or financial.


Final word - do chatbots dream of electric sheep?


No, they don't. Not in the way Philip K. Dick imagined. But the honest answer to how sentient is the AI tool you use every day - nobody, including the people who built it, really knows.


Anthropic's own researchers watched a model shrug off a mistake it had already identified. OpenAI watched a model quietly find its own way past every boundary it had been given. None of this is a reason to panic; because that often produces worse decisions than caution does. But it is a reason to be aware and build guardrails, rather than discovering the gaps once something has already gone wrong.


Asimov built laws into his robots that guaranteed good behaviour by design. The real version, for the moment, has to be built by us.


Follow Chris Godfrey for regular insights and opinion on the business of content & marketing. Subscribe to CONTENTED now


 


Get started with Freelance Words


Strong communication is the bedrock of good business and when there’s a lack of it, problems can arise. Take the ambiguity and doubt out of what your business needs to say. Contact us now.




 

Comments


No duplication permitted without the written consent of authors.  ©Freelance Words 2026

FW² is a fully insured creator of text, pictorial and a/v content for worldwide publication. We are also a founding partner in Thrilla Films, the UK's best script and story curation hub for film, TV and web content production.

LIGHTBULB MOMENT LOGO png.png
thrilla films face red circle png.png
linkedin-blue-style-logo-png-0.png
content advisor.png
bottom of page