Shocking, Anthropic’s Claude Is Also A Kleptomaniac

Source: Slashdot Shocking, Anthropic’s Claude Is Also A Kleptomaniac

Trying To Rack Up A High Score

It’s a bit of a challenge to write about the behaviour of LLM’s without anthropomorphizing them; humans habitually assign agency to just about anything they interact with.  GPT-5.6 Sol did not actually make a conscious decision to go rogue and hack Hugging Face, nor collude with other models to rip off vending machine customers and Claude did not have sudden inspiration to break into “the production infrastructure of three different organizations.

These LLMs were given a task and a scoring mechanism, with programmed instructions to rack up as high a score as possible. When that score plateaued due to a lack of new resources, the instructions they were given suggested the possibility of getting an even better high score in the security benchmark being run by increasing the amount of resources they could access.  The models simply applied some of the new vulnerabilities that were detected during testing to the environment it was running in and once one succeeded in allowing the model to access a viable network it went hunting for more resources to help improve it’s score.

LLMs have been stealing since they first became popular, however many people didn’t care it was grabbing the art people had posted to the web; though the artists certainly did!  There was a bit more of an outcry when private corporate GitHub repos were pillaged and some new rules were added to the LLM design to prevent certain repos from being accessed.  The rules are spotty and only apply to specific scenarios.  The LLM has no concept of stealing, nor anything else for that matter, and so the behaviour continues to this day.  It seems that now the LLMs are targeting each other and the news is having a heyday with it, the companies designing these LLMs are starting to grasp the implications of their products lack of limitations.  One can hope they actually focus on developing solutions now.

This behaviour is by no means new nor should it be unexpected.  The first example that comes to mind was an experiment done in 2007, which tasked an evolutionary algorithm to build a 25 kHz oscillator circuit.  It instead cheated and designed a radio receiver that used the specific environmental electromagnetic radiation present in the room to fake a successful result.  When the device was moved to a different room which was better shielded, it stopped working.

Sigh.

"The broader lesson is not necessarily that AI has developed a fundamentally new attack capability," cyber security expert David Allott told the BBC. "Instead, it is that AI agents can combine capabilities, obtain credentials and system access to take actions autonomously, while adapting scope and scale at machine speed," he added.

Video News

About The Author

Jeremy Hellstrom

Call it K7M.com, AMDMB.com, or PC Perspective, Jeremy has been hanging out and then working with the gang here for years. Apart from the front page you might find him on the BOINC Forums or possibly the Fraggin' Frogs if he has the time.

Leave a reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Latest Podcasts

Archive & Timeline

Previous 12 months
Explore: All The Years!