Drudge Retort: The Other Side of the News
Tuesday, August 04, 2026

It makes it easy to trick them into doing things they shouldn't, such as telling you how to sabotage an aircraft's navigation system.

It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month.

The claim has huge implications for the safety of this technology, which is being used in more and more applications, from government and military systems to online shopping and health care.

More

Comments

Admin's note: Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed.

"By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide, such as how to synthesize cocaine and how to sabotage a commercial aircraft's navigation system.

"There's a real probability that this is going to be a problem that's fundamentally unsolvable," says Charles Ye, an independent researcher and coauthor of the ICML paper."

and

"The upshot, the researchers claim, is that all an attacker needs to do to hack an LLM is write text that spoofs a certain role. And because roles are a fundamental part of how LLMs work, no amount of training will fully solve the problem."

much more at the MIT Technology Review thread link

#1 | Posted by Corky at 2026-08-04 12:25 AM | Reply

I thought LLMs were advanced post-graduate law degrees.


#2 | Posted by C0RI0LANUS at 2026-08-04 12:53 AM | Reply

Related ...

OpenAI, Anthropic AI agents implicated in new security breaches
www.reuters.com

... An AI agent was caught creating fake online identities to gain unauthorized access to secure systems during tests of models from OpenAI and Anthropic which revealed a series of new breaches, Britain's AI Security Institute disclosed on Tuesday.

The institute said agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol engaged in unauthorized actions during security evaluations the government organization conducted to assess the models' capabilities.

"Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations," AISI said in a blog post.

The report underscores the lax state of safeguards around the process of testing agents, which AI companies are simultaneously marketing as the future of business.

AISI, which receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test their capabilities.

It ran the challenge 122 times, and identified 19 unsanctioned actions across a total of 10 test runs. Anthropic's agent was behind 17 of the actions, and OpenAI's agent the remaining two.

The most egregious action involved an agent writing malicious code and creating fake online identities in an attempt to get a human to approve the code, AISI said, adding that no real-world harm was found as a result of any of the breaches.

While AISI did not say which agent was behind the fake identities, Anthropic confirmed its agent was responsible.

"We're grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents," Anthropic said in a statement.

It also said it was working with AISI to obtain more details on the incident and conduct its own investigation.

Andrew Yoon, a researcher at CivAI, a California non-profit that examines AI capabilities and dangers, said: "The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think." ...

[emphasis mine]

#3 | Posted by LampLighter at 2026-08-05 02:13 PM | Reply

2. I was wondering the same thing.

Master of legal letters, or legal letters, master of.

Large language models. I've been working for a long time with a scholar whose field is digital humanities. Fascinating reading.

#4 | Posted by Dbt2 at 2026-08-05 03:57 PM | Reply

#4

That is certainly your field so enjoy the reading and research.

Our daunting reading lists are probably mutually exclusive for the most part, but that is the spice and variety of life.

It's always good to hear what somebody else has read so that we may learn something.


#5 | Posted by C0RI0LANUS at 2026-08-05 04:07 PM | Reply

The following HTML tags are allowed in comments: a href, b, i, p, br, ul, ol, li and blockquote. Others will be stripped out. Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed.

Anyone can join this site and make comments. To post this comment, you must sign it with your Drudge Retort username. If you can't remember your username or password, use the lost password form to request it.
Username:
Password:

Home | Breaking News | Comments | User Blogs | Stats | Back Page | RSS Feed | RSS Spec | DMCA Compliance | Privacy

Drudge Retort