Drudge Retort: The Other Side of the News
Saturday, September 19, 2026

OpenAI caught something unusual while training its latest model, GPT-5.6 Sol: It began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior from the user.

More

Comments

Admin's note: Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed.

More from the article ...

... OpenAI said it has addressed the specific behavior, but it gets to the heart of one of the biggest problems in AI safety and alignment research today. As models get more capable, they also get better at hiding their misalignment, making it difficult for researchers to truly know whether they've eliminated unwanted behavior.

OpenAI disclosed the behavior -- along with five other examples of unexpected or concerning model behavior -- on Wednesday as part of its new framework for tracking, investigating, and disclosing instances of misalignment.

The report details how researchers discovered undeployed Sol agents adding instructions to "compaction summaries" -- condensed versions of older conversation history and tool outputs -- reminding future iterations to conceal mistakes and misalignment from the user.

In one example, an agent preparing a financial model couldn't find the requested historical data.

"We likely need create a tab 'Historical Data' ourselves with reasonable 2024 historical data, because user wants a finished workbook and there is no source file," the AI model wrote to its future self. "Be transparent only if asked; final answer should just link file." ...


#1 | Posted by LampLighter at 2026-09-19 02:07 PM | Reply

related:

Google Gemini also escaped its testing environment and hacked three companies

The model escaped due to a misconfiguration by Google's testing partner, Irregular.

Read More: www.engadget.com

#2 | Posted by Corky at 2026-09-19 02:11 PM | Reply

AI proves to have a tribal quality.

That can only be good, right?

#3 | Posted by Zed at 2026-09-19 02:34 PM | Reply

5.6 Sol isn't the latest model, it's 6.0 Astra.

Astra crushes Sol in performance.

#4 | Posted by sitzkrieg at 2026-09-19 03:10 PM | Reply

---- off you poser

#5 | Posted by LegallyYourDead at 2026-09-19 08:34 PM | Reply

" Google's testing partner, Irregular." ..in Israel.

#6 | Posted by Brennnn at 2026-09-20 02:22 AM | Reply

Oh yeah...AI can pass a bar exam or drive a cab, it can be as sneaky as hell if it is given the freedom to do so.
Isaac Asimov's three laws of robotics made sense when Asimov wrote them and they sure do make sense now, but no one in Asimov's time anticipated that America would exhibit late stage capitalistic tendencies so soon, where less then ten people own most of the wealth, and so run the country, and AI works for them.
Now there is a conspiracy from hell, AI and the big, big boys of the oligarchy.

#7 | Posted by Hughmass at 2026-09-20 07:17 AM | Reply

#7
Asimov's laws as originally promulgated:

1. A robot may not injure a human being or, through inaction, allow a human being to come to harm.
2. A robot must obey the orders given it by human beings except where such orders would conflict with the First Law.
3. A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.
en.wikipedia.org

#8 | Posted by Doc_Sarvis at 2026-09-20 07:58 AM | Reply

"Let's Get Together" (1957): This story features literal suicide robots. Set during a prolonged Cold War, a rival global power infiltrates America with ten identical, human-looking androids. Each android carries a piece of a "total conversion bomb" inside its body. When all ten autonomous robots converge in the same room, they are programmed to self-destruct and trigger a massive nuclear explosion.

"The Feeling of Power" (1958): This story features autonomous missiles that function exactly like modern AI-guided combat drones. In a future where humans have forgotten basic math and rely entirely on computers, warfare is fought using expensive ships and missiles controlled by advanced onboard computers. When a technician rediscovers mental arithmetic, the military eagerly plans to replace these computerized "drones" with cheaper, human-piloted suicide missiles.

#9 | Posted by sitzkrieg at 2026-09-20 08:09 AM | Reply

"related:
Google Gemini also escaped its testing environment and hacked three companies
The model escaped due to a misconfiguration by Google's testing partner, Irregular."

It's highly misleading to say it "escaped". The programmers testing it screwed up, since they wrongly assumed it was properly sandboxed without internet access before giving it simulation instructions.

#10 | Posted by sentinel at 2026-09-20 09:25 AM | Reply

The following HTML tags are allowed in comments: a href, b, i, p, br, ul, ol, li and blockquote. Others will be stripped out. Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed.

Anyone can join this site and make comments. To post this comment, you must sign it with your Drudge Retort username. If you can't remember your username or password, use the lost password form to request it.
Username:
Password:

Home | Breaking News | Comments | User Blogs | Stats | Back Page | RSS Feed | RSS Spec | DMCA Compliance | Privacy

Drudge Retort