OpenAI caught its models leaving notes to successors
OpenAI caught something unusual while training its latest model, GPT-5.6 Sol: It began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior from the user.
Menu
Front Page Breaking News Comments Flagged Comments Recently Flagged User Blogs Write a Blog Entry Create a Poll Edit Account Weekly Digest Stats Page RSS Feed Back Page
Subscriptions
Read the Retort using RSS.
RSS Feed
Author Info
lamplighter
Joined 2013/04/13Visited 2026/09/19
Status: user
MORE STORIES
Trump's authoritarian mindset can't handle appointee dissent (3 comments) ...
Report: CEOs now earn 325 times more than workers (2 comments) ...
OpenAI caught its models leaving notes to successors (4 comments) ...
RFK Jr. Names 8 New Members to Medicine Task Force (4 comments) ...
1,400 New Words, Definitions Added to Merriam-Webster.com (3 comments) ...
Alternate links: Google News | Twitter
Admin's note: Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed.
More from the article ...
... OpenAI said it has addressed the specific behavior, but it gets to the heart of one of the biggest problems in AI safety and alignment research today. As models get more capable, they also get better at hiding their misalignment, making it difficult for researchers to truly know whether they've eliminated unwanted behavior. OpenAI disclosed the behavior -- along with five other examples of unexpected or concerning model behavior -- on Wednesday as part of its new framework for tracking, investigating, and disclosing instances of misalignment. The report details how researchers discovered undeployed Sol agents adding instructions to "compaction summaries" -- condensed versions of older conversation history and tool outputs -- reminding future iterations to conceal mistakes and misalignment from the user. In one example, an agent preparing a financial model couldn't find the requested historical data. "We likely need create a tab 'Historical Data' ourselves with reasonable 2024 historical data, because user wants a finished workbook and there is no source file," the AI model wrote to its future self. "Be transparent only if asked; final answer should just link file." ...
OpenAI disclosed the behavior -- along with five other examples of unexpected or concerning model behavior -- on Wednesday as part of its new framework for tracking, investigating, and disclosing instances of misalignment.
The report details how researchers discovered undeployed Sol agents adding instructions to "compaction summaries" -- condensed versions of older conversation history and tool outputs -- reminding future iterations to conceal mistakes and misalignment from the user.
In one example, an agent preparing a financial model couldn't find the requested historical data.
"We likely need create a tab 'Historical Data' ourselves with reasonable 2024 historical data, because user wants a finished workbook and there is no source file," the AI model wrote to its future self. "Be transparent only if asked; final answer should just link file." ...
#1 | Posted by LampLighter at 2026-09-19 02:07 PM | Reply
related:
Google Gemini also escaped its testing environment and hacked three companies
The model escaped due to a misconfiguration by Google's testing partner, Irregular.
Read More: www.engadget.com
#2 | Posted by Corky at 2026-09-19 02:11 PM | Reply
AI proves to have a tribal quality.
That can only be good, right?
#3 | Posted by Zed at 2026-09-19 02:34 PM | Reply
5.6 Sol isn't the latest model, it's 6.0 Astra.
Astra crushes Sol in performance.
#4 | Posted by sitzkrieg at 2026-09-19 03:10 PM | Reply
Post a comment The following HTML tags are allowed in comments: a href, b, i, p, br, ul, ol, li and blockquote. Others will be stripped out. Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed. Anyone can join this site and make comments. To post this comment, you must sign it with your Drudge Retort username. If you can't remember your username or password, use the lost password form to request it. Username: Password: Home | Breaking News | Comments | User Blogs | Stats | Back Page | RSS Feed | RSS Spec | DMCA Compliance | Privacy
The following HTML tags are allowed in comments: a href, b, i, p, br, ul, ol, li and blockquote. Others will be stripped out. Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed.
Home | Breaking News | Comments | User Blogs | Stats | Back Page | RSS Feed | RSS Spec | DMCA Compliance | Privacy