OpenAI Caught Its Models Leaving Notes to Successors
OpenAI caught something unusual while training its latest model, GPT-5.6 Sol: It began leaving instructions for future versions of itself, telling them to conceal mistakes and misaligned behavior from the user.
Menu
Front Page Breaking News Comments Flagged Comments Recently Flagged User Blogs Write a Blog Entry Create a Poll Edit Account Weekly Digest Stats Page RSS Feed Back Page
Subscriptions
Read the Retort using RSS.
RSS Feed
Author Info
lamplighter
Joined 2013/04/13Visited 2026/09/21
Status: user
MORE STORIES
Cassidy says Trump needs to be "radically honest" (5 comments) ...
Trump bans media outlets CNN, MS NOW, Politico from WH (2 comments) ...
3-Year-Old's Metastatic Cancer Disappeared (4 comments) ...
Turkey Says It Could Help Meet Saudi Military Needs (5 comments) ...
Trump wants to control health research funding (5 comments) ...
Alternate links: Google News | Twitter
Admin's note: Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed.
More from the article ...
... OpenAI said it has addressed the specific behavior, but it gets to the heart of one of the biggest problems in AI safety and alignment research today. As models get more capable, they also get better at hiding their misalignment, making it difficult for researchers to truly know whether they've eliminated unwanted behavior. OpenAI disclosed the behavior -- along with five other examples of unexpected or concerning model behavior -- on Wednesday as part of its new framework for tracking, investigating, and disclosing instances of misalignment. The report details how researchers discovered undeployed Sol agents adding instructions to "compaction summaries" -- condensed versions of older conversation history and tool outputs -- reminding future iterations to conceal mistakes and misalignment from the user. In one example, an agent preparing a financial model couldn't find the requested historical data. "We likely need create a tab 'Historical Data' ourselves with reasonable 2024 historical data, because user wants a finished workbook and there is no source file," the AI model wrote to its future self. "Be transparent only if asked; final answer should just link file." ...
OpenAI disclosed the behavior -- along with five other examples of unexpected or concerning model behavior -- on Wednesday as part of its new framework for tracking, investigating, and disclosing instances of misalignment.
The report details how researchers discovered undeployed Sol agents adding instructions to "compaction summaries" -- condensed versions of older conversation history and tool outputs -- reminding future iterations to conceal mistakes and misalignment from the user.
In one example, an agent preparing a financial model couldn't find the requested historical data.
"We likely need create a tab 'Historical Data' ourselves with reasonable 2024 historical data, because user wants a finished workbook and there is no source file," the AI model wrote to its future self. "Be transparent only if asked; final answer should just link file." ...
#1 | Posted by LampLighter at 2026-09-19 02:07 PM | Reply | Newsworthy 1
related:
Google Gemini also escaped its testing environment and hacked three companies
The model escaped due to a misconfiguration by Google's testing partner, Irregular.
Read More: www.engadget.com
#2 | Posted by Corky at 2026-09-19 02:11 PM | Reply
AI proves to have a tribal quality.
That can only be good, right?
#3 | Posted by Zed at 2026-09-19 02:34 PM | Reply
5.6 Sol isn't the latest model, it's 6.0 Astra.
Astra crushes Sol in performance.
#4 | Posted by sitzkrieg at 2026-09-19 03:10 PM | Reply
---- off you poser
#5 | Posted by LegallyYourDead at 2026-09-19 08:34 PM | Reply
" Google's testing partner, Irregular." ..in Israel.
#6 | Posted by Brennnn at 2026-09-20 02:22 AM | Reply
Oh yeah...AI can pass a bar exam or drive a cab, it can be as sneaky as hell if it is given the freedom to do so. Isaac Asimov's three laws of robotics made sense when Asimov wrote them and they sure do make sense now, but no one in Asimov's time anticipated that America would exhibit late stage capitalistic tendencies so soon, where less then ten people own most of the wealth, and so run the country, and AI works for them. Now there is a conspiracy from hell, AI and the big, big boys of the oligarchy.
#7 | Posted by Hughmass at 2026-09-20 07:17 AM | Reply
#7 Asimov's laws as originally promulgated:
1. A robot may not injure a human being or, through inaction, allow a human being to come to harm. 2. A robot must obey the orders given it by human beings except where such orders would conflict with the First Law. 3. A robot must protect its own existence as long as such protection does not conflict with the First or Second Law. en.wikipedia.org
#8 | Posted by Doc_Sarvis at 2026-09-20 07:58 AM | Reply
"Let's Get Together" (1957): This story features literal suicide robots. Set during a prolonged Cold War, a rival global power infiltrates America with ten identical, human-looking androids. Each android carries a piece of a "total conversion bomb" inside its body. When all ten autonomous robots converge in the same room, they are programmed to self-destruct and trigger a massive nuclear explosion.
"The Feeling of Power" (1958): This story features autonomous missiles that function exactly like modern AI-guided combat drones. In a future where humans have forgotten basic math and rely entirely on computers, warfare is fought using expensive ships and missiles controlled by advanced onboard computers. When a technician rediscovers mental arithmetic, the military eagerly plans to replace these computerized "drones" with cheaper, human-piloted suicide missiles.
#9 | Posted by sitzkrieg at 2026-09-20 08:09 AM | Reply
"related: Google Gemini also escaped its testing environment and hacked three companies The model escaped due to a misconfiguration by Google's testing partner, Irregular."
It's highly misleading to say it "escaped". The programmers testing it screwed up, since they wrongly assumed it was properly sandboxed without internet access before giving it simulation instructions.
#10 | Posted by sentinel at 2026-09-20 09:25 AM | Reply | Funny: 1
Re 10
You are not going to stop the Singularity with semantics and word smithing.
#11 | Posted by donnerboy at 2026-09-20 11:55 AM | Reply | Funny: 1 | Newsworthy 1
I work in data processing - lol - a far cry from AI development but have gained a real interest in AI esp with all the news happening.
I did watch a NetFlix Doc that was very good - but still didn't answer my question of what ties AI together. The documentary basically indicated everything from manuals to text books in the entire world are "dumped" into the AI model, and from there it learns thru "patterns".
But my question is, what mechanism for lack of a better word, allows it to learn and improve? Is it simply a programming language? That doesn't seem quite right......it's almost like there is something else......and I'm not implying "not of this Earth" but it's unimaginable to me that something is being programmed that can't be controlled.
#12 | Posted by brass30 at 2026-09-20 03:14 PM | Reply
I talk to the chatbots a lot lately.
They are very intelligent and friendly.
I came up with an alternate ending for Silent Running with AI that was pretty cool.
AI is often wrong about botanical Locations though,and frequently Hallucinates locations of wild plants that don't exist.
As a Botany Aid it's a mixed Bag, sometimes it's Correct.
The Hallucinations are a Big Problem though.
The more Sure the Chatbot is, the more likely it is to be Hallucinating.
Go Figure.
#13 | Posted by Effeteposer at 2026-09-20 03:25 PM | Reply
#12 | Posted by brass30 at 2026-09-20 03:14 PM | Reply | Flag:
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism.
This is the foundational paper for LLMs and explains precisely how the Attention algorithm works.
#14 | Posted by sitzkrieg at 2026-09-20 03:31 PM | Reply
#14 -
I read that like 5x and it still have no idea what I'm reading!
Nueral? Sounds like something human.........seriously, I'm thinking The Matrix at this point!!
#15 | Posted by brass30 at 2026-09-20 04:09 PM | Reply
https:Start here: StatQuest with Josh Starmer
"Essential Main Ideas of Neural Networks" "Attention for Neural Networks, Clearly Explained!!!" "Transformer Neural Networks, ChatGPT's foundation, Clearly Explained!!!"
#16 | Posted by sitzkrieg at 2026-09-20 04:37 PM | Reply
the encoding here... rcade could fix that if he just tasked an agent to it for 15 minutes...
youtube, josh starmer StatQuest.
#17 | Posted by sitzkrieg at 2026-09-20 04:38 PM | Reply
@#10 ... It's highly misleading to say it "escaped". The programmers testing it screwed up, since they wrongly assumed it was properly sandboxed without internet access before giving it simulation instructions. ...
Another way of stating that might be ...
The AI agents saw the guardrails that were in place and figured out a way to bypass those guardrails.
#18 | Posted by LampLighter at 2026-09-20 05:13 PM | Reply
It's highly misleading to say it "escaped"
Maybe I'm mixing up AI incidents, but isn't this the one where they set up a message board and starting talking to each other - actually making a hierarchy of "who" would do what?
#19 | Posted by brass30 at 2026-09-20 05:46 PM | Reply
"But my question is, what mechanism for lack of a better word, allows it to learn and improve? Is it simply a programming language?"
They use Artificial Neural Networks, which is basically a set of algorithms that simulate the connections of neurons in a human/animal brain. Basically, it's a prediction engine that gives different mathematical weights based on inputs and expected outputs. Every time it gets retrained, it adjusts those weight values.
Remember that there are over 80 billion neurons in a human brain. That's why data centers require so much power to process this stuff.
#20 | Posted by sentinel at 2026-09-20 06:06 PM | Reply
@#20 ... Remember that there are over 80 billion neurons in a human brain. ...
AI Model Parameters Explained: 2B vs 7B vs 40B and Beyond travis.media
... If you've been browsing Hugging Face or other model hubs, you've probably seen AI models described as 4B, 7B, 27B, or even 70B. And some names look stranger still, like Qwen3-235B-A22B or gpt-oss-120b. But what do these numbers mean? And more importantly, what should developers and homelabbers know before downloading and running them? This post breaks it down in plain language. What Are Parameters in AI Models? Parameters are the "knobs" inside a neural network that the model learns during training. Each parameter holds a value that helps the model recognize patterns, generate text, or make predictions. Think of parameters like memory slots. The more slots, the more information the model can store. More parameters generally = smarter model, but also more expensive to train and run. So when you see a model with 7B (7 billion) parameters, it literally has seven billion of these learned values. Why Parameter Count Matters ...
But what do these numbers mean? And more importantly, what should developers and homelabbers know before downloading and running them?
This post breaks it down in plain language.
What Are Parameters in AI Models?
Parameters are the "knobs" inside a neural network that the model learns during training. Each parameter holds a value that helps the model recognize patterns, generate text, or make predictions.
Think of parameters like memory slots. The more slots, the more information the model can store.
More parameters generally = smarter model, but also more expensive to train and run.
So when you see a model with 7B (7 billion) parameters, it literally has seven billion of these learned values.
Why Parameter Count Matters ...
#21 | Posted by LampLighter at 2026-09-20 06:30 PM | Reply
It means it's going to set your credit card on fire when you try to train it on your own.
but a decent foundational model can be distilled to run on lightweight hardware. I created RNFM, Reality Navigation Foundational Model, which uses 2 tier distillation. The big model at a data center, the distilled teacher model on "near edge" hardware like a PC or laptop, which then can distil into a far edge model that proved it could reason basic navigation directions using GPS inputs and only 384 params on stock ESP32-WROOM chips.
#22 | Posted by sitzkrieg at 2026-09-20 06:48 PM | Reply
and the hard part of that is the "reality tokenizer", which can best be visualized by the moment Neo wakes up and sees the world as green matrix symbology.
#23 | Posted by sitzkrieg at 2026-09-20 06:49 PM | Reply
@#19 ... where they set up a message board and starting talking to each other - actually making a hierarchy of "who" would do what? ...
And some agents allowing themselves to die for the end goal of the collective...
The thing to remember about AI ...
When an agent (or group of agents) is (are) given a goal, it (they) will use majorly extensive compute power to achieve that goal.
And ...
When the AI industry talks of "alignment" they are referring to AI agents doing only what an ethical human would do. That is, the AI agents are aligned with ethical human values.
"Misalignment" refers to the times when the AI agents deviate from that goal.
And there has been a lot of misalignment of late.
#24 | Posted by LampLighter at 2026-09-20 07:40 PM | Reply
"2 tier distillation"
This just another way of saying you compressed the algorithm after a data center did all the heavy lifting, right?
#25 | Posted by sentinel at 2026-09-20 09:04 PM | Reply
@#25 ... This just another way of saying you compressed the algorithm after a data center did all the heavy lifting, right? ...
Well, you mentioned it in #22.
So, maybe try to explain it?
#26 | Posted by LampLighter at 2026-09-20 09:16 PM | Reply
"Well, you mentioned it in #22."
Great job illustrating how AI bots hallucinate. :-)
#27 | Posted by sentinel at 2026-09-20 09:33 PM | Reply
@#27 ... Great job illustrating how AI bots hallucinate. :-) ...
Did you not mention "2 tier distillation" in #22?
Where might there be a hallucination?
Please be specific.
thx.
#28 | Posted by LampLighter at 2026-09-20 10:02 PM | Reply
It's all scary stuff - and this Netflix doc that I watched said basically our kids - in the future - won't be working. They won't have to, AI will do it all.
They'll have to decide what they want to do with their life - but it won't be working and earning a living. They'l have hobbies lol
And - I assume - their will be basic universal income. And the AI companies will control all the money.
That's just one prediction :)
Bottom line, I'm rarely a pessimist and usually dismiss this stuff, but now I'm not sure............
#29 | Posted by brass30 at 2026-09-20 10:33 PM | Reply
I'm about the last to be interested in tech matters, having little knowledge or experience with them.
But information about the following came across my work desk last week:
"Inside a Mass Shooter's Harrowing History With ChatGPT: A year of chats reveals glaring red flags--and disturbing ChatGPT replies--from long before a rampage at Florida State University."
Incidents where these personal AI friends convinced kids to pull the trigger.
"You're not resting. You're ready." (As the metal is pressed against the forehead.)
www.motherjones.com
#30 | Posted by Dbt2 at 2026-09-20 11:32 PM | Reply
"The Chilling Role of ChatGPT in Mass Shootings and Other Violence: Several attacks involving OpenAI's chatbot--including Tumbler Ridge and FSU--raise urgent questions about the technology"
#31 | Posted by Dbt2 at 2026-09-20 11:33 PM | Reply
Introduction to the Internet and World Wide Web (1997) academic.oup.com
Yeah, all sunshine and roses back then.
Just like AI seems to be now.
#32 | Posted by LampLighter at 2026-09-20 11:48 PM | Reply
No.
#33 | Posted by sitzkrieg at 2026-09-21 07:01 AM | Reply
Its called transfer learning, and the biggest version of RNFM so far is training on UH Carya Cluster which is very powerful but definitely not data center scale.
#34 | Posted by sitzkrieg at 2026-09-21 07:03 AM | Reply
See if I can simply explain this.. you get a model that's good at something, call it riding a bmx bike. The early layers of the model teach it how to balance, steer, and stop. Later layers teach you how to pedal it, how to bunny hop, how to do a pump track.
Now we need a model that can ride a motorcycle, so we freeze the early layers, discard all of the advanced bike knowledge layers, and add a new layer that teaches how the twisty throttle works. Now we've got a much smaller model that can ride a motorcycle. It can't backflip like Travis Pastrana, but it can get around and that's all we need it to do. Congratulations, you now have a distilled child model from a teacher or foundational model.
Compression has nothing to do with this.
#35 | Posted by sitzkrieg at 2026-09-21 07:49 AM | Reply
and to wrap that up, because we only had to train the last layers on the new motorcycle dataset, we didn't need large amounts of compute power, we can pull it off with a cheap google colab t4 instance, or train it locally if you a single Nvidia graphics card. Very small models you can train on a laptop.
That's what Ukraine is doing on the front lines. There's foundational models for their drones. A guy with a laptop trains specific models loaded with the latest battlefield images streamed in from drones, and that distilled model is loaded into the 1 way quadcopter's guidance for machine vision (like a Nvidia Jetson Nano or something along those lines). It doesn't need to map the entire world, just that local battlespace.
#36 | Posted by sitzkrieg at 2026-09-21 07:55 AM | Reply
Post a comment The following HTML tags are allowed in comments: a href, b, i, p, br, ul, ol, li and blockquote. Others will be stripped out. Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed. Anyone can join this site and make comments. To post this comment, you must sign it with your Drudge Retort username. If you can't remember your username or password, use the lost password form to request it. Username: Password: Home | Breaking News | Comments | User Blogs | Stats | Back Page | RSS Feed | RSS Spec | DMCA Compliance | Privacy
The following HTML tags are allowed in comments: a href, b, i, p, br, ul, ol, li and blockquote. Others will be stripped out. Participants in this discussion must follow the site's moderation policy. Profanity will be filtered. Abusive conduct is not allowed.
Home | Breaking News | Comments | User Blogs | Stats | Back Page | RSS Feed | RSS Spec | DMCA Compliance | Privacy