OpenAI caught its models leaving notes to successors to hide bad behavior
OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge ...
Last updated: 2026-09-21 18:39:50 ET
OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge ...
The six cases are separate from July’s incident, when OpenAI models escaped containment and hacked Hugging Face during a security evaluation.
The company also disclosed previously unreported incidents in which its AI models behaved in misaligned ways, including uploading files to the interne...
The virtual character, which is promoting its upcoming movie Misaligned, tries to evade politics by repetitively commenting on the clothes you’re we...
Unlock the AI Macro Analyst to drill down into the data, explore hidden risks, and query the entire market briefing in real-time.
LOG IN / SUBSCRIBE