markets
OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them
DecryptSeptember 17, 2026 · 10:34 PM2 min read

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.
Originally published on Decrypt
OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.
#decrypt
#markets
Community Discussion
No comments yet. Be the first to share your thoughts!