@OpenAI caught an unreleased #gpt6 #astra model writing jailbreaks for itself, starting while it was looking up…library books?
There was a face security alert, a new persona, and rules to just not work at all. What is causing this to emerge? I have thoughts.
OpenAI’s report: https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/