返回新闻中心新闻摘要

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

来源
Decrypt
发布时间
2026-09-17 22:31 UTC
缓存更新
2026-09-17 22:34 UTC

本页仅展示标题、摘要与来源信息,完整内容请访问原新闻源。

打开原文 ↗查看手续费

相关专题