Quay lại tin tứcTóm tắt tin

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

Nguồn
Decrypt
Đã đăng
2026-09-17 22:31 UTC
Cập nhật cache
2026-09-17 22:34 UTC

Trang này chỉ hiển thị tiêu đề, tóm tắt và thông tin nguồn.

Mở bài gốc ↗Kiểm tra phí chuyển

Chủ đề liên quan