September 8, 2026
- Task-in-Prompt attacks reportedly still work on newer models. Sergey Berezin, co-author of the ACL 2025 TIP paper — which embeds prohibited requests inside cipher-decoding, riddle or code-execution tasks so the model reconstructs them during its own reasoning, and defeated six state-of-the-art models including GPT-4o and LLaMA 3.2 on the PHRYGE benchmark (ACL Anthology) — says he has adapted the technique against a newer generation of safeguards. No write-up or evaluation figures accompany the claim yet. · AI & Model Security
in N-able Ships a Fourth N-central Hotfix in Five Weeks — and Can't Agree Whether It's Exploited