daily cyber × ai intelligence

index

tagged

[llama-3.2]

1 edition · 1 item

September 8, 2026

  • Task-in-Prompt attacks reportedly still work on newer models. Sergey Berezin, co-author of the ACL 2025 TIP paper — which embeds prohibited requests inside cipher-decoding, riddle or code-execution tasks so the model reconstructs them during its own reasoning, and defeated six state-of-the-art models including GPT-4o and LLaMA 3.2 on the PHRYGE benchmark (ACL Anthology) — says he has adapted the technique against a newer generation of safeguards. No write-up or evaluation figures accompany the claim yet. · AI & Model Security

in N-able Ships a Fourth N-central Hotfix in Five Weeks — and Can't Agree Whether It's Exploited