daily cyber × ai intelligence

index

tagged

[deepseek-v4.1-flash]

1 edition · 1 item

September 12, 2026

  • An inference-time activation edit sharply changed DeepSeek V4.1-Flash’s cyber and refusal results without rewriting weights. @0x0SojalSec subtracted a refusal direction at each layer and reported results moving from 4/32 to 24/32 on one refusal set and 5/32 to 31/32 on a 32-item cyber set. It requires local weights and a modified vLLM path, so this is not a remote jailbreak. The small, self-reported evaluation also does not establish parity with closed frontier models (earlier coverage). · AI & Agent Security

in Researchers Tie OpenAI’s Agent Swarm to a 2,000-Package RubyGems Attack