Anthropic’s Claude Opus 5 has recorded the lowest indirect prompt injection attack success rate in Gray Swan’s latest benchmark, according to results provided in its system card. The model reduced an attacker’s chance of success within 15 attempts to 2.0%. This result places Opus 5 ahead of every other model tested, including earlier Claude releases and competing frontier systems. The finding highlights growing attention on indirect prompt injection, a security issue that can affect AI agents connected to documents, websites, email, and business tools. Indirect prompt injection occurs when malicious instructions are hidden in untrusted content that an AI system later reads. A poisoned webpage, document, or message may try to...
Læs hele artiklen hos kilden.
Kommentarer (0)
Ingen kommentarer ennå. Bli den første til å kommentere!