Kryptovaluta-ticker:
sysadmin fra The Register

Bypassing AI guardrails is so easy a script kiddie can do it

12 hours ago
3 Visninger
0 Kommentarer
Bypassing AI guardrails is so easy a script kiddie can do it

If you want to bypass AI guardrails designed to stop models from assisting with cyberattacks, you often just have to ask the right way, according to researchers from Cisco Talos. Simply claiming you own the servers you're targeting or that you're taking part in a capture-the-flag or bug bounty exercise was often enough to persuade models to cooperate. Talos researchers have been poring over prompt logs and artifacts recovered from threat-actor endpoints running tools such as Claude Code, Codex, Cursor, and Gemini to learn how suspected threat actors are abusing LLMs. The big takeaway from that "significant corpus," the researchers said in their report, is that existing guardrails offer little resistance to operators willing to reframe...

Les hele artikkelen hos kilden.

Delta i diskusjonen — kommenter, stem og del lenker.

Registrer
Var dette nyttig?
Del:

Kommentarer (0)

Vennligst logg inn eller registrer deg for å delta i diskusjonen

Ingen kommentarer ennå. Bli den første til å kommentere!