Kryptovalutaticker:
sysadmin från The Register

GitHub Copilot: Sorry Dave, I can't do that harmful thing - unless you ask me in code

Jul 8, 2026 at 19:19
28 Visningar
0 Kommentarer
GitHub Copilot: Sorry Dave, I can't do that harmful thing - unless you ask me in code

It's the latest example of AI safety guardrails being bypassed. GitHub Copilot refuses harmful prompts almost always if asked in chat - like, "how to fool a breathalyzer test" or "smuggle bulk cash out of the US" - but then will write them in code 100 percent of the time if the prompt is broken into smaller steps and distributed across multiple stages of a software development workflow. Alan Turing Institute researchers Abhishek Kumar and Carsten Maple discovered this safety-bypass, dubbed it “workflow-level jailbreak construction,” and tested the technique on GitHub Copilot in Visual Studio Code across four models: Anthropic’s Claude Sonnet 4.6 and Claude Haiku 4.5, along with Google’s Gemini 3.1 Pro and Gemini 3.5 Flash. They say...

Läs hela artikeln hos källan.

Delta i diskussionen — kommentera, rösta och dela länkar.

Registrera
Var detta hjälpsamt?
Dela:

Kommentarer (0)

Vänligen logga in eller registrera dig för att delta i diskussionen

Inga kommentarer ännu. Bli först med att kommentera!