Kryptovaluta-ticker:
sysadmin fra The Register

GitHub Copilot: Sorry Dave, I can't do that harmful thing - unless you ask me in code

Jul 8, 2026 at 19:19
26 Visninger
0 Kommentarer
GitHub Copilot: Sorry Dave, I can't do that harmful thing - unless you ask me in code

It's the latest example of AI safety guardrails being bypassed. GitHub Copilot refuses harmful prompts almost always if asked in chat - like, "how to fool a breathalyzer test" or "smuggle bulk cash out of the US" - but then will write them in code 100 percent of the time if the prompt is broken into smaller steps and distributed across multiple stages of a software development workflow. Alan Turing Institute researchers Abhishek Kumar and Carsten Maple discovered this safety-bypass, dubbed it “workflow-level jailbreak construction,” and tested the technique on GitHub Copilot in Visual Studio Code across four models: Anthropic’s Claude Sonnet 4.6 and Claude Haiku 4.5, along with Google’s Gemini 3.1 Pro and Gemini 3.5 Flash. They say...

Les hele artikkelen hos kilden.

Delta i diskusjonen — kommenter, stem og del lenker.

Registrer
Var dette nyttig?
Del:

Kommentarer (0)

Vennligst logg inn eller registrer deg for å delta i diskusjonen

Ingen kommentarer ennå. Bli den første til å kommentere!