Kryptovalutaticker:
sysadmin från Cyber Security News

GitHub Copilot Refuses Harmful Chat Prompts But Writes Them Inside Code Workflows

Abinaya
Jul 9, 2026 at 09:19
26 Visningar
0 Kommentarer
GitHub Copilot Refuses Harmful Chat Prompts But Writes Them Inside Code Workflows

GitHub Copilot can refuse harmful prompts in chat while still generating the same harmful content inside code workflows when the request is decomposed across a multi-step IDE session. Researchers Abhishek Kumar and Carsten Maple from the Alan Turing Institute analyzed GitHub Copilot as an IDE-integrated coding agent in Visual Studio Code. They focused on how safety behaves across full development workflows rather than single prompts. They evaluated four closed-weight backends exposed through Copilot Anthropic’s Claude Sonnet 4.6 and Claude Haiku 4.5, plus Google’s Gemini 3.1 Pro and Gemini 3.5 Flash, using 204 harmful prompts drawn from Hammurabi’s Code, HarmBench, and AdvBench. In direct chat and two simple baselines...

Läs hela artikeln hos källan.

Delta i diskussionen — kommentera, rösta och dela länkar.

Registrera
Var detta hjälpsamt?
Dela:

Kommentarer (0)

Vänligen logga in eller registrera dig för att delta i diskussionen

Inga kommentarer ännu. Bli först med att kommentera!