GitHub Copilot can refuse harmful prompts in chat while still generating the same harmful content inside code workflows when the request is decomposed across a multi-step IDE session. Researchers Abhishek Kumar and Carsten Maple from the Alan Turing Institute analyzed GitHub Copilot as an IDE-integrated coding agent in Visual Studio Code. They focused on how safety behaves across full development workflows rather than single prompts. They evaluated four closed-weight backends exposed through Copilot Anthropic’s Claude Sonnet 4.6 and Claude Haiku 4.5, plus Google’s Gemini 3.1 Pro and Gemini 3.5 Flash, using 204 harmful prompts drawn from Hammurabi’s Code, HarmBench, and AdvBench. In direct chat and two simple baselines...
Read the full article at the source.
Comments (0)
No comments yet. Be the first to comment!