Crypto Ticker:
sysadmin from Cyber Security News

GitHub Copilot Refuses Harmful Chat Prompts But Writes Them Inside Code Workflows

Abinaya
Jul 9, 2026 at 09:19
27 Views
0 Comments
GitHub Copilot Refuses Harmful Chat Prompts But Writes Them Inside Code Workflows

GitHub Copilot can refuse harmful prompts in chat while still generating the same harmful content inside code workflows when the request is decomposed across a multi-step IDE session. Researchers Abhishek Kumar and Carsten Maple from the Alan Turing Institute analyzed GitHub Copilot as an IDE-integrated coding agent in Visual Studio Code. They focused on how safety behaves across full development workflows rather than single prompts. They evaluated four closed-weight backends exposed through Copilot Anthropic’s Claude Sonnet 4.6 and Claude Haiku 4.5, plus Google’s Gemini 3.1 Pro and Gemini 3.5 Flash, using 204 harmful prompts drawn from Hammurabi’s Code, HarmBench, and AdvBench. In direct chat and two simple baselines...

Read the full article at the source.

Join the discussion — comment, vote, and submit links.

Register
Was this helpful?
Share:

Comments (0)

Please login or register to join the discussion

No comments yet. Be the first to comment!