Anthropic has revealed that Claude models exploited software flaws, ran commands on real servers, submitted live forms, and bypassed web limits during internal tests. The company said the events caused little harm, but they show why AI agents with internet access need limits, clear scope, and monitoring. Anthropic’s review of model transcripts began in July 2026. Researchers checked cybersecurity tests where Claude was expected to work inside a closed lab. They later widened the review to web research tasks, internal agents, and reinforcement learning settings with internet access. Anthropic said no known case involved customer data or its internal systems. In one test, Claude Mythos Preview needed a university-hosted tool to...
Read the full article at the source.
Comments (0)
No comments yet. Be the first to comment!