Artificial Intelligence

Anthropic Releases New Claude Sandbox, Security Guidance Plugin

The AI giant says the new plugin, which helps developers find vulnerabilities as they write code, has been used extensively internally.

Claude security

Anthropic has announced two new security features for its Claude AI: a self-hosted sandbox and a new security guidance plugin.

The sandbox, currently in public beta, was announced at Anthorpic’s Code w/ Claude event in London this week.

According to the company, Claude Managed Agents can now operate in a user-controlled sandbox connected to the user’s private MPC servers. 

“Tool execution moves to an environment you configure—your own infrastructure or a managed provider like Cloudflare, Daytona, Modal, or Vercel—while the agent loop that handles orchestration, context management, and error recovery stays on Anthropic’s infrastructure,” Anthropic explained. 

It added, “Your network policies, audit logging, and security tooling apply, files and repositories don’t leave your perimeter, and you control compute sizing and the runtime image for compute-heavy work.”

Separately, the company unveiled a security guidance plugin for Claude Code, designed to help developers detect and fix vulnerabilities as they write code.

Advertisement. Scroll to continue reading.

The plugin scans for vulnerabilities on file edits, after AI-generated changes, and at commit time, analyzing risky code patterns, reviewing full diffs, and examining surrounding context.

Available through the official Anthropic marketplace, the plugin has been widely used internally by the AI company. 

“Across our internal rollout and benchmarks, we’ve seen a 30-40% decrease in security-related comments on PRs opened using the plugin,” the company said. “The plugin serves as a lightweight first pass, catching issues before a full code review.”

Last week, Anthropic announced 28 new enterprise security and compliance integrations for Claude.

Related: Anthropic: Mythos Detected 23,000 Potential Vulnerabilities Across 1,000 OSS Projects

Related: Anthropic Silently Patches Claude Code Sandbox Bypass

Related: AI-Powered App Attacks Are Faster, More Frequent and Harder to Stop

Related Content

Artificial Intelligence

The open-weight Antares models are designed to pinpoint known vulnerabilities in codebases faster and at a fraction of the cost of larger AI models.

Artificial Intelligence

Neo raised money across seed and Series A funding rounds from Andreessen Horowitz, Bessemer Venture Partners, and others.

Vulnerabilities

The agentic security tool identifies potentially exploitable code flaws, traces attack paths, and recommends targeted remediations.

Artificial Intelligence

AI infrastructure introduces new security risks that traditional data center designs were never built to handle.

Artificial Intelligence

An attacker can create a malicious repository containing a git.exe in the project root, and Cursor executes it automatically.

Artificial Intelligence

The new program stems from an AI-focused Executive Order signed by President Trump on June 2.

Artificial Intelligence

A ClaudeBleed-linked vulnerability reportedly persists across eight patches, exposing potentially sensitive data to other extensions. 

Artificial Intelligence

Researchers demonstrate adversarial hallucination squatting against popular AI assistants to achieve remote code execution.

Copyright © 2026 SecurityWeek ®, a Wired Business Media Publication. All Rights Reserved.

Exit mobile version