Docker's Agent Sandbox Strategy: What Developers Need to Know

Docker's Agent Sandbox Strategy: What Developers Need to Know

I’ve been watching Docker’s moves in the AI space closely, and their announcement at WeAreDevelopers World Congress North America feels significant. They’re not just building another container tool. They’re addressing a real problem that most teams still haven’t figured out: how do you give AI agents meaningful work without losing control of your infrastructure?

The core idea is elegant. Docker Cloud Sandboxes let developers start agent work locally on their laptop, then move that sandbox to managed cloud compute when tasks get longer or more resource-intensive. When you’re done, you bring the results back. Same command-line tool, same workflow, different execution layer.

The Real Problem They’re Solving

Mark Cavage laid out the four requirements for what he calls an “agent factory”: containment, control, choice, and capacity. That framework resonated with me because it maps to actual pain points I hear from teams building with agents.

Containment means knowing where an agent can act. Control means seeing and stopping its activity. Choice means not being locked into one model or tool. Capacity means running work beyond a single developer’s laptop. Docker Sandboxes addresses all four, but the containment piece is where the real innovation sits.

The keynote demo was telling. An agent running in a container with the host Docker socket mounted could read secrets from the host. Not because of a vulnerability, but because the configuration granted that access. The same agent attempt inside a sandbox failed. The microVM boundary enforced by infrastructure, not configuration, stopped it cold. That’s the difference between permission systems and architectural guarantees.

Kits and Governance at Scale

The Docker Sandbox Kit specification is the piece that could matter most for teams. A Kit is an OCI image containing an agent, its tools, and declarations of what access it needs. Teams can build, push, pull, and scan them like any container image. More importantly, they can review changes to permissions as part of code review.

This is where governance becomes practical. If a Kit suddenly asks for network access to a new destination or a new credential, that change shows up in your manifest. Reviewers see it. The runtime enforces the policy, not the agent.

Tushar Jain’s keynote focused on governing at the runtime layer rather than the agent layer. Different teams use different models and tools. Governing each tool separately creates a fractured security posture. A common runtime layer that enforces execution, tools, credentials, and permissions gives you visibility across all your agents, regardless of what model or framework is running underneath. His demo showed an agent’s request to delete a GitHub repository getting blocked by a default-deny rule at the runtime level. That enforcement came from infrastructure, not the agent.

What This Means for Your Workflow

If you’re building with AI agents today, you’re probably running them locally or managing your own cloud deployments. Cloud Sandboxes collapse that complexity. You iterate locally, scale to the cloud when you need it, and maintain the same environment across both contexts.

The pricing model (billed by the second) aligns incentives nicely. You’re not paying for idle compute. Longer agent tasks that would drain your laptop battery now run on Docker-managed infrastructure. Interactive work stays local. That’s a practical division of labor.

But here’s what caught my attention: Docker’s committing to bring the Kit specification to the CNCF for neutral governance. That’s not a small decision. It signals that this isn’t a Docker-specific solution. Other runtimes can implement the spec. Developers get a common language for describing agent environments. That’s how an idea becomes infrastructure.

The Honest Limitations

Mark Lechner, Docker’s CISO, was refreshingly candid about what these controls can’t do. Blocking an unapproved network destination is straightforward. Recognizing that an otherwise permitted email is going to the wrong customer is not. Narrow permissions matter. An agent that can read and draft messages without sending them is still safer than one that can send. But runtime controls are not a substitute for intent verification or human review.

Agent identity remains an open industry challenge, Tushar acknowledged in the Q&A. You can have team-specific policies that give different access to different users running the same agent. But proving which agent actually took an action still needs work.

Where This Leads

I see three implications for developers. First, sandboxed agent execution is becoming standardized infrastructure, not a security research project. Second, governance is moving to the runtime layer, which means less friction for teams trying to enforce policy at scale. Third, the ecosystem is beginning to coalesce around common specifications. That maturity matters.

The workshops and customer sessions showed real workloads: Spectro Cloud running repeatable agent tasks at the edge, J.P. Morgan Payments starting with docker compose, ClickHouse hardening images, agents exchanging work in separate sandboxes. These aren’t hypotheticals.

If you want to understand [https://mgks.dev/tags/ai-agents/](AI agents) as infrastructure rather than experimentation, this is the direction the industry is moving. And if you’re building systems that need to govern those agents, understanding the [https://mgks.dev/tags/security/](runtime security layer) they run in is becoming essential.

The question isn’t whether agents will become more autonomous. It’s whether your infrastructure will let you trust them when they do.

Read Next