Why a runc container's isolation rests on the one host kernel, and the runtimes that put another kernel in between — gVisor's Sentry, a kernel in Go in user space, with its gofers, two layers of namespaces and the systrap platform; Kata Containers' VM per pod and Firecracker — chosen per pod with RuntimeClass, side by side on one Kubernetes node, with Agent Sandbox's Sandbox, SandboxTemplate, SandboxWarmPool and SandboxClaim on top.
A container's isolation is only as strong as the code that serves its syscalls.
- With runc, that code is the host kernel, shared by every container on the node:
- namespaces, cgroups and seccomp are checks inside it;
- one bug in a syscall path the container can reach is code execution in the host kernel.
- A sandboxed runtime puts another kernel in between:
- gVisor: the Sentry, a kernel written in Go, running as an ordinary process;
- Kata Containers: a small VM per pod, with a Linux kernel of its own.
- AI agents make this urgent: they run code they wrote themselves, and anything they read can steer it.
The drawing builds on
Linux Containers: the Runtime Stack
, which shows how runc
makes a container.
The gVisor details were checked on gVisor release-20260928.0, run on a Linux host: the platform, the
processes and namespaces the host sees, and the RPCs that add a container to a sandbox.
The Kata details follow Kata Containers 4.2, and Agent Sandbox v1.0.4, whose API is v1beta1.
Reading the drawing
- The drawing is one Kubernetes node, as bands, top to bottom:
- Kubernetes API: our
kubectl, the RuntimeClasses, the pods, and later Agent Sandbox's objects; - The node: the kubelet, and containerd with its runtime handlers.
- Under them, three lanes, one per runtime: runc, gVisor, Kata. In each lane:
- Shims, runtimes: the programs containerd starts for the pod;
- Inside the containers: the processes, as they see themselves;
- Serving the syscalls: nothing for runc, the Sentry and its gofers for gVisor, a VM for Kata;
- Host kernel: the host processes, the real namespaces and the guards; under the three, the node's one
kernel and who calls it from each lane.
- The lane heads place each runtime on the isolation spectrum, with its trade-offs.
- A click on a pod in the Kubernetes API band lights everything it is made of; a second click lets go.
- The book on a layer's rail opens that layer's doc: the spectrum's trade-offs, RuntimeClass and Agent
Sandbox's objects, containerd's config, the shims, the Sentry and the gVisor platforms, the host's view.
Any click closes it.
The three runtimes side by side
| runc | gVisor | Kata Containers |
|---|
| What serves a syscall | the host kernel | the Sentry, in user space | a guest kernel, in a VM |
| Host processes per pod | the shim, every container process | the shim, the Sentry and its stubs, a gofer per container | the shim, QEMU, virtiofsd |
| What the pod reaches in the host kernel | every syscall seccomp allows | the Sentry's syscalls, under 100 | KVM and the VMM's devices |
| RuntimeClass | none: the default | gvisor → handler runsc | kata-qemu-runtime-rs (kata-deploy) |
| Shim | containerd-shim-runc-v2 | containerd-shim-runsc-v1 | containerd-shim-kata-v2 |
Links