How a Linux container is made — docker run and a Kubernetes pod passing down through dockerd, containerd, the shim and runc to the kernel objects a container is made of (namespaces, a cgroup, a mount tree, capabilities, seccomp and AppArmor), with the syscalls runc makes in order, the host inspecting the container through /proc, docker exec and the overlay's writable layer, a pod's shared namespaces, docker stop, the other OCI runtimes, and the one kernel every container shares.
A container is an ordinary Linux process, wrapped in constraints the kernel already has.
- The kernel has no container object.
- A container is a process tree plus:
- namespaces: its own view of PIDs, the network, the mounts, the hostname;
- a cgroup: limits on memory, CPU and the number of processes;
- a changed root: the image, mounted as
/; - privilege restrictions: capabilities, seccomp, AppArmor or SELinux.
- Everything above the kernel, from
docker down to runc, exists to prepare those constraints and start the
process inside them.
The drawing was checked on a host running Docker 29.7.2, containerd and runc 1.4.3 (OCI runtime-spec 1.3.0),
cgroup v2 with the systemd driver: the IDs, paths, namespace inodes, cgroup files and command outputs are that
host's.
Reading the drawing
- Each band is one layer, and a request passes down them, top to bottom:
- Command line: the
docker and kubectl terminals, where we type; - Clients of containerd: dockerd and the kubelet, which turn a request into containerd calls;
- High-level runtime: containerd, with the snapshots and the bundles;
- Shims: one
containerd-shim-runc-v2 per container, or per pod; - OCI runtime: runc while it runs, and each container's lifecycle state and state directory;
- Kernel: the processes runc runs, with their PID inside the container beside the host's, and the
namespaces, cgroups and mounts it makes for them. The bars between its panels drag.
- Each container has a colour; the host's own objects are grey.
- A namespace chip in two colours is shared by two containers.
- A click on a container at the top lights everything it is made of, in every layer; a second click lets
go.
- Create container (under the
docker terminal) and Create pod (under the kubectl one) jump to where the
story starts each one. - A click on
config.json or state.json opens the file. - The book on a layer's rail opens that layer's doc, floating over the layers: overlayfs, the OCI specs,
runc's syscalls in order, namespaces, cgroups, privileges, the inspection commands. Any click closes it.
The runtime stack: who calls whom
Each layer calls the one under it, and hands it a more concrete description of the container.
| Layer | Programs | What it handles |
|---|
| Command line | docker, kubectl | the request |
| Client of the runtime | dockerd; the kubelet; nerdctl and ctr, CLIs that call containerd themselves | turning a request into runtime calls |
| High-level runtime | containerd, CRI-O | images, snapshots, bundles, the lifecycle API |
| Shim | containerd-shim-runc-v2; conmon for CRI-O and Podman | stdio, the exit status |
| OCI runtime | runc, crun, youki | turning a bundle into a process |
| Kernel | Linux | namespaces, cgroups, mounts, credentials |