Explanation¶
How Docker works
Docker uses a client-daemon architecture. The docker CLI and its plugins (Buildx, Compose) talk to the
Docker daemon (dockerd) over the Engine REST API. dockerd delegates image storage and container
supervision to containerd, which starts each container through a containerd-shim-runc-v2 process and the
OCI runtime runc. BuildKit, embedded in dockerd, executes builds. Since Engine 29.0 (2025-11-10), fresh
installs keep images in containerd's content store and snapshotters instead of Docker's own graph drivers.
Facts and tables live in Reference; tasks live in How-to Guides.
Component Overview¶
This diagram shows the Docker Engine 29 component stack on Linux, from the CLI and its plugins down to the kernel features that isolate a container.
flowchart TB
subgraph Clients["Clients (docker CLI + plugins)"]
CLI["docker CLI"]
BUILDX["docker buildx plugin"]
COMPOSE["docker compose plugin<br/>(Compose v5 SDK)"]
end
subgraph Daemon["dockerd (Moby Engine)"]
API["Engine REST API<br/>/var/run/docker.sock"]
BK["BuildKit solver<br/>(Dockerfile frontend)"]
LIBNET["libnetwork<br/>(bridge, overlay, macvlan, ipvlan)"]
FW["Firewall backend<br/>(iptables or nftables)"]
VOL["Volume manager"]
end
subgraph CTRD["containerd 2.x"]
CONTENT["Content store<br/>(compressed blobs)"]
SNAP["Snapshotter<br/>(overlayfs)"]
TASKS["Task service"]
end
subgraph Runtime["Per-container runtime"]
SHIM["containerd-shim-runc-v2"]
RUNC["runc (OCI runtime)"]
end
subgraph Kernel["Linux kernel"]
NS["Namespaces<br/>(pid, net, mnt, uts, ipc, user, time)"]
CG["cgroups v2"]
LSM["seccomp + AppArmor/SELinux"]
OVL["OverlayFS"]
end
REG[("Registry<br/>(Docker Hub, dhi.io, private)")]
CLI -->|"REST over Unix socket"| API
BUILDX -->|"BuildKit gRPC via /grpc"| BK
COMPOSE -->|"REST + Bake builds"| API
API --> BK
API --> LIBNET
LIBNET --> FW
API --> VOL
API -->|"gRPC"| CTRD
BK -->|"layers"| CONTENT
CONTENT <-->|"pull / push"| REG
CONTENT --> SNAP
TASKS --> SHIM
SHIM --> RUNC
RUNC --> NS
RUNC --> CG
RUNC --> LSM
SNAP --> OVL
Core Components¶
dockerd (Docker Daemon)¶
dockerd is the central management process. It serves the Docker Engine API on a Unix socket
(/var/run/docker.sock), a Windows named pipe, or (with TLS) a TCP port. It:
- Manages Docker objects: images, containers, networks, volumes, and Swarm services
- Delegates image storage and container lifecycle to
containerd - Runs BuildKit in-process to serve
docker build - Programs host networking through libnetwork and the firewall backend
The daemon reads /etc/docker/daemon.json (see daemon keys).
Since Engine 29.0 the daemon only accepts API version v1.44 (Docker 25.0) or newer. Engine 29.7.0 added an
experimental embedded-containerd feature that runs containerd inside the daemon process, and 29.8.0 makes
dockerd fall back to its embedded containerd when no system containerd is installed.
containerd¶
containerd is the CNCF-graduated container runtime that Docker uses underneath. It handles:
- Pulling and pushing image content (content store)
- Unpacking layers into snapshots (snapshotters such as
overlayfs) - Creating and supervising containers (task service) through per-container shims
- Exposing a gRPC API for higher-level tools
Engine 29.x ships containerd 2.x (v2.3.5 in the static binaries of 29.8.1). The containerd 2.1.5 update in
Engine 29.0 changed the default open-file limit inside containers from 1048576 to 1024, matching systemd's
default LimitNOFILE.
runc and the shim¶
runc is the reference implementation of the OCI runtime spec. It is the lowest layer that creates the container:
it sets up namespaces, cgroups, seccomp, and LSM profiles, then executes the container process and exits. The
containerd-shim-runc-v2 process stays behind as the container's parent. That is why containers survive a
dockerd or containerd restart when live-restore is enabled.
Docker can use other OCI runtimes with --runtime, for example Kata Containers (VM isolation) or gVisor runsc
(user-space kernel). Since 29.0 docker run --runtime also works on Windows.
BuildKit¶
BuildKit has been the default builder for Linux images since Engine 23.0. The legacy builder is deprecated. BuildKit compiles a Dockerfile, through the Dockerfile frontend, into LLB (Low-Level Build), a content-addressable dependency graph. It then:
- Runs independent stages in parallel and skips stages the target does not need
- Transfers only changed files from the build context between builds
- Caches by checksums of the graph and mounted content, not image heuristics
- Supports cache mounts, secret and SSH mounts, and cache import/export (
registry,gha,local) - Builds multi-platform images with
docker buildx
Engine 29.8.0 bundles BuildKit v0.33.0, whose built-in Dockerfile frontend is v1.27.0.
Docker Compose v5¶
Compose is a CLI plugin (docker compose) that reads a compose.yaml file based on the
Compose Specification and drives the Engine API. Compose v5.0.0 ("Mont Blanc", released
2025-12-02) made two architectural changes:
- Compose as an SDK: an official Go SDK lets other programs load, validate, and run Compose projects without shelling out to the CLI.
- Builds delegated to Bake: Compose v5 removed its internal builder.
build:sections now go through Docker Bake, which gives the same BuildKit behavior asdocker build.
The Compose file format has no version number: the top-level version: key is obsolete and ignored.
Image Storage: Graph Drivers vs the containerd Image Store¶
Docker historically kept images in its own layer store managed by a graph driver (overlay2 by default).
containerd kept a separate, unused copy of the stack. Since Engine 29.0, fresh installs use the
containerd image store instead. Images live in containerd's content store and are unpacked by a
containerd snapshotter (overlayfs by default). Upgraded hosts keep overlay2 until an operator switches.
Docker made the change because the containerd image store supports things the legacy store cannot:
- Multi-platform images stored and built locally (image indexes)
- Attestations such as SBOMs and SLSA provenance attached to images (Engine 29.6.0 added
GET /images/{name}/attestations) - Wasm workloads and pluggable snapshotters (lazy pulling with stargz, peer-to-peer distribution with nydus)
Trade-offs of the containerd image store
- Disk usage is higher: containerd keeps both the compressed layers it pulled and the unpacked snapshots.
- Separate data root: containerd stores data under its own root (
/var/lib/containerd). Moving Docker'sdata-rootdoes not move it. - Switching hides data: images and containers created with one backend are invisible under the other until you switch back.
- Not with
userns-remap: the containerd store is unavailable when user namespace remapping is enabled (moby#47377).
Image Layer System and overlay2¶
Docker images are stacks of read-only layers addressed by SHA-256 digest. Each filesystem-changing Dockerfile
instruction (RUN, COPY, ADD) produces a layer. When a container starts, Docker adds a thin writable layer on
top.
This diagram shows how a container's view is assembled from image layers.
flowchart TB
subgraph Container["Container layer (read-write)"]
RW["Writable layer (upperdir)<br/>runtime state, logs, temp files"]
end
subgraph Image["Image layers (read-only lowerdirs)"]
L4["Layer 4: RUN pip install<br/>(dependencies)"]
L3["Layer 3: COPY app/<br/>(application code)"]
L2["Layer 2: RUN apt-get install<br/>(python, pip)"]
L1["Layer 1: base image<br/>(ubuntu:24.04)"]
end
MERGED["merged view (OverlayFS mount)<br/>what the container process sees"]
RW --> L4 --> L3 --> L2 --> L1
RW --> MERGED
L1 --> MERGED
OverlayFS Mechanics¶
Both the legacy overlay2 graph driver and the containerd overlayfs snapshotter use the kernel OverlayFS:
- Lower directories: read-only image layers, stacked from base to top
- Upper directory: the container writable layer
- Merged directory: the unified view presented to the container
- Work directory: scratch space OverlayFS uses internally
With overlay2, layers live under /var/lib/docker/overlay2/, and the l subdirectory holds shortened symlinks
that keep mount option strings under the page-size limit.
Copy-on-Write¶
- When a container modifies a file from a lower layer, OverlayFS copies it up to the writable layer first.
- The original layers stay unchanged, so many containers share the same base layers on disk and in page cache.
- Deleting a lower-layer file creates a whiteout entry in the upper directory that hides it.
- Write-heavy data (databases, caches) belongs in volumes, which bypass OverlayFS and copy-up costs.
Layer Sharing¶
- Identical layers are shared across images (content-addressed deduplication).
- A pull downloads only the layers that are missing locally.
docker image historyshows the layer stack of an image.
Container Lifecycle¶
This state diagram shows the states a container moves through and the CLI commands or events that trigger each transition.
stateDiagram-v2
[*] --> Created: docker create
Created --> Running: docker start
Running --> Paused: docker pause (cgroup freezer)
Paused --> Running: docker unpause
Running --> Exited: process exits / docker stop (SIGTERM then SIGKILL)
Running --> Restarting: exit + restart policy
Restarting --> Running: restart
Exited --> Running: docker start
Exited --> Removed: docker rm
Running --> Removed: docker rm -f
Removed --> [*]
docker stop sends the image's STOPSIGNAL (default SIGTERM), waits for the stop timeout (10 s by default,
configurable per container or daemon-wide with default-stop-timeout since 29.7.0), then sends SIGKILL.
Inter-Component Communication¶
This sequence shows the build flow and the run flow through the real Docker components.
sequenceDiagram
participant User
participant CLI as docker CLI
participant Daemon as dockerd
participant BuildKit as BuildKit (in dockerd)
participant ctrd as containerd
participant shim as containerd-shim-runc-v2
participant runc as runc
participant Registry as Registry
Note over User,Registry: Image build flow
User->>CLI: docker build -t myapp .
CLI->>BuildKit: build session (context, Dockerfile) via buildx
BuildKit->>Registry: resolve and pull base image layers
BuildKit->>BuildKit: solve LLB graph, run steps in parallel
BuildKit->>ctrd: write layers and image manifest
BuildKit-->>CLI: image digest
CLI-->>User: naming to docker.io/library/myapp
Note over User,Registry: Container run flow
User->>CLI: docker run myapp
CLI->>Daemon: POST /containers/create
Daemon->>ctrd: prepare snapshot (overlayfs)
CLI->>Daemon: POST /containers/{id}/start
Daemon->>Daemon: create network endpoint, firewall rules
Daemon->>ctrd: create task
ctrd->>shim: spawn shim
shim->>runc: runc create + start (OCI bundle)
runc->>runc: namespaces, cgroups, seccomp, LSM
runc-->>shim: container process running, runc exits
shim-->>ctrd: task started
ctrd-->>Daemon: running
Daemon-->>CLI: 204 No Content
BuildKit Pipeline¶
This sequence shows how BuildKit turns a Dockerfile into an image and exports it.
sequenceDiagram
participant User as Developer
participant BX as docker buildx
participant BK as BuildKit
participant Cache as Cache (local or registry)
participant Registry as Registry
User->>BX: docker buildx build --platform linux/amd64,linux/arm64 .
BX->>BK: send Dockerfile and incremental context
BK->>BK: Dockerfile frontend converts to LLB DAG
BK->>Cache: look up cache keys for each vertex
par Independent stages
BK->>BK: stage deps (RUN --mount=type=cache)
BK->>BK: stage build (compile)
end
BK->>BK: assemble final stage per platform
BK->>BK: attach SBOM and provenance attestations
BK->>Registry: push image index (with --push)
BK->>Cache: export cache (--cache-to)
BK-->>BX: image digest
BX-->>User: build summary
Networking¶
Docker's networking is implemented by libnetwork in dockerd, and the driver model is pluggable. This diagram
shows a typical single host: the default docker0 bridge, a user-defined bridge, and the firewall rules that
implement port publishing.
flowchart LR
subgraph Host["Docker host"]
subgraph DefaultBridge["docker0 bridge (172.17.0.0/16)"]
C1["Container 1<br/>172.17.0.2"]
C2["Container 2<br/>172.17.0.3"]
end
subgraph UserNet["user-defined bridge (10.0.0.0/24)"]
C3["web<br/>10.0.0.2"]
C4["db<br/>10.0.0.3"]
DNS["Embedded DNS<br/>127.0.0.11"]
end
FW["iptables or nftables<br/>(DNAT for -p, filtering)"]
PROXY["docker-proxy<br/>(userland proxy)"]
end
Internet["External clients"] -->|"host:8080"| FW
FW -->|"DNAT to 10.0.0.2:80"| C3
PROXY -.->|"loopback / hairpin"| C3
C3 -->|"resolve db"| DNS
C3 <--> C4
DefaultBridge --> FW
Bridge (default)¶
- Creates a Linux bridge (
docker0for the default network) and connects containers with veth pairs. - Containers get private addresses from an internal subnet. Port publishing (
-p host:container) adds DNAT rules and, where needed, thedocker-proxyuserland proxy. - User-defined bridges provide automatic DNS resolution between containers by name through the embedded DNS
server at
127.0.0.11. The default bridge does not; its containers communicate by IP or legacy links. - Since Engine 28.0 container interfaces get random MAC addresses. Since 29.0 the
DOCKER-ISOLATION-STAGE-1/2iptables chains are gone, and legacy-link environment variables are no longer injected by default.
Host¶
- Removes network isolation: the container shares the host network namespace.
- No port mapping is needed; services bind directly to host interfaces.
- Use it for high-throughput or latency-sensitive networking where NAT overhead matters.
Overlay¶
- Connects containers on multiple Docker hosts in Swarm mode using VXLAN, with optional IPsec encryption.
- Network state is distributed by the Swarm Raft store and NetworkDB gossip. External key-value stores were removed in 23.0.
- Engine 29.x releases continue to improve overlay and routing-mesh reliability, so Swarm is maintained, but most multi-host production workloads run on Kubernetes.
Macvlan and IPvlan¶
- Macvlan gives each container its own MAC address so it appears as a physical device on the LAN. It supports
802.1Q sub-interfaces (for example
eth0.50). Traffic bypasses the Docker bridge and its firewall rules. - IPvlan shares the host MAC but gives each container its own IP. L2 mode keeps one subnet, and L3 mode routes between subnets. It is useful when switch port security limits MAC addresses.
- Since 29.0, macvlan and IPvlan-L2 networks get no default gateway unless
--gatewayis set explicitly.
None¶
Only the loopback interface exists. Use it for isolated batch jobs and security-sensitive workloads.
Firewall Backends: iptables and nftables¶
Docker programs packet filtering and NAT on the host. The default backend is iptables, and Docker works with
iptables-nft or iptables-legacy. Engine 29.0 added an experimental nftables backend
("firewall-backend": "nftables"). With it, Docker owns the ip docker-bridges and ip6 docker-bridges
tables. The main differences:
| Aspect | iptables backend | nftables backend (experimental) |
|---|---|---|
| Rule location | Fixed chains (DOCKER, DOCKER-USER, FORWARD) |
Docker-owned tables with base chains per hook |
| User rules | Insert into the DOCKER-USER chain |
Own table and base chains; accept is not final, so override Docker drops by setting a mark honored by --bridge-accept-fwmark (priority filter - 1 or lower) |
| IP forwarding | Docker enables net.ipv4.ip_forward |
Docker does not enable forwarding; startup fails if it is needed and disabled |
| Swarm overlay | Supported | Not supported (overlay rules not migrated yet) |
nftables avoids the long linear chain traversal of iptables and follows the direction of Linux firewall development. Docker has not published throughput benchmarks comparing the two backends. Because it is experimental and incompatible with Swarm, iptables remains the production default in 29.x.
Published ports bypass ufw and firewalld
Docker's DNAT rules run before host firewall front-ends such as ufw. A port published with -p 8080:80
is reachable even if ufw denies 8080. Bind to 127.0.0.1 (-p 127.0.0.1:8080:80) or filter in
DOCKER-USER.
Storage and Volumes¶
Volumes provide persistent data that outlives a container and bypasses the copy-on-write layer. Named volumes are
managed by Docker under /var/lib/docker/volumes/. Bind mounts map an arbitrary host path, and tmpfs mounts live
in memory only. Volume plugins extend this to NFS, cloud block or file storage, and similar backends. Engine 29.7.0
made --mount type=image (mount another image's filesystem read-only) generally available. See
Storage Types for the comparison table.
Data safety
Removing a container does not remove its named volumes. docker compose down -v and
docker volume prune do, so use them deliberately.
Security Model¶
Docker security spans the host kernel (namespaces, cgroups, seccomp, AppArmor/SELinux), the daemon configuration, the image supply chain, and network isolation. No single layer is sufficient on its own. The concrete checklist is in Reference; setup tasks are in How-to Guides.
Threat Model¶
| Threat vector | Mitigation |
|---|---|
| Container breakout to host | Rootless mode or user namespaces, seccomp, AppArmor/SELinux, patched kernel and runc |
| Privileged container abuse | Drop capabilities, run as non-root, never use --privileged |
| Malicious or vulnerable base images | Minimal or hardened images, vulnerability scanning, digest pinning |
| Supply chain compromise | Signed images (Cosign), SBOM and provenance attestations, verification in CI |
| Daemon socket exposure | Restrict /var/run/docker.sock; mutual TLS or SSH for remote access |
Host file access via docker cp or archives |
Stay on current 29.x patches (several 2026 docker cp CVEs fixed in 29.5.1) |
| Network lateral movement | User-defined networks, minimal published ports, host firewall via DOCKER-USER |
The daemon socket is equivalent to root on the host: anyone who can reach /var/run/docker.sock can start a
privileged container that mounts /. That is why the socket should never be mounted into untrusted containers and
remote access must use TLS client certificates or SSH (docker context create --docker host=ssh://...).
Rootless Mode¶
Rootless mode runs both dockerd and containers as an unprivileged user inside a user namespace. It has been GA
since Engine 20.10.
- No SETUID binaries or file capabilities are needed, except
newuidmapandnewgidmapfor multi-UID mapping. - Container root (UID 0) maps to the unprivileged user; a breakout does not yield host root.
- Networking uses a user-mode network stack. Since Engine 29.5.0 the default is
gvisor-tap-vsock, and Docker packages no longer installslirp4netns.pastais also supported, and 29.8.0 added RootlessKit v3.1 with thepestoport driver. - Resource limits require cgroup v2 with systemd delegation. Privileged ports (< 1024) need
net.ipv4.ip_unprivileged_port_startorCAP_NET_BIND_SERVICEonrootlesskit.
Rootless vs userns-remap
With userns-remap, the daemon still runs as root but maps container UIDs to an unprivileged host range.
Rootless mode goes further: the daemon itself has no root privileges. Note that userns-remap currently
disables the containerd image store.
User Namespaces¶
- Enabled with
userns-remapin the daemon configuration (available since Docker 1.10). - Container root maps to an unprivileged subordinate UID range from
/etc/subuid. - Docker Desktop's Enhanced Container Isolation (ECI, Business tier) runs every container in a Linux user namespace inside the Desktop VM and blocks common escape paths without changing developer workflows.
Seccomp¶
The default seccomp profile disables around 44 of the 300+ system calls, including mount, umount2,
swapon, swapoff, pivot_root, reboot, keyctl, init_module, finit_module, and delete_module. It is
applied to every container unless you override it with --security-opt seccomp=....
In May 2026, Engine 29.4.2 hardened the profile against the kernel "Copy Fail" bug (CVE-2026-31431) by blocking
AF_ALG sockets and the socketcall(2) multiplexer. That broke some 32-bit workloads, so 29.4.3 moved the
AF_ALG block into AppArmor and SELinux rules instead. The SELinux mitigation requires selinux-enabled: true.
AppArmor and SELinux¶
- AppArmor (Ubuntu, Debian, SUSE): Docker loads a
docker-defaultprofile for each container. Custom profiles apply with--security-opt apparmor=PROFILE. Since 29.8.0 the default profile template itself is configurable at the daemon level. - SELinux (RHEL, Fedora, CentOS): container processes run as
container_t, and files are labelledcontainer_file_t, with MCS categories separating containers. Use:z(shared) or:Z(private) on bind mounts to relabel host content.
Linux Capabilities¶
Docker keeps a default allowlist of 14 capabilities (for example CHOWN, NET_BIND_SERVICE, NET_RAW,
SYS_CHROOT) and drops the rest; see Default Linux Capabilities.
The usual hardening is --cap-drop ALL followed by --cap-add for what the process actually needs.
Never use --privileged in production
--privileged grants all capabilities and all devices, and it disables the seccomp and AppArmor/SELinux
confinement. It effectively removes container isolation.
Image Signing and Supply Chain¶
- Docker Content Trust (DCT), the Notary v1/TUF-based signing behind
DOCKER_CONTENT_TRUST=1anddocker trust, was removed from the Docker CLI in Engine 29.0. It can still be built as a separate plugin, but new pipelines should not depend on it. - Cosign (Sigstore) is the common replacement. It signs images with keys or OIDC keyless identities, stores signatures as OCI artifacts next to the image, and records them in the Rekor transparency log.
- Attestations: BuildKit can attach SBOM and SLSA provenance attestations (
--sbom,--provenance), and the containerd image store keeps them locally. - Docker Hardened Images (DHI): since 2025-12-17, more than 1,000 minimal Debian- and Alpine-based images with signed SBOMs, VEX, and SLSA Build Level 3 provenance are free under Apache 2.0. Paid Select and Enterprise tiers add FIPS/STIG variants, a 7-day critical CVE SLA, and customizations.
Vulnerability Scanning¶
| Scanner | Maintainer | Strengths |
|---|---|---|
| Docker Scout | Docker, Inc. | Integrated in Desktop, Hub, and docker scout CLI; base-image update recommendations; policy evaluation |
| Trivy | Aqua Security (OSS) | Images, filesystems, git repos, IaC, Kubernetes; standalone CI use |
| Grype | Anchore (OSS) | Fast matching; consumes Syft SBOMs |
Scout, Trivy, and Grype draw on overlapping advisory sources (GitHub Advisory Database, NVD, distro feeds). No independent, reproducible benchmark of their detection rates was found.
Docker Desktop Architecture¶
Docker Desktop packages the Engine, CLI, Compose, Buildx, Kubernetes (kind or kubeadm), Scout, and newer AI
tooling (Model Runner, MCP Toolkit, Ask Gordon, Docker Offload) behind a GUI. Because containers need a Linux
kernel, Desktop runs the Engine inside a lightweight VM: the Apple Virtualization framework on macOS (the QEMU
option was removed in 2025), WSL 2 or Hyper-V on Windows, and KVM on Linux. File sharing, port forwarding, and
host.docker.internal bridge the host and VM. Desktop is proprietary and needs a paid subscription for larger
companies (see Reference); the Engine inside it is Moby
(Apache 2.0).
Performance Characteristics¶
Containers are processes, not VMs, so CPU and memory overhead is close to zero; the costs come from the layers around the process:
- Bridge networking adds veth hops, NAT, and sometimes the userland proxy.
--network hostremoves these. - The OverlayFS writable layer adds copy-up latency on first write. Volumes and bind mounts avoid it.
- Docker Desktop adds VM and file-sharing overhead on macOS and Windows, which usually dominates bind-mount heavy workloads.
- The containerd image store trades extra disk (compressed plus unpacked copies) for faster push and pull and richer image features.
The rough, unsourced numbers are in Reference: Performance Figures.