Skip to content

Explanation

How Docker works

Docker uses a client-daemon architecture. The docker CLI and its plugins (Buildx, Compose) talk to the Docker daemon (dockerd) over the Engine REST API. dockerd delegates image storage and container supervision to containerd, which starts each container through a containerd-shim-runc-v2 process and the OCI runtime runc. BuildKit, embedded in dockerd, executes builds. Since Engine 29.0 (2025-11-10), fresh installs keep images in containerd's content store and snapshotters instead of Docker's own graph drivers.

Facts and tables live in Reference; tasks live in How-to Guides.

Component Overview

This diagram shows the Docker Engine 29 component stack on Linux, from the CLI and its plugins down to the kernel features that isolate a container.

flowchart TB
    subgraph Clients["Clients (docker CLI + plugins)"]
        CLI["docker CLI"]
        BUILDX["docker buildx plugin"]
        COMPOSE["docker compose plugin<br/>(Compose v5 SDK)"]
    end

    subgraph Daemon["dockerd (Moby Engine)"]
        API["Engine REST API<br/>/var/run/docker.sock"]
        BK["BuildKit solver<br/>(Dockerfile frontend)"]
        LIBNET["libnetwork<br/>(bridge, overlay, macvlan, ipvlan)"]
        FW["Firewall backend<br/>(iptables or nftables)"]
        VOL["Volume manager"]
    end

    subgraph CTRD["containerd 2.x"]
        CONTENT["Content store<br/>(compressed blobs)"]
        SNAP["Snapshotter<br/>(overlayfs)"]
        TASKS["Task service"]
    end

    subgraph Runtime["Per-container runtime"]
        SHIM["containerd-shim-runc-v2"]
        RUNC["runc (OCI runtime)"]
    end

    subgraph Kernel["Linux kernel"]
        NS["Namespaces<br/>(pid, net, mnt, uts, ipc, user, time)"]
        CG["cgroups v2"]
        LSM["seccomp + AppArmor/SELinux"]
        OVL["OverlayFS"]
    end

    REG[("Registry<br/>(Docker Hub, dhi.io, private)")]

    CLI -->|"REST over Unix socket"| API
    BUILDX -->|"BuildKit gRPC via /grpc"| BK
    COMPOSE -->|"REST + Bake builds"| API
    API --> BK
    API --> LIBNET
    LIBNET --> FW
    API --> VOL
    API -->|"gRPC"| CTRD
    BK -->|"layers"| CONTENT
    CONTENT <-->|"pull / push"| REG
    CONTENT --> SNAP
    TASKS --> SHIM
    SHIM --> RUNC
    RUNC --> NS
    RUNC --> CG
    RUNC --> LSM
    SNAP --> OVL

Core Components

dockerd (Docker Daemon)

dockerd is the central management process. It serves the Docker Engine API on a Unix socket (/var/run/docker.sock), a Windows named pipe, or (with TLS) a TCP port. It:

  • Manages Docker objects: images, containers, networks, volumes, and Swarm services
  • Delegates image storage and container lifecycle to containerd
  • Runs BuildKit in-process to serve docker build
  • Programs host networking through libnetwork and the firewall backend

The daemon reads /etc/docker/daemon.json (see daemon keys). Since Engine 29.0 the daemon only accepts API version v1.44 (Docker 25.0) or newer. Engine 29.7.0 added an experimental embedded-containerd feature that runs containerd inside the daemon process, and 29.8.0 makes dockerd fall back to its embedded containerd when no system containerd is installed.

containerd

containerd is the CNCF-graduated container runtime that Docker uses underneath. It handles:

  • Pulling and pushing image content (content store)
  • Unpacking layers into snapshots (snapshotters such as overlayfs)
  • Creating and supervising containers (task service) through per-container shims
  • Exposing a gRPC API for higher-level tools

Engine 29.x ships containerd 2.x (v2.3.5 in the static binaries of 29.8.1). The containerd 2.1.5 update in Engine 29.0 changed the default open-file limit inside containers from 1048576 to 1024, matching systemd's default LimitNOFILE.

runc and the shim

runc is the reference implementation of the OCI runtime spec. It is the lowest layer that creates the container: it sets up namespaces, cgroups, seccomp, and LSM profiles, then executes the container process and exits. The containerd-shim-runc-v2 process stays behind as the container's parent. That is why containers survive a dockerd or containerd restart when live-restore is enabled.

Docker can use other OCI runtimes with --runtime, for example Kata Containers (VM isolation) or gVisor runsc (user-space kernel). Since 29.0 docker run --runtime also works on Windows.

BuildKit

BuildKit has been the default builder for Linux images since Engine 23.0. The legacy builder is deprecated. BuildKit compiles a Dockerfile, through the Dockerfile frontend, into LLB (Low-Level Build), a content-addressable dependency graph. It then:

  • Runs independent stages in parallel and skips stages the target does not need
  • Transfers only changed files from the build context between builds
  • Caches by checksums of the graph and mounted content, not image heuristics
  • Supports cache mounts, secret and SSH mounts, and cache import/export (registry, gha, local)
  • Builds multi-platform images with docker buildx

Engine 29.8.0 bundles BuildKit v0.33.0, whose built-in Dockerfile frontend is v1.27.0.

Docker Compose v5

Compose is a CLI plugin (docker compose) that reads a compose.yaml file based on the Compose Specification and drives the Engine API. Compose v5.0.0 ("Mont Blanc", released 2025-12-02) made two architectural changes:

  • Compose as an SDK: an official Go SDK lets other programs load, validate, and run Compose projects without shelling out to the CLI.
  • Builds delegated to Bake: Compose v5 removed its internal builder. build: sections now go through Docker Bake, which gives the same BuildKit behavior as docker build.

The Compose file format has no version number: the top-level version: key is obsolete and ignored.

Image Storage: Graph Drivers vs the containerd Image Store

Docker historically kept images in its own layer store managed by a graph driver (overlay2 by default). containerd kept a separate, unused copy of the stack. Since Engine 29.0, fresh installs use the containerd image store instead. Images live in containerd's content store and are unpacked by a containerd snapshotter (overlayfs by default). Upgraded hosts keep overlay2 until an operator switches.

Docker made the change because the containerd image store supports things the legacy store cannot:

  • Multi-platform images stored and built locally (image indexes)
  • Attestations such as SBOMs and SLSA provenance attached to images (Engine 29.6.0 added GET /images/{name}/attestations)
  • Wasm workloads and pluggable snapshotters (lazy pulling with stargz, peer-to-peer distribution with nydus)

Trade-offs of the containerd image store

  • Disk usage is higher: containerd keeps both the compressed layers it pulled and the unpacked snapshots.
  • Separate data root: containerd stores data under its own root (/var/lib/containerd). Moving Docker's data-root does not move it.
  • Switching hides data: images and containers created with one backend are invisible under the other until you switch back.
  • Not with userns-remap: the containerd store is unavailable when user namespace remapping is enabled (moby#47377).

Image Layer System and overlay2

Docker images are stacks of read-only layers addressed by SHA-256 digest. Each filesystem-changing Dockerfile instruction (RUN, COPY, ADD) produces a layer. When a container starts, Docker adds a thin writable layer on top.

This diagram shows how a container's view is assembled from image layers.

flowchart TB
    subgraph Container["Container layer (read-write)"]
        RW["Writable layer (upperdir)<br/>runtime state, logs, temp files"]
    end
    subgraph Image["Image layers (read-only lowerdirs)"]
        L4["Layer 4: RUN pip install<br/>(dependencies)"]
        L3["Layer 3: COPY app/<br/>(application code)"]
        L2["Layer 2: RUN apt-get install<br/>(python, pip)"]
        L1["Layer 1: base image<br/>(ubuntu:24.04)"]
    end
    MERGED["merged view (OverlayFS mount)<br/>what the container process sees"]
    RW --> L4 --> L3 --> L2 --> L1
    RW --> MERGED
    L1 --> MERGED

OverlayFS Mechanics

Both the legacy overlay2 graph driver and the containerd overlayfs snapshotter use the kernel OverlayFS:

  • Lower directories: read-only image layers, stacked from base to top
  • Upper directory: the container writable layer
  • Merged directory: the unified view presented to the container
  • Work directory: scratch space OverlayFS uses internally

With overlay2, layers live under /var/lib/docker/overlay2/, and the l subdirectory holds shortened symlinks that keep mount option strings under the page-size limit.

Copy-on-Write

  • When a container modifies a file from a lower layer, OverlayFS copies it up to the writable layer first.
  • The original layers stay unchanged, so many containers share the same base layers on disk and in page cache.
  • Deleting a lower-layer file creates a whiteout entry in the upper directory that hides it.
  • Write-heavy data (databases, caches) belongs in volumes, which bypass OverlayFS and copy-up costs.

Layer Sharing

  • Identical layers are shared across images (content-addressed deduplication).
  • A pull downloads only the layers that are missing locally.
  • docker image history shows the layer stack of an image.

Container Lifecycle

This state diagram shows the states a container moves through and the CLI commands or events that trigger each transition.

stateDiagram-v2
    [*] --> Created: docker create
    Created --> Running: docker start
    Running --> Paused: docker pause (cgroup freezer)
    Paused --> Running: docker unpause
    Running --> Exited: process exits / docker stop (SIGTERM then SIGKILL)
    Running --> Restarting: exit + restart policy
    Restarting --> Running: restart
    Exited --> Running: docker start
    Exited --> Removed: docker rm
    Running --> Removed: docker rm -f
    Removed --> [*]

docker stop sends the image's STOPSIGNAL (default SIGTERM), waits for the stop timeout (10 s by default, configurable per container or daemon-wide with default-stop-timeout since 29.7.0), then sends SIGKILL.

Inter-Component Communication

This sequence shows the build flow and the run flow through the real Docker components.

sequenceDiagram
    participant User
    participant CLI as docker CLI
    participant Daemon as dockerd
    participant BuildKit as BuildKit (in dockerd)
    participant ctrd as containerd
    participant shim as containerd-shim-runc-v2
    participant runc as runc
    participant Registry as Registry

    Note over User,Registry: Image build flow
    User->>CLI: docker build -t myapp .
    CLI->>BuildKit: build session (context, Dockerfile) via buildx
    BuildKit->>Registry: resolve and pull base image layers
    BuildKit->>BuildKit: solve LLB graph, run steps in parallel
    BuildKit->>ctrd: write layers and image manifest
    BuildKit-->>CLI: image digest
    CLI-->>User: naming to docker.io/library/myapp

    Note over User,Registry: Container run flow
    User->>CLI: docker run myapp
    CLI->>Daemon: POST /containers/create
    Daemon->>ctrd: prepare snapshot (overlayfs)
    CLI->>Daemon: POST /containers/{id}/start
    Daemon->>Daemon: create network endpoint, firewall rules
    Daemon->>ctrd: create task
    ctrd->>shim: spawn shim
    shim->>runc: runc create + start (OCI bundle)
    runc->>runc: namespaces, cgroups, seccomp, LSM
    runc-->>shim: container process running, runc exits
    shim-->>ctrd: task started
    ctrd-->>Daemon: running
    Daemon-->>CLI: 204 No Content

BuildKit Pipeline

This sequence shows how BuildKit turns a Dockerfile into an image and exports it.

sequenceDiagram
    participant User as Developer
    participant BX as docker buildx
    participant BK as BuildKit
    participant Cache as Cache (local or registry)
    participant Registry as Registry

    User->>BX: docker buildx build --platform linux/amd64,linux/arm64 .
    BX->>BK: send Dockerfile and incremental context
    BK->>BK: Dockerfile frontend converts to LLB DAG
    BK->>Cache: look up cache keys for each vertex
    par Independent stages
        BK->>BK: stage deps (RUN --mount=type=cache)
        BK->>BK: stage build (compile)
    end
    BK->>BK: assemble final stage per platform
    BK->>BK: attach SBOM and provenance attestations
    BK->>Registry: push image index (with --push)
    BK->>Cache: export cache (--cache-to)
    BK-->>BX: image digest
    BX-->>User: build summary

Networking

Docker's networking is implemented by libnetwork in dockerd, and the driver model is pluggable. This diagram shows a typical single host: the default docker0 bridge, a user-defined bridge, and the firewall rules that implement port publishing.

flowchart LR
    subgraph Host["Docker host"]
        subgraph DefaultBridge["docker0 bridge (172.17.0.0/16)"]
            C1["Container 1<br/>172.17.0.2"]
            C2["Container 2<br/>172.17.0.3"]
        end
        subgraph UserNet["user-defined bridge (10.0.0.0/24)"]
            C3["web<br/>10.0.0.2"]
            C4["db<br/>10.0.0.3"]
            DNS["Embedded DNS<br/>127.0.0.11"]
        end
        FW["iptables or nftables<br/>(DNAT for -p, filtering)"]
        PROXY["docker-proxy<br/>(userland proxy)"]
    end
    Internet["External clients"] -->|"host:8080"| FW
    FW -->|"DNAT to 10.0.0.2:80"| C3
    PROXY -.->|"loopback / hairpin"| C3
    C3 -->|"resolve db"| DNS
    C3 <--> C4
    DefaultBridge --> FW

Bridge (default)

  • Creates a Linux bridge (docker0 for the default network) and connects containers with veth pairs.
  • Containers get private addresses from an internal subnet. Port publishing (-p host:container) adds DNAT rules and, where needed, the docker-proxy userland proxy.
  • User-defined bridges provide automatic DNS resolution between containers by name through the embedded DNS server at 127.0.0.11. The default bridge does not; its containers communicate by IP or legacy links.
  • Since Engine 28.0 container interfaces get random MAC addresses. Since 29.0 the DOCKER-ISOLATION-STAGE-1/2 iptables chains are gone, and legacy-link environment variables are no longer injected by default.

Host

  • Removes network isolation: the container shares the host network namespace.
  • No port mapping is needed; services bind directly to host interfaces.
  • Use it for high-throughput or latency-sensitive networking where NAT overhead matters.

Overlay

  • Connects containers on multiple Docker hosts in Swarm mode using VXLAN, with optional IPsec encryption.
  • Network state is distributed by the Swarm Raft store and NetworkDB gossip. External key-value stores were removed in 23.0.
  • Engine 29.x releases continue to improve overlay and routing-mesh reliability, so Swarm is maintained, but most multi-host production workloads run on Kubernetes.

Macvlan and IPvlan

  • Macvlan gives each container its own MAC address so it appears as a physical device on the LAN. It supports 802.1Q sub-interfaces (for example eth0.50). Traffic bypasses the Docker bridge and its firewall rules.
  • IPvlan shares the host MAC but gives each container its own IP. L2 mode keeps one subnet, and L3 mode routes between subnets. It is useful when switch port security limits MAC addresses.
  • Since 29.0, macvlan and IPvlan-L2 networks get no default gateway unless --gateway is set explicitly.

None

Only the loopback interface exists. Use it for isolated batch jobs and security-sensitive workloads.

Firewall Backends: iptables and nftables

Docker programs packet filtering and NAT on the host. The default backend is iptables, and Docker works with iptables-nft or iptables-legacy. Engine 29.0 added an experimental nftables backend ("firewall-backend": "nftables"). With it, Docker owns the ip docker-bridges and ip6 docker-bridges tables. The main differences:

Aspect iptables backend nftables backend (experimental)
Rule location Fixed chains (DOCKER, DOCKER-USER, FORWARD) Docker-owned tables with base chains per hook
User rules Insert into the DOCKER-USER chain Own table and base chains; accept is not final, so override Docker drops by setting a mark honored by --bridge-accept-fwmark (priority filter - 1 or lower)
IP forwarding Docker enables net.ipv4.ip_forward Docker does not enable forwarding; startup fails if it is needed and disabled
Swarm overlay Supported Not supported (overlay rules not migrated yet)

nftables avoids the long linear chain traversal of iptables and follows the direction of Linux firewall development. Docker has not published throughput benchmarks comparing the two backends. Because it is experimental and incompatible with Swarm, iptables remains the production default in 29.x.

Published ports bypass ufw and firewalld

Docker's DNAT rules run before host firewall front-ends such as ufw. A port published with -p 8080:80 is reachable even if ufw denies 8080. Bind to 127.0.0.1 (-p 127.0.0.1:8080:80) or filter in DOCKER-USER.

Storage and Volumes

Volumes provide persistent data that outlives a container and bypasses the copy-on-write layer. Named volumes are managed by Docker under /var/lib/docker/volumes/. Bind mounts map an arbitrary host path, and tmpfs mounts live in memory only. Volume plugins extend this to NFS, cloud block or file storage, and similar backends. Engine 29.7.0 made --mount type=image (mount another image's filesystem read-only) generally available. See Storage Types for the comparison table.

Data safety

Removing a container does not remove its named volumes. docker compose down -v and docker volume prune do, so use them deliberately.

Security Model

Docker security spans the host kernel (namespaces, cgroups, seccomp, AppArmor/SELinux), the daemon configuration, the image supply chain, and network isolation. No single layer is sufficient on its own. The concrete checklist is in Reference; setup tasks are in How-to Guides.

Threat Model

Threat vector Mitigation
Container breakout to host Rootless mode or user namespaces, seccomp, AppArmor/SELinux, patched kernel and runc
Privileged container abuse Drop capabilities, run as non-root, never use --privileged
Malicious or vulnerable base images Minimal or hardened images, vulnerability scanning, digest pinning
Supply chain compromise Signed images (Cosign), SBOM and provenance attestations, verification in CI
Daemon socket exposure Restrict /var/run/docker.sock; mutual TLS or SSH for remote access
Host file access via docker cp or archives Stay on current 29.x patches (several 2026 docker cp CVEs fixed in 29.5.1)
Network lateral movement User-defined networks, minimal published ports, host firewall via DOCKER-USER

The daemon socket is equivalent to root on the host: anyone who can reach /var/run/docker.sock can start a privileged container that mounts /. That is why the socket should never be mounted into untrusted containers and remote access must use TLS client certificates or SSH (docker context create --docker host=ssh://...).

Rootless Mode

Rootless mode runs both dockerd and containers as an unprivileged user inside a user namespace. It has been GA since Engine 20.10.

  • No SETUID binaries or file capabilities are needed, except newuidmap and newgidmap for multi-UID mapping.
  • Container root (UID 0) maps to the unprivileged user; a breakout does not yield host root.
  • Networking uses a user-mode network stack. Since Engine 29.5.0 the default is gvisor-tap-vsock, and Docker packages no longer install slirp4netns. pasta is also supported, and 29.8.0 added RootlessKit v3.1 with the pesto port driver.
  • Resource limits require cgroup v2 with systemd delegation. Privileged ports (< 1024) need net.ipv4.ip_unprivileged_port_start or CAP_NET_BIND_SERVICE on rootlesskit.

Rootless vs userns-remap

With userns-remap, the daemon still runs as root but maps container UIDs to an unprivileged host range. Rootless mode goes further: the daemon itself has no root privileges. Note that userns-remap currently disables the containerd image store.

User Namespaces

  • Enabled with userns-remap in the daemon configuration (available since Docker 1.10).
  • Container root maps to an unprivileged subordinate UID range from /etc/subuid.
  • Docker Desktop's Enhanced Container Isolation (ECI, Business tier) runs every container in a Linux user namespace inside the Desktop VM and blocks common escape paths without changing developer workflows.

Seccomp

The default seccomp profile disables around 44 of the 300+ system calls, including mount, umount2, swapon, swapoff, pivot_root, reboot, keyctl, init_module, finit_module, and delete_module. It is applied to every container unless you override it with --security-opt seccomp=....

In May 2026, Engine 29.4.2 hardened the profile against the kernel "Copy Fail" bug (CVE-2026-31431) by blocking AF_ALG sockets and the socketcall(2) multiplexer. That broke some 32-bit workloads, so 29.4.3 moved the AF_ALG block into AppArmor and SELinux rules instead. The SELinux mitigation requires selinux-enabled: true.

AppArmor and SELinux

  • AppArmor (Ubuntu, Debian, SUSE): Docker loads a docker-default profile for each container. Custom profiles apply with --security-opt apparmor=PROFILE. Since 29.8.0 the default profile template itself is configurable at the daemon level.
  • SELinux (RHEL, Fedora, CentOS): container processes run as container_t, and files are labelled container_file_t, with MCS categories separating containers. Use :z (shared) or :Z (private) on bind mounts to relabel host content.

Linux Capabilities

Docker keeps a default allowlist of 14 capabilities (for example CHOWN, NET_BIND_SERVICE, NET_RAW, SYS_CHROOT) and drops the rest; see Default Linux Capabilities. The usual hardening is --cap-drop ALL followed by --cap-add for what the process actually needs.

Never use --privileged in production

--privileged grants all capabilities and all devices, and it disables the seccomp and AppArmor/SELinux confinement. It effectively removes container isolation.

Image Signing and Supply Chain

  • Docker Content Trust (DCT), the Notary v1/TUF-based signing behind DOCKER_CONTENT_TRUST=1 and docker trust, was removed from the Docker CLI in Engine 29.0. It can still be built as a separate plugin, but new pipelines should not depend on it.
  • Cosign (Sigstore) is the common replacement. It signs images with keys or OIDC keyless identities, stores signatures as OCI artifacts next to the image, and records them in the Rekor transparency log.
  • Attestations: BuildKit can attach SBOM and SLSA provenance attestations (--sbom, --provenance), and the containerd image store keeps them locally.
  • Docker Hardened Images (DHI): since 2025-12-17, more than 1,000 minimal Debian- and Alpine-based images with signed SBOMs, VEX, and SLSA Build Level 3 provenance are free under Apache 2.0. Paid Select and Enterprise tiers add FIPS/STIG variants, a 7-day critical CVE SLA, and customizations.

Vulnerability Scanning

Scanner Maintainer Strengths
Docker Scout Docker, Inc. Integrated in Desktop, Hub, and docker scout CLI; base-image update recommendations; policy evaluation
Trivy Aqua Security (OSS) Images, filesystems, git repos, IaC, Kubernetes; standalone CI use
Grype Anchore (OSS) Fast matching; consumes Syft SBOMs

Scout, Trivy, and Grype draw on overlapping advisory sources (GitHub Advisory Database, NVD, distro feeds). No independent, reproducible benchmark of their detection rates was found.

Docker Desktop Architecture

Docker Desktop packages the Engine, CLI, Compose, Buildx, Kubernetes (kind or kubeadm), Scout, and newer AI tooling (Model Runner, MCP Toolkit, Ask Gordon, Docker Offload) behind a GUI. Because containers need a Linux kernel, Desktop runs the Engine inside a lightweight VM: the Apple Virtualization framework on macOS (the QEMU option was removed in 2025), WSL 2 or Hyper-V on Windows, and KVM on Linux. File sharing, port forwarding, and host.docker.internal bridge the host and VM. Desktop is proprietary and needs a paid subscription for larger companies (see Reference); the Engine inside it is Moby (Apache 2.0).

Performance Characteristics

Containers are processes, not VMs, so CPU and memory overhead is close to zero; the costs come from the layers around the process:

  • Bridge networking adds veth hops, NAT, and sometimes the userland proxy. --network host removes these.
  • The OverlayFS writable layer adds copy-up latency on first write. Volumes and bind mounts avoid it.
  • Docker Desktop adds VM and file-sharing overhead on macOS and Windows, which usually dominates bind-mount heavy workloads.
  • The containerd image store trades extra disk (compressed plus unpacked copies) for faster push and pull and richer image features.

The rough, unsourced numbers are in Reference: Performance Figures.

Sources