Isolation and privilege
Namespaces, cgroups, and capabilities make a process a container. Root inside it is still a risk.
A container is a process with restrictions
Start with the claim Container Security makes early: a container is an ordinary Linux process, started with a set of restrictions. Run ps on the host and you will see it. What makes it "contained" is three kernel mechanisms applied together.
Namespaces: what the process can see
Each namespace gives the process a private copy of one kind of system resource:
| Namespace | Isolates |
|---|---|
| PID | Process IDs. The entrypoint is PID 1 inside, and host processes are invisible |
| Mount | The filesystem tree. The image's root replaces the host's |
| Network | Interfaces, routes, ports, and firewall rules |
| UTS | Hostname and domain name |
| IPC | Shared memory and semaphores |
| User | User and group IDs. Root inside can map to an unprivileged UID outside |
| Cgroup | The view of the cgroup hierarchy |
The network namespace matters most day to day. A container has its own localhost. The “Container networking and composition” notes build on that.
cgroups: what the process can use
Control groups account for and limit resources: memory, CPU shares and quotas, I/O, and the number of processes. A container's --memory=512m is a cgroup limit. When the processes inside exceed it, the kernel's OOM killer acts within that cgroup, whatever the host has free. That is the same mechanism described in the “Memory, limits, and a method for performance” notes, now per container. The pids limit also stops fork bombs inside one container from exhausting the host.
Capabilities: root in pieces
Traditional Unix has one powerful user. Linux splits root's power into capabilities: binding ports below 1024 (CAP_NET_BIND_SERVICE), changing file ownership (CAP_CHOWN), and the catch-all CAP_SYS_ADMIN. Docker starts containers with a reduced default set and applies a seccomp profile that blocks dangerous system calls. That reduces risk but does not remove it.
Why root in a container still matters
Without user-namespace remapping, UID 0 in the container is UID 0 on the host, restricted by the mechanisms above. Any of these turns a container compromise into a host compromise: a kernel or runtime escape bug, a mounted Docker socket, a sensitive host path mounted in, or --privileged. The defence is layered least privilege:
- Run as a non-root user. Add
USER 10001(or a named user) in the image, and give that user only the files it needs. - Drop capabilities.
--cap-drop=ALL, then add back only what is proven necessary. - Read-only root filesystem.
--read-only, with explicit volumes ortmpfsfor writable paths. - No
--privileged, no Docker socket mounts, and no host paths beyond what is needed. - Rootless mode or user namespaces where the platform supports them.
A web service listening on 8080 and reading its config needs none of root's powers. That is the situation in the lab.
Verify, do not assume
docker exec waybill id # uid=10001 ...
docker inspect waybill --format '{{.Config.User}} {{.HostConfig.CapDrop}} {{.HostConfig.ReadonlyRootfs}}'
grep Cap /proc/$(docker inspect -f '{{.State.Pid}}' waybill)/status # effective capability bitmaps
Confirm the service still works after each restriction. The lab's validator checks both that privilege dropped and that health is unchanged.
Key terms
- Namespace
- A kernel feature that gives a process its own view of one resource type, such as PIDs, mounts, network, hostname (UTS), IPC, users, or cgroups.
- Capability
- One slice of root's privileges, such as
CAP_NET_BIND_SERVICEorCAP_SYS_ADMIN, that can be granted or dropped independently. - Rootless container
- A container whose root user maps to an unprivileged user on the host through a user namespace.
- Read-only root filesystem
- Running with the image layers mounted read-only, so the process can write only to declared volumes or tmpfs.
Read further
- Container Security, Chapters "Linux System Calls, Permissions, and Capabilities", "Control Groups", and "Container Isolation" (Purchase)
Build a container by hand withunshareandchrootas the book does. Note which namespace provides which isolation, and how capabilities split root's power into pieces. - Container Security, Chapter "Strengthening Container Isolation" (Purchase)
seccomp, AppArmor and SELinux, and the discussion of user namespaces and rootless containers. Note which defaults Docker applies for you.