A penetration tester has gained code execution inside one of your production containers. They can run arbitrary commands as root within the container. Your job: prove that SELinux prevents them from breaking out to the host or accessing other containers, even with root privileges inside the compromised container. Then fix a misconfigured volume mount that is weakening the isolation.
Namespaces and cgroups isolate containers at the process and resource level, but they are not a security boundary on their own. Kernel vulnerabilities and misconfigured mounts can provide escape routes. SELinux adds mandatory access control that the kernel enforces regardless of whether the process has root inside its namespace.
When Podman or Docker launches a container, the runtime assigns each process an SELinux type and an MCS (Multi-Category Security) label:
The container_t type restricts what the process can access on the host. The s0:c123,c456 portion is the MCS label - a pair of randomly assigned categories unique to each container. This is what prevents container A from reading container B's files, even though both run as container_t.
Every container gets a unique MCS category pair at launch. Files created by or mounted into a container inherit that container's categories. The kernel enforces that a process with categories c123,c456 cannot access files labeled with categories c789,c012, even if the type matches. This provides per-container isolation without needing separate types for each container.
Mounting host directories into containers requires careful labeling. Podman and Docker support two volume label suffixes:
:z (shared) relabels the volume content with a label accessible by all containers:
:Z (private) relabels with the specific container's MCS categories. Only that container can access the files:
Without either suffix, host file labels remain unchanged. If labeled default_t, the container gets access denials. If labeled too broadly, you may expose host data unintentionally.
For high-security workloads, you can override the default container_t type:
The container_loopback_t type allows only loopback network access. Other specialized types include container_init_t for containers running systemd. You can also write your own type that extends container_t with specific additional permissions.
Container file access goes through the svirt_sandbox_file_t type (aliased in newer policies to container_file_t). The policy allows container_t to read and write this type, but only when the MCS categories match. Host files with types like etc_t or shadow_t remain off-limits regardless of what happens inside the container.
If an attacker exploits a kernel vulnerability to bypass the namespace boundary, their process is still labeled container_t:s0:c123,c456. The kernel denies writes to bin_t, etc_t, or sshd_key_t regardless of UID or capabilities:
You have a test environment with three containers:
:zThe penetration tester has root inside the webapp container. Your tasks:
ps -eZ:z mount is a problem - it allows all containers to access the shared data:Z or implementing a dedicated file typeBonus objective: Write a custom container type that restricts the webapp container to only accept inbound connections and blocks all outbound network access.
You can configure and verify SELinux container isolation including MCS categories and volume labels. +300 XP
Next: Level 6 - The Boss Level