Blog

Firecracker agent sandboxes on nested KVM

· by

On September 25, 2026 at 01:39 UTC, plori started to run new agent sessions in Firecracker microVMs instead of gVisor pods. Each node is itself a KVM guest, so our microVMs use nested KVM. From the switch to September 27, 19:38 UTC, the fleet cold-booted 305 microVMs. The median time from the create request to a healthy agent process was 1.38 s, and the 95th percentile was 1.76 s. Most of that time is the agent process start. Firecracker itself configured and started each guest in a median of 10 ms.

One part did not work as the Firecracker documentation describes. On these hosts, two of the documented ways to correct a guest's wall clock fail. Our guests use a third documented way, NTP, with the node as their time server.

What does the setup look like?

Each node is a KVM virtual machine with 2 vCPUs, 8 GB of memory and a 120 GB local disk. The nodes run Ubuntu 24.04 with kernel 6.8, and the kvm_amd module runs with nested virtualization enabled. Before we started, we checked that /dev/kvm works inside a node. We pin Firecracker v1.17.0 and a 6.18.44 guest kernel from Firecracker's CI builds by checksum.

A small node daemon starts and stops the microVMs. The control plane calls it over HTTP to create a VM, change its labels, stop it, and read its state. Inside each guest, an init process fetches its configuration from the node and then starts the agent process. An agent session VM has 2 vCPUs and 1024 MiB of memory.

How long do create, boot and health take?

The node daemon logs three phases for every VM that becomes ready:

  • create: from the accepted request to a Firecracker API socket that answers. This includes the network slot, the root disk clone and the jailer start.
  • boot: from the first Firecracker configuration call to InstanceStart.
  • health: from InstanceStart to the first successful response from the agent process's health endpoint.
Phase Production, 305 cold boots, Sept 25 to 27 First test node, one VM at a time, Sept 24
create median 19 ms, p95 32 ms, max 42 ms 7 to 13 ms
boot median 10 ms, p95 18 ms, max 40 ms 9 to 12 ms
health median 1.34 s, p95 1.72 s, max 3.28 s not measured separately
total median 1.38 s, p95 1.76 s, p99 2.96 s, max 3.36 s 1.41 to 1.47 s (1 vCPU, 512 MiB guest)

The test node ran before we added the jailer. On a development host, the jailer increased the create phase from 21 to 26 ms to 37 to 53 ms. The time to a ready VM stayed at about 1.05 s, because the agent process start takes about 0.95 s there.

The five slowest production boots had a health phase of 2.5 to 3.3 s. All five became ready within three seconds of each other on September 26. That is our only production evidence about boots in parallel. In the full week, samples at 5-minute intervals show at most five VMs in use at the same time across the fleet. We have not measured boot time under higher concurrency.

We also measured snapshots on the test node. A full snapshot of a 512 MiB guest took 300 to 404 ms to create. A restore took 3 to 6 ms to load and resume, and 58 ms to a healthy agent process. Production did not create or restore a snapshot in this period, because the pool always cold-boots new VMs.

Why did the guest clock break, and how did we correct it?

A guest restored from a snapshot continues its wall clock from the moment of the snapshot. Firecracker documents three ways to correct this. The snapshot documentation describes clock_realtime: true in the snapshot load request, which advances the guest clock at restore time. The Firecracker FAQ names NTP in the guest as the standard solution. For accurate time at scale, the FAQ recommends the KVM PTP device, /dev/ptp0, as a chrony reference clock.

On our nodes, clock_realtime and KVM PTP both fail for the same reason. KVM provides them only when the host kernel uses a TSC-based clock source. Each of our nodes is itself a KVM guest without an invariant TSC. At boot, its kernel logs tsc: Marking TSC unstable due to TSCs unsynchronized and then uses the kvm-clock clock source.

  • clock_realtime: the snapshot load fails with clock_realtime requested but not present in the snapshot state. Firecracker returns this error when the saved KVM clock does not have the KVM_CLOCK_REALTIME flag. The Linux kernel sets that flag in KVM_GET_CLOCK only when it can read the wall time and the TSC together. That read requires a TSC-based clock source on the host that takes the snapshot.
  • KVM PTP: the guest kernel has CONFIG_PTP_1588_CLOCK_KVM=y, but the guest has no /dev/ptp0, and chronyd stops with Could not open /dev/ptp0. The device depends on the KVM_HC_CLOCK_PAIRING hypercall, which returns KVM_EOPNOTSUPP when the host does not use a TSC-based clock source.

Without a correction, a restored guest was behind by the full time it spent in the snapshot: 4 s in our first test.

Our correction is NTP from the node. The node already keeps its own clock correct with chronyd. We configured the node's chronyd to answer the VM network slots. It also serves local time if its own upstream sources become unavailable. Each slot forwards UDP port 123 at a fixed guest-side address to the node. The guest thus needs no route to the public internet for time. The guest's chrony configuration is:

server 192.0.2.1 iburst minpoll 0 maxpoll 2 prefer
makestep 1 -1
driftfile /run/chrony.drift

minpoll 0 maxpoll 2 sets a poll interval of 1 to 4 s, so the guest corrects its clock within seconds of a restore. With makestep 1 -1, chronyd steps the clock for every correction larger than 1 s, with no limit on the number of steps. The chrony.conf documentation describes the limit argument: a negative value disables it. A pool VM can be restored many times, so a limit on steps does not fit. The guest init also sets the clock from a timestamp in its boot configuration when the difference is more than 2 s.

We measured the result on the test node. A fresh guest followed the node within 0.2 ms. A guest restored from a 120 s old snapshot read 120 s behind for the first 3 s. Its offset was 0 s at 5 s after it became ready.

The guest init checks for /dev/ptp0 at every boot. If the device is present, it adds the PTP reference clock to the chrony configuration. On our development host, which has native KVM, the device is present. On our nodes it is not, and the guest uses only the node server.

One operational note: after you write a new chronyd configuration file, restart chronyd. systemctl enable --now does not restart a daemon that already runs, so it continues with the old configuration.

The node daemon starts every VM through the Firecracker jailer, never Firecracker directly. The Firecracker production host guide recommends a unique uid and gid for each Firecracker instance. Each VM receives a uid and gid of 900000 plus its slot number. No user account exists for these IDs. Firecracker runs with no capabilities. It cannot open a file that belongs to another slot, and it cannot send a signal to another VM's process. The slot's TAP device belongs to the same uid. An unprivileged Firecracker process cannot attach to a TAP device that it does not own.

The jailer builds a chroot for each VM. The node daemon puts hard links to the VM's files into it. These are the root disk, the console log and, for a restore, the snapshot state and memory files. Every VM sees its root disk at the same path, /rootfs.ext4. A snapshot records that path, so a snapshot taken in one jail loads in another jail.

Hard links only work in one filesystem, and the root disk clone must be fast. The local disk of our nodes has one ext4 partition. Ext4 does not support reflinks, so the FICLONE call failed. The clone then fell back to a full copy of the 4 GiB disk. That copy took 2.1 s for a cold boot and 3.2 s for a restore. We now create an XFS filesystem with reflink enabled inside an image file on that disk. All VM state is in that filesystem. The clone is now part of the create phase, which was 7 to 13 ms on the test node. All 305 production boots used a reflink clone.

Two more jailer details:

  • The node daemon creates each VM's cgroup (cgroup v2) itself and starts the jailer directly in that cgroup. It passes --parent-cgroup with no --cgroup values, so the jailer uses that cgroup and creates no other.
  • We do not use --daemonize or --new-pid-ns. With --new-pid-ns, the jailer forks and exits, and Firecracker becomes the init process of its PID namespace. A SIGTERM from the daemon then does not stop it. Without these options, the process that the daemon starts and waits for is Firecracker itself.

How is each VM's network isolated?

Each VM receives a network slot. A slot is a network namespace with a veth pair to the host and a TAP device for the guest. The veth pair has a /30 subnet of its own. Every guest has the same address, 169.254.0.21, and the same MAC address. NAT rules in the slot's namespace translate between the guest address and the slot address. The guest configuration then contains no node-specific address, and a restored guest can continue in any slot. Firecracker's network for clones guide describes the same namespace-and-NAT method.

The guest reaches three node services at one fixed address, 192.0.2.1, from the documentation range in RFC 5737:

  • the boot configuration, which the guest init fetches once
  • NTP from the node's chronyd
  • a relay to the control plane

The node serves the boot configuration only once per boot. Every later request receives 410 Gone. The agent's own code runs as an unprivileged user and starts after the init process, so it cannot read the secrets in that configuration. On the host, the firewall drops these service ports on every interface that is not a slot interface.

Guest egress rules apply in this order:

  1. Accept replies on existing connections.
  2. Accept the three node services above.
  3. Reject traffic to the node's own addresses and to all other slots.
  4. Accept an allow list. This list is empty in production.
  5. Reject all private ranges from RFC 1918, including the private network between nodes. Also reject all link-local addresses, including the metadata address 169.254.169.254, and the carrier-grade NAT range.
  6. Accept all other traffic, which is the public internet.

The rules reject and do not drop. A TCP connection receives a reset, and other packets receive an ICMP "administratively prohibited" reply. A connection to a denied address fails in milliseconds and does not wait for a timeout. A separate rule drops guest packets that do not come from the guest address. The guest reaches the control plane only through the relay.

The host also limits traffic in the other direction. The control plane connects to a VM at the node's private address and a port for that slot. The host translates that address only for packets from the control plane's sources on the private interface.

Two host problems took time to find:

  • A NAT rule in the prerouting hook does not see packets that the host itself sends. The node daemon's own health checks to a slot need a second NAT rule in the output hook.
  • The Ubuntu image has ufw on, with a default input policy of DROP. A drop in ufw's tables overrides an accept in our own nftables table. A health check through loopback worked, and the same check from the private network timed out with no reply.

How does a node decide whether a new VM fits?

The node daemon admits VMs by memory. The node budget is its total memory minus the larger of 1024 MiB and 10 percent of the total. On our 8 GB nodes, the total is 7935 MiB, so the budget is about 6.9 GiB. Each VM counts its memory reservation plus 64 MiB for host memory outside the guest. On the test node, the Firecracker process of an idle 2 vCPU, 1 GiB guest used 102 MB of RSS. With one core busy, it used 109 MB. The VM's cgroup used 150 MB and 164 MB. At a reservation of 1024 MiB, a node holds six VMs. The daemon refuses a create that does not fit with 503 and a message that names the failed check.

Each VM also receives host-side limits from its priority:

Priority Cgroup memory limit oom_score_adj
pool (not yet claimed) memory.max, guest size plus 128 MiB 1000
user service memory.max, guest size plus 128 MiB 800
agent session memory.high, guest size plus 128 MiB -500

memory.high throttles and reclaims but does not kill, so a session's own cgroup does not stop it. If the host runs out of memory, the kernel OOM killer takes pool VMs first and sessions last. The node daemon runs at -900. In a pressure test on a node of the production type on September 26, the daemon ran at 0. It was the first process that the OOM killer stopped, four times in one run, and its memory pressure loop stopped with it.

Swap is off on every node. The Firecracker production host guide recommends this so that guest memory does not go to disk. The node's Ubuntu image has an 8 GiB swap file. In a test on September 25 with that swap file on, the host swapped session memory. The slowest 15-second session write rate fell from 320 MiB/s to 108 MiB/s.

In the week to September 27, the largest Firecracker RSS of one production agent VM was 772 MiB.

What does memory overcommit add, and is it proven?

With overcommit, each VM reserves its guest size divided by a ratio. At a ratio of 1.5, a 1024 MiB session reserves 683 MiB. The sum of guest sizes can be up to 1.5 times the budget. A node can then hold nine sessions instead of six. The host must then recover memory when guests use more than it has.

Every VM receives a balloon device with free page reporting at boot, because Firecracker accepts these only before boot. The daemon uses a low mark of 1536 MiB and a clear mark of 2048 MiB of available host memory. Below the low mark, it refuses new VMs and inflates balloons in priority order, pool VMs first. If that does not free enough memory, it stops pool VMs and then user services. It never stops a session.

We tested this on a disposable node of the production type with nine sessions:

  • Available host memory fell by about 260 MiB/s. Our first marks were 512 and 1024 MiB with a 2 s check interval. With them, the daemon saw the pressure only 2 to 3 s before the OOM killer ran. At 1536 and 2048 MiB with a 1 s interval, every test phase passed, and the lowest available host memory was 229 MiB.
  • In the first run, no balloon deflated. The Firecracker API server answers an eleventh open connection with HTTP 503, and our client opened a new connection for each call. The daemon now keeps one client for each VM.
  • One balloon target was the full difference between guest size and reservation. It starved the guest, and one session ran 9 times slower (7.6 to 0.83 iterations per second). The daemon now leaves at least 128 MiB of available memory in each guest.

Since September 26, production nodes run at a ratio of 1.5. In the week to September 27, the memory pressure loop took zero steps. The lowest available memory on any node was 5778 MiB, far above the low mark. The balloon steps and the stop order worked in tests on one node. They are not proven in production.

What is not proven yet?

Snapshot restore and memory pressure handling ran only on test nodes. Three more questions are open:

  • Boot times with more than five VMs in use at the same time.
  • Moving a snapshot to another node. This needs the same Firecracker version and CPU features on both nodes.
  • The CPU cost of nested virtualization. Our first comparison between the host and a guest used different software builds, so the numbers cannot be compared.

Any host whose kernel does not use a TSC-based clock source has the same clock problem. This is common when the host is itself a KVM guest. It is not certain, because a nested host that receives an invariant TSC can still use the TSC. To check a host, read /sys/devices/system/clocksource/clocksource0/current_clocksource. With tsc, both methods can work. With kvm-clock, as on our nodes, clock_realtime restores and KVM PTP in the guest fail. In that case, serve NTP from the node.