When several VMs on a Proxmox host want more CPU than the physical cores can give, each guest’s virtual CPUs get paused while other VMs run. Inside the guest that shows up as steal time, st in top or vmstat. The guest’s scheduler doesn’t know which of its vCPUs are being starved, so it keeps putting work on them: latency spikes, lock holders get preempted, databases slow down far more than the lost CPU time alone suggests.
The steal governor is a new kernel feature aimed at exactly that. As of 28 September 2026 it was queued for the Linux 7.4 merge window, so it isn’t in a released kernel yet.
What the steal governor does
Per the Phoronix report and the patch series, it pairs a scheduler mechanism with a virtualisation driver:
- it tracks steal time per vCPU,
- and dynamically adjusts a preferred CPU mask, so the scheduler avoids placing tasks on heavily contended vCPUs: the VM “backs off” from the CPUs that are being stolen from.
The patch authors target workloads where preemption costs more than the lost time, such as lock-holder preemption, critical sections, TLB and cache misses, and database workloads in particular. They also note that pure CPU-crunching workloads may regress slightly. It’s a tool for latency-sensitive guests on oversubscribed hosts, not a free speed-up.
Details such as how it’s enabled and its defaults are best taken from the kernel documentation once 7.4 is released.
When you’ll get it
It runs in the guest, so the VM’s kernel needs it. That means 7.4 or newer inside the VM: distributions that follow mainline closely (Fedora, Arch) will get it soonest. Debian 13 ships 6.12 and LTS distributions stay on their kernel series, so on those it won’t arrive through normal updates.
Measure steal time today
vmstat 1 10 # "st" column = % of time stolen
mpstat -P ALL 1 5 # per-vCPU %steal (sysstat package)
A few percent now and then is normal on a shared host. Sustained double digits during your busy periods means your VMs are competing for cores.
What helps on Proxmox now
The real fix for steal time is less contention:
- Don’t give every VM many vCPUs. Idle vCPUs are cheap, but a VM with more vCPUs than it needs competes harder when it gets busy. Size VMs to their real load.
- Prioritise important VMs with CPU units (relative weight under contention, default 100):
qm set <vmid> --cpuunits 200. - Cap noisy VMs with a CPU limit (in cores):
qm set <vmid> --cpulimit 2. - Pin latency-critical VMs to dedicated cores with
qm set <vmid> --affinity 4-7, and keep other VMs off those cores. - Watch the host: if the node itself sits near 100% CPU, no guest-side scheduler change can create cores that aren’t there.
Containers (LXC) don’t have steal time in this sense: they’re processes on the host kernel and share its scheduler directly. Use CPU limits and units on them in the same way.
[discussion]
Comments are powered by Giscus — backed by GitHub Discussions. Sign in with GitHub to join the conversation.