vLLM Preemption: Swapping vs Recompute | Generated by AI

Home 2026.07

Question: What does preemption mean in vLLM’s scheduler?

Answer: When KV cache blocks run out because too many sequences are running concurrently, vLLM has to kick a sequence out of the running batch to free GPU memory for others. That’s preemption — it’s the scheduler’s answer to the fact that you can’t know upfront how many tokens a generation will need, so admission control alone can’t prevent overcommitment.

Two recovery modes, set via --preemption-mode:

# vllm/core/scheduler.py (V0) — simplified logic
class PreemptionMode(Enum):
    SWAP = "swap"
    RECOMPUTE = "recompute"

def _preempt(self, seq_group, blocks_to_swap_out):
    if self.user_specified_preemption_mode is None:
        # beam search / multi-seq groups can't just redo prefill
        # cleanly -> use swap; otherwise recompute is cheaper
        preemption_mode = (PreemptionMode.RECOMPUTE
                            if seq_group.get_max_num_running_seqs() == 1
                            else PreemptionMode.SWAP)
    else:
        preemption_mode = self.user_specified_preemption_mode

    if preemption_mode == PreemptionMode.RECOMPUTE:
        self._preempt_by_recompute(seq_group)
    else:
        self._preempt_by_swap(seq_group, blocks_to_swap_out)

Which one is default depends on version: In vLLM V1, the default preemption mode is RECOMPUTE rather than SWAP, since recomputation has lower overhead in the V1 architecture. V0 defaulted to recompute for single-sequence groups but fell back to swap for beam search / multi-seq groups where recompute isn’t valid. If not specified, recomputation is used by default since it incurs lower overhead than swapping — but when the sequence group has multiple sequences (e.g., beam search), recomputation is not supported, so swapping is used instead.

Which sequence gets picked as victim: vLLM typically enforces a LIFO policy and selects the most recently scheduled request as the victim, and the preempted request is prepended to the front of the WAITING queue rather than discarded — so it doesn’t lose its place in line, just re-pays prefill cost.

When it fires, you’ll see this in logs:

WARNING scheduler.py:1057 Sequence group 0 is preempted by PreemptionMode.RECOMPUTE
mode because there is not enough KV cache space.

This can affect end-to-end performance; increasing gpu_memory_utilization or tensor_parallel_size provides more KV cache memory. Track it via the Prometheus vllm:num_preemptions_total counter — if it grows steadily under nominal load, reduce –max-model-len, raise –gpu-memory-utilization, or add capacity.

Practically for your setup: if you’re serving long-context requests on MI300X and see preemption thrashing, first check gpu_memory_utilization (raise it), then consider --enable-chunked-prefill (prioritizes decode over new prefill so running sequences aren’t starved), and if you have adversarial/bursty traffic mixing short+long requests, that’s exactly the DoS pattern in that DoS paper — attacker keeps nudging memory over the threshold to force expensive recompute thrashing on victims.

References:


Generated by AI. Curating and sharing still takes effort. If you find it useful, feel free to donate. WeChat: @lzwjavaWeChat QR · X: @lzwjava · Say hi 👋

Back Donate