llm-d-router v0.10.0 flow control · one H200 · 2026-09-30
Priority holdback alongside eviction
Flow control can protect interactive requests from batch work by evicting in-flight batch requests, or by holding back part of the pool so batch never fills it. Is it worth running both?
It depends on how the EPP measures saturation, and at scale the saturation measure matters more than holdback. Nothing tested protects against bursts larger than the reserve.
Request counts
Optional while slots are the limit; eviction alone keeps TTFT near the floor. Not enough once KV cache is the limit (32B, 128 slots): TTFT ~50 s with or without holdback.
Token counts
Required. Without it, TTFT p50 is 2.6 s under steady load (1.5B) and 3.8 s at 32B; with it, 0.03 s and 0.39 s.
KV utilization
Helps, not enough. With default thresholds it overshoots at startup; holdback halves the wait under steady load, but TTFT p50 stays at 1–1.5 min at 32B.
Hybrid
Mixed. Fixes steady load, makes bursts worse (p95 1.5 s → 5.3 s).
Bursts, any mode
No help. Bursts larger than the reserve still evict or queue.
Resume off
No help. Bursts then redo about as much decode as the batch needs (+74% batch time), with or without holdback.
1Setup
One vLLM 0.30 replica on an H200 with --max-num-seqs 32, behind llm-d-router v0.10.0 with flow control and eviction enabled.
Batch requests (priority −1) go through a queue processor that resumes an evicted stream from its saved tokens. Interactive requests (priority 10) go straight to the gateway.
Holdback is priority-holdback-policy. A batch ceiling of 0.9 or 0.8 caps batch at 29 or 26 of 32 slots; interactive keeps ceiling 1.0 and evicts only once the pool is full.
Batch: 3 jobs × 64 requests × 1,500 output tokens. 3 runs per arm; charts show the mean with min–max whiskers.
2Request-count saturation
With eviction alone, batch refills every freed slot, so about 9 in 10 interactive requests evict a batch request. A 20% reserve absorbs steady arrivals (0 evictions at 1 and 2 req/s, 18 vs 219 at 4 req/s). Bursts of 32 overwhelm any reserve.
Qwen2.5-1.5B, short prompts. Holdback lowers TTFT p95 to the 0.13 s floor at 1 req/s, where waiting on an eviction showed up, and costs up to ~3 s of batch time. Holdback without eviction never gets TTFT near the floor (40% reserve: p50 1.05 s vs 0.22 s under bursts).
All numbers
Steady load: one interactive request every 1/rate seconds for 60 s. Bursts: 15 bursts of 32 requests, 6 s apart; this table includes the arms without eviction. Batch deadlines were 45, 75 and 100 s (jobs a, b, c).
steady 1 req/s
arm
evictions
TTFT p50
TTFT p95
batch done
job a met
job b met
job c met
eviction
54
0.12 s
0.57 s
62.4 s
87%
100%
100%
eviction + 10% holdback
11
0.12 s
0.13 s
64.6 s
96%
100%
100%
eviction + 20% holdback
0
0.12 s
0.13 s
64.8 s
100%
100%
100%
steady 2 req/s
arm
evictions
TTFT p50
TTFT p95
batch done
job a met
job b met
job c met
eviction
108
0.12 s
0.18 s
65.0 s
86%
100%
100%
eviction + 10% holdback
34
0.12 s
0.13 s
64.3 s
88%
100%
100%
eviction + 20% holdback
1
0.12 s
0.13 s
65.1 s
98%
100%
100%
steady 4 req/s
arm
evictions
TTFT p50
TTFT p95
batch done
job a met
job b met
job c met
eviction
219
0.12 s
0.18 s
70.1 s
80%
100%
100%
eviction + 10% holdback
68
0.11 s
0.12 s
65.9 s
78%
100%
100%
eviction + 20% holdback
18
0.12 s
0.13 s
72.9 s
89%
100%
100%
bursts of 32 every 6 s
arm
evictions
TTFT p50
TTFT p95
batch done
job a met
job b met
job c met
no eviction
0
1.80 s
6.72 s
75.6 s
100%
100%
100%
10% holdback
0
1.46 s
6.29 s
75.1 s
100%
100%
100%
40% holdback
0
1.05 s
4.93 s
80.5 s
100%
100%
100%
eviction
350
0.22 s
0.43 s
69.1 s
16%
100%
100%
eviction + 10% holdback
329
0.22 s
0.41 s
69.9 s
0%
100%
100%
eviction + 20% holdback
324
0.21 s
0.39 s
69.7 s
0%
100%
100%
3Token and hybrid saturation
Counting tokens, an interactive request weighs about a ninth of a batch request, so one eviction lets ~8 interactive requests in while vLLM has one free sequence. The rest queue inside vLLM, out of flow control's reach. Holdback keeps real slots free and fixes steady load; under bursts it does not help, and with hybrid counting it hurts.
concurrency-detector can count requests, estimated tokens, or both (hybrid: per endpoint, the larger ratio). Output tokens are estimated as min(1.5 × prompt, max_tokens), so these runs use long prompts: ~4.9k-token batch prompts and 500-token interactive prompts × 128 tokens on Qwen2.5-1.5B, with maxTokenConcurrency set to 32 batch requests' worth (204,800). Load is driven in-cluster with a per-run cache_salt, so TTFT is lower here than in section 2; compare only within this section.
All numbers
requests · steady 2/s
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
125
0.03 s
0.16 s
85.5 s
eviction + 20% holdback
0
0.03 s
0.16 s
87.3 s
tokens · steady 2/s
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
159
2.64 s
4.53 s
90.7 s
eviction + 20% holdback
0
0.03 s
0.31 s
85.4 s
hybrid · steady 2/s
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
132
0.27 s
0.53 s
84.2 s
eviction + 20% holdback
0
0.03 s
0.13 s
86.0 s
requests · bursts
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
451
0.18 s
0.34 s
90.5 s
eviction + 20% holdback
428
0.16 s
0.30 s
95.6 s
tokens · bursts
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
44
4.70 s
9.73 s
96.1 s
eviction + 20% holdback
4
3.60 s
9.38 s
93.1 s
hybrid · bursts
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
442
0.15 s
1.52 s
96.2 s
eviction + 20% holdback
231
0.64 s
5.33 s
95.7 s
4At scale: Qwen2.5-32B, 128 slots
When KV cache runs out before sequence slots do, the saturation detector matters far more than holdback. Request counts admit ~5× what fits and interactive requests wait ~50 s whatever the holdback. Token counts capped at the real KV size are the only setup close to usable, and holdback is what gets them there: TTFT p50 3.8 s → 0.39 s under steady load. The KV-utilization detector with default thresholds is the worst: interactive requests wait about two minutes; holdback halves that under steady load (115 s → 58 s p50) and trims it under bursts (112 s → 94 s).
One H200, --max-num-seqs 128, 267k tokens of KV cache, which fits ~26 batch requests at full length. Batch: 2 jobs × 64 requests, ~8k-token prompts × 2,000 tokens. Interactive: ~2k-token prompts × 256 tokens, steady 1 req/s for 180 s or 6 bursts of 32 every 30 s. The token cap is the measured KV size (267,472); the KV detector uses its defaults (queue depth 5, KV usage 0.8). In-cluster driver, per-run cache_salt. 2 runs per arm, 1 for request counts.
Request counts: the EPP sees 128 in flight as full while vLLM runs ~28 and queues ~100 on KV capacity. An interactive request that evicts one joins that queue.
KV utilization: at the start of a run the cache is empty, so the EPP dispatches ~46 batch requests before its metrics (50 ms) show the pool full. Queue depth then holds saturation at 3–4 (16 waiting / 5), and an interactive request waits for ~11 evictions. This is the reactive-detector failure its documentation describes.
Token counts: admission matches what fits, so evictions do real work. Under bursts, though, a burst of 32 counts as only ~28% of the cap, so most of it is admitted and queues inside vLLM (p50 5.8–9.2 s).
All numbers
requests · steady 1/s
arm
runs
evictions
TTFT p50
TTFT p95
batch done
eviction
1
177
49.80 s
334.92 s
931.5 s
eviction + 20% holdback
1
116
48.79 s
183.73 s
873.3 s
tokens · steady 1/s
arm
runs
evictions
TTFT p50
TTFT p95
batch done
eviction
2
298
3.83 s
9.82 s
1013.5 s
eviction + 20% holdback
2
7
0.39 s
10.45 s
973.5 s
kv · steady 1/s
arm
runs
evictions
TTFT p50
TTFT p95
batch done
eviction
2
188
115.02 s
279.26 s
909.5 s
eviction + 20% holdback
2
234
58.43 s
117.43 s
937.0 s
requests · bursts
arm
runs
evictions
TTFT p50
TTFT p95
batch done
eviction
1
160
46.95 s
96.49 s
946.6 s
eviction + 20% holdback
1
132
46.95 s
96.48 s
926.8 s
tokens · bursts
arm
runs
evictions
TTFT p50
TTFT p95
batch done
eviction
2
26
9.23 s
26.11 s
894.6 s
eviction + 20% holdback
2
9
5.79 s
24.30 s
972.7 s
kv · bursts
arm
runs
evictions
TTFT p50
TTFT p95
batch done
eviction
2
162
111.60 s
249.35 s
916.4 s
eviction + 20% holdback
2
128
94.49 s
126.14 s
910.8 s
5Without resume
Turning resume off (evicted requests restart from scratch after backoff) barely matters under steady load but nearly doubles batch work under bursts: ~288k tokens redone against 288k needed, and the batch takes 74% longer. Interactive TTFT does not change. Holdback does not rescue it: under bursts it trims redone work ~8% and still finishes later.
Request-count saturation, same long-prompt mix and in-cluster driver as section 3. Without a render_url the processor cannot resume, so an evicted request is retried from the start with exponential backoff; resume and immediate requeue cannot be switched off separately.
All numbers
steady 2/s · resume
arm
evictions
decode redone
TTFT p95
batch done
eviction
125
0k
0.16 s
85.5 s
eviction + 20% holdback
0
0k
0.16 s
87.3 s
steady 2/s · restart
arm
evictions
decode redone
TTFT p95
batch done
eviction, restart
134
7k
0.12 s
87.9 s
eviction + 20% holdback, restart
1
1k
0.19 s
87.7 s
bursts · resume
arm
evictions
decode redone
TTFT p95
batch done
eviction
451
0k
0.34 s
90.5 s
eviction + 20% holdback
428
0k
0.30 s
95.6 s
bursts · restart
arm
evictions
decode redone
TTFT p95
batch done
eviction, restart
483
288k
0.34 s
157.7 s
eviction + 20% holdback, restart
437
264k
0.31 s
168.9 s
6A larger model: Qwen2.5-14B
Holdback removes evictions here too (166 → 11 → 2) but TTFT p95 stays at ~0.92 s in every arm, and batch takes 4–5% longer. Interactive latency comes from sharing the GPU with long-context batch work, not from waiting on evictions.
Keep resume on. Without it, each burst eviction throws away work and the batch takes 74% longer; holdback does not win that back.
Keep request-count saturation with eviction as the baseline. It keeps interactive TTFT near the floor under both steady load and bursts.
Add holdback, sized to steady interactive concurrency, when evictions themselves cost something (lost prefix cache, churn). Expect 1–6% more batch time.
If saturation is counted in tokens on a server with a sequence cap, run holdback: token counts alone admit more requests than the server can run.
When KV cache binds before sequence slots (long prompts and outputs, high concurrency), request counts cannot protect interactive traffic. Count tokens with the cap set to the real KV capacity, and run holdback with it. The KV-utilization detector needs thresholds tuned for the workload; its defaults overshoot.
Neither setting absorbs bursts larger than the reserve. That needs request-count backpressure or a larger reserve.
8Caveats and open questions
Saturation is counted in requests except in sections 3 and 4. A different token cap would move the token-mode results.
The 45 s batch job misses deadlines only under eviction (100% met without it on the same build). Leading suspect, not yet confirmed: all 192 batch requests reach the gateway at once, so the processor's earliest-deadline-first order never applies and a resumed request rejoins the gateway's FCFS queue behind later-deadline work.
Why hybrid counting with holdback gets worse under bursts is not yet explained.
At 32B, request-count arms have 1 run each (each takes ~16 min) and the other arms 2. The KV detector ran only with its default thresholds. An earlier 32B run at 2 req/s for 300 s thrashed (777 evictions in 14 minutes) and was stopped; the interactive load was cut to 1 req/s for 180 s for every arm.
The 14B runs did not use a per-run cache_salt, so prefill comparisons are left out; its 20% holdback arm has 2 runs.
One interactive request out of 1,440 (tokens mode, bursts, eviction only) got a 503.
Every completed batch output was token-identical to a direct vLLM run (121 runs).