From Simulation to Production Part II: Soft-Reflective Flow Control for llm-d
Part 2 of the sim2real series. A case study in closing the AI-native loop across a second problem class: dispatch prioritization.
Introduction
The first article in this series traced a complete traversal of the AI-native loop for llm-d's admission control layer: an AI agent framework (Nous) used a high-fidelity simulator (BLIS) to discover a probabilistic, proactive admission control policy that cuts critical-traffic TTFT by up to 97% at near-saturation workloads. That article restates the core methodology: simulation lets the loop run at machine speed, reserving expensive GPU time for validating the few candidates that actually merit real-hardware testing.
This article applies the same loop to a different problem: flow control. The distinction matters. Admission control acts at the gate: it decides whether an incoming request enters the system at all, and may reject it. Flow control acts at the dispatch queue: it decides which already-admitted request gets sent to a backend server next. The two mechanisms protect different points in the pipeline and neither replaces the other.
The concrete outcome is the soft-reflective ceiling policy: a new flow control plugin for llm-d-router that reduces critical-class TTFT p99 by up to 98% at near-capacity load without rejecting a single request. It is now merged into llm-d-router, and users can enable it today with a single YAML change.