Signal-driven
Classify intent, safety, and domain signals — then route each request to the right model.

System-level intelligence for heterogeneous LLM inference
Deploy fast, route by signal, and keep every decision observable.
Classify intent, safety, and domain signals — then route each request to the right model.
Deploy as Envoy ExtProc or local vllm-sr. No client changes for existing integrations.
From rules to reinforcement learning — every routing decision is configurable and measurable.
Route each request to the best model pool through one OpenAI-compatible API.

Heterogeneous models
Unify a fragmented model landscape across four dimensions.
Explore how it worksModels specialize in different work.
Compose personalized model paths.
GPUs, accelerators, edge, and cloud coexist.
Route across heterogeneous compute.
Inference spans edge, private, and cloud.
Keep data within its boundaries.
“Best” changes by user and workload.
Make every preference executable.
16 signal families across heuristic and learned detectors, from knowledge base routing to history-aware reasks.
12 routing strategies spanning rules, latency heuristics, reinforcement learning, and ML selection.
18 research papers spanning routing, systems, safety, and multimodality.
One supported local path. Copy the installer, run it, then open the dashboard. The supported first-run path is a single installer that sets up the CLI and local serve flow on macOS and Linux.
Downloads the installer, prepares Docker, and writes vllm-sr to your PATH.
curl -fsSL https://vllm-sr.ai/install.sh | bashRemoves ~/.local/share/vllm-sr and ~/.local/bin/vllm-sr. Stop any running serve session first.
rm -rf ~/.local/share/vllm-sr && rm -f ~/.local/bin/vllm-srRouting Blueprint
Explore how the architecture extracts signals, composes decisions, and executes the selected path.
Structural mapping from communication theory to the routing pipeline.
The user request is the raw source message before encoding.
vLLM Semantic Router keeps the public surface as vllm-sr/auto, then coordinates closed, open, and hybrid model pools inside the serving layer.
Route by task shape, risk, confidence, and model capability; run bounded collaboration; return one OpenAI-compatible response.

92.6 vs Fugu Ultra 92.0

96.0 vs Fugu Ultra 95.5

50.0 matches Fugu Ultra

See how signals, policies, and models connect for every request.
Classify intent and complexity, then route each request to the best model in your fleet.
Combine embeddings, domain, PII, jailbreak, preference, and more into executable routing decisions.
Match prompts to model descriptions with embedding similarity for Cursor-style Auto routing.
Balance quality, latency, cost, and load without reading prompt content.
Chain lightweight models for triage and escalate hard prompts to frontier models.
Run on gateways, Kubernetes, or locally — with 16 signal families for every request.
Same router, any infrastructure
16 heuristic and learned detectors
Research threads that trace the router's evolving ideas across safety, multimodality, orchestration, and system design.
vLLM Semantic Router Team
arXiv Technical Report
We introduce vLLM Semantic Router, a signal-driven decision routing framework for Mixture-of-Modality deployments that composes heterogeneous signals into deployment-specific routing policies across cost, privacy, latency, and safety constraints.
Huamin Chen, Xunzhuo Liu, Bowei He, Fuyuan Lyu, Yankai Chen, Xue Liu, Yuhan Liu, Junchen Jiang
arXiv Technical Report
We synthesize the project’s recent routing, fleet, multimodal, and governance results into the Workload-Router-Pool (WRP) architecture, connecting signal-driven routing to a full-stack inference optimization framework and outlining future research directions across workload, router, and pool design.
Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, Huamin Chen
arXiv Technical Report
We formalize the visual confused deputy as a security failure mode in computer-using agents and introduce a dual-channel guardrail that independently checks click targets and action reasoning before execution.
Huamin Chen, Xunzhuo Liu, Junchen Jiang, Bowei He, Xue Liu
arXiv Technical Report
We introduce Outcome-Aware Tool Selection (OATS), an offline embedding refinement method that improves semantic-router tool ranking under single-digit millisecond CPU budgets without adding serving-time model inference.
Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, Huamin Chen
arXiv Technical Report
We propose Adaptive VLM Routing (AVR), which estimates action difficulty and routes computer-use agent steps to the cheapest model that still satisfies a target reliability threshold.
Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, Huamin Chen
arXiv Technical Report
We combine Flash Attention, prompt compression, and near-streaming body processing to cut routing latency from seconds to tens of milliseconds while keeping the router lightweight enough to share hardware with serving.
Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang, Bowei He, Xue Liu
arXiv Technical Report
We present a queueing-theory-grounded fleet planner and discrete-event simulator for sizing multi-pool LLM GPU fleets against P99 TTFT targets, without requiring hardware profiling runs up front.
Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang, Bowei He, Xue Liu
arXiv Technical Report
We derive the minimum-cost two-pool LLM fleet directly from the workload CDF and P99 TTFT target, then use Compress-and-Route to make the optimal boundary deployable in practice.
Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang, Bowei He, Xue Liu
arXiv Technical Report
We derive the 1/W law showing that tokens per watt roughly halve whenever the serving context window doubles, making context-length routing topology a larger energy-efficiency lever than a pure GPU generation upgrade.
Xunzhuo Liu, Hao Wu, Huamin Chen, Bowei He, Xue Liu
arXiv Technical Report
We show how probabilistic ML predicates in policy languages can silently co-fire on the same query, and implement conflict detection plus a softmax-based prevention mechanism in the Semantic Router DSL.
Huamin Chen, Xunzhuo Liu, Bowei He, Xue Liu
arXiv Technical Report
We extend the Semantic Router DSL from stateless, per-request routing to multi-step agent workflows, emitting verified decision nodes for orchestration frameworks, Kubernetes artifacts, YANG/NETCONF payloads, and protocol-boundary gates from a single declarative source file.
Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, Huamin Chen
arXiv Technical Report
We show that conversational memory and retrieval-grounded routing let a lightweight 8B model recover most of a 235B model’s performance on persistent user-specific queries while cutting effective inference cost by 96%.
Xunzhuo Liu, Bowei He, Xue Liu, Haichen Zhang, Huamin Chen
SIGIR 2026 Industry Track
We present a real-time verification component for long-document RAG that processes contexts up to 32K tokens, balancing latency and grounding coverage so interactive systems can detect unsupported answers without falling back to truncated checks.
Huamin Chen, Xunzhuo Liu, Junchen Jiang, Bowei He, Xue Liu
arXiv Technical Report
We propose token-budget-aware pool routing, which estimates each request’s total token budget using a self-calibrating bytes-per-token ratio and dispatches it to short or long vLLM pools to cut fleet cost while avoiding KV-cache failures.
Chen Wang, Xunzhuo Liu, Yuhan Liu, Yue Zhu, Xiangxi Mo, Junchen Jiang, Huamin Chen
NeurIPS - MLForSys
We present a semantic router that classifies queries based on their reasoning requirements and selectively applies reasoning only when beneficial.
Chen Wang, Xunzhuo Liu, Yue Zhu, Alaa Youssef, Priya Nagpurkar, Huamin Chen
We present a category-aware semantic caching where similarity thresholds, TTLs, and quotas vary by query category, with a hybrid architecture separating in-memory HNSW search from external document storage.
Huamin Chen, Luay Jalil
Internet Engineering Task Force (IETF)
This document specifies the Semantic Inference Routing Protocol (SIRP), a framework for content-level classification and semantic routing in AI inference systems.
H. Chen, L. Jalil, N. Cocker
Internet Engineering Task Force (IETF) - Network Management Research Group
This document specifies multi-provider extensions for agentic AI inference APIs. Published: 20 October 2025. Intended Status: Informational. Expires: 23 April 2026.
Maintainers across research, infrastructure, and model systems shape the project together.
vLLM Semantic Router is a community project. Development and testing compute are supported by the organizations below. Thank you for your support.
Donations are collected through the vLLM project to support development, maintenance, and adoption across the ecosystem. Learn more on vllm.ai
Shape every model path with signals, preferences, and policy.