OpenAI-compatible inference gateway that turns a single-node setup into a fault-tolerant, observable, multi-node LLM serving system.
- Routes traffic across multiple vLLM nodes in real time.
- Supports three routing strategies:
round_robin,least_loaded,latency_based. - Handles failures with retries + failover (no single-node bottleneck).
- Enforces backpressure (
semaphore + queue) to stay stable under load. - Exposes production-ready Prometheus metrics and Grafana dashboard.
- Benchmarks include node distribution visibility (
X-Routed-Node).
Client -> Router Gateway -> vLLM Node Pool
Example pool:
http://localhost:8001http://localhost:8002http://localhost:8003
The gateway keeps API compatibility with POST /v1/chat/completions, so existing clients continue to work.
./scripts/setup_linux.shEdit gateway/.env:
VLLM_NODES=http://localhost:8001,http://localhost:8002,http://localhost:8003
ROUTING_STRATEGY=least_loaded
NODE_TIMEOUT_SECONDS=120
MAX_RETRIES=2
HEALTH_CHECK_INTERVAL_SECONDS=15
HEALTH_FAILURE_THRESHOLD=3
GATEWAY_PORT=8000
AUTH_TOKEN=your-secret-token-here
MODEL_NAME=mistralai/Mistral-7B-Instruct-v0.3
CONNECT_TIMEOUT_SECONDS=10
MAX_CONCURRENT_REQUESTS=16
MAX_QUEUE_SIZE=32./scripts/start_cluster.shsource .venv/bin/activate
uvicorn gateway.main:app --host 0.0.0.0 --port 8000./scripts/smoke_test.shPOST /v1/chat/completions- OpenAI-compatible routed inferenceGET /health- gateway status + router config + node listGET /stats- per-node load/latency/errors/health snapshotGET /metrics- Prometheus metrics
Reference (Colab T4, single node, Qwen2.5-3B-Instruct): ~7× higher aggregate throughput concurrent vs naive; sweep shows throughput leveling off once client concurrency exceeds server limits. Colab reproduction.
Run from the repository root with PYTHONPATH=.. Set --model to the same model ID as vLLM and MODEL_NAME in gateway/.env.
| Harness | Purpose |
|---|---|
benchmark.run_benchmark |
Compares sequential (naive) vs overlapping (concurrent) requests; reports throughput and latency. |
benchmark.run_load_sweep |
Varies client concurrency to observe scaling and saturation under gateway / vLLM limits. |
Throughput vs sequential (--mode both):
PYTHONPATH=. python -m benchmark.run_benchmark \
--mode both \
--num-requests 60 \
--concurrency 12 \
--base-url http://localhost:8000 \
--auth-token your-secret-token-here \
--model mistralai/Mistral-7B-Instruct-v0.3 \
--output benchmark/results/run_both.jsonAdd --log-node to summarize routing across nodes via X-Routed-Node.
Concurrency sweep (writes JSON + CSV under benchmark/results/):
PYTHONPATH=. python -m benchmark.run_load_sweep \
--num-requests 60 \
--base-url http://localhost:8000 \
--auth-token your-secret-token-here \
--model mistralai/Mistral-7B-Instruct-v0.3 \
--output-json benchmark/results/load_sweep.json \
--output-csv benchmark/results/load_sweep.csvOptional: --concurrencies 1,2,4,8 to align sweep steps with MAX_CONCURRENT_REQUESTS and vLLM --max-num-seqs.
Prometheus + Grafana track both gateway and node behavior:
- traffic and latency:
requests_total,request_latency_seconds - overload safety:
queue_depth,requests_rejected_total - node-level routing:
requests_per_node_total,node_active_requests - resiliency:
node_failures_total,upstream_errors_total
Grafana:
- http://localhost:3000
- dashboard:
grafana/dashboards/inference.json
- If 3 full replicas do not fit one GPU, use fewer nodes or multiple GPUs.
- Every request needs
Authorization: Bearer <AUTH_TOKEN>.
The inference_aware strategy extends the existing routing logic with inference-specific telemetry. Instead of selecting nodes by pure load or historical wall-clock latency, it estimates the expected time to deliver the response based on per-node TTFT and token generation throughput.
output_estimate = max_tokens * 0.5 # treat max_tokens as an upper bound, not exact
estimated_latency = avg_ttft + (output_estimate / avg_tokens_per_second)
score = estimated_latency * (1 + active_requests * load_penalty)
The node with the lowest score is selected. avg_ttft and avg_tokens_per_second are updated via exponential moving average (alpha=0.2) after each request completes.
When a node has fewer than INFERENCE_AWARE_MIN_SAMPLES completed requests, or when its last update is older than INFERENCE_AWARE_STALENESS_SECONDS, it is treated as having no data and assigned a score of +inf. If all nodes have infinite scores (no inference data yet), the strategy falls back to least_loaded.
This means the strategy is safe to enable from a cold start — it behaves identically to least_loaded until enough telemetry is collected.
Use inference_aware when:
- You have multiple vLLM backends with different GPU generations or batch sizes.
- You see uneven latency across backends (one node is consistently slower to produce the first token).
- Requests have varying
max_tokensvalues and you want to avoid queuing long jobs behind short ones on a slow node.
All settings are read from environment variables (or a gateway/.env file).
| Variable | Default | Description |
|---|---|---|
ROUTING_STRATEGY |
least_loaded |
Set to inference_aware to enable the new strategy. Also supports round_robin and latency_based. |
INFERENCE_AWARE_MIN_SAMPLES |
3 |
Minimum completed requests before a node is scored on inference metrics. Below this, the node falls back to +inf score. |
INFERENCE_AWARE_STALENESS_SECONDS |
60.0 |
Discard inference telemetry older than this many seconds. Prevents stale data from biasing routing after a backend restart or load spike. |
INFERENCE_AWARE_LOAD_PENALTY |
0.1 |
Per-active-request penalty multiplier applied to the estimated latency. Increase this value to route more aggressively away from busy nodes. |
All existing metrics are preserved. The following metrics have been updated with a node label to enable per-backend breakdowns.
Note: Adding a node label to an existing metric is a breaking change for existing Prometheus queries and alert rules that aggregate these metrics without a label filter. Update any dashboards or recording rules accordingly.
| Metric | Labels (updated) | Change |
|---|---|---|
time_to_first_token_seconds |
node |
Added node label; buckets changed from latency range to (0.05, 0.1, 0.25, 0.5, 1.0, 2.0, 5.0) seconds |
tokens_per_second |
node |
Added node label; buckets changed to (1, 5, 10, 25, 50, 100, 200, 500) tokens/sec |
tokens_generated_total |
model, node |
Added node label |
tokens_prompt_total |
model, node |
Added node label |
upstream_errors_total |
code, node |
Added node label; node="" for errors before node selection |
A new helper function observe_inference(node, ttft_seconds, tokens_per_second_value, prompt_tokens, completion_tokens, model) is available for callers that need to record per-node inference metrics without going through observe_request.
pip install -e ".[dev]"
# or, without editable install:
pip install pytest>=8 pytest-asyncio>=0.23 httpx>=0.27cd /path/to/Distributed-LLM-Router
python -m pytest tests/ -vAll tests use mocked backends — no running vLLM instance is required.
| Test file | What it covers |
|---|---|
tests/test_router.py |
All four routing strategies, unhealthy node exclusion, inference-aware fallback, stale data handling |
tests/test_node_manager.py |
Active request tracking, EMA latency, inference stats (TTFT + TPS), health state machine |
tests/test_metrics.py |
Counter increments, TTFT histogram, per-node inference observations |
tests/test_streaming.py |
Semaphore release on completion and exception, TTFT recording, record_inference_stats called |
Because the gateway supports only one routing strategy at runtime, strategy comparison requires running the benchmark once per strategy:
# Step 1: run with least_loaded (default)
ROUTING_STRATEGY=least_loaded python gateway/main.py &
python benchmark/run_strategy_comparison.py --output benchmark/results/least_loaded.json
# Step 2: restart with inference_aware
kill %1
ROUTING_STRATEGY=inference_aware python gateway/main.py &
python benchmark/run_strategy_comparison.py --output benchmark/results/inference_aware.json
# Step 3: compare JSON filespython benchmark/run_strategy_comparison.py \
--gateway-url http://localhost:8000 \
--requests 50 \
--concurrency 5 \
--model mistralai/Mistral-7B-Instruct-v0.3 \
--output benchmark/results/my_run.json
| Flag | Default | Description |
|---|---|---|
--gateway-url |
http://localhost:8000 |
Gateway endpoint |
--requests |
50 |
Total number of requests per mode |
--concurrency |
5 |
Max concurrent requests |
--model |
mistralai/Mistral-7B-Instruct-v0.3 |
Model name to request |
--output |
benchmark/results/strategy_comparison.json |
Output JSON path |
--auth-token |
your-secret-token-here |
Bearer token |
--no-streaming |
— | Skip streaming mode, run non-streaming only |
- Latency p50 / p95 / p99 (ms)
- TTFT p50 / p95 (ms, streaming only)
- Average tokens per second
- Request throughput (req/s)
- Error rate
DISCLAIMER: Results from this script reflect a combination of gateway routing overhead and backend inference performance. The inference_aware strategy requires a warm-up period (INFERENCE_AWARE_MIN_SAMPLES requests) before it can score nodes on inference metrics. For meaningful comparisons, use identical hardware, identical backend configurations, and run each strategy with the same prompt mix and concurrency settings.