Skip to content

About

Distributed LLM Router: an OpenAI-compatible FastAPI gateway that intelligently routes requests across multiple vLLM nodes with load balancing, failover retries, backpressure, and real-time Prometheus/Grafana observability.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

Distributed LLM Router

OpenAI-compatible inference gateway that turns a single-node setup into a fault-tolerant, observable, multi-node LLM serving system.

Project Goals

  • Routes traffic across multiple vLLM nodes in real time.
  • Supports three routing strategies: round_robin, least_loaded, latency_based.
  • Handles failures with retries + failover (no single-node bottleneck).
  • Enforces backpressure (semaphore + queue) to stay stable under load.
  • Exposes production-ready Prometheus metrics and Grafana dashboard.
  • Benchmarks include node distribution visibility (X-Routed-Node).

What it is

Client -> Router Gateway -> vLLM Node Pool

Example pool:

  • http://localhost:8001
  • http://localhost:8002
  • http://localhost:8003

The gateway keeps API compatibility with POST /v1/chat/completions, so existing clients continue to work.

Quick start (5 steps)

1) Setup environment

./scripts/setup_linux.sh

2) Configure router

Edit gateway/.env:

VLLM_NODES=http://localhost:8001,http://localhost:8002,http://localhost:8003
ROUTING_STRATEGY=least_loaded
NODE_TIMEOUT_SECONDS=120
MAX_RETRIES=2
HEALTH_CHECK_INTERVAL_SECONDS=15
HEALTH_FAILURE_THRESHOLD=3
GATEWAY_PORT=8000
AUTH_TOKEN=your-secret-token-here
MODEL_NAME=mistralai/Mistral-7B-Instruct-v0.3
CONNECT_TIMEOUT_SECONDS=10
MAX_CONCURRENT_REQUESTS=16
MAX_QUEUE_SIZE=32

3) Start vLLM cluster

./scripts/start_cluster.sh

4) Start gateway

source .venv/bin/activate
uvicorn gateway.main:app --host 0.0.0.0 --port 8000

5) Smoke test

./scripts/smoke_test.sh

Core endpoints

  • POST /v1/chat/completions - OpenAI-compatible routed inference
  • GET /health - gateway status + router config + node list
  • GET /stats - per-node load/latency/errors/health snapshot
  • GET /metrics - Prometheus metrics

Benchmark

Reference (Colab T4, single node, Qwen2.5-3B-Instruct): ~7× higher aggregate throughput concurrent vs naive; sweep shows throughput leveling off once client concurrency exceeds server limits. Colab reproduction.

Run from the repository root with PYTHONPATH=.. Set --model to the same model ID as vLLM and MODEL_NAME in gateway/.env.

Harness Purpose
benchmark.run_benchmark Compares sequential (naive) vs overlapping (concurrent) requests; reports throughput and latency.
benchmark.run_load_sweep Varies client concurrency to observe scaling and saturation under gateway / vLLM limits.

Throughput vs sequential (--mode both):

PYTHONPATH=. python -m benchmark.run_benchmark \
  --mode both \
  --num-requests 60 \
  --concurrency 12 \
  --base-url http://localhost:8000 \
  --auth-token your-secret-token-here \
  --model mistralai/Mistral-7B-Instruct-v0.3 \
  --output benchmark/results/run_both.json

Add --log-node to summarize routing across nodes via X-Routed-Node.

Concurrency sweep (writes JSON + CSV under benchmark/results/):

PYTHONPATH=. python -m benchmark.run_load_sweep \
  --num-requests 60 \
  --base-url http://localhost:8000 \
  --auth-token your-secret-token-here \
  --model mistralai/Mistral-7B-Instruct-v0.3 \
  --output-json benchmark/results/load_sweep.json \
  --output-csv benchmark/results/load_sweep.csv

Optional: --concurrencies 1,2,4,8 to align sweep steps with MAX_CONCURRENT_REQUESTS and vLLM --max-num-seqs.

Observability highlights

Prometheus + Grafana track both gateway and node behavior:

  • traffic and latency: requests_total, request_latency_seconds
  • overload safety: queue_depth, requests_rejected_total
  • node-level routing: requests_per_node_total, node_active_requests
  • resiliency: node_failures_total, upstream_errors_total

Grafana:

Notes

  • If 3 full replicas do not fit one GPU, use fewer nodes or multiple GPUs.
  • Every request needs Authorization: Bearer <AUTH_TOKEN>.

Inference-Aware Routing

The inference_aware strategy extends the existing routing logic with inference-specific telemetry. Instead of selecting nodes by pure load or historical wall-clock latency, it estimates the expected time to deliver the response based on per-node TTFT and token generation throughput.

Scoring formula

output_estimate = max_tokens * 0.5        # treat max_tokens as an upper bound, not exact
estimated_latency = avg_ttft + (output_estimate / avg_tokens_per_second)
score = estimated_latency * (1 + active_requests * load_penalty)

The node with the lowest score is selected. avg_ttft and avg_tokens_per_second are updated via exponential moving average (alpha=0.2) after each request completes.

Fallback behavior

When a node has fewer than INFERENCE_AWARE_MIN_SAMPLES completed requests, or when its last update is older than INFERENCE_AWARE_STALENESS_SECONDS, it is treated as having no data and assigned a score of +inf. If all nodes have infinite scores (no inference data yet), the strategy falls back to least_loaded.

This means the strategy is safe to enable from a cold start — it behaves identically to least_loaded until enough telemetry is collected.

When to use it

Use inference_aware when:

  • You have multiple vLLM backends with different GPU generations or batch sizes.
  • You see uneven latency across backends (one node is consistently slower to produce the first token).
  • Requests have varying max_tokens values and you want to avoid queuing long jobs behind short ones on a slow node.

Configuration

All settings are read from environment variables (or a gateway/.env file).

New env vars for inference-aware routing

Variable Default Description
ROUTING_STRATEGY least_loaded Set to inference_aware to enable the new strategy. Also supports round_robin and latency_based.
INFERENCE_AWARE_MIN_SAMPLES 3 Minimum completed requests before a node is scored on inference metrics. Below this, the node falls back to +inf score.
INFERENCE_AWARE_STALENESS_SECONDS 60.0 Discard inference telemetry older than this many seconds. Prevents stale data from biasing routing after a backend restart or load spike.
INFERENCE_AWARE_LOAD_PENALTY 0.1 Per-active-request penalty multiplier applied to the estimated latency. Increase this value to route more aggressively away from busy nodes.

New and Updated Metrics

All existing metrics are preserved. The following metrics have been updated with a node label to enable per-backend breakdowns.

Note: Adding a node label to an existing metric is a breaking change for existing Prometheus queries and alert rules that aggregate these metrics without a label filter. Update any dashboards or recording rules accordingly.

Metric Labels (updated) Change
time_to_first_token_seconds node Added node label; buckets changed from latency range to (0.05, 0.1, 0.25, 0.5, 1.0, 2.0, 5.0) seconds
tokens_per_second node Added node label; buckets changed to (1, 5, 10, 25, 50, 100, 200, 500) tokens/sec
tokens_generated_total model, node Added node label
tokens_prompt_total model, node Added node label
upstream_errors_total code, node Added node label; node="" for errors before node selection

A new helper function observe_inference(node, ttft_seconds, tokens_per_second_value, prompt_tokens, completion_tokens, model) is available for callers that need to record per-node inference metrics without going through observe_request.


Testing

Install dev dependencies

pip install -e ".[dev]"
# or, without editable install:
pip install pytest>=8 pytest-asyncio>=0.23 httpx>=0.27

Run the test suite

cd /path/to/Distributed-LLM-Router
python -m pytest tests/ -v

All tests use mocked backends — no running vLLM instance is required.

Test coverage

Test file What it covers
tests/test_router.py All four routing strategies, unhealthy node exclusion, inference-aware fallback, stale data handling
tests/test_node_manager.py Active request tracking, EMA latency, inference stats (TTFT + TPS), health state machine
tests/test_metrics.py Counter increments, TTFT histogram, per-node inference observations
tests/test_streaming.py Semaphore release on completion and exception, TTFT recording, record_inference_stats called

Benchmarking

Strategy comparison

Because the gateway supports only one routing strategy at runtime, strategy comparison requires running the benchmark once per strategy:

# Step 1: run with least_loaded (default)
ROUTING_STRATEGY=least_loaded python gateway/main.py &
python benchmark/run_strategy_comparison.py --output benchmark/results/least_loaded.json

# Step 2: restart with inference_aware
kill %1
ROUTING_STRATEGY=inference_aware python gateway/main.py &
python benchmark/run_strategy_comparison.py --output benchmark/results/inference_aware.json

# Step 3: compare JSON files

Benchmark options

python benchmark/run_strategy_comparison.py \
  --gateway-url http://localhost:8000 \
  --requests 50 \
  --concurrency 5 \
  --model mistralai/Mistral-7B-Instruct-v0.3 \
  --output benchmark/results/my_run.json
Flag Default Description
--gateway-url http://localhost:8000 Gateway endpoint
--requests 50 Total number of requests per mode
--concurrency 5 Max concurrent requests
--model mistralai/Mistral-7B-Instruct-v0.3 Model name to request
--output benchmark/results/strategy_comparison.json Output JSON path
--auth-token your-secret-token-here Bearer token
--no-streaming — Skip streaming mode, run non-streaming only

Metrics reported

  • Latency p50 / p95 / p99 (ms)
  • TTFT p50 / p95 (ms, streaming only)
  • Average tokens per second
  • Request throughput (req/s)
  • Error rate

DISCLAIMER: Results from this script reflect a combination of gateway routing overhead and backend inference performance. The inference_aware strategy requires a warm-up period (INFERENCE_AWARE_MIN_SAMPLES requests) before it can score nodes on inference metrics. For meaningful comparisons, use identical hardware, identical backend configurations, and run each strategy with the same prompt mix and concurrency settings.

About

Distributed LLM Router: an OpenAI-compatible FastAPI gateway that intelligently routes requests across multiple vLLM nodes with load balancing, failover retries, backpressure, and real-time Prometheus/Grafana observability.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages