SLA-Aware Inference Gateway
A gateway that keeps image-classification latency inside a 300 ms SLA by routing each request between an accurate ResNet-50 (int8 ONNX) and a fast MobileNetV3, based on live p95 latency. Built from scratch and measured end to end under load.
Problem
At twice its capacity a single accurate model collapses: p95 latency of 2.7 s, 27% of requests failing, and 0.4% answered within the SLA. Always using the small model avoids that at a permanent 8-point accuracy cost.
What I built
The whole system: ONNX export and int8 quantization of 8 candidate models evaluated on 9,500 held-out images, the model servers, the gateway and its control loop, canary rollouts, an open-loop load generator, Docker Compose, kind and GKE deployments, and the benchmark suite (34 tests).
Result
2,721 → 192 msp95 at 2× capacity, with 99.8% of requests within the SLA and zero errors, while keeping accuracy 4 points above the always-fast option. A faulty canary was rolled back in 16 s with no client-visible errors. Found and fixed a CFS-throttling issue in ONNX Runtime that cut p95 by 3.8×. On GKE, a 6-minute 4× spike stayed 99.8% within SLA while the HPA and cluster autoscaler added pods and a node in 81 s. Measured energy per inference with powermetrics: ONNX Runtime int8 ResNet-50 at 0.79 J vs 3.75 J for Core ML fp32 on the same CPU, and 0.15 J for the fast tier's MobileNetV3.
- Python
- FastAPI
- ONNX Runtime
- int8 quantization
- Kubernetes (GKE)
- Prometheus
- OpenTelemetry