Scoring Server - Design a Low-Latency Real-Time Scoring Service
Interview Experience
Round 1 System Design
Problem
Design a scoring server that receives raw feature vectors in real time and returns model scores within 20ms p99. The model is a gradient boosted tree (100 MB serialized). The server handles 50K requests/sec at peak.
Requirements
- Latency: p99 < 20ms end-to-end (network + inference).
- Throughput: 50K RPS peak, 10K RPS average.
- The model is updated daily; zero-downtime rollout required.
- Feature input: JSON payload, ~50 float fields.
Design Points
Load Balancer -> Scoring Fleet (stateless workers)
Workers: deserialize JSON -> validate -> run model ->
**return** score
Model loaded in-process (no subprocess call)
Blue/Green deploy: new model warmed up, traffic shifted atomically
Discussion Questions
- How do you manage model warm-up time when spinning up new instances?
- How do you validate incoming features for schema drift before scoring?
- What metrics do you instrument: latency histogram, score distribution, error rate?
Follow-ups
- How do you A/B test two model versions in production with consistent user assignment?
- What happens when a feature is missing in the payload — impute, reject, or score with default?
- How do you handle a latency spike caused by a single slow feature transformation?
- How would the design differ for a deep learning model that requires a GPU?
Full Details
Round 1 System Design
Problem
Design a scoring server that receives raw feature vectors in real time and returns model scores within 20ms p99. The model is a gradient boosted tree (100 MB serialized). The server handles 50K requests/sec at peak.
Requirements
- Latency: p99 < 20ms end-to-end (network + inference).
- Throughput: 50K RPS peak, 10K RPS average.
- The model is updated daily; zero-downtime rollout required.
- Feature input: JSON payload, ~50 float fields.
Design Points
Load Balancer -> Scoring Fleet (stateless workers)
Workers: deserialize JSON -> validate -> run model ->
**return** score
Model loaded in-process (no subprocess call)
Blue/Green deploy: new model warmed up, traffic shifted atomically
Discussion Questions
- How do you manage model warm-up time when spinning up new instances?
- How do you validate incoming features for schema drift before scoring?
- What metrics do you instrument: latency histogram, score distribution, error rate?
Follow-ups
- How do you A/B test two model versions in production with consistent user assignment?
- What happens when a feature is missing in the payload — impute, reject, or score with default?
- How do you handle a latency spike caused by a single slow feature transformation?
- How would the design differ for a deep learning model that requires a GPU?
About This Question
This is a candidate experience report from a ixl interview during the phone round.
It covers the following topics: Phone, System Design, System Design, Coding, Onsite .