Design a Load Balancing for Goldman Sachs
Last updated: May 11, 2026
Quick Overview
Design a low-latency load balancing system that handles millions of requests. Discuss trade-offs in consistency, availability, and performance.
Goldman Sachs
May 11, 202649
4
3,930 solved
Design a low-latency load balancing system that handles millions of requests. Discuss trade-offs in consistency, availability, and performance.
ML system design at Goldman Sachs goes beyond model selection. This Onsite question evaluates your ability to design end-to-end ML pipelines, from data collection to model serving, while considering production constraints like latency and reliability.
What the Interviewer Expects
- Map the business problem to a concrete ML objective
- Propose reasonable features and a baseline model
- Discuss basic model evaluation metrics
- Outline a simple serving architecture
Key Topics to Cover
How to Approach This
- Start by clarifying functional and non-functional requirements with the interviewer.
- Estimate the scale: QPS, storage, bandwidth. This drives your design decisions.
- Draw a high-level architecture first, then deep dive into 1-2 critical components.
- Discuss trade-offs explicitly (e.g., consistency vs availability, SQL vs NoSQL).
- Address failure scenarios, monitoring, and how the system handles 10x traffic spikes.
Possible Follow-up Questions
- How would you debug a model that works well offline but poorly online?
- How would you handle a 10x increase in prediction requests?
- How would you run A/B tests on different model versions?
Practice a Similar Problem on Codemia
Solve a related problem with our interactive workspace, get AI feedback, and view detailed solutions.
Solve on CodemiaSample Answer
Requirements
Functional Requirements
- Request Handling: The system must efficiently distribute incoming requests to multiple backend servers, ensuring low latency (< 100ms response time).
- **Health Che...
Capacity Estimation
To estimate capacity, assume we expect to handle 10 million requests per day.
-
Requests per second (RPS):
- 10 million requests / 86400 seconds = ~115.74 RPS.
-
Concurrent Users: Assum...