Design a Data Pipeline Service
Last updated: April 17, 2026
Quick Overview
Design a fault-tolerant data pipeline system that handles millions of requests. Discuss trade-offs in consistency, availability, and performance.
Atlassian
April 17, 2026152
7
1,930 solved
Design a fault-tolerant data pipeline system that handles millions of requests. Discuss trade-offs in consistency, availability, and performance.
System design interviews at Atlassian typically last 45-60 minutes. You are expected to drive the conversation, starting from requirements gathering through to a detailed architecture. The interviewer will evaluate your ability to handle ambiguity and make practical engineering decisions.
What the Interviewer Expects
- Drive the design discussion proactively with minimal interviewer guidance
- Perform detailed capacity estimation and use it to inform design decisions
- Design for global scale with multi-region deployment and data consistency
- Deep dive into 2-3 critical components with implementation-level detail
- Address complex trade-offs: CAP theorem, eventual consistency, conflict resolution
- Discuss operational excellence: deployment strategy, chaos engineering, SLOs/SLIs
- Propose a phased rollout plan from MVP to full-scale system
Key Topics to Cover
How to Approach This
- Start by clarifying functional and non-functional requirements with the interviewer.
- Estimate the scale: QPS, storage, bandwidth. This drives your design decisions.
- Draw a high-level architecture first, then deep dive into 1-2 critical components.
- Discuss trade-offs explicitly (e.g., consistency vs availability, SQL vs NoSQL).
- Address failure scenarios, monitoring, and how the system handles 10x traffic spikes.
Possible Follow-up Questions
- How would you handle a region-wide outage?
- How would you optimize costs as the system scales?
- How would you migrate from a monolithic to a microservices architecture?
Practice a Similar Problem on Codemia
Solve a related problem with our interactive workspace, get AI feedback, and view detailed solutions.
Solve on CodemiaSample Answer
Requirements
Functional Requirements
- Data Ingestion: The service must accept millions of events per second from various Atlassian products (e.g., Jira, Confluence).
- Data Processing: Capable of tran...
Capacity Estimation
Assuming each Atlassian product generates around 1 million events per minute:
- Total Events per Second: 1 million events / 60 seconds = ~16,667 events/second.
- Data Size: If each event avera...