100M active user
10M+ DAU
Search = 100K/sec
Peak Search = 1M/sec (Read)
Write: 10K orders/sec
Peak orders will be 30K orders/sec
Read:Write ratio varies from 10:1 to 30:1 (normal to peak)
Storage:
Product Details + metadata = 1KB / product
Total SKU = 500M items
Storage = 500M * 1KB = 0.5TB
Image storage = 500KB.
Each items has on an average 3 images
Total image storage = 500 * 3 * 500M = 750TB
User data = 100M users * 1 KB = 100GB
Orders (including payment metadata) = 10K * 10^5 * 2KB /day = 2TB/day
Order data in a year = 750TB/year (This is without replication)
For strong availability we will need a replication factor of 3 times atleast.
User:
POST v1/users
Header:
Authentication-Key: Bearer
Idempotence-Key: uuid
{
name, city, address, email, phone_no. etc
}
Response:
{
user_id: uuid, created_at: timestamp, success: 200, cart_id: uuid
}
Products:
Search by product desc:
GET v1/products/q=Bluetooth+Headphone&min_price=1000&max_price=10000&limit=10
Response:
{
products: [{
product_id: uuid, price: 1290, currency: "INR", title: string, description: string, rating: 4.5, variants: [{variant_id: uuid, color: red, image: url_string}, {variant_id: uuid, color: blue, image: url_string}]
}, {}, {}, ....],
has_more: True
cursor: uuid
}
GET v1/products/{product_id}
{
product_id: uuid, price: 1290, currency: "INR", title: string, description: string, rating: 4.5, variants: [{variant_id: uuid, color: red, image: url_string}, {variant_id: uuid, color: blue, image: url_string}]
}
Cart:
POST /v1/pre-order/{cart_id}/{product_id}
Header:
Authentication-Key: Bearer
Idempotence-Key: uuid
{
product_id: uuid, qty: 2
}
Response:
{
{
products: [{product_id: uuid, qty: 2, item_price: 1000, total_price: 2000, currency: INR}, {}, {}]
}
}
GET /v1/pre-order/{cart_id}/
{
products: [{product_id: uuid, qty: 2, item_price: 1000, total_price: 2000, currency: INR}, {}, {}]
}
DELETE /v1/orders/{cart_id}/{product_id}
Header:
Authentication-Key: Bearer
Idempotence-Key: uuid
{
products: [{product_id: uuid, qty: 2, item_price: 1000, total_price: 2000, currency: INR}, {}, {}]
}
Response:
{
products: [{product_id: uuid, qty: 2, item_price: 1000, total_price: 2000, currency: INR}, {}, {}]
}
Order Placement:
POST v1/orders/{cart_id}/place_order
Header:
Authentication-Key: Bearer
Idempotency-Key: uuid
Response:
{order_id: uuid, payment_url: string, amount: int, currency: string}
DELETE v1/orders/{order_id}
Response: 204
Payment:
POST v1/orders/{cart_id}/{order_id}/payment
Header:
Authentication-Key: Bearer
Idempotency-Key: uuid
{amount: int, currency: string, desc: string}
Response:
{payment_id: uuid, status: PROCESSING, created_at: timestamp}
Track Order:
GET v1/orders/{order_id}
Response:
{order_id: uuid, order_placed_at: timestamp, status:[{scan_id: uuid, scan_time: timestamp, location: string}], eta: timestamp}
Users request is routed through Load Balancer and API Gateway, where we also handle Rate Limiter
Product Service:
When user adds a product, we store the metadata and info in a MongoDB, users are provided a pre-signed URL to upload the images/videos.
Once the products are added, an event is emitted to Kafka, where a consumer adds the products to Elastic Search for quick search functionality based on filters provided.
Auth Service: Handles the authentication, user login, password reset etc
Cart Service: It handles the functionality of adding products to a cart.
When an item is added, we first check the validation service, which checks the availability of the product before adding the details to Redis. It is then async to Postgres for durability
Order Service: It handled the fucntionality of placing orders. Before placing an order, we check the validation service, then call the Payment service, get a payment link, and redirect user to payment page.
We rely on webhooks to fetch the status of the payment, while showing processing to the user.
Order Placement and Payment are done using idempotency keys, so that retry do not cause double order/payment
Search Functionality: Search is powered by ES, which can search based on title, description, price_range etc.
Once we have the list of the products_ids, we can score them based on user_relevance, search relevance (title are ranked higher than description matching), business rules (promoted products, profit margin). The product ids are hydrated and returned to the user.
This is also cached in Redis and CDN, as these are very static data.
Redis: We use Redis for product details, product rankings, idempotency key checks, cart functionality, Validation etc
We will choose the following DB for the following services:
Users Table:
(
user_id: uuid [PK]
name, city, address, cart_id...
)
Cart:
(
id: uuid,
cart_id: uuid [FK]
user_id: uuid [FK]
product_id: uuid [FK]
qty: int,
amount: int
base_price: int,
currency: string,
created_at: timestamp
updated_at: timestamp
)
Index at: (user_id, cart_id)
Order Table:
(
user_id: uuid, [FK]
order_id: uuid [PK]
payment_status: string
payment_id: uuid [FK]
order_status: string
)
Payment Table:
(
user_id, order_id, payment_id, status, created_at, payment_method...
)
Mongo for product id, as the product catalog is highly unstructured. Tshirts and laptop have very different specification:
{product_id: uuid, description: string, features: {RAM, CPU, Storage}, price, qty}
Elastic Search will be used for product search, fuzzy matching (since typos are expected in product search), autocomplete
Cart:
We store the info about a user's cart in Redis (for logged in user) and Browser/application memory for non-logged in user and merge them when they log in.
Cart is async stored in Postgres for durability. We do accept the trade off of losing cart data in case of Redis outage for faster reads/writes.
When an user adds a product in the cart, we check the validation service for product availability. This info can be preloaded to Redis when we can anticipate high traffic (prime day sale, black friday) and use Lua scripts to atomically handle it. For regular period, we can store it in Postgres. Either way, we will durably store it.
If the product is available, we will move it from available to blocked section. We will have a periodic check on the cart update times and if it crosses a threshold time (24 hours without any cart activity for the product), we will release the blocked status and move it available. This is to prevent, abandoned carts to hold a product for long time, while legitimate users can't buy it.
Order Placement:
While placing an order, the same mechanism happens for checking the availability and holding the product. We will call the payment gateway, and redirect users to the payment page. We will show PROCESSING to the user, while check the payment status through webhook.
For handling retries, we will use idempotency keys.
Once the order is placed, a Kafka event is added, which is listened by various services like Notification, warehouse, ML.
We also decrement the blocked section by qty and update the Postgres.
Failure Scenarios:
The biggest risk here is payment related.
We will follow a SAGA pattern, where we will do a correction step everytime something goes wrong.
Payment failed -> Retry with idempotency key
Payment failed due to wrong/declined card -> show error to the user
Payment charged -> Add a refund
Order Cancelled/returned -> Add a refund based on business rules
All payment related logs are strictly append only for audit purposes.
To handle flash sales, we can use Redis to keep the counts, add users to the queue.
Show the product details through CDN and heavily cache the product details, ratings etc.
We have separated out the services using Kafka topics so that producers and consumers can independently scale.
We have ML models to rank the products based on user's behaviour. During heavy regression, we can stop this service, and use a simpler metrics like no. of products sold in last month.
For read heavy sections like products, we can also add Read replicas.
For cold cache, we can use request coalescing, while other requests serve stale cache (product details is fine, showing the number of available qty is also fine, as it will be checked in the validation service). We can also have a worker actively identify and re-cache hot keys