List functional requirements for the system (Ask the chat bot for hints if stuck.)...
This is essential for ensuring only registered users can publish advertisements.
Users can search for advertisements by keywords and view the found advertisements.
The users can edit their advertisements.
The user can delete their advertisements
The admin can add the sections to structure the advertisements.
List non-functional requirements for the system...
This service should be highly available.
This service should be highly scalable.
When users publish new advertisements, they appear within 10 milliseconds.
The service should be highly reliable.
Only registered users can publish, edit, and delete the advertisements; others may only read.
The advertisements will be shown in chronological order in the board messages.
The advertisements can have only text.
Estimate the scale of the system you are going to design...
The section of advertisement information:
id long (8 bytes)
parent_id long (8 bytes)
name string (200 bytes)
description string (1024 bytes)
createdAt timestamp (4 bytes)
changedAt timestamp (4 bytes)
The advertisement information:
id long (8 bytes)
parent_id (8 bytes)
content string (2048 bytes)
changedAt timestamp (4 bytes)
createdAt timestamp (4 bytes)
user_id long (8 bytes)
The user information:
id long (8 bytes)
email string (200 bytes)
nickname string (100 bytes)
createdAt timestamp (4 bytes)
role string(1 byte)
status string (1 byte)
Let's have 1 MAU. Every user can read the advertisement 5 times daily and publish a new advertisement 2 times a week.
Every advertisement takes 2 kilobytes, and it takes 4 gigabytes per year to store added advertisements, or 40 Gigabytes for 5 years.
In order to search the advertisement we introduce the reverse index.
The reversed index contains a word and a list of unique IDs, where each unique ID takes 8 bytes. Let's count the every word take 100 unique ids and it has 1 million words and every word takes 500 bytes. In total, it takes 8 PetaBytes.
Define what APIs are expected from the system...
The Craigslist can benefit from a well-defined set of API to facilitate communication between different component and potentially integrate with external systems. Here's a breakdown of potential APIis:
User API:
Admin API:
Defining the system data model early on will clarify how data will flow among different components of the system. Also you could draw an ER diagram using the diagramming tool to enhance your design...
We use PostgreSQL, which we favor for strong consistency over availability.PostgreSQL is scalable and fault-tolerant.
The section of advertisement information:
id long (8 bytes)
parent_id long (8 bytes)
name string (200 bytes)
visible byte (1 byte)
description string (1024 bytes)
createdAt timestamp (4 bytes)
changedAt timestamp (4 bytes)
The advertisement information:
id long (8 bytes)
parent_id (8 bytes)
content string (2048 bytes)
changedAt timestamp (4 bytes)
createdAt timestamp (4 bytes)
user_id long (8 bytes)
The user information:
id long (8 bytes)
email string (200 bytes)
nickname string (100 bytes)
createdAt timestamp (4 bytes)
role string(1 byte)
status string (1 byte)
You should identify enough components that are needed to solve the actual problem from end to end. Also remember to draw a block diagram using the diagramming tool to augment your design. If you are unfamiliar with the tool, you can simply describe your design to the chat bot and ask it to generate a starter diagram for you to modify...
We will take microservice approach. Advertisement service manages the ads API. If ads is updated, Advertisement service sends an update event to Search service to update index.
Dig deeper into 2-3 components and explain in detail how they work. For example, how well does each component scale? Any relevant algorithm or data structure you like to use for a component? Also you could draw a diagram using the diagramming tool to enhance your design...
Advertisement service:
Client sends a request to get advertisement by id, Gateway API picks up the less loaded(we uses least connection load balancing) Advertisement Service instance to process the request. Advertisement Service attempts to invoke the Cache (Redis) if the requested advertisement is there. If so Advertisement Service returns the found advertisement to the client otherwise, it invokes PostgreSQL to get the requested advertisement, after the persisted advertisement is added to the Cache. Advertisement Service returns it to the client. Suppose the client requests to search advertisements by words, the Advertisement Service invokes the Search service to find the advertisements by words. In that case, The Search service uses Elastic search and its index to find the advertisements' ids and if the Search service returns the advertisements' ids,they are returned to the Advertisement Service,the Advertisement Service request advertisement object from PostgreSQL, updates/adds them to Cache and returns advertisement object to clients. If the advertisement service updates the advertisement in PostgresSQL,the cache, also the advertisement service sends the update advertisement event to the Search service with Kafka.
We'going to deploy a few instances Advertisement service, Search service, and User service to AWS by using Kubernates to orchestrate service instances (containers) pods. Every of them are stateless, and Kubernates manages the full pod lifecycle. PostgreSQL uses master and slave instances, master instance processes the write requests and replicates the updates to slave instances, the slave instances caters the read requests.
Explain any trade offs you have made and why you made certain tech choices...
The big decision was between NoSQL and RDB.
RDB comes on top because:
Pro:
Con:
Because RDB doesn't have good availability, otherwise, instances would have different data versions.
In order to the clients can update data simultaneously, we can use the RDB pessimistic data lock, if client try to publish data update, it checked the data version in memory and in RDB and if they don't match, the client reads the last data from RDB and merges it and try to update the data again.
The search service would use the index that will be built by the Elastic Search, when a client adds/updates the advertisement,the advertisement service send the updated data to Search service with Kafka. The Search service receives the advertisement data from Kafka, tokenized this data and adds them to the reverse index: word, word weight, and list of advertisement primary keys. When a client searches the advertisement by words, the Search service uses the reverse index in the Elasttic Search to find the appropriate advertisement primary keys and returns them to the Advertisment Service.
The admin can add the sections to structure the advertisement. When the admin adds new sections, they aren't visible at once to the user. if the admin enables the root section, they become available for clients to add new advertisements there.
We can partition the advertisement table by primary key, which guarantees even data distribution.
In order to offload the RDB, we can use Redis as a cache with an eviction policy as LRU or LFU. Redis gives good performance and fault-tolerance.
Try to discuss as many failure scenarios/bottlenecks as possible.
The advertisement service is a crucial service for the system if it fails, the system can't work properly. In order to prevent this service downtime, we can replicate the instances of this service and deploy them to different AZ in AWS, The K8S can manage the service instances efficiently, also we can keep a copy of this service in master-to-master replication so that this copy takes over to process client requests.
PostgreSQL is a mandatory part of the system. We can have the number of replicas that can be promoted to master if the current master is down. We can replicate data to the user PostgreSQL cluster to not loose client data.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?