Storage Requirements for Tagging Service
Here's how to estimate the storage requirements for your tagging service, considering the below assumptions:
Digital Items:
Tags:
Calculations:
Storage for digital items:
item_storage = num_items * avg_item_size
item_storage = 100,000,000 * 500 KB
item_storage = 50,000,000,000 KB
item_storage = 50 TB
Storage for tags:
tag_storage = num_items * max_tags_per_item * tag_length
tag_storage = 100,000,000 * 10 * 128
tag_storage = 128,000,000,000 bytes
tag_storage = 128 GB
Total storage:
total_storage = item_storage + tag_storage
total_storage = 50 TB + 128 GB
Here's a breakdown of potential APIs for your tagging service, addressing both user interaction and internal functionalities:
User-facing APIs:
1. Upload Digital Item
2. Get Digital Item
3. Add Tags to Item
4. Remove Tags from Item
5. Search for Items
Internal APIs (optional):
1. Get Tag Suggestions
2. Get Popular Tags
3. Normalize Tag
User Data, Data Item Metadata
Tags
Search Index
Best Strategy: Hash Partitioning by User ID
Partitioning Algorithm: Consistent Hashing is a popular choice for its ability to distribute data evenly across shards and handle node addition/removal efficiently.
Best Strategy: Vertical Sharding
Here's a breakdown of the main components needed for your tagging service:
1. User Management:
2. Item Upload Service:
3. Tag Management Service:
4. Search Service:
5. API Gateway:
6. Monitoring and Logging:
7. Queueing System (Optional):
8. Administration Panel (Optional):
Communication and Data Flow:
This high-level design provides a solid foundation for your tagging service. You can further refine the components and their functionalities based on your specific requirements and chosen technologies.
Below diagram shows sequence diagram considering a scenario where user searches for data items, adds tags and saves tags.
The Tag Management Service plays a crucial role in your tagging system by handling all aspects of tag creation, modification, and retrieval associated with uploaded digital items. Here's a closer look at its functionalities:
Responsibilities:
Normalization Goals:
Normalization Techniques:
Performance Considerations:
Implementation Options:
1. Entity-Attribute-Value (EAV) Model:
+---------------+-----------------+--------------------+
| item_id | attribute 1 | attribute 2 |
|---------------|-----------------|--------------------|
| 1 | dog | playful |
| 2 | cat | lazy |
| 3 | car | red |
+---------------+-----------------+--------------------+
Advantages:
Disadvantages:
2. Document-oriented Databases:
{
"item_id": 1,
"metadata":
{
"filename": "image.jpg",
"size": 1024
},
"tags": ["cat", "funny"]
}
Advantages:
Disadvantages:
Explain any trade offs you have made and why you made certain tech choices...
Try to discuss as many failure scenarios/bottlenecks as possible.
What are some future improvements you would make? How would you mitigate the failure scenario(s) you described above?