Big Data questions and answers
Spark, Hadoop and batch processing at volume. Page 9 of 9.
System Design practice on Codemia
Work through 120+ system design problems with detailed solutions, from rate limiters to multi-region storage.
Answers 481-495
- What ways can a Consumer consume message in Kafka?
- What's the most elegant/right way to stop a spark job running on a Kubernetes cluster?
- When add hdfs file to hive and use in udf, it comes an error
- When is a Kafka connector preferred over a Spark streaming solution?
- Why are my Airflow tasks queued but not running?
- Why do we use distributed cache in hadoop?
- Why does Spark's OneHotEncoder drop the last category by default?
- Why doesn't Hadoop file system support random I/O?
- Why is Ceph and its CRUSH algorithm less used for big data analytics?
- Why Kafka so fast
- Why mutual exclusion is required in MapReduce distributed system?
- Why spark.ml don't implement any of spark.mllib algorithms?
- Why using apache kafka in real-time processing
- Why we require Apache Kafka with NoSQL databases?
- Yarn Distributed cache, no mapper/reducer

Course
Beginner
27 lessons
10 hours
System Design Fundamentals
Build a strong foundation in designing scalable, reliable distributed systems.
View the course