Big Data questions and answers
Spark, Hadoop and batch processing at volume. Page 2 of 9.
System Design practice on Codemia
Work through 120+ system design problems with detailed solutions, from rate limiters to multi-region storage.
Answers 61-120
- Counting Number of messages stored in a kafka topic
- Create Spark DataFrame in Spark Streaming from JSON Message on Kafka
- CSV Connector For Kafka
- Data aggregation using apache flink
- data structure for indexing big file
- Database synchronization time in cassandra
- Datanode denied communication with namenode because hostname cannot be resolved
- Dealing with unbalanced datasets in Spark MLlib
- Decision tree implementation issue in apache spark with java
- Delaying Kafka Streams consuming
- Difference between kafka and nifi
- Direct Kafka Stream with PySpark (Apache Spark 1.6)
- Distibuted Cache in Reduce Hadoop
- Distribute messages equally into partitions in kafka
- Distributed alternatives to hadoop
- Distributed cache with Pig and Python
- Distributed Hash Tables Preventing nodes from storing petabytes of data?
- Distributed Kafka Connect topic configuration
- Distributed logs in Cassandra
- DistributedCache Hadoop - FileNotFound
- DistributedCache in Hadoop 2.x
- Does a file need to be in HDFS in order to use it in distributed cache?
- Does an EMR master node know its cluster ID?
- Does Spark Structured Streaming maintain the order of Kafka messages?
- DStream filtering and offset management in Spark Streaming Kafka
- Dynamically update topics list for spark kafka consumer
- Efficiency of Querying 10 Billion Rows (with High Cardinality) in ScyllaDB
- Efficiently convert edge list to adjacency list using MapReduce
- Encrypting the Hadoop Distributed Cache file
- End-to-end Exactly-once processing in Apache Flink
- Error Compiling Hadoop WordCount MapReduce Example
- Error handling in hadoop map reduce
- Error when Spark 2.2.0 standalone mode write Dataframe to local single-node Kafka
- ETL in Java Spring Batch vs Apache Spark Benchmarking
- Eventsourcing in Apache Kafka
- Exception while accessing KafkaOffset from RDD
- External shuffle shuffling large amount of data out of memory
- Extract the time stamp from kafka messages in spark streaming?
- Extremely slow S3 write times from EMR/ Spark
- Fail to create SparkContext
- Feature Selection in PySpark
- Fetch all rows in cassandra
- FileNotFound Exception when trying to store file in hadoop distributed cache
- Finding median of large set of numbers too big to fit into memory
- Fluentd vs Kafka
- Flume use case reading from HTTP and push to HDFS via Kafka
- General techniques to work with huge amounts of data on a non-super computer
- Get context from Pod launched with Airflow KubernetesPodOperator
- Get Kafka compressed message size
- get topic from kafka message in spark
- Geting exception while using DistributedCache in Hadoop
- Geting messages of Offset is getting reset in structured streaming mode in Spark
- Getting Multiple messages on spark-submit seeking to EARLIEST and Resetting offset for partition topic-partition
- Guava version while using spark-shell
- Hadoop - Directory Structure and Distributed Cache
- Hadoop - Large files in distributed cache
- Hadoop 1.2.1 - using Distributed Cache
- Hadoop cache file for all map tasks
- Hadoop Distributed Cache - modify file
- Hadoop Distributed Cache don't work

Course
Beginner
27 lessons
10 hours
System Design Fundamentals
Build a strong foundation in designing scalable, reliable distributed systems.
View the course