Big Data questions and answers
Spark, Hadoop and batch processing at volume. Page 6 of 9.
System Design practice on Codemia
Work through 120+ system design problems with detailed solutions, from rate limiters to multi-region storage.
Answers 301-360
- Multiple windows of different durations in Spark Streaming application
- NoSuchMethodError with Spark Streaming 2.2.0. and Kafka 0.8
- Not able to import the spark packages
- Not Serializable exception when reading Kafka records with Spark Streaming
- Object not serializable (org.apache.kafka.clients.consumer.ConsumerRecord) in Java spark kafka streaming
- One Kafka consumer in a group consistently rejects coordinator, but only when Spark and Kafka are both in EC2
- org.apache.spark.SparkException Task not serializable
- org.apache.spark.sql.AnalysisException Can't extract value from probability
- Out of memory exception during TFIDF generation for use in Spark's MLlib
- Output Dstream of Apache Spark in Python
- Parallelize a collection with Spark
- Parsing one terabyte of text and efficiently counting the number of occurrences of each word
- Partition By Multiple Nested Fields in Kafka Connect HDFS Sink
- Passing bigger data in a service-oriented architecture
- Performance Benchmarks for Kafka KTables
- Performance decrease for huge amount of columns. Pyspark
- Persisting Spark Streaming output
- Pig Distributed cache
- Pod template for specifying tolerations when running Spark on Kubernetes
- Poor performance with Spark streaming, Kafka and multiple topics
- Process parquet file row-wise
- Production architecture for big data real time machine learning application?
- Py4JJavaError An error occurred while calling None.org.apache.spark.api.java.JavaSparkContext
- PySpark 2.x Programmatically adding Maven JAR Coordinates to Spark
- PySpark Can only call getServletHandlers on a running MetricsSystem
- pyspark NameError name 'spark' is not defined
- Query on Hadoop High Availability
- Re-use files in Hadoop Distributed cache
- Read and process a batch of messages from Kafka
- Read from Kafka and write to hdfs in parquet
- Read Kafka topic in a Spark batch job
- Read sharded output from Hadoop job from DistributedCache
- Read timed out Httpfs HDFS
- Reading Avro messages from Kafka with Spark 2.0.2 (structured streaming)
- Reading file inside driver Hadoop
- Reading file inside main function - Hadoop
- Reading HAR file from DistributedCache in mapreduce
- real time log processing using apache spark streaming
- Regarding Apache nifi - Distrubuted Cache
- Relationship between number of subtasks in Flink and resource usage
- Remove Airflow Scheduler logs
- Retaining data in Apache Kafka
- Retrieve history of past kafka consumers
- Retrieve Timestamp based data from Kafka
- Retrieving the top 100 numbers from one hundred million of numbers
- Reuse static variable in Hadoop
- Right database for machine learning on 100 TB of data
- Run ML algorithm inside map function in Spark
- Running dependent hadoop jobs in one driver
- Running Tensorflow on big data
- Save null Values in Cassandra using DataStax Spark Connector
- Scaling out with 200+ Kafka topics
- Schedule each Apache Spark Stage to run on a specific Worker Node
- Sending Large CSV to Kafka using python Spark
- Should programmers use SSIS, and if so, why?
- Simplest way to go about transforming data from kafka
- Slow Performance with Apache Spark Gradient Boosted Tree training runs
- Solr shard distribution data not distributed evenly
- Sorting 1 trillion integers
- spark-streaming-kafka-0-10 auto.offset.reset is always set to none

Course
Beginner
27 lessons
10 hours
System Design Fundamentals
Build a strong foundation in designing scalable, reliable distributed systems.
View the course