Big Data questions and answers
Spark, Hadoop and batch processing at volume. Page 5 of 9.
System Design practice on Codemia
Work through 120+ system design problems with detailed solutions, from rate limiters to multi-region storage.
Answers 241-300
- is it possible to use apache mahout without hadoop dependency?
- Is scikit-learn suitable for big data tasks?
- Is there a .NET equivalent to Apache Hadoop?
- Is there a way to dynamically stop Spark Structured Streaming?
- java.lang.IllegalArgumentException Wrong FS , expected hdfs//localhost9000
- java.lang.IllegalStateException Error reading delta file, spark structured streaming with kafka
- java.lang.NoClassDefFoundError org/apache/spark/Logging
- Join Vs Reduce In Batch Processing
- Json file data into kafka topic
- Kafka->Spark->Cassandra forcing data locality
- Kafka -> Flink DataStream -> MongoDB
- Kafka and hotspots in a partition
- Kafka as an Akka-persistence journal
- kafka cluster configuration
- kafka connect hdfs sink connector is failing even when json data contains schema and payload field
- Kafka connect HDFS sink ERROR failed creating a WAL
- Kafka Connect How can I send protobuf data from Kafka topics to HDFS using hdfs sink connector?
- Kafka connect or Kafka Client
- Kafka Consumer Assignment returns Empty Set
- Kafka consumer in Spark Streaming
- Kafka data types of messages
- Kafka KStream-KTable join race condition
- Kafka multiple consumers for a partition
- Kafka number of topics vs number of partitions
- Kafka producer to read data files
- Kafka RecordMetadata use?
- Kafka Spark streaming unable to read messages
- Kafka Stream Suppress session-windowed-aggregation
- kafka streams - joining partitioned topics
- Kafka Streams - Send on different topics depending on Streams Data
- Kafka Streams How to ensure offset is committed after processing is completed
- Kafka Streams use case
- Kafka Streams with lookup data on HDFS
- Kafka to Elasticsearch, HDFS with Logstash or Kafka Streams/Connect
- Kafka to Pandas dataframe without Spark
- Kafka topic partition and Spark executor mapping
- Kafka topic partitions to Spark streaming
- Kafka vs StreamSets
- KafkaUtils class not found in Spark streaming
- Kubernetes executor do not parallelize sub DAGs execution in Airflow
- Lambda Architecture with Apache Spark
- Large data, workflows using pandas
- Large data workflows using pandas
- Large scale Machine Learning
- Lazy Method for Reading Big File in Python?
- Learning Kafka 0.8.2
- Limit kafka batch size when using Spark Structured Streaming
- Limit Kafka batches size when using Spark Streaming
- Limit on the number of topics in Kafka
- Loading a pyspark ML model in a non-Spark environment
- Loading data from RDBMS to Hadoop with multiple destinations
- Machine Learning Big Data
- make avro schema from a dataframe - spark - scala
- MapReduce alternatives
- MapReduce atomic renames
- Mapreduce with third party API
- Max number of tuple replays on Storm Kafka Spout
- Message routing in kafka
- Metadata requests in Kafka producer
- Moving data from Snowflake to Kafka

Course
Beginner
27 lessons
10 hours
System Design Fundamentals
Build a strong foundation in designing scalable, reliable distributed systems.
View the course