Big Data questions and answers
Spark, Hadoop and batch processing at volume. Page 8 of 9.
System Design practice on Codemia
Work through 120+ system design problems with detailed solutions, from rate limiters to multi-region storage.
Answers 421-480
- Spark Structured Streaming Kafka Offset Management
- Spark Structured Streaming program that reads from non-empty Kafka topic (starting from earliest) triggers batches locally, but not on EMR cluster
- Spark Structured Streaming with Hbase integration
- Spark Structured Streaming with Kafka - How to repartition the data and distribute the processing among worker nodes
- Spark Structured Streaming with Kafka SASL/PLAIN authentication
- Spark Structured Streaming with secured Kafka throwing Not authorized to access group exception
- Spark submit to kubernetes packages not pulled by executors
- Spark unable to download kafka library
- Spark What is the time complexity of the connected components algorithm used in GraphX?
- Spark Word2vec vector mathematics
- Spark write Dataset in kafka, enable KryoSerializer
- Spark/k8s How to run spark submit on Kubernetes with client mode
- Split single DStream into multiple Hive tables
- stopping spark streaming after reading first batch of data
- Store TreeSet on Hadoop DistributedCache
- Store your events directly from kafka into database?, when or why using S3/HDFS before?
- Storing Avro schema in schema registry
- Strange delays in spark streaming
- Streaming data from Kafka into Cassandra in real time
- Streaming messages from one Kafka Cluster to another
- Structured Streaming - Foreach Sink
- structured streaming Kafka 2.1->Zeppelin 0.8->Spark 2.4 spark does not use jar
- Submit Spark Application on Kubernetes in Cluster mode Configured service account doesn't have access
- System Design of Google Trends?
- Technically what is the difference between s3n, s3a and s3?
- Tensorflow Dataset API with HDFS
- The benefits of Flink Kafka Stream over Spark Kafka Stream? And Kafka Stream over Flink?
- Tracking an expected set of Kafka events
- Unable to create spark session
- Unable to create SparkApplications on Kubernetes cluster using SparkKubernetesOperator from Airflow DAG
- Understand Kafka replication factor
- Understand Kafka write speed
- updating file in distributed cache in hadoop
- Use Distributed Cache - HIVE STREAMING
- Use hdfs as backend storage for kafka, is it doable?
- Use kafka to detect changes on values
- Use Kafka topics to store data for many years
- Use schema to convert AVRO messages with Spark to DataFrame
- Use schema to convert ConsumerRecord value to Dataframe in spark-kafka
- Using Apache Kafka for log aggregation
- using AWS Glue with Apache Avro on schema changes
- Using Kafka to import data to Hadoop
- Using Silhouette Clustering in Spark
- Using Spark Structured Streaming to Read Data From Kafka, Issue of Over-time is Always Occured
- What actions does job.commit perform in aws glue?
- What are the differences between airflow and Kubeflow pipeline?
- What causes unknown resolver null in Spark Kafka Connector?
- What do columns ‘rawPrediction’ and ‘probability’ of DataFrame mean in Spark MLlib?
- What do foreachBatches contain in a streaming query from multiple Kafka topics?
- What do you use Apache Kafka for?
- What environment do I need for Testing Big Data Frameworks?
- What is most efficient way to write from kafka to hdfs with files partitioning into dates
- What is rank in ALS machine Learning Algorithm in Apache Spark Mllib
- What is the biggest Couchbase cluster nodes number?
- What is the difference between Big Data and Data Mining?
- What is the difference between DistributedCache.getCacheFiles() and DistributedCache.getLocalCacheFiles()
- What is the differences between Apache Spark and Apache Apex?
- What is the optimal way to read from multiple Kafka topics and write to different sinks using Spark Structured Streaming?
- what is this topic __consumer_offsets in Kafka
- What is transformation_ctx used for in aws glue?

Course
Beginner
27 lessons
10 hours
System Design Fundamentals
Build a strong foundation in designing scalable, reliable distributed systems.
View the course