Big Data questions and answers
Spark, Hadoop and batch processing at volume. Page 1 of 9.
System Design practice on Codemia
Work through 120+ system design problems with detailed solutions, from rate limiters to multi-region storage.
Answers 1-60
- Accessing a File from Distributed Cache in Pig UDF Java class, Amazon EMR
- Accessing data on distributed database on OrientDB
- Accessing Distributed Cache in Pig StoreFunc
- Accessing file in Pig through Distributed Cache
- Accessing Hadoop Distributed Cache in UDF
- Accessing Maxmind Geo API in Hadoop using Distributed Cache
- Add JAR files to a Spark job - spark-submit
- Adding external file for use in MapReduce driver Class
- Airflow/k8s How do I correctly set permissions for DAGs stored in a persistent volume?
- Algorithm for counting common group memberships with big data
- Algorithm for detecting duplicates in a dataset which is too large to be completely loaded into memory
- Alpakka kafka vs Kafka streams
- alternative solutions for Hadoop/Hive distributed-cache for handling very large dictionary file?
- apache- kafka with 100 millions of topics
- Apache Beam over Apache Kafka Stream processing
- Apache Flink connect versus join
- Apache Flink vs Apache Spark as platforms for large-scale machine learning?
- Apache Kafka for Time Series Data Persistence
- Apache Kafka KRaft - Kafka Storage Tool
- Apache Kafka Mirroring vs. Replication
- Apache Spark-Kafka.TaskCompletionListenerException & KafkaRDD$KafkaRDDIterator.close NPE on local cluster(Client Mode)
- Apache Spark + Delta Lake concepts
- Apache Spark ALS recommendations approach
- Apache Spark Getting a InstanceAlreadyExistsException when running the Kafka producer
- Apache Spark MLLib for real time analytics
- apache spark MLLib how to build labeled points for string features?
- apache spark streaming - kafka - reading older messages
- Are all distributed database designed to process data in parallel?
- Are directories handled by Hadoop cache symlinks?
- Avro schema versioning
- AWS DynamoDB and MapReduce in Java
- Best of breed indexing data structures for Extremely Large time-series
- Best practice for integrating Kafka and HBase
- Best way to join two (or more) kafka topics in KSQL emiting changes from all topics?
- BrokerNotAvailableError Could not find the leader Exception while Spark Streaming
- Bucket records based on time(kafka-hdfs-connector)
- Build stateful chain for different events and assign global ID in spark
- Can a model be created on Spark batch and use it in Spark streaming?
- Can I create an RDD from a kafka topic if I do not know the until offset?
- Can I extract fp-tree (any format) in spark?
- Can Kafka streams deal with joining streams efficiently?
- Cannot launch SparkPi example on Kubernetes Spark 2.4.0
- Cannot process data using Spark Continuous Streaming
- Capture Kubernetes Spark driver and executor logs in S3 and view in History Server
- Caret train rf model - how long it takes to execute big data?
- Cassandra + kafka for event sourcing
- Cassandra selective copy
- Changing replication of existing files in HDFS
- Clarification of use-cases for Hadoop versus RabbitMQ+Celery
- Classify data using Apache Mahout
- Cloud solution to parse and process 1M+ rows
- clustering very large dataset in R
- Combine results from batch RDD with streaming RDD in Apache Spark
- Conditional routing in Apache NiFi
- Configure hadoop/hbase in fully-distributed mode
- Configuring external data source for Elastic MapReduce
- Confusion about distributed cache in Hadoop
- Connecting Pyspark with Kafka
- Copy Files from NFS or Local FS to HDFS
- Count number of columns in pyspark Dataframe?

Course
Beginner
27 lessons
10 hours
System Design Fundamentals
Build a strong foundation in designing scalable, reliable distributed systems.
View the course