Highlights
Question A
1. Please describe the concept of elasticity in the context of cloud computing. Please use examples.
2. Please describe an example scenario where it would be good to run a Hadoop cluster on the cloud. Also describe another example scenario where it would be good to run a Hadoop cluster on premise.
3. Please explain why MapReduce does schema-on-read instead of schema-on-write.
4. Please explain how to group data by a certain attribute using the MapReduce framework.
5. What are the advantages of using Hive over traditional relational databases (RDBMS)? Please give an example application where it is better to use Hive rather than RDBMS. Explain why in this case it is better to use Hive rather than RDBMS.
6. What are the main reasons for introducing YARN?
7. Please describe the function of the application master in the YARN framework.
8. Please describe the advantages of using the lambda architecture of storm to process large data streams
9. Please describe the different types of big data problems that Apache Spark can be used to solve.
10. Please describe what the map and the reduce functions in Apache Spark can be used for.
11. Name two advantages of using SparkSQL datasets over using RDDs.
12. What are the benefits of using Apache Spark Streaming’s structured API?
13. Why would a company that used to only support disaster recovery now consider supporting high availability when moving their infrastructure onto the cloud? In your answer you should include an explanation of the difference between disaster recovery and high availability.
14. Please describe the advantages of using the following AWS services:
a. Relational Database Service (RDS)
b. Elastic Load Balancer
c. Cloud Formation
Question B
Answer the sub-questions of question B using the data described below

You will answer the questions either using the Spark RDD API or Spark SQL dataframes/datasets. Please complete the question using spark RDDs for the questions marked as [Spark RDD], please complete the questions using dataframes/datasets for questions marked as [Spark SQL]. When you program Spark SQL you can use the dataset/dataframe operations or use SQL syntax.
E.g. dataset/dataframe operations syntax:
df.filter($"age" > 21).show()
df.select($"ID", $"age" + 1).show()
E.g. SQL syntax:
val sqlDF = spark.sql("SELECT * FROM animals")
sqlDF.show()
Write spark code to do the following. Your solution can consist of one or more lines of Spark code. You do not need to make the output format look good. Marks will be awarded for more efficient code. For example, code that results in less data shuffles.
This IT Computer Science Assignment has been solved by our IT Computer Science Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing Style. Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered.
You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turn tin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.