CSE5BDC : Gain in Depth Experience Playing Around with Big Data Tools - IT/Computer Science Assignment Help

Download Solution Order New Solution
Assignment Task:

Task:

Objectives

1. Gain in depth experience playing around with big data tools (Hive, SparkRDDs, and Spark SQL).

2. Solve challenging big data processing tasks by finding highly efficient solutions.

3. Experience processing three different types of real data a. Standard multi-attribute data (Bank data) b. Time series data (Twitter feed data) c. Bag of words data.

4. Practice using programming APIs to find the best API calls to solve your problem. Here are the API descriptions for Hive, Spark (especially spark look under RDD. There are a lot of really useful API calls). a) [Hive] https://cwiki.apache.org/confluence/display/Hive/LanguageManual

b) [Spark] http://spark.apache.org/docs/latest/api/scala/index.html#package

c) [Spark SQL] https://spark.apache.org/docs/latest/sql-programming-guide.html https://spark.apache.org/docs/latest/api/scala/index.html#org.apache.spark.sql.Datase t https://spark.apache.org/docs/latest/api/sql/index.html - If you are not sure what a spark API call does, try to write a small example and try it in the spark shell

Assignment structure:

• A script which puts all of the data files into HDFS automatically is provided for you. Whenever you start the docker container again you will need to run the following script to upload the data to HDFS again, since HDFS state is not maintained across docker runs: $ bash put_data_in_hdfs.sh The script will output the names of all of the data files it copies into HDFS. If you do not run this script, solutions to the Spark questions will not work since they load data from HDFS.

• For each Hive question a skeleton .hql file is provided for you to write your solution in. You can run these just like you did in labs: $ hive -f Task_XX.hql

• For each Spark question, a skeleton project is provided for you. Write your solution in the .scala file in the src directory. Build and run your Spark code using the provided scripts: $ bash build_and_run.sh

Tips:

1. Look at the data files before you begin each task. Try to understand what you are dealing with!

2. For each subtask we provide small example input and the corresponding output in the assignment specifications below. These small versions of the files are also supplied with the assignment (they have “-small” in the name). It’s a good idea to get your solution working on the small inputs first before moving on to the full files.

3. In addition to testing the correctness of your code using the very small example input. You should also use the large input files that we provide to test the scalability of your solutions.

4. It can take some time to build and run Spark applications from .scala files. So for the Spark questions it’s best to experiment using spark-shell first to figure out a working solution, and then put your code into the .scala files afterwards.

This CSE5BDC IT Assignment has been solved by our IT Experts at onlineassignmentbank. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.

Be it a used or new solution, the quality of the work submitted by our assignment Experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.