Exploratory Data Analysis (EDA) & Data Cleaning, PySpark, MySQL Database - IT Assignment Help

Download Solution Order New Solution
Assignment Task


1. Load the downloaded data into HDFS 
3. Create an internal table in Hive to store the data 
a. Create the table structure.

b. Load the data from HDFS into the Hive table 

4. Create an internal table in Hive with partitions 
a. Create a Partition Table in Hive using “work class” as the Partition Key 


4. Access the following two tables created as part of Problem 1 (HDFS and Hive) and perform the steps as mentioned below: a. Access Hive External Table with partition 
i. Query the table to get the number of adults based on income and gender 
ii. Query the table to get the number of adults based on income and work class 


b. Access Hive Internal Table with Partition.

i. Query the table to get the number of adults based on income and gender.
ii. Query the table to get the number of adults based on income and work class.


Make a note of the time taken for getting the result in comparison with the time taken to get results with Hive. 


5. Comment on the time taken for executing these commands using Spark as compared to the time taken for execution in Hive (Problem Statement 1). 

3. INCOME CLASSIFIER 


Problem Statement 1
Income Classifier is an application that will be used to classify individuals based on the annual income. An individual’s annual income may be influenced by various factors such as age, gender, occupation, education level, and so on. 
Write a program to build classification models using PySpark. Explore the possibility of classifying income based on an individual’s personal information. Perform the following steps to build and compare different classifiers
 

Load data from the staging table (Table created in Step 3) into this table 


5. Create an external table in Hive to hold the same data stored in HDFS 

6. Create an external table in Hive with partitions using “work class” as Partition Key 

7. For each of the four tables created above, perform the following operations 

• Find out the number of adults based on income and gender. Note the time taken for getting the result 
• Find out the number of adults based on income and work class. Note the time taken for getting the result 
• Write your observations by comparing the time taken for executing the commands between: 
a. Internal & External Tables 
b. Partitioned & Non-partitioned Tables 

8. Delete the internal as well as external tables. Comment on the effect on data and metadata after the deletion is performed for both internal and external tables. 

2. DATA INGESTION 
Problem Statement 2 
In a similar scenario as above, the data is available in a MySQL database. Due to the inefficiency of RDBMS systems to store and analyze Big Data, it is recommended that we move the data to the Hadoop Ecosystem. 
Ingest the data from MySQL database into Hive using Sqoop. A data pipeline needs to be created to ingest data from an RDBMS into Hadoop Cluster and then load data into Hive. To make the analysis faster, use Spark on top of Hive after getting data into the Hadoop cluster. Using Spark, query different tables from Hive to analyze the dataset. 

Steps to be performed: 
1. Create the necessary structure in a MySQL database using the steps mentioned below: 
a. Create a new database in MySQL with the name mid-project. 
b. Create a table in this database with the name census_adult to store the input dataset. 
c. Load the dataset into the tabled. Verify whether data is loaded properly.

e. Verify the table for unwanted data such as ‘?’,’ Nan’ and ‘Null’.
f. Get the counts for the columns which contain unwanted data.

g. Clean the data by replacing the unwanted data with others 


2. Import the above data from MySQL into a Hive table using Sqoop 

3. Connect to PySpark using a web console to access the created Hive table. Perform the following queries and note the time taken for execution in each of the queries. 
a. Query the table to get the number of adults based on income and gender 
b. Query the table to get the number of adults based on income and workplace.

 


4. Access the following two tables created as part of Problem 1 (HDFS and Hive) and perform the steps as mentioned below: a. Access Hive External Table with partition 
i. Query the table to get the number of adults based on income and gender 
ii. Query the table to get the number of adults based on income and workplace 


b. Access Hive Internal Table with Partition i. Query the table to get the number of adults based on income and gender 
ii. Query the table to get the number of adults based on income and workplace 


Make a note of the time taken for getting the result in comparison with the time taken to get results with Hive. 


5. Comment on the time taken for executing these commands using Spark as compared to the time taken for execution in Hive (Problem Statement 1). 

3. INCOME CLASSIFIER 


Problem Statement 3 
Income Classifier is an application that will be used to classify individuals based on the annual income. An individual’s annual income may be influenced by various factors such as age, gender, occupation, education level, and so on. 
Write a program to build classification models using PySpark. Explore the possibility of classifying income based on an individual’s personal information. Perform the following steps to build and compare different classifiers
 

 

 

This IT Assignment has been solved by our IT Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.

Be it a used or new solution, the quality of the work submitted by our assignment Experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.