ICT707 - Data Science Practice - Data Analysis - Machine Learning - Computer Science Assignment Help

Download Solution Order New Solution
Assignment Details:

Task 1
In this component, we need to utilize Python 3 and PySpark to complete the following data analysis tasks:

  1. Exploratory data analysis
  2. Recommendation engine
  3. Classification
  4. Clustering

You need to choose a dataset to complete these tasks. Explain the data set you’ve chosen, including its source URL. Demonstrate your exploratory data analysis in this section. 

Task I.1: Exploratory Data Analysis
This subtask requires you to explore your dataset by

  • Telling its number of rows and columns,
  • Doing the data cleaning (missing values or duplicated records) if necessary
  • Selecting 3 columns, and drawing 1 plot for each to summarise it

Task I.2: Recommendation engine
This subtask requires you to implement a recommender system on Collaborative filtering with Alternative Least Squares Algorithm. You need to include

  • Model training and predictions
  • Model evaluation using MSE

Task I.3: Classification
This subtask requires you to implement a classification system with Logistic regression with LBFGS class. You need to include

  • Logistic Regression model training
  • Model evaluation

Task I.4: Clustering
This subtask requires you to implement a clustering system with K-means. You need to include

  • Model training
  • Model evaluation

Task 2

You are required to write a report to explain your design and implementation of the machine learning parts in your code, including the following topics:

  • Introduction/summary/explanation to the ML algorithm/concepts.
  • The learning settings, such as how to prepare training/testing set, what are the key parameters and how you set them up.
  • Comments/evaluation for the models learned
     

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.