Machine Learning Assignment - IT Assignment Help

Download Solution Order New Solution
Assignment Task

 

1. Data visualization is an essential part of machine learning. In this question, you will use python libraries (such as matplotlib) to create different types of plots.
You have to include the legends in the plots to denote the class labels and other relevant information.

Datasets: COVID-19| Face mask dataset|
1. Find the two CSVs present in the COVID-19 folder and explore the files. Specifically, check the number of columns, their names, type of columns (categorical or
continuous), and the possible data values/range for each column. Illustrate these in your report.
2. Use the file covid_19_india.csv and show the trend of confirmed, and cured cases along with the number of deaths in India for the time period of the data in one plot.
3. Use the file covid_vaccine_statewise.csv and plot the total number of doses administered of Covishield, Covaxin, and Sputnik in the following cities: Kerala, Delhi, Rajasthan, Haryana, Uttar Pradesh, Tamil Nadu for the time period of the data.
4. Use the Face mask dataset for the following tasks:
(a) Randomly select 5 images from each class and visualise them as images.
(b) We can visualize only 2D or 3D data using scatter plots. For features dimensions higher than three, we may use T-distributed Stochastic Neighbor Embedding (t-SNE) to reduce the number of features. Use the t-SNE to reduce the dataset to 2 dimensions and visualize the scatter plot. What is your inference regarding the class separation?

2. For this question, you can use the decision tree classifier from sklearn.
Dataset: data (Use bank-additional-full.csv)
Target variable: As mentioned in the dataset description
Use the first 70% of the samples for training and the remaining 30% for testing. Implement the function to split the dataset and calculate accuracy. You cannot use any
inbuilt version. Though, you can use Numpy, Pandas, Random, etc.
1. Take the Complexity parameter as a hyperparameter, and perform a grid search for finding its optimal value. You have to perform grid search for at least 10 values
of the Complexity parameter. You have to implement calculating sum of impurity value of all leaf nodes for each complexity parameter value.
Plot a curve between Complexity parameter and testing accuracy. Plot a curve between Complexity parameter and sum of impurity of leaf nodes. Comment on the effect of the Complexity parameter on the performance of the classifier. You have to implement grid search and cannot use any inbuilt implementation for the algorithm.

 

This IT Assignment has been solved by our IT experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.