COSC2670 - Practical Data Science With Python Assignment

Download Solution Order New Solution

Assignment Task

1. Introduction

In this assignment, you will examine a data file in CSV and carry out the first steps of the data science process, including the cleaning and exploring of data. You will need to develop and implement appropriate steps, in Jupyter Notebook (available in Anaconda), to load a data file into memory, clean, process, and analyse it. In this assignment, you may need to use Python packages/libraries such as Pandas, NumPy, Matplotlib, and/or Seaborn. This assignment is intended to give you practical experience with the typical first steps of the data science process.

Task 1: Data Preparation

This dataset contains data about the digital connectivity information of children in a school attendance age that have internet connection at home. Check the file Readme-A1data.txt for details about the dataset.

Your task is to prepare the provided data for analysis. You will start by loading the CSV data from the file (using appropriate Pandas functions) and then clean the data (using Python, not manually).

Note that there are at least four types of errors/issues of the data. You can presume the first 4 columns (ISO3, Countries and areas, Region, Sub-region) don’t contain any error.

Task 2: Data Exploration

Use the cleaned data YourStudentNumber-cleaned-A1data.csv (which you obtained in the above Task) and complete the following subtasks.

2.1. Consider the overall/total percentage of children in a school attendance age that have internet connection at home (i.e., column Total). Create (and display) the side-by-side boxplot (as one graph/chart) having the data separated/grouped by Region. Compute the Median (of the total percentage) for each Region.

2.2. Compute the Mean (of the percentage of school-age children who have an internet connection at home) for the Wealth quintile (Poorest) and Wealth quintile (Richest), respectively. Display/list the top 10 countries with the highest percentages for Wealth quintile (Poorest) and Wealth quintile (Richest), respectively.

2.3. Consider the data about children that are from the Lower middle income (LM) group. Compare the percentages of different categories of Residence (Rural versus Urban), using at least three statistics measures (of your choice).

Task 3: Written Report

The report should comprise the following sections:

  • Data Preparation: Provide a concise explanation of how you addressed Task 1. Create a sub-section for each type of errors. Explain and justify the approach taken to address each kind of errors.
  • Data Exploration: For each subtask in Task 2, create a sub-section with corresponding numbering; explain and justify how you explored the data as required, and summarise the results (i.e., findings).

This Data Science has been solved by our PhD Experts at My Uni Paper. Our Assignment Writing Experts are efficient in providing a fresh solution to this question. We are serving more than 10000+ Students in Australia, the UK, and the US by helping them to score HD in their academics. Our Experts are well-trained to follow all marking rubrics and referencing styles.

Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.