Highlights
Related learning outcomes:
Overview
COVID-19 (previously known as "2019 novel coronavirus") is dominating the news. As you are well aware, this novel virus emerged a few months ago, and is spreading around the world.
While it is a major public health emergency, it also provides us with an excellent learning opportunity. Scientists and researchers across the world are collecting and analysing data on the virus. Most of it is publicly available.
In this assignment, you will work with a portion of that data. You will use this partial dataset to demonstrate your skills on a number of tasks, detailed below.
The data was made available by Johns Hopkins University, strictly for educational and academic research purposes. They rely on publicly available data from multiple sources, that do not always agree. This is worth keeping in mind in your analysis.
You can download the data from Blackboard, in the Assignment Item 1 folder, under Assessment.
You will get information on all reported cases up until March 8th. The data is organised in daily CSV files, from January 22nd to March 8th.
To ensure that you all work with the same dataset, it is important that you download it from Blackboard, rather than trying to access the data directly from Johns Hopkins. After the assignment is completed, you will still be able to extend your analysis by accessing new data if you want to. The data on Blackboard is simply a snapshot.
This is the first time we are experiencing COVID-19, and obviously the first time we are running this assignment. While the intent and objectives of the assessment item will not change, details may evolve as new information comes to light. Any changes to the details will be communicated via Blackboard and Slack. This may be a release of additional data, or some tweaks to specific tasks.
Tasks
1. Data manipulation and cleansing
You have two options to access the data.
In the first option, the dataset is structured in a way that makes updates very convenient: every day, a new file is simply added to the repository. This means working with multiple separate CSV files. This is not suitable for an efficient analysis.
In the second option, the dataset is structured around three separate files (confirmed cases, deaths, recovered patients), but the number of columns increases with each update of the dataset.
You need to select one option, and then represent your data in memory in a way that will make your analysis possible. In doing so, you need to keep in mind that:
There may be inconsistencies in the formats used, or errors in the data. How are you handling that? It is good practice to separate code and data. Your solution has to work for these files, but should also be usable if you had more files (in option 1) or more columns (in option 2). This is especially important as we may end up using an expanded dataset, as explained earlier.
This Engineering Assignment has been solved by our Engineering Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.