COSC2670 - Practical Data Science with Python - Computer Science Assignment Help

Download Solution Order New Solution
Assignment Task

 

Objective The key objectives of this assignment are to learn how to process messy text data using python. Data is always going to start in a form you may not want. So, you must massage the data into a form you do want. Python is a perfect tool for processing large volumes of raw data. Australians love sports, so it seems fitting to begin our data adventure by processing a bit of statistical data from Australian Rules Football. 

 

The Data The data you will be processing is data that has been crawled from the web to produce a set of markdown files. The data is raw statistical information for Australian Rules Football. You should study several of the input files in a text editor to get a feel for what you will have to parse. You will quickly start to recognize clear patterns in the line structure that you will exploit to quickly process thousands of lines of raw data. The two main file types are team-based statistics. The data is not clean, but it is well-formed. So this makes it very amenable to python data wrangling to extract and aggregate all of the data into a more usable form – a dataframe. pandas is a popular Python package that is used regularly in data science. So you will find that dataframes are a very useful way to organize and process columns of data similar to a database table – without all the overhead of a full rbdms system. However, as many of you learning python still, we will not use a dataframe in this project. It would be relatively easy to convert the data array we are using in this project into pandas as you will have done all the hard work of cleaning the data, but we will save that challenge for another day.

 

Program output When your program is combined with the datasets supplied, your program should produce the output as specified in the code template. The primary output format is actually a tsv file which is similar to a csv file, but uses tabs instead of commas to separate fields. When working with long strings of text, you are more likely to avoid collisions as tabs can easily be removed from a text file by replacing them with spaces, but not commas. Once you have tsv files (or even csv files), it is easy to serialize them out to a file for storage and reload them when you need the data again. You can also easily load some and not all of the data if there is a lot to process too. Do not change 2 the output functions provided, or you will fail to pass the automated harness tests. All you need to do is to implement the functions in the skeleton code. Once you get each of these functions to work, it will output the answers automatically. Each function is worth a subset of the the 30 possible points you can get on the project.

 

 


This (COSC2670) Computer Science Assignment has been solved by our Computer Science experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinctio.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.