INFO411/911 - Data Mining and Knowledge Discovery Assignment

Download Solution Order New Solution

Assignment Task

Overview

This assignment consists of two tasks. There are several questions to be answered for each of the two tasks. You may need to do some research on background information for this assignment. For example, you may need to develop a deeper understanding of writing code in R, or study the general characteristics of GPS, obtain general geographic information about Rome, and study other topics that are related to the tasks in this assignment.

Task

1. Preface:

The analysis of results from urban mobility simulations can provide very valuable information for the identification and addressing of problems in an urban road network. Public transport vehicles such as busses and taxis are often equipped with GPS location devices and the location data is submitted to a central server for analysis. The metropolitan city of Rome, Italy collected location data from 320 taxi drivers that work in the center of Rome. Data was collected during the period from 01/Feb/2014 until 02/March/2014. An extract of the dataset is found in taxi.csv.

The dataset contains 4 attributes:

1. ID of a taxi driver. This is a unique numeric ID.

2. Date and time in the format Y:m:d H:m:s.msec+tz, where msec is micro-seconds, and tz is a timezone adjustment. (You may have to change the format of the date into one that R can understand).

3. Latitude

4. Longitude

Purpose of this task

Perform a general analysis of this dataset. Learn to work with large datasets. Obtain general information of the behaviour of some taxi drivers. Analyse and interpret results. This task also serves as a preparation for projects that will be based on this dataset.

Questions

1. By using the data in taxi.csv perform the following tasks:

(a) Plot the location points (2D plot using all of the latitude,longitude value pairs in the dataset). Clearly define and justify what you consider is invalid, noise, and outlier. Clearly indicate the points that are invalid, outliers or noise points in your plot. The plot should be informative! Clearly explain the rationale that you used when identifying invalid points, noise points, and outliers. Remove invalid points, outliers and noise points before answering the subsequent questions.

(b) Compute the minimum, maximum, and mean location values.

(c) Obtain the most active, least active, and average activity of the taxi drivers (most time driven, least time driven, and mean time driven) . Explain the rationale of your approach and explain your results.

(d) Look at the file Student_Taxi_Mapping.txt. The file contains two columns. The first column is a 4- digit code, the 2nd column is the ID of a taxi driver. Use the first and last three digits of your student number to optain a 4-digit code. Locate that code in the first column of the file Student_Taxi_Mapping.txt then use the corresponding ID of the taxi driver listed in column 2. Thus, for example, if your student number is 7671167 then you would look up 7167 in file Student_Taxi_Mapping.txt to find that the corresponding taxi ID is 12. Use the taxi ID that is listed next to your 4-digit student code to answer the following questions:

i. Plot the location points for taxi=ID

ii. Compare the mean, min, and max location value of taxi=ID with the global mean, min, and max.

iii. Compare total time driven by taxi=ID with the global mean, min, and max values.

iv. Compute the distance traveled by taxi=ID. To compute the distance between two points on the surface of the earth use the following method

2. Preface

Banks are often posed with a problem to whether or nor a client is credit worthy. Banks commonly employ data mining techniques to classify a customer into risk categories such as category A (highest rating) or category C (lowest rating). A bank collects data from past credit assessments. The file "creditworthiness.csv" contains 2500 records. 1962 of these records have been assessed for credit worthyness. Each assessment lists 46 attributes of a customer. The last attribute (the 47-th attribute) is the result of the assessment. Open the file and study its contents. You will notice that the columns are coded by numeric values. The meaning of these values is defined in the file "definitions.txt". For example, a value 3 in the 47-th column means that the customer credit worthiness is rated "C". Any value of attributes not listed in definitions.txt is "as is". This poses a "prediction" problem. A machine is to learn from the outcomes of past assessments and, once the machine has been trained, to assess any customer who has not yet been assessed. For example, the value 0 in column 47 indicates that this customer has not yet been assessed.

Purpose of this task

You are to start with an analysis of the general properties of this dataset by using suitable visualization and clustering techniques (i.e. Such as those introduced during the lectures), and you are to obtain an insight into the degree of difficulty of this prediction task. Then you are to design and deploy an appropriate supervised prediction model (i.e. MLP) to obtain a prediction of customer ratings.

Analyze the general properties of the dataset and obtain an insight into the difficulty of the prediction task. Create a statistical analysis of the attributes and their values, then list 5 of the most interesting (most valuable) attributes. Explain the reasons that make these attributes interesting. Use the Self-Organizing Map method to obtain further insights of value into the properties of the dataset and to obtain an insight into the difficulty of the supervised learning problem (i.e. from the results that you obtained from the SOM, can it be expected that a prediction model will be able to achieve a 100% prediction accuracy?).

Explain how these insights affect your choice and design of the supervised classification method. (i.e. from the results that you obtained, can it be expected that a prediction model will be able to achieve a 100% prediction accuracy?). Always explain your answers.

2. Deploy a suitable prediction model based on MLP to predict the credit worthiness of customers which have not yet been assessed. The prediction capabilities of the MLP in lab4 was poor. Your task is to:

a.) Describe a valid strategy that maximises the accuracy of predicting the credit rating. Explain why your strategy can be expected to maximize the prediction capabilities.

b.) Use your strategy to train MLP(s) then report your results. Give an interpretation of your results. What is the best classification accuracy (expressed in % of correctly classified data) that you can obtain for data that were not used during training (i.e. the test set)?

c.) You will find that 100?curacy cannot be obtained on the test data. Explain reasons to why a 100?curacy could not be obtained on this test dataset. What would be needed to get the prediction accuracy closer to 100%?

d.) To be answered by INFO911 students only: Deploy another prediction model (other than MLP) to obtain a 2nd set of results. Analyse and compare the two sets of results. Explain the strength and weaknesses of each of the two prediction models.
 

This INFO411/911Engineering has been solved by our PHD Experts at My Uni Paper.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.