Highlights
Task:
This assignment is a practical data analytics project that follows on from the data exploration you did in Assignment 2.
You will be acting as a data scientist at a consultant company and you need to make a prediction on a dataset. The dataset can be found below.
You need to build classifiers using the techniques covered in the lectures to predictthe class attribute. At the very minimum, you need to produce a classifier for each method we have covered. However, if you explore the problem very thoroughly (as you should do in industry), preprocessing the data, looking at different methods, choosing their best parameters settings and identifying the best classifier in a principled and explainable way, then you should be able to get a better mark. If you choose to use KNIME and you show 'expert' use (i.e. exploring multiple classifiers, with different settings, choosing the best in a principled way and being able to explain why you built the model the way you did), this will attract a better mark. If you choose to use R or Python to build, optimise and test different models, this will also attract better marks.
Kaggle
Competition
For this assignment you will use the Kaggle website (kaggle.com) to submit your assignment solution. The report itself will be submitted through Canvas as for the other assignments. Go to this link to sign up to the competition on Kaggle: https://www.kaggle.com/t/fa3d108d56cb44b38c81b44e27f89e49 (Links to an external site.). You need to use the link to access the project, because it is a private project for students in 31250 and 32130 only. Sharing the competition with anyone not relevant to the subject is strictly prohibited. To submit to Kaggle you will need to make a Kaggle login using your UTS email address, and set your display name (in My Profile -> Edit Profile -> Display Name) as UTS_32130_xxxxxxxx where xxxxxxxx is your student ID . Submissions will not be considered if they don't meet these criteria.
Classification task
Build a classifier that classifies the "Salary" attribute - with 0 if the salary is predicted to be <=50K, and 1 for >50K. The Kaggle competition will be assessed using area under the ROC curve (AUC). Therefore, the output of your classifier needs to be a probability, i.e. a number between 0 and 1, representing the confidence that Salary is >50K for each instance.
You can do different data pre-processing and transformations (e.g. grouping values of attributes, converting them to binary, etc.), providing explanations for why you have chosen to do that. You may need to spit the training set into training, validation and test sets to accurately set the parameters and evaluate the quality of the classifier.
This IT Assignment has been solved by our IT Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.