Highlights
Task:
PROJECT TITLE 1.
Introduction
- A brief narrative as to how your organisation currently performs data analytics & how the CRISP-DM methodology may help your organisation to better meet its strategic priorities with respect to data analytics and business intelligence.
- An overview of the data analytic task you are going to do.
- A clear justification as to why the task you are attempting is of value to your business/ community etc. Please justify with references to the literature. - State the insight you intend to gain
2. Data Selection and Pre-processing
- Select a data set consisting of at least 2,000 observations/records and preferably above 10,000. Identify an anonymised data set, the strategic objectives of the business. Please select a data set from one of the following sources: https://archive.ics.uci.edu/ml/datasets.html http://www.cs.waikato.ac.nz/ml/weka/datasets.html
- Describe your data set and reference its origin.
- If you have 15 or less attributes, table your attributes with attribute name, description and data type and then show the minimum/average/maximum and stdev values for the training set and test set. For nominal variables, then show the most and least frequently occurring nominal value(s). If you have more than 15 attributes, then group attributes into themes (e.g. customer, orders, employees) and describe the type of information and data types in each theme including number of each variable type (e.g. nominal, interval, ratio etc). You may want to highlight significant variables identified by some attribute selection algorithm.
- Briefly table the following characteristics of the entire data set: number of instances, patterns per target class (if classification), limitations such as possible conflicting patterns, missing values, outliers/erroneous values.
- Explain how you have sampled your data to create the ‘in sample’ and ‘out of sample’ data sets. If you have used instance weightings to balance your data set(s) then explain how the weightings were determined.
2 - Provide a statistical summary in tabular form for the resulting ‘in sample’ (training/validation set) and ‘out of sample’ (test set). State whether or not there was any overlap in training and test set instances and if so, justify why your test set is not compromised.
- What pre-processing and transformation was performed on the variables and why? (e.g. standardising numerical variables and/or using scaling, taking logs to reduce skewness, or log differences to reduce non-stationarity; converting numerical variables to discrete ones; converting numerical or symbolic patterns into bit patterns; removing patterns with missing or outlier values; adding noise or jitter to patterns to expand the data set; adding instance weightings or replicating certain pattern classes to improve class distributions; transforming time-series data into static training/test patterns)
- How did you ensure that your pre-processing did not compromise your test set (e.g. use of standardisation)
- Consideration will also be given to the ‘curse of dimensionality’, its issues and how its impact can be reduced.
The above IT Assignment has been solved by our IT Assignment Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.