Highlights
TASK:
Q1. For each of the following meetings, explain in detail which phase in the CRISP-DM process is represented: (2.5 x 5 = 12.5 marks)
In a company, the senior executives were enquiring whether the project deployment can take place next week. Therefore, business analysts, data analysts, and data scientists meet to discuss the usefulness and accuracy of the model.
The business analyst meets the data engineer to brainstorm how the data can be collected and explored further.
The business analyst meets the CEO to discuss the feasibility of data mining in customer relationship management.
The data mining project manager meets the data engineer and data scientist to discuss various process changes or other improvements that can be made for better results towards business decisions.
In the weekly meeting for a data mining project, business analysts, data analysts, and data scientists meet to confirm whether logistic regression or decision tree algorithm would be suitable for the given task.
Q2. Explain whether the task required in each of the below cases is supervised or unsupervised learning. If it is supervised, then explain whether it is a classification or prediction. (2.5 x 5 = 12.5 marks)
Detecting whether the blip on the radar is a flock of geese or an incoming nuclear missile.
The strategist of a political party wants to identify the best group(s) in a county to canvas for political donations.
A Wall Street analyst has to determine the movement of the stock prices for some companies based on the price/ earnings ratio.
A retailer has three types of discount coupons – Coupon A, Coupon B, and Coupon C. Based on prior purchase behaviour with respect to discount coupons, the retailer wants to distribute each coupon to the right customer so as to increase the sales.
Which service package (S1, S2, or none) will a customer likely purchase if given incentive I?
Q3a. Identify the nominal and ordinal variables in the below dataset. Justify your answer.
(4 marks)
College Name State Ranking Public or Private Number of applications received Number of applications accepted New student enrolled Percent of faculty with PhD Stud faculty ratio Graduation rate
A AK low Private 193 146 55 76 11.9 15
B AK high Public 1852 1427 928 67 10 C AK low Public 146 117 89 39 9.5 39
D AK medium Public 2065 1598 1162 48 13.7 E AL high Public 2817 1920 984 53 14.3 40
F AL high Private 345 320 179 52 32.8 55
G AL medium Public 1351 892 570 72 18.9 51
Q3b. Considering only two variables (Percent of faculty with PhD and Graduation rate) from the dataset in Q3a, we get the covariance matrix (the diagonal elements are the variances of the variables) as follows:
What is the percentage of variance explained by each of the two variables? How much variation will we lose if we drop the variable Percent of faculty with PhD to reduce the dimension from 2D to 1D? Clearly show the calculation steps (4 marks)
Q3c. We performed Principal Component Analysis (PCA) without normalization on the two variables from Q3b and got the following results:
Comment on how PCA has done a better job in data reduction (from 2D to 1D) than dropping any one of the original variables. (3 marks)
Calculate the value of the first and third observation (i.e., for College A and C) corresponding to PC1 (Given average Percent of faculty with PhD = 69.65 and average Graduation rate = 60.60). Clearly show the calculation steps. (5 marks)
Q4. Answer the following questions based on the below customer dataset.
Sl. No Fever Age Sex Headache Sore Throat Covid +ve1 Yes 24 Male Yes Yes Yes
2 No 21 Female No No No
3 No 28 Female No Yes No
4 No 25 Male Yes Yes Yes
Frame a supervised learning problem statement based on the above dataset. Which variables can be the input variables, and which can be the target variable? (3 + 3 = 6 marks)
Can we use linear regression for the data mining problem related to the above dataset? Justify your answer (1 + 2 = 3 marks)
This Statistics Assignment has been solved by our Statistics experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.