Highlights
1. Chapter 1
a. (2 pts.) Jane developed two different regression models. She calculated the MSE of these models on the same data that she used to fit the models. She, then,
recommended her supervisor to adopt Model
#1 because it has a lower MSE than Model
#2. Is her recommendation correct? Why?
b. (2 pts.) At a large regional hospital, the Data Analytics department is working on a project to create a model that predicts the length-of-stay for each patient at the time of admission to the hospital. The team collects data on 15000 inpatients that were discharged within the last 12 months. The data includes 15 demographics variables, 30 vital measurements, length-of-stay, and 25 other medical treatment plan variables. An accurate length-of-stay prediction will allow the hospital to better manage the staffing and bed utilization. In addition, the hospital hopes to come up with a strategy to shorten the average length-of-stay as it can potentially reduce the risk of healthcare-acquired infection.
i. Is this project scenario a classification or regression problem?
ii. Are they interested in inference or prediction or both?
iii. Provide and. (Do not include the response variable in .)
c. (2 pts.) Rank these four classification models from the most flexible model to the least flexible model: KNN with K=1, KNN with K=2, LDA, QDA. Which one among these four models has the most bias with respect to the bias-variance tradeoff concept?
d. (2 pts.) In a certain classification problem, the number of predictors ?? is extremely large, and the number of observations ?? is small. Do you think LDA or QDA would perform better? Why?
e. (2 pts.) What is the purpose of vif() function in-car package? What value of vif is bad? What do you do with the variables that have a bad value of vif?
2. Chapter 2
This dataset features the salaries of 612 NHL players for the 2016/2017 season. We use some predictors to build a linear regression model to predict NHL player's salaries.
a. (2 pts.) Write the model in an equation form where is salary, is handedness dummy variables for right handedness, and and for the other variables. Please substitute all parameters with the values from the above code.
b. (2 pts.) Provide an interpretation of the coefficient of “GP” in the model in the context of this problem.
c. (2 pts.) Provide an interpretation of R-squared (.6048) in the context of this problem.
d. (2 pts.) Are all the five predictors significant? Explain.
e. (2 pts.) Estimate the salary of a left-handed hockey player who was drafted in 2014, 2 nd round, played 80 games and, while this player was on the ice, the team has 600 shots on goal. Please show your calculation.
3. Chapter 3.
We used R to fit a logistic regression model to predict the probability of default from balance" and student status (Yes, No).
a. (2 pts.) Write the model in an equation form. Please substitute all parameters with the values from the above code.
b. (2 pts.) Calculate the probability of default for a student who has a credit card balance of $2200. Please show your calculation.
c. (2 pts.) Suppose that an individual has a 25% chance of defaulting on her credit card payment. What are the odds that she will default? Please show your calculation.
d. (2 pts.) Calculate the error rate based on the given confusion table in the last part of the above code. Please show your calculation.
e. (2 pts.) If the accuracy is not as important as the sensitivity, how can you improve the sensitivity of this model?
4. Chapter 4
A team used the Ping Man robot to test golf balls. At one point, there was a mixed up that caused the team to end up with 20 unmarked balls in a basket. They knew that 7 balls were Snell MTB-X, 3 balls were Volvik S4, and 10 balls were Callaway Soft X, but they could not tell from looking at the balls. The team randomly got one of the balls and fed it to the robot. The ball went 284 yards. The team wants to know the probability that the ball is Volvik S4.
a. (2 pts.) If we want to use an LDA model to perform a prediction, what is the assumption that we need to satisfy in terms of the variance of the predictor?
b. (3 pts.) Assume that all necessary assumptions for LDA are satisfied and follows a normal distribution. Please fill in the blanks for these parameters for LDA:
c. (2 pts.) Write the mathematical expression (in terms of and )) for the probability that a
golf ball that goes yards is Volvik S4. That is, _______.
d. (3 pts.) Predict the probability that a golf ball that goes 284 yards is Volvik S4. Please show your calculation.
5. Chapter 5
a. (3pt.) Two engineers were independently testing a cubic polynomial regression model on the same dataset. Both of them used leave-one-out cross-validation. They had almost identical R code, except that one of the engineers had set.seed(3) on the first line of his R code, while the other one had set.seed(7). Would these engineers end up with different results? Why?
b. (3pt.) If one of the two engineers in Problem 1a used the validation set approach, while the other one used 10-fold cross-validation. Both of them performed the test multiple times with different seed numbers. What would you expect to see in terms of the mean square errors? Why?
c. (4pt., three points on the code and one point on the answer.) In this question, you will analyze the Default data set. Perform logistics regression to
predict default using 2 models.
default~balance+student
default~balance*student
Use 10-fold cross-validation to evaluate the model. Also check whether the predictors are significant. Which model is the best? Explain.
6. Chapter 6:
a. (2pt.) Should we use the bootstrap to estimate prediction error, instead of k-fold cross-validation? If yes, how? If no, why not?
b. (2pt.) A certain dataset contains 20 qualitative variables, each has more than 10 classes. In performing subset selection for a regression problem, would you use best subset selection or forward stepwise selection? Why?
c. (2pt.) In subset selection, what is the common purpose of BIC, Cp, and adjusted? Explain their similarities.
d. (2pts) A certain large dataset contains 50 independent variables. Can a LASSO model be used to reduce the number of predictors? If no, why? If yes, how?
e. (2pts) Why do we need cross-validation when working with a LASSO model? How does cross-validation is used in LASSO model?
7. Chapter 7
a. (4pt.) Suppose we fit a curve with basis functions and . We fit the linear regression model and obtain coefficient estimates .
b. (6pts) This question uses the variables dis (the weighted mean of distances to five Boston employment centers) and nox (nitrogen oxides concentration in parts per 10 million) from the Boston data (library MASS). We will treat dis as the predictor and nox as the response.
Use the ns() function to fit a natural cubic spline to predict nox using dis. Perform a 5-fold cross-validation in order to select the best degrees of freedom upto 10 degrees.
o Plot degrees of freedom and its cross-validation error.
o Plot the resulting fits of the best degrees of freedom.
Please use set.seed(1) on the first line of your R code.
This POMS.6120: Statistics Assignment has been solved by our Statistics Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.