Descriptive Statistics and the Normal Distribution Assignment

Download Solution Order New Solution

Assignment Task

Questions

1. Get to know your scientific question

(a) Identify the variable of interest.

(b) Identify the population(s) and sample(s).

(c) Identify the parameter(s) and statistic(s).

(d) What is the scientific question?

2. Get to know your data

(a) Identify the types of your data: nominal data, ordinal data or quantitative data.

3. Calculate descriptive statistics in Excel

(a) Calculate the statistics for your variable of interest, such as sample mean (x¯), median, mode, variance (s 2 ), standard deviation (s), and 1st and 3rd quartiles.

(b) Use your qualitative variable to split your full data into two subsets: one set per category on your qualitative variable (Hint: The categorical variable has two categories/labels). Calculate the above statistics for each subset to compare.

4. Display your data with charts and graphs in Excel

(a) Construct displays that best describe ONLY your qualitative variable (e.g. bar chart, pie chart); and describe the distribution (Hint: How do we display qualitative data? which of all your variables are qualitative? 

(b) Construct displays that best describe your variable of interest and describe its distribution. (Use: Frequency distribution tables, histograms, etc.

(c) Construct displays that best describe the variable of interest by category (Use: Boxplot.

5. Building a Simple Linear Regression Model: Preprocess: Covariance, Correlation, Coefficient of Determination and SLR.

(a) Identify all quantitative variables from the dataset. 

(b) Construct a Scatter Plot to show the relationship between Tip $ (y) and each independent variable. Calculate the sample correlation coefficients for all pairs. Describe the association.

(c) Which pair has the strongest linear association?

(d) Write down the general formula for the Simple Linear Regression Model between y and x. (Write the formula using general parameters notation β0 and β1, what should be capitalize or lowercase? what should be added, if any? )

6. Describe the linear relationship between Tip $ (y) and the variable you answered in 1(c) (above) as x.

(a) Calculate the slope and y-intercept of the least squares regression line using Excel. Write down the linear equation.

(b) Interpret the regression slope.

(c) What percentage of the total variation in y can be explained by this independent variable x?

7. Use the regression model to predict Tip (y). (a) What is the predicted tip per table with 185 ? (Fill in the blank with units and name of the independent variable you chose.) 4. Is there a linear relationship between y and x?

(a) Test the significance of the slope of the regression equation. Use α = 0.08. i. Write down the hypotheses. ii. What is the p-value? iii. Describe your decision.

(b) Develop a 96% confidence interval for the population slope. Does this confidence interval include 0?

(c) State your conclusion Data → Data Analysis → Regression → Confidence level.)

8: Develop a multiple regression model to predict the Tip (y) using all the other variables of interest as listed above. (Round all numerical answers to two decimal places as needed.)

(a) Identify qualitative variable(s) from the list of variables of interest, if there is any, and create a dummy variable in Excel. (Note: use Excel function =IF() and use alphabetical order to assign values 0 and 1)

(b) Perform a multiple regression with the Data Analysis Toolpak in Excel, and write down the regression equation for Model 1. (Enter in Excel the confidence level given in question 1(e). 

(c) Explain the variation of the dependent variable after accounting for the effects of the other independent variables:

i. What percentage of total variation in the Tip (y) can be explained by Model 1? 

ii. What is the value of the adjusted multiple coefficient of determination, R2 A?

(d) Is the overall regression model significant using α = 0.02? State the hypotheses and your conclusion.

(e) Which independent variables are signifcant predictors using α = 0.4 or confidence level 60%? Which are not significant? (After accounting for the effects of the other independent variables)

9. Develop a second multiple regression model (Model 2) using ONE step of the “backward elimination method”. (Remember: variables should be removed one at the time and regression analysis i.e. coefficients, R2 , p-values, etc must be re-calculated at each step) (Round all numerical answers to two decimal places as needed.)

(a) Which variable should you remove from Model 1? Why?

(b) Perform a multiple regression with the Data Analysis Toolpak in Excel, and write down the regression equation for Model 2. (Enter in Excel the confidence level given in question 2(e).

(c) Explaining the variation of the dependent variable:

i. What percentage of total variation in the Tip (y) can be explained by Model 2? How does this compare with the percentage you obtained with Model 1?

ii. What is the value of the adjusted multiple coefficient of determination, R2 A? How does this compare with the one you obtained with Model 1?

(d) Is the overall regression model (Model 2) significant using α = 0.08?

(e) Are all the independent variables in Model 2 significant predictors using α = 0.1 or confidence level 90 ?ter accounting for the effects of the other independent variables?

(f) Prediction:

i. Is Model 2 better than Model 1?

ii. Predict the tip per table(y) with Day = Weekend; Bill ($) = 126; Diners() = 3; Table Minutes (minute) = 57 using “the best” model (between Model 1 and Model 2). 

(g) Interpret regression coefficients.

i. Interpret the coefficient of Diners.

3. Check the assumptions for regression analysis for the model you have chosen. Make necessary plots in Excel to justify.

(a) Is the relationship between the dependent and independent variables linear?

(b) Do the residuals exhibit some patterns across values of the independent variables?

(c) Are the variations of the dependent variable the same across all values of the independent variables?

(d) Do the residuals follow the normal probability distribution?

(e) Conclusion: Are the results from the regression analysis reliable?

This Statistics has been solved by our PHD Experts at My Uni Paper.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.