Advanced Quantitative Methods - Logistic Regression - Drugtreatment-Spring

Download Solution Order New Solution
This data set is from the UMARU IMPACT study. The purpose of the study was to detect the predictors of staying  drug free and, in particular, whether the duration of a drug treatment program matters. This kind of test is sometimes called a “dose, response” test and the process of determining the proper “dose” is known in medicine as “titration.” Randomization occurred at two sites which used slightly different treatment modalities. At site A, 444 individuals were randomized to test the difference between a 3 and 6 “modified therapeutic communities” approach which incorporated elements of health education and relapse prevention. Clients at this site were taught how to recognize “high risk” situations that are triggers to relapse and were taught the skills to enable them to cope with these situations without using drugs. In the trial at site B, 184 clients were randomized to receive either a 6- or 12-month therapeutic community program involving a highly structured lifestyle in a communal setting. You are to answer the questions below using R and the data set “drugtreatment.csv”. A description of the data set is included below. I’ll provide more overview in class.   Part I (Preliminary steps) 1. Before we do any statistical analysis we should get to know our dataset and our sample. This typically involves running some frequencies and means.   a. Summarize the variables in the dataset. Discuss measures of central tendency and variation.   b. Run frequencies on the following variables: Race Site Treat Dfree IVHX. Report the basic frequencies. Use the description of the dataset above to help you interpret the output. (This will explain what the “0” and “1” codes mean).   c. Analyze how the various variables differ by levels of the DFREE variable. (This will give you an overview for the modeling).   Part II (Core Assignment) Develop a logistic regression model to determine whether treatment duration matters and what else predicts likelihood of staying drug free. Write a 2-3 page front memo with a short literature review motivating the problem and analysis and highlighting your argument, actionable findings. In the appendix portion (approx 5 pages) you should be sure to explain the substantive findings and their meaning, as well as all of the key statistical attributes (for example the odds ratios, the various measures of goodness of fit, what was significant and what was not in predicting the outcome).   Additional Questions to answer What do the results imply to you? (include all output from final model!!). What are the main predictors of staying drug free? Who should be targeted for extra treatment? Does the longer treatment seem more effective than the shorter one? Does this vary by site? What interactions do you see and what do they imply? (Please include the code in the appendix/conclusion section of the document).   Be sure to interpret all of the logistic statistics (I want to see your final model, but I do want to see about variables which you tried which were not statistically significant). Remember, there are no standardized coefficients in logistic regression. Use the absolute value of the Z statistic to describe the strongest relationships. Use graphics to reinforce your arguments if you wish (not required)   Part III (Advanced extensions) Please try (2) of these (provide a small summary in the appendix with findings) 1. Try to fit an ROC curve and try some interactions (many are included on the dataset). The ROC curve is described in the Logistic Regression with R materials.   2. Analyze the interactions and consider creating more of them. Remember from regular regression that interactions involves multiplying two dummy variables to create a new variable and then adding this interaction variable to see if the two variables depend on each other to create differential relationships—it is a way of testing if the interaction is more than the sum of the two variables. If you try the interactions then be sure to test for multicollinearity by running your regression as a regular (OLS not Logistic) regression. This is because logistic regression doesn’t have the VIF. Often interactions cause multi-collinearity. If this happens, remove the individual variables that you interacted from the model and just put the interaction in.   3. Run the model separately by site. Then run the full model with the treat site interaction. How do the results differ?   4. Test for threshold effects or nonlinearities on the age and/or Beck variables. To do this, create dummy variables which indicate “high” values on these variables (a common approach is to consider the top quarterly to be “high”.   5. Analyze the false positive, false negative rate, the sensitivity and the specificity. What happens to the classification accuracy and these statistics if you move down the classification cutoff? What happens to them if you move it up?   6. Give everyone a probability of remaining drug free by saving the probability from your final model. Report on the distribution of that probability. Is it “spread out” such that the difference in probability across the interquartile range is at least 10 percentage points? Give those in the lowest quartile a “1” on a “highrisk” dummy variable. How do these “highrisk” individuals differ on the variables in the dataset? These individuals would be targeted for special treatment in a risk management system.   7. Compare the effects in site A and site B. This comparison is confounded if you do not factor in the duration of treatment. Resolve this issue by looking at a constant length of treatment and then comparing whether the approach taken in site A is more or less effective than site B.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.