Highlights
QUESTION B: Estimating Constituency-Level Results from the EU Referendum
In the 2016 UK referendum on leaving the EU, the results of the vote were not released for individual electoral constituencies. However, many scholars would like to know why people voted to leave the EU, and how support for leaving differed across constituencies. One previous study has already estimated constituency-level support for ‘leave’ in an authoritative way. Your tasks in this question are (i) to produce estimates of the percentage of voters that voted ‘leave’ in every
constituency using multilevel modeling and post-stratification that are as close as possible to this existing set of estimates, as measured by the Mean Absolute Error (MAE), and (ii) to use your results to explain why people voted to leave.
You need to:
i)
Estimate an appropriate logistic multilevel model explaining voting for leave, using the predictors in the dataset. 1
ii)
Present the multilevel model results and interpret how the variables affect voting to leave the EU (Note: you do not need to discuss statistical significance).
iii)
Produce post-stratified estimates of the percentage of people who voted ‘leave’ in all 631 constituencies in England, Scotland and Wales
iv)
Compare your results to the existing estimates using the Mean Absolute Error You should present and explain your approach and results in a brief report, explaining why your estimates do or not perform well compared to the existing estimates. Note: if you cannot get very close to the existing results, do not worry. Your grade depends on the quality of your analysis,
presentation, and interpretation, not how close your results are to the existing estimates. The survey data is called “e” and is in the file “eusurvey.Rda”. It comes from 2017 British Election Study and it contains the following variables:
QUESTION C: Describing and Classifying Tweets
Many companies monitor social media posts in order to gauge how customers feel about their company and their competitors. For this question, imagine that you have been hired as a consultant by one of the major American airline companies to analyze tweets about airlines. They want to find
out how people talk about airlines on Twitter and then build a predictive tool that can classify tweets in future into ‘negative’ or ‘positive’ sentiment toward airlines, to help them respond better to their customers in real time. They have provided you with a dataset of 11,541 tweets about airlines that have been labeled as ‘negative’ or ‘positive’ by their staff. The dataset also identifies
which airline each tweet is talking about.
Your task is to prepare a brief report that describes the tweets and recommends a classification method for future tweets. You need to:
1. Use appropriate tools to describe the tweets. In particular, what words are associated with negative or positive sentiment? How does word usage differ across different airlines?
2. Use your analysis from (1) to build a short dictionary of negative and positive words describing airlines, then use it to classify tweets as ‘negative’ if they contain more negative than positive language, and ‘positive’ otherwise [code for creating your own dictionary is provided below]
3. Use an appropriate supervised machine-learning method to classify the tweets into ‘negative’ and ‘positive’
4. Compare the performance of your classifiers from (2) and (3), and use this analysis to decide which one would be the better classifier for the company to use for future tweets Here is some advice for part (2):
This Data Analysis Assignment has been solved by our Data Analysis experts at onlineassignmentbank. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.