Statistics and Data Modelling Assignment2

Download Solution Order New Solution

Assignment Task

Assessment 2 will be an individually written empirical project. You must select from one of the datasets listed on Blackboard and provide an analysis using this dataset. You are then to write up your analysis using Microsoft Word in a style format suggested (i.e. in the style of an empirical research paper). Your project must be no longer than words. You must only choose ONE of the datasets to perform your analysis on. The datasets available are the same as those in Assessment 1. You may wish to/can extend upon the analysis you did for Assessment 1. However, this time you may wish to source your own supplementary data or datasets. This is up to you and relates to the question you wish to ask (and your broader interests). Your project must cover the following points in a way you feel is most appropriate for your project:

  1. Statement of your research problem
  2. Descriptive analysis
  3. Inferential analysis
  4. Discussion and conclusions

For those aiming for a higher classification in this task should in addition consider covering:

  1. The robustness of your results in point 3
  2. Considerations of functional forms
  3. How your research relates to the More detail of the above points is provided below.

1. Statement of your research problem

In this section you are required to consider a research problem. Having spent more time on the module, you may now wish to pursue the same research question as before, but this time with an adapted methodology. Or, you may simply want to pursue a completely different question altogether. Finally, you may wish to take this opportunity to `try out’ a potential research question that you may want to pursue in your dissertation. Once you have decided on a topic or research question, you can specify this as a hypothesis, or you may wish to simply state a question. These will usually be derived from some ‘research problem’ as defined by the current literature on this area and a theoretical framework.

2. Description of the data and Descriptive Analysis

For this section you must ensure you tell the audience about.

  • The nature of the dataset (i.e. cross-sectional, time-series)
  • The nature of the variables (i.e. continuous, ordinal, categorical)

This section will require you to do some work in R. Each of the datasets is different, so the descriptive analysis will also be very different. Think about what the most appropriate way to present your data is. You can provide a range of tables/plots/histograms to explain what the data shows and the composition of the data and its variables, as covered in Topic 2.

3. Inferential analysis

In this section you may wish to use any of the techniques discussed in the module. This could include t-tests, simple linear regression, multiple linear regression and/or ARCH/GARCH models. You may

also wish to use multiple methodologies. Remember, the purpose of inferential analysis is to build statistical evidence around your question and its answer. The more evidence you have, the more rounded your answer will be. You should think more strongly about how to present your results in an appropriate manner (e.g. in the format of regression tables).

4. Discussion and conclusions

Have you answered your question? What was the answer? What are the implications? Think about the question you have asked and the result you got. If you identify that X causes Y, why is this important to individuals, policy makers or other researchers.

5. Robustness of the result

When consider the robustness of the result you need to make sure that your analysis doesn’t suffer from heteroscedasticity or autocorrelation. Have you tested for violation of assumptions? What was the result? Did you need to correct for them? If so, how did you do this? You need to make sure your model(s) do not violate assumptions. It is not enough to just have a model, but the model needs to be robust and efficient; otherwise there potentially is no validity in your statistical approach.

6. Consideration of alternative functional forms

We have seen that the default linear regression may not always be the most helpful way to specify our regression models. This depends on the nature of our variables, our motivations and theoretical reasoning a-prior the research. Should you specify an alternative functional form? Some discussion on the models functional form should be considered and discussed. This naturally has implications for your results and implications.

7. How your research relates to the literature

Was the set of research questions and empirical model defined by present literature? Are you attempting to conduct a set of analysis on something nobody has done before? Do your results contribute evidence to interesting topics in finance, accounting and economics. Fulfilling this aspect of your assignment involves discussing other research which is similar to yours. What did other people find? Does this support your work or contradict it? Why might this be the case?

Datasets

1. Wages

(Economics, Accounting & Finance based dataset)

This type of dataset is familiar from IT labs. We have provided an extended version here which has additional variables. A common question here would be to try and test hypotheses related to how much people earn. This could be as a simple t – test (Do X earn more than Y) or something more encompassing. You could wish to construct a multiple regression which allows you to test wages for multiple variables. This may be of more interest to students who have a more general interest in the wider economy.

2. Ethereum 

(Finance & Investment based dataset)

This dataset is very simple. There are only two columns, dates and closing prices of Ethereum (the 2 nd largest Cryptocurrency after Bitcoin). However, do not let this fool you. A typical question here could be to look at the average return over different periods (filtering) or you could simply try to test if the mean return is equal to zero (Remember this is what the theory says it should be!), similar to what was achieved in IT Lab 3.1 with Bitcoin data. This is an exceptionally interesting area of research, but challenging for those who wish to push their skills in R.

3. Capital Asset Pricing Model ("capm.csv")

(Investment & Finance based dataset)

This dataset is a CLASSIC in finance research. The capital asset pricing model (CAPM) is an important model in the field of finance. Attempting to model stock returns is perhaps the ONLY thing people are really interested in finance (both in research and in industry). In this dataset you can test the CAPM model that you may have learned about in your other modules. You will develop a much greater understanding of how risk metrics (Beta) are calculated, and this will no doubt be helpful for you in the future. However, bear in mind this can also be a tricky dataset to work with.

4. CEO Salaries & Firm Performance 

(Accounting based dataset)

CEO salaries are hotly discussed in news and media, alongside those of football players. However, what determines a CEO salary? Is it biographical information? Is it the person’s sex? Education? OR could it be how their company performs? Are CEO salaries simply a rational reward for running a successful company? This is an interesting alternative to the wage’s dataset, with a more precise focus. This topic may be of interest to accounting students or those with a keen interest in evaluating company performance. This can inform management accounting decision making. It requires an understanding of some key accountancy ratios as these can often be important variables to consider!

5. R&D Spending and Firm Size

(Accounting based dataset)

The relation between firm size and innovation is a topic that has been much studied in accounting literature, however theoretical and empirical studies are still inconclusive. Does the firm size, profits and/or profit margin determine a firms R&D (research and development) spending and/or R&D intensity (R&D spending as a percentage of sale), or vice versa? Does R&D intensity decrease with firm size? This is particularly pertinent to management accounting decision making within large companies, particularly when deciding an optimal level of R&D spending relative to the size of the firm.

6. Twitter 

(Investment & Finance based dataset)

This dataset gives share price data for both Twitter and the market (S&P 500). Twitter is very interesting at present the day, with the proposed takeover of Twitter by Elon Musk. There could be some interesting ways to use this dataset, such as beta analysis, subset analysis, amongst others. Your choice to how you use this data!

In Blackboard, there is a folder which contains the “.csv” file for each of these datasets. Accompanying this file is another file which explains what each of the variables represents. You are perfectly welcome to look at each of the datasets before deciding over which one you will choose. We have provided for you a range of data sets and you may select one of these for your assessment. We have split these data sets into subject specialisms, however if you feel interested by a data set outside of your specialism please feel free to use whichever you prefer.

This Statistics has been solved by our PhD Experts at My Uni Paper.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.