Highlights
This item of coursework will contribute to 50% of the overall module marks.
The solutions of all the following exercises need to be submitted into the module assessment area of the Blackboard, as a lab-based assignment, by the end of the day on Friday in the ninth week, contributing to your portfolio of evidence relating to Data Validation exercises. You may like to use this file to present your functioning code along with program outputs through this R Markdown document.
Exercise 1. The dataset mpg is part of the R datasets package. It contains a subset of the fuel economy data that the Environment Protection Agency (EPA) makes available. It contains only car models which had a new release every year between 1999 and 2008 - this was used as a proxy for the popularity of the car. Applying an appropriate R data visualization method on the mpg data, perform the following tasks:
(a). Write code that displays a graph which plots in the order of decreasing medians of the vehicle's miles-per-gallon on the highway (hwy) against their manufacturers.
(b). Plot the graph and list the manufacturers in the order of fuel efficiency of their vehicles.
(c). Using the graph, find out which companies produce the most and the least fuel-efficient cars.
Exercise 2. The diamonds dataset within R's ggplot2 contains 10 columns (price, carat, cut, color, clarity, length(x), width(y), depth(z), depth
percentage, top width) for 53940 different diamonds. Using this dataset, carry out the following tasks.
(a). Write code to plot histograms for cut, carat, and price. Plot the histograms and comment on their shapes.
(b). Write code to display an appropriate graph that facilitates the investigation of a three-way relationship between cut, carat, and price. Plot the graph.
(c). Based on the above graph, what are your conclusions regarding the three-way relationship?
Exercise 3. Before deciding about selecting a particular machine learning technique for a data science problem, it is important to study the data
distribution, particularly through visualization. However, visualizing multivariate data with two or more variables is difficult in a two-dimensional plot. In this exercise, you are required to study the R's iris dataset which is multivariate data consisting of four features or properties (Sepal.Length, Sepal.Width, Petal.Length, Petal.Width) characterizing three species of an iris flower (setosa, versicolor, and virginica). The principal component analysis (PCA) is a technique that can help facilitate the visualization of multivariate data distribution. The first two principal components (PC1 and PC2) obtained after applying PCA,
can explain the majority of variation in the data. In order to study the data variability in iris data-set, perform the following tasks.
(a). Write code to obtain PC scores.
(b). Write code to obtain a scatter plot representing PC1 vs. PC2, wherein data clusters corresponding to three flower types are clearly marked using possibly an elipsoid.
(c). Run the codes to make the scatter plot and comment on the feature distribution.
Exercise 4. In this task, you are required to analyze the Animals dataset from the MASS package. This dataset contains brain weight (in grams) and body weight (in kilograms) for 28 different animal species. The three largest animals are dinosaurs, whose measurements are obviously the result of scientific modeling rather than precise measurements. A scatter plot given below fails to describe any obvious relationship between brain weight and body weight variables. You are required to apply
appropriate power transformations to the variables to obtain a more interpretable plot and describe the obtained relationship. To this end, undertake the following tasks.
Task-1. Check whether each of the variables has a normal distribution. Your response should be based on an appropriate statistical test as well as smoothed histogram plots.
Task-2. A power transformation of a variable X consists of raising X to the power lambda. Using an appropriate statistical test and/or plot, find the best lambda values needed for transforming each of the variables requiring power transformation.
Task-3. Apply power transformation and verify whether transformed variables have a normal distribution through the statistical tests as well as smoothed histogram plots.
Task-4. Create a scatter plot of the transformed data. Based on the visual inspection of the plot, provide your interpretation of the relationship between brain weight and body weight variables. You may like to add an appropriate smoothed line curve to your plot to help in interpretation.
This IT/Computer Science Assignment has been solved by our IT/Computer Science Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.