Highlights
Assessment Overview
In Week 2, we started using basic R functions (Lab 4), tried implementing regression (Lab 5) and PCA (Lab 6).
In this assessment, you will apply the data analytics and visualization skills to further analyse the provided TwitterSpam dataset. Table 1 shows the features description of the dataset .
Table 1. Account-based and Content-based Features Description
Feature Name Description Account-based Features account age no_follower no_following no_userfavorites no_lists no_tweets
The age of an account # of followers # of followings # of favourites the user received # of lists in which the user is a member of # of tweets that has been posted by the user Content-based Features no retweets no_tweetfavorites no_hashtag no_usermention no_urls no_char no_digits of times this tweet has been retweeted # of favourites this tweet received # of hashtags in this tweet # of times this tweet being mentioned # of URLs contained in this tweet # of characters in this tweet # of digits in this tweet
Assessment Requirements and Instructions
Follow the instructions below and write an essay that covers the following tasks. R script, R screenshot, your results and explanations should be covered for each question.
1 Load TwitterSpam dataset into R studio, use ggplot function to make density plot of Tweets' number (column: no_tweets) to
compare spam and non-spam. Please clarify on how you make the plot, and what's your observation from the density plot?
2 Load TwitterSpam dataset into R studio, use ggplot function to make density plot of retweets' number (column: no_retweets) to compare spam and non-spam. Please clarify on how you make the plot, and what's your observation from the density plot?
3 Compared with the result from task 1 and task 2, for the 'no_tweets' and 'no_retweets' column, which one is more valuable for detecting Twitter spam? Why?
4 Use ggplot function to make scatterplots to present the relation of posted tweets number (column: no_tweets) and the number of followers (column: no_follower). Please explain how you make the scatterplot and the relationship between the number of posted tweets and the number of followers you observe from the scatterplot.
Add regression line to the scatterplot generated in Step 2. Compared with scatterplot only, what are the advantages of adding a regression line?
Data exploration is the initial step in data analysis, where users explore a data set in an unstructured way to uncover initial patterns, characteristics, and points of interest. Except approaches mentioned in Assessment 1 and Assessment 2 (1)(2)(3) and characteristics observed from these approaches, what other characteristics, points of interest, or initial patterns you find from the dataset? Please describe one of your findings and give a detailed description of how you achieve it.
This NIT3202 - IT Assignment has been solved by our IT Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment Experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.