Cybersecurity - K-Means and Self Organising Maps (SOM) - IT Assessment Answer

Download Solution Order New Solution
Internal Code: 1AGBJH

IT Assessment Answer

  • Part A:
  1.  Work in groups of maximum 4 people. o Your solution must use a large dataset of tweets. This dataset will be made available to you on the Data Science Cluster of the Department Computer Science. o The tweet attributes and contents of the tweets should be used.
  • Part B
  1.  Individual assignments
Part 1 
  1. Data Acquisition / Generation / Fabrication
  2.  You access the Twitter data set on your platform at the following Google
  3.  For access on the data science cluster please contact TechTeam who will make the Twitter data available under a certain folder.
  • Data Cleaning: Your data needs to be reliable. Dirty data is data that has values, which are missing, incorrect or inconsistent (Krishnan et al., 2015). Do data cleaning on your dataset with a program/script.
  1. Do exploratory data analysis on the tweet dataset. 
Additional to your EDA also do the following: 3.1 Extract URL’s out of the text. Determine if some of the URLs are Phishing. To do this you can use normal regular expressions. Regular expression in this context is a pattern describing some text. You can read more on regular expressions Extend your data set to include a feature indicating possible phishing. 3.2 Do language detection by using the contents of tweets. Once you  have identified the language then link it back to the language specified by the user account holder in the user profile. This will not help us to identify if a person deceives, but it might be interesting to correlate the tweet contents with the user profiles. This can possibly also be then linked to the country/land of the user profile. 3.3 Emoticons. Investigate whether to which extend users in general, in our Twitter data set use emoticons.
  1. Identify attributes from tweets and contents that can play a role in the discovery of Identity Deception. Do feature engineering.
  1. Unsupervised machine learning for the detection of identity Deception
  • Twitter as a social media platform has millions of users of which a large percentage can be users who lie about their identities, also referred to identity deception. Because of the big volumes of data, such as the case with Twitter, it is impossible to manually determine which identities deceive and which are true. Therefore we focus on machine learning as an intelligent automated way to assist us in identifying identity deception. Broadly, machine learning can be classified into supervised machine
  • The problem with supervised machine learning is that you have to tell the algorithm what to do. So basically a human has to first decide which identities are deceitful and which are not. Training data sets need to be developed that label certain identities as deceitful, this on its own is not an easy process because very rarely you will find users themselves that will indicate that their identities are fake. To overcome this problem we will use unsupervised machine learning to discover clusters (subsets) of data that may be give us an indication of certain types of user profiles that appears to be deceitful. In this case we do not, upfront determine who is deceitful or not, but we rather let the algorithm do the clustering itself and hopefully we can then learn something about indicators of deception in user profiles from the different clusters.
  • Data clustering is a common problem in various domains. The intention is to categorise similar data into one cluster based on features of the data that is similar. In our case raw Twitter data is taken and then mapped onto a new feature space that will assist in the clustering of the data into different categories. Common algorithms that do this are k-means and Self Organising maps (SOM).
PART 2
  1. This part of the assignment can only be completed on an individual basis i.e. each student submit 100% his/her own work.
  2. Write an essay on state-of-the-art developments in automatic Deception Detection.
  3. Your assignment should be maximum 1500 words including headings and bibliography.
  4. Each student will present his own Part B assignment as part of the practical demonstration. More details will be provided.
This Computer Science Assessment has been solved by our Computer Science experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.