Highlights
Task:
Background.
Social data analytics have become a vital asset for organizations and governments. For example, over the last few years, governments started to extract knowledge and derive insights from vastly growing open/social data to personalize the advertisements in elections, improve government services, predict intelligence activities, as well as to improve national security and public health. A key challenge in analyzing social data is to transform the raw data generated by social actors into curated data, i.e., contextualized data and knowledge that is maintained and made available for use by end-users and applications. In this assignment you will explore Big Data Technologies for analysing the data generated on social networks. Reference. Beheshti et al., "DataSynapse: A Social Data Curation Foundry". Distributed Parallel Databases 37(3): 351-384 (2019). Download: https://doi.org/10.1007/s10619-018-7245-1 Dataset. The Twitter dataset, including 10k tweets, is available on iLearn. Twitter1 serves many objects as JSON2 , including Tweets and Users. These objects all encapsulate core attributes that describe the object. Each Tweet has an author, a message, a unique ID, a timestamp of when it was posted, and sometimes geo metadata shared by the user. Each User has a Twitter name, an ID, a number of followers, and most often an account bio. With each Tweet, Twitter generates 'entity' objects, which are arrays of common Tweet contents such as hashtags, mentions, media, and links. If there are links, the JSON payload can also provide metadata such as the fully unwound URL and the webpage’s title and description. 1 https://developer.twitter.com/en/docs/tweets/data-dictionary/overview/intro-to-tweet-json 2 JSON is based on key-value pairs, with named attributes and associated values. These attributes, and their state, are used to describe objects. © https://data-science-group.github.io/ So, in addition to the text content itself, a Tweet can have over 140 attributes associated with it. Let’s start with an example Tweet: The following JSON illustrates the structure for these objects and some of their attributes: Source: https://developer.twitter.com/en/docs/twitter-api/v1/data-dictionary/overview
Part 1. Extraction (30%) You will use a technology to provide a customizable feature extraction to harness desired features from each tweet. These features should include: ? Schema-based features (10%). This category is related to the properties of a social item. For example, according to the Twitter schema, a tweet may have attributes such as text, source and language; and a user may have attributes such as username, description and timezone. ? Lexical-based features (10%). This category is related to: o the words or vocabulary of a language such as keyword, topic, phrase, abbreviation, special characters (e.g., ‘#’ in a tweet), slangs, informal language and spelling errors. o entities that can be extracted by the analysis and synthesis of natural language (NL) and speech, such as part-of-speech (e.g., verb, noun, etc), named entity type (e.g., person, organization, product, etc), and named entity (i.e., an instance of an entity type such as ‘Malcolm Turnbull’ as an instance of entity type Person).
The above IT Assignment has been solved by our IT Assignment Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.