COMP-SCI-3306 - Frequent Itemsets, Clustering, Advertising - IT Assignment Help

Download Solution Order New Solution
Assignment Task

 

Overview
Read the following carefully as it differs from the last assignment.
For students who are enrolled in the course COMP SCI 3306 (i.e. undergraduate students), the assignment must be done in groups consisting of TWO students. Please use A3-groups on MyUni to organise yourselves into groups.
If you have problems/questions regarding grouping or require assistance, please contact the teaching assistant Mahdi (mahdi.kazemimoghaddam@adelaide.edu.au).
For other students who are enrolled in the course COMP SCI 7306 (i.e. postgraduate students), this assignment must be done individually. Do not join a group in this case. References to sections, examples, etc. refer to the book of “Leskovec, Rajaraman and Ullman: Mining Massive Datasets (Second Edition)”.

Assignment
Exercise 1 Frequent Itemsets
For this exercise, you have to read Section 6.4 up to 6.4.3.
1. Implement the simple, randomized algorithm given in 6.4.1
2. Implement the algorithm of Savasere, Omiecinski, and Navathe (SON algorithm) in 6.4.3
3. Compare the two algorithms on the datasets T10I4D100K, T40I10D100K, chess, connect, mushroom, pumsb, pumsb star and report the outcomes.
4. Experiment with different sample sizes in the simple randomized algorithm such as 1, 2, 5, 10% and compare your results (including the result
produced by the SON algorithm).

Your approach should be as efficient as possible in terms of runtime and memory requirements.
Report on challenges that you might have observed in the implementation and by running experiments.

Exercise 2 Clustering
1. Perform a hierarchical clustering on the one-dimensional set of points
1, 4, 9, 16, 25, 36, 49, 64, 81. assuming the clusters are represented by their centroid (average), and at each step the clusters with the closest centroids are merged. (Exercise 7.2.1)

2. Implement the K-means algorithm and carry out experiments on the Iris
dataset (note that you are not allowed to use the libraries such as scikitlearn to implement the algorithm itself, but you are free to compare your results with such). The dataset can be accessed from scikit-learn library.

a) Plot the K-means clustering results by plotting the first 2 dimensions of the input data as well as the converged centroids.
b) Provide some discussions about how you picked the value of K in the K-means algorithm.

 

This COMP-SCI-3306 - IT Assignment has been solved by our IT experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.