Artificial Neural Network, Parameters of The Network, Multiclass Data, Partial Decision Tree & Different Decision Tree - IT Assignment Help

Download Solution Order New Solution
Assignment Task -                 
 


Question 1: Artificial Neural Network 
Ruben is in the process of designing an Artificial Neural Network to predict the pitch-value of the head pose given the vertical positions of the nose and chin as input attributes. Below is given an incomplete diagram of the architecture he has chosen. For the hidden units he has chosen a linear activation function f(a) = a, and for the output layer he has mistakenly used the suboptimal activation function f(a) = a2 + 2a. 

network

a. Copy the diagram from Figure 3 and complete it by drawing in the missing arrows, and in each hidden node, the fundamental computation steps taken. Add symbols to all nodes to identify them, as well as to the edges, inputs, and outputs. Use standard notation where possible and include index ranges of summations and products. 
b. Ruben will use backpropagation to find the optimal values of the intrinsic parameters of the network. Assuming that the error propagating from the output nodes to the hidden layers is δy, give the error term δz1 of the top hidden node. Provide step by step calculations. 
c. Assuming you use the standard loss function for regression, provide the partial derivative of the total network error E with respect to the weight connecting the first unit of the hidden layer to the output layer. Provide step by step calculations. 
d. You now want a network that is capable of encoding temporal dynamics. What type of neural network can do this? Sketch two different architectures that have the capacity of remembering past states by extending the network sketched in Figure 3 in two different ways. 
e. You now feed the first frame of a time series to the network capable of encoding temporal dynamics. Explain for both of your networks what problems you may encounter when making a prediction for this frame and how you could deal with this. 
f. Explain how batch normalization works and what it is used for.


Question 2: Decision Tree 
In Figure 4, there are data points with two feature values, x1, and x2, and in addition, label attributes indicated by their shape S in the set of shapes S ∈ {circle, square, triangle} and color 
c in the set of colors c ∈ {red, green, blue}. 

data

a. You must simplify this dataset into a binary set of labels of crosses and naughts, by creating a small set of rules that turn a particular shape and color data point to either naught or cross, losing all information about their previous shape and color in the process. Your set of rules must cover all possible combinations of shapes and colors, turning each colored shape into either naught or across, and should result in a balanced dataset of roughly equal numbers of naughts and crosses. Multiple solutions are possible. Provide the answer as a list of rules. 
b. Draw the new diagram of noughts and crosses. In this, draw the decision boundaries created by a monothetic decision tree with a perfect classification rate on the training data. Number the hyperplanes in the order they are used. 
c. Place a new datapoint indicated by an asterisk-shape at a random location in your new diagram. Then provide the class that the decision tree will infer for this data point and provide the sequence of decisions made by the decision tree to get to that prediction. Write down the decisions as equations. 
d. Draw the decision tree that created the decision boundaries of (b). Number the nodes to indicate what decision boundary the node is responsible for, and annotate for each edge what feature is compared against, and whether going down that edge means the feature value is greater or smaller than (or equal to) the threshold. 
e. In learning decision trees, one has to find for every node the query that maximizes the change in entropy, where this change in entropy is given by: 
δi(N) = i(N) − PLi(NL) − (1 − PL)i(NR)
For the partially learned tree shown in Figure 5, with two classes named ω1 and ω2, calculate the values of PL, I (NL), I (NR), and the change in entropy for the split made at the root node. The numbers in each node indicate how many examples of that class there are. 

Partial decision tree
f. Figure 6 shows two decision trees, with equivalent functionality (that is, every data point passed through the two trees results in exactly the same classification). Argue whether it is possible to choose one tree over the other, referring to fundamental Machine Learning concepts where possible.

DIFFERENT DECISION TREES

 

This IT Assignment has been solved by our IT Experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.

Be it a used or new solution, the quality of the work submitted by our assignment Experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.