Reinforcement Learning using a Look-Up-Table (LUT) - on-Policy Learning vs off-Policy Learning - IT Assignment Help

Download Solution Order New Solution

Assignment Help
 

Part 2 - Reinforcement Learning using a Look-Up-Table (LUT):

 

Instructions

In this second part you are to implement Reinforcement Learning for your robot tank. For this you will need to use the Robocode environment. Before doing this, you need to decide which actions you would like to support and when & how to generate rewards. Once your learning mechanism is working, you should capture the contents of a trained LUT to file. This will be later helpful when replacing the LUT component
with a neural net.

Here are some tips:

 You will find it helpful to attempt to battle only a single enemy. Pick one that comes pre-installed with Robocode.  It might help if rewards are not just terminal.? You will have to apply some quantization or dimensionality reduction. If your state space is large you’ll, run out of memory when declaring your look up table.

For now you should implement RL, specifically Q-learning, using a look-up table. Furthermore since this is going to be replaced with the neural net from part one, the class headers defining the LUT should be consistent with the class headers of the neural net class from part 1. Again, to help you, I have suggested java interface definitions for you to start with:

 Submission instructions

(2) Once you have your robot working, measure its learning performance as follows:

a) Draw a graph of a parameter that reflects a measure of progress of learning and comment on the convergence of learning of your robot.
b) Using your robot, show a graph comparing the performance of your robot using on-policy learning vs off-policy learning.
c) Implement a version of your robot that assumes only terminal rewards and show & compare its behaviour with one having intermediate rewards.
(3) This part is about exploration. While training via RL, the next move is selected randomly with probability e and greedily with probability 1−e
a) Compare training performance using different values of e including no exploration at all. Provide graphs of the measured performance of your tank vs e

As for part 1, your submission should be a brief document clearly showing the graphs requested about. Please number your graphs as above and also include in your report an appendix section containing your source code.

 


This  IT Assignment has been solved by our IT experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+Students in Australia, UK & US by helping them to score HD in their academics. Our Experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.

Get It Done! Today

Country
Applicable Time Zone is AEST [Sydney, NSW] (GMT+11)
+

Every Assignment. Every Solution. Instantly. Deadline Ahead? Grab Your Sample Now.