Highlights
Task Description:
Question 1
Assume a known MDP as above: In all states except B & E, your agent can perform 4 actions: Up, Down, Left, Right. In squares B & E your agent can only take the Exit action and an immediate reward of +10 & -10 respectively. You agent’s actions are successful 90% of the time; 10% of the time the agent moves in one of the two orthogonal (perpendicular) directions, with equal probability. If the movement is blocked by a wall (the blue squares, or the perimeter of the grid), the agent stays in place. For example, if your agent is in square C, and picks the Right action, 90% of the time it stays in place, 5% of the time in A, and 5% of the time in E.
Given the policy, p, as specified by the arrows, evaluate it for three iterations using the Policy Evaluation algorithm - fill in the values in the table below:
Question 2
Consider the same Gridworld as in question 1, but now assume that the MDP is unknown, i.e. you do not know the Reward Function, R, nor do you know the Transition Function, T. Your agent interacts with the Gridworld environment for four episodes as below.
(a) Using Model-based Reinforcement Learning, estimate the values below:
T(A, Up, C)=
T(A, Up, A)=
T(C, Left, A)=
T(C, Up, E)=
----
R(E, Exit, T) =
R(A, Up, C) =
R(B, Exit, T)=
R(C, Up, E)=
(b) After all Reward & Transition functions are learned as per (a), what methods could you use to solve the MDP?
(c) Using Q-Learning (Model-free), and assuming that all Q-Values for every pair have been initialized to 0, and a learning rate of (a = 0.6), fill in the Q-Table below after your agent has experienced the first three episodes above. The red cells show that the pair is not available. You can ignore these. Note that you will have to create the intermediate tables (after Episodes 1 and 2) for yourselves to get to the last one (after Episode 3), but only this last one is required here.
This IT Assessment has been solved by our IT experts at My Uni Paper. Our Assignment Writing Experts are efficient to provide a fresh solution to this question. We are serving more than 10000+ Students in Australia, UK & US by helping them to score HD in their academics. Our experts are well trained to follow all marking rubrics & referencing style.
Be it a used or new solution, the quality of the work submitted by our assignment experts remains unhampered. You may continue to expect the same or even better quality with the used and new assignment solution files respectively. There’s one thing to be noticed that you could choose one between the two and acquire an HD either way. You could choose a new assignment solution file to get yourself an exclusive, plagiarism (with free Turnitin file), expert quality assignment or order an old solution file that was considered worthy of the highest distinction.
© Copyright 2026 My Uni Papers – Student Hustle Made Hassle Free. All rights reserved.