Portrait of Reza Alvandi is unavailable

Reza Alvandi

Research Intern - McGill University
Supervisor
Research Topics
Reinforcement Learning
Stochastic Control

Publications

Confidence Intervals for the Return Process in Markov Decision Processes
In this work, we derive confidence intervals for the return process in discounted reward Markov Decision Processes with continuous state and… (see more) action spaces. These confidence bounds depend only on the statistics of the value function, which may be derived using dynamic programming. In the special case of MDPs with uniformly bounded value functions, simpler confidence intervals are provided for the return process. Finally, we study the effect of epistemic uncertainty on the derived confidence intervals. Numerical examples are provided to show how these bounds may be used in practice.