Portrait de Reza Alvandi n'est pas disponible

Reza Alvandi

Stagiaire de recherche - McGill
Superviseur⋅e principal⋅e
Sujets de recherche
Apprentissage par renforcement
Contrôle stochastique

Publications

Confidence Intervals for the Return Process in Markov Decision Processes
In this work, we derive confidence intervals for the return process in discounted reward Markov Decision Processes with continuous state and… (voir plus) action spaces. These confidence bounds depend only on the statistics of the value function, which may be derived using dynamic programming. In the special case of MDPs with uniformly bounded value functions, simpler confidence intervals are provided for the return process. Finally, we study the effect of epistemic uncertainty on the derived confidence intervals. Numerical examples are provided to show how these bounds may be used in practice.