Publications

Distributional reinforcement learning with linear function approximation

Subhodeep Moitra

Despite many algorithmic advances, our theoretical understanding of practical distributional reinforcement learning methods remains limited.… (see more) One exception is Rowland et al. (2018)'s analysis of the C51 algorithm in terms of the Cramer distance, but their results only apply to the tabular setting and ignore C51's use of a softmax to produce normalized distributions. In this paper we adapt the Cramer distance to deal with arbitrary vectors. From it we derive a new distributional algorithm which is fully Cramer-based and can be combined to linear function approximation, with formal guarantees in the context of policy evaluation. In allowing the model's prediction to be any real vector, we lose the probabilistic interpretation behind the method, but otherwise maintain the appealing properties of distributional approaches. To the best of our knowledge, ours is the first proof of convergence of a distributional algorithm combined with function approximation. Perhaps surprisingly, our results provide evidence that Cramer-based distributional methods may perform worse than directly approximating the value function.

2019-02-08

ArXiv (preprint)

Distributional reinforcement learning with linear function approximation

Subhodeep Moitra

Despite many algorithmic advances, our theoretical understanding of practical distributional reinforcement learning methods remains limited.… (see more) One exception is Rowland et al. (2018)'s analysis of the C51 algorithm in terms of the Cramer distance, but their results only apply to the tabular setting and ignore C51's use of a softmax to produce normalized distributions. In this paper we adapt the Cramer distance to deal with arbitrary vectors. From it we derive a new distributional algorithm which is fully Cramer-based and can be combined to linear function approximation, with formal guarantees in the context of policy evaluation. In allowing the model's prediction to be any real vector, we lose the probabilistic interpretation behind the method, but otherwise maintain the appealing properties of distributional approaches. To the best of our knowledge, ours is the first proof of convergence of a distributional algorithm combined with function approximation. Perhaps surprisingly, our results provide evidence that Cramer-based distributional methods may perform worse than directly approximating the value function.

2019-02-08

ArXiv (preprint)

Dendritic solutions to the credit assignment problem

Blake Richards

Timothy P. Lillicrap

2019-02-01

Current Opinion in Neurobiology (published)

doi.org

A Geometric Perspective on Optimal Representations for Reinforcement Learning

Will Dabney

Robert Dadashi

Adrien Ali Taiga

Dale Eric. Schuurmans

Tor Lattimore

Clare Lyle

We propose a new perspective on representation learning in reinforcement learning based on geometric properties of the space of value functi… (see more)ons. We leverage this perspective to provide formal evidence regarding the usefulness of value functions as auxiliary tasks. Our formulation considers adapting the representation to minimize the (linear) approximation of the value function of all stationary policies for a given environment. We show that this optimization reduces to making accurate predictions regarding a special class of value functions which we call adversarial value functions (AVFs). We demonstrate that using value functions as auxiliary tasks corresponds to an expected-error relaxation of our formulation, with AVFs a natural candidate, and identify a close relationship with proto-value functions (Mahadevan, 2005). We highlight characteristics of AVFs and their usefulness as auxiliary tasks in a series of experiments on the four-room domain.

2019-01-31

ArXiv (preprint)

Author Correction: Why rankings of biomedical image analysis competitions should be interpreted with care

Lena Maier-Hein

Matthias Eisenmann

Annika Reinke

Sinan Onogur

Marko Stankovic

Patrick Scholz

Tal Arbel

Hrvoje Bogunovic

Andrew P. Bradley

Aaron Carass

Carolin Feldmann

Alejandro F. Frangi

Peter M. Full

Bram van Ginneken

Allan Hanbury

Katrin Honauer

Michal Kozubek

Bennett Landman

Keno März

Oskar Maier … (see 18 more)

Klaus Maier-Hein

Bjoern Menze

Henning Müller

Peter F. Neher

Wiro Niessen

NASIR RAJPOOT

Gregory C. Sharp

Korsuk Sirinukunwattana

Stefanie Speidel

Christian Stock

Danail Stoyanov

Abdel Aziz Taha

Fons van der Sommen

Ching-Wei Wang

Marc-André Weber

Guoyan Zheng

Pierre Jannin

Annette Kopp-Schneider

2019-01-30

Nature Communications (published)

doi.org

Session-Based Social Recommendation via Dynamic Graph Attention Networks

Weiping Song

Zhiping Xiao

Yifan Wang

Laurent Charlin

Ming Zhang

Jian Tang

Online communities such as Facebook and Twitter are enormously popular and have become an essential part of the daily life of many of their … (see more)users. Through these platforms, users can discover and create information that others will then consume. In that context, recommending relevant information to users becomes critical for viability. However, recommendation in online communities is a challenging problem: 1) users' interests are dynamic, and 2) users are influenced by their friends. Moreover, the influencers may be context-dependent. That is, different friends may be relied upon for different topics. Modeling both signals is therefore essential for recommendations. We propose a recommender system for online communities based on a dynamic-graph-attention neural network. We model dynamic user behaviors with a recurrent neural network, and context-dependent social influence with a graph-attention neural network, which dynamically infers the influencers based on users' current interests. The whole model can be efficiently fit on large-scale data. Experimental results on several real-world data sets demonstrate the effectiveness of our proposed approach over several competitive baselines including state-of-the-art models.

2019-01-30

Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining (published)

doi.org

What comes next? Extractive summarization by next-sentence prediction

Jingyun Liu

Jackie Cheung

Annie Priyadarshini Louis

Existing approaches to automatic summarization assume that a length limit for the summary is given, and view content selection as an optimiz… (see more)ation problem to maximize informativeness and minimize redundancy within this budget. This framework ignores the fact that human-written summaries have rich internal structure which can be exploited to train a summarization system. We present NEXTSUM, a novel approach to summarization based on a model that predicts the next sentence to include in the summary using not only the source article, but also the summary produced so far. We show that such a model successfully captures summary-specific discourse moves, and leads to better content selection performance, in addition to automatically predicting how long the target summary should be. We perform experiments on the New York Times Annotated Corpus of summaries, where NEXTSUM outperforms lead and content-model summarization baselines by significant margins. We also show that the lengths of summaries produced by our system correlates with the lengths of the human-written gold standards.

2019-01-12

ArXiv (preprint)

A Geometric Perspective on Optimal Representations for Reinforcement Learning

Will Dabney

Robert Dadashi

Adrien Ali Taiga

Dale Schuurmans

Tor Lattimore

Clare Lyle

We propose a new perspective on representation learning in reinforcement learning based on geometric properties of the space of value functi… (see more)ons. We leverage this perspective to provide formal evidence regarding the usefulness of value functions as auxiliary tasks. Our formulation considers adapting the representation to minimize the (linear) approximation of the value function of all stationary policies for a given environment. We show that this optimization reduces to making accurate predictions regarding a special class of value functions which we call adversarial value functions (AVFs). We demonstrate that using value functions as auxiliary tasks corresponds to an expected-error relaxation of our formulation, with AVFs a natural candidate, and identify a close relationship with proto-value functions (Mahadevan, 2005). We highlight characteristics of AVFs and their usefulness as auxiliary tasks in a series of experiments on the four-room domain.