Publications

Transnational conservation to anticipate future plant shifts in Europe
Yohann Chauvier-Mendes
Peter H. Verburg
Dirk N. Karger
Loïc Pellissier
Sébastien Lavergne
Niklaus E. Zimmermann
Wilfried Thuiller
Deployable Reinforcement Learning with Variable Control Rate
Yong Wang
METhodological RadiomICs Score (METRICS): a quality scoring tool for radiomics research endorsed by EuSoMII
Burak Kocak
Tugba Akinci D’Antonoli
Nathaniel Mercaldo
Angel Alberich-Bayarri
Bettina Baessler
Ilaria Ambrosini
Anna E. Andreychenko
Spyridon Bakas
Regina G. H. Beets-Tan
Keno Bressem
Irene Buvat
Roberto Cannella
Luca Alessandro Cappellini
Armando Ugo Cavallo
Leonid L. Chepelev
Linda Chi Hang Chu
Aydin Demircioglu
Nandita M. deSouza
Matthias Dietzel
Salvatore Claudio Fanni … (see 40 more)
Andrey Fedorov
Laure S. Fournier
Valentina Giannini
Rossano Girometti
Kevin B. W. Groot Lipman
Georgios Kalarakis
Brendan S. Kelly
Michail E. Klontzas
Dow-Mu Koh
Elmar Kotter
Ho Yun Lee
Mario Maas
Luis Marti-Bonmati
Henning Müller
Nancy Obuchowski
Fanny Orlhac
Nikolaos Papanikolaou
Ekaterina Petrash
Elisabeth Pfaehler
Daniel Pinto dos Santos
Andrea Ponsiglione
Sebastià Sabater
Francesco Sardanelli
Philipp Seeböck
Nanna M. Sijtsema
Arnaldo Stanzione
Alberto Traverso
Lorenzo Ugga
Lisanne V. van Dijk
Joost J. M. van Griethuysen
Robbert W. van Hamersvelt
Peter van Ooijen
Federica Vernuccio
Alan Wang
Stuart Williams
Jan Witowski
Zhongyi Zhang
Alex Zwanenburg
Renato Cuocolo
Abstract B049: Pancreatic beta cell stress pathways drive pancreatic ductal adenocarcinoma development in obesity
Cathy C. Garcia
Aarthi Venkat
Sherry Agabiti
Lauren Lawres
Rebecca Cardone
Richard G. Kibbey
Mandar Deepak Muzumdar
Closing the Gap between TD Learning and Supervised Learning - A Generalisation Point of View
Matthieu Geist
Benjamin Eysenbach
Some reinforcement learning (RL) algorithms can stitch pieces of experience to solve a task never seen before during training. This oft-soug… (see more)ht property is one of the few ways in which RL methods based on dynamic-programming differ from RL methods based on supervised-learning (SL). Yet, certain RL methods based on off-the-shelf SL algorithms achieve excellent results without an explicit mechanism for stitching; it remains unclear whether those methods forgo this important stitching property. This paper studies this question for the problems of achieving a target goal state and achieving a target return value. Our main result is to show that the stitching property corresponds to a form of combinatorial generalization: after training on a distribution of (state, goal) pairs, one would like to evaluate on (state, goal) pairs not seen together in the training data. Our analysis shows that this sort of generalization is different from i.i.d. generalization. This connection between stitching and generalisation reveals why we should not expect SL-based RL methods to perform stitching, even in the limit of large datasets and models. Based on this analysis, we construct new datasets to explicitly test for this property, revealing that SL-based methods lack this stitching property and hence fail to perform combinatorial generalization. Nonetheless, the connection between stitching and combinatorial generalisation also suggests a simple remedy for improving generalisation in SL: data augmentation. We propose a temporal data augmentation and demonstrate that adding it to SL-based methods enables them to successfully complete tasks not seen together during training. On a high level, this connection illustrates the importance of combinatorial generalization for data efficiency in time-series data beyond tasks beyond RL, like audio, video, or text.
Consciousness-Inspired Spatio-Temporal Abstractions for Better Generalization in Reinforcement Learning
Mingde Zhao
Harm van Seijen
Romain Laroche
Inspired by human conscious planning, we propose Skipper, a model-based reinforcement learning framework utilizing spatio-temporal abstracti… (see more)ons to generalize better in novel situations. It automatically decomposes the given task into smaller, more manageable subtasks, and thus enables sparse decision-making and focused computation on the relevant parts of the environment. The decomposition relies on the extraction of an abstracted proxy problem represented as a directed graph, in which vertices and edges are learned end-to-end from hindsight. Our theoretical analyses provide performance guarantees under appropriate assumptions and establish where our approach is expected to be helpful. Generalization-focused experiments validate Skipper's significant advantage in zero-shot generalization, compared to some existing state-of-the-art hierarchical planning methods.
A Data-Driven Measure of Relative Uncertainty for Misclassification Detection
Eduardo Dadalto Câmara Gomes
Marco Romanelli
Georg Pichler
Decoupling regularization from the action space
Sobhan Mohammadpour
Regularized reinforcement learning (RL), particularly the entropy-regularized kind, has gained traction in optimal control and inverse RL. W… (see more)hile standard unregularized RL methods remain unaffected by changes in the number of actions, we show that it can severely impact their regularized counterparts. This paper demonstrates the importance of decoupling the regularizer from the action space: that is, to maintain a consistent level of regularization regardless of how many actions are involved to avoid over-regularization. Whereas the problem can be avoided by introducing a task-specific temperature parameter, it is often undesirable and cannot solve the problem when action spaces are state-dependent. In the state-dependent action context, different states with varying action spaces are regularized inconsistently. We introduce two solutions: a static temperature selection approach and a dynamic counterpart, universally applicable where this problem arises. Implementing these changes improves performance on the DeepMind control suite in static and dynamic temperature regimes and a biological design task.
Efficient Dynamics Modeling in Interactive Environments with Koopman Theory
Arnab Kumar Mondal
Sai Rajeswar
The accurate modeling of dynamics in interactive environments is critical for successful long-range prediction. Such a capability could adva… (see more)nce Reinforcement Learning (RL) and Planning algorithms, but achieving it is challenging. Inaccuracies in model estimates can compound, resulting in increased errors over long horizons. We approach this problem from the lens of Koopman theory, where the nonlinear dynamics of the environment can be linearized in a high-dimensional latent space. This allows us to efficiently parallelize the sequential problem of long-range prediction using convolution while accounting for the agent's action at every time step. Our approach also enables stability analysis and better control over gradients through time. Taken together, these advantages result in significant improvement over the existing approaches, both in the efficiency and the accuracy of modeling dynamics over extended horizons. We also show that this model can be easily incorporated into dynamics modeling for model-based planning and model-free RL and report promising experimental results.
Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
Brady Neal
Vasilis Syrgkanis
Ensemble Distillation for Unsupervised Constituency Parsing
Behzad Shayegh
Yanshuai Cao
Xiaodan Zhu
Jackie CK Cheung
Lili Mou
Evaluating Representation Learning on the Protein Structure Universe
Arian R. Jamasb
Alex Morehead
Chaitanya K. Joshi
Kieran Didi
Simon Mathis
Charles Harris
Jianlin Cheng
Pietro Lio
Tom L. Blundell
We introduce ProteinWorkshop, a comprehensive benchmark suite for representation learning on protein structures with Geometric Graph Neural … (see more)Networks. We consider large-scale pre-training and downstream tasks on both experimental and predicted structures to enable the systematic evaluation of the quality of the learned structural representation and their usefulness in capturing functional relationships for downstream tasks. We find that: (1) large-scale pretraining on AlphaFold structures and auxiliary tasks consistently improve the performance of both rotation-invariant and equivariant GNNs, and (2) more expressive equivariant GNNs benefit from pretraining to a greater extent compared to invariant models. We aim to establish a common ground for the machine learning and computational biology communities to rigorously compare and advance protein structure representation learning. Our open-source codebase reduces the barrier to entry for working with large protein structure datasets by providing: (1) storage-efficient dataloaders for large-scale structural databases including AlphaFoldDB and ESM Atlas, as well as (2) utilities for constructing new tasks from the entire PDB. ProteinWorkshop is available at: github.com/a-r-j/ProteinWorkshop.