Publications

R-MelNet: Reduced Mel-Spectral Modeling for Neural TTS
A guided multiverse study of neuroimaging analyses
Jessica Dafflon
Pedro F. da Costa
František Váša
Ricardo Pio Monti
Peter J. Hellyer
Federico Turkheimer
Jonathan Smallwood
Emily Jones
Robert Leech
For most neuroimaging questions the range of possible analytic choices makes it unclear how to evaluate conclusions from any single analytic… (see more) method. One possible way to address this issue is to evaluate all possible analyses using a multiverse approach, however, this can be computationally challenging and sequential analyses on the same data can compromise predictive power. Here, we establish how active learning on a low-dimensional space capturing the inter-relationships between pipelines can efficiently approximate the full spectrum of analyses. This approach balances the benefits of a multiverse analysis without incurring the cost on computational and predictive power. We illustrate this approach with two functional MRI datasets (predicting brain age and autism diagnosis) demonstrating how a multiverse of analyses can be efficiently navigated and mapped out using active learning. Furthermore, our presented approach not only identifies the subset of analysis techniques that are best able to predict age or classify individuals with autism spectrum disorder and healthy controls, but it also allows the relationships between analyses to be quantified.
Integrating Equity, Diversity, and Inclusion throughout the lifecycle of Artificial Intelligence in health
Milka Nyariro
Elham Emami
Samira Abbasgholizadeh Rahimi
Reproducible between-person brain-behavior associations do not always require thousands of individuals
Colin G. DeYoung
Tyler A. Sassenberg
Rany Abend
Timothy A. Allen
Roger E. Beaty
Mark A. Bellgrove
Scott D. Blain
Robert S. Chavez
Stephen A. Engel
Ma Feilong
Alex Fornito
Erhan Genç
Vina M. Goghari
Rachael Grazioplene
Jamie L. Hanson
James V. Haxby
Kirsten Hilger
Philipp Homan
Keanan J. Joyner … (see 12 more)
Antonia N. Kaczkurkin
Robert D. Latzman
Elizabeth A. Martin
Luca Passamonti
Alan D. Pickering
Adam Safron
Michelle N. Servaas
Luke D. Smillie
R. Nathan Spreng
Jeggan Tiego
Essi Viding
Jan Wacker
Marek et al. analyzed three very large magnetic resonance imaging (MRI) datasets and concluded that thousands of participants are necessary … (see more)to ensure replicable results in “brain-wide associations studies,” which they defined as “studies of the associations between common inter-individual variability in human brain structure/function and cognition or psychiatric symptomatology.” This conclusion overgeneralizes the implications of their findings and is likely to have an unwarranted chilling effect on neuroimaging research focused on individual differences, preventing good research with samples in the hundreds from being funded and conducted. To fend off these negative consequences, we explain why their conclusion is not fully justified, discuss methods that can yield larger effects, and suggest practical guidelines for sample size, recognizing the potential utility of samples in the hundreds.
Annotation Cost-Sensitive Deep Active Learning with Limited Data (Student Abstract)
Disentanglement via Mechanism Sparsity Regularization: A New Principle for Nonlinear ICA
Pau Rodríguez
Yash Sharma
Katie E Everett
Rémi Le Priol
Alexandre Lacoste
This work introduces a novel principle we call disentanglement via mechanism sparsity regularization, which can be applied when the latent f… (see more)actors of interest depend sparsely on past latent factors and/or observed auxiliary variables. We propose a representation learning method that induces disentanglement by simultaneously learning the latent factors and the sparse causal graphical model that relates them. We develop a rigorous identifiability theory, building on recent nonlinear independent component analysis (ICA) results, that formalizes this principle and shows how the latent variables can be recovered up to permutation if one regularizes the latent mechanisms to be sparse and if some graph connectivity criterion is satisfied by the data generating process. As a special case of our framework, we show how one can leverage unknown-target interventions on the latent factors to disentangle them, thereby drawing further connections between ICA and causality. We propose a VAE-based method in which the latent mechanisms are learned and regularized via binary masks, and validate our theory by showing it learns disentangled representations in simulations.
Estimating Social Influence from Observational Data
Caterina De Bacco
David Blei
We consider the problem of estimating social influence, the effect that a person's behavior has on the future behavior of their peers. The k… (see more)ey challenge is that shared behavior between friends could be equally explained by influence or by two other confounding factors: 1) latent traits that caused people to both become friends and engage in the behavior, and 2) latent preferences for the behavior. This paper addresses the challenges of estimating social influence with three contributions. First, we formalize social influence as a causal effect, one which requires inferences about hypothetical interventions. Second, we develop Poisson Influence Factorization (PIF), a method for estimating social influence from observational data. PIF fits probabilistic factor models to networks and behavior data to infer variables that serve as substitutes for the confounding latent traits. Third, we develop assumptions under which PIF recovers estimates of social influence. We empirically study PIF with semi-synthetic and real data from Last.fm, and conduct a sensitivity analysis. We find that PIF estimates social influence most accurately compared to related methods and remains robust under some violations of its assumptions.
Fair Representation Learning through Implicit Path Alignment
Qi CHEN
Jiaqi Li
Boyu Wang
A Generalized Bootstrap Target for Value-Learning, Efficiently Combining Value and Feature Predictions
Anthony GX-Chen
Blake A. Richards
Estimating value functions is a core component of reinforcement learning algorithms. Temporal difference (TD) learning algorithms use bootst… (see more)rapping, i.e. they update the value function toward a learning target using value estimates at subsequent time-steps. Alternatively, the value function can be updated toward a learning target constructed by separately predicting successor features (SF)--a policy-dependent model--and linearly combining them with instantaneous rewards. We focus on bootstrapping targets used when estimating value functions, and propose a new backup target, the
Head2Toe: Utilizing Intermediate Representations for Better Transfer Learning
Utku Evci
Michael Curtis Mozer
Improving Robustness against Real-World and Worst-Case Distribution Shifts through Decision Region Quantification
Leon Bungert
A. Nguyen
Ren'e Raab
Falk Pulsmeyer
B. Eskofier
Dario Zanca
The reliability of neural networks is essential for their use in safety-critical applications. Existing approaches generally aim at improvin… (see more)g the robustness of neural networks to either real-world distribution shifts (e.g., common corruptions and perturbations, spatial transformations, and natural adversarial examples) or worst-case distribution shifts (e.g., optimized adversarial examples). In this work, we propose the Decision Region Quantification (DRQ) algorithm to improve the robustness of any differentiable pre-trained model against both real-world and worst-case distribution shifts in the data. DRQ analyzes the robustness of local decision regions in the vicinity of a given data point to make more reliable predictions. We theoretically motivate the DRQ algorithm by showing that it effectively smooths spurious local extrema in the decision surface. Furthermore, we propose an implementation using targeted and untargeted adversarial attacks. An extensive empirical evaluation shows that DRQ increases the robustness of adversarially and non-adversarially trained models against real-world and worst-case distribution shifts on several computer vision benchmark datasets.
3D Infomax improves GNNs for Molecular Property Prediction
Hannes Stärk
Gabriele Corso
Christian Dallago
Stephan Günnemann
Pietro Lio
Molecular property prediction is one of the fastest-growing applications of deep learning with critical real-world impacts. Including 3D mol… (see more)ecular structure as input to learned models improves their predictions for many molecular properties. However, this information is infeasible to compute at the scale required by most real-world applications. We propose pre-training a model to understand the geometry of molecules given only their 2D molecular graph. Using methods from self-supervised learning, we maximize the mutual information between a 3D summary vector and the representations of a Graph Neural Network (GNN) such that they contain latent 3D information. During fine-tuning on molecules with unknown geometry, the GNN still generates implicit 3D information and can use it to inform downstream tasks. We show that 3D pre-training provides significant improvements for a wide range of molecular properties, such as a 22% average MAE reduction on eight quantum mechanical properties. Crucially, the learned representations can be effectively transferred between datasets with vastly different molecules.