Portrait de Hugo Larochelle

Hugo Larochelle

Directeur scientifique, Équipe de direction
Professeur associé, Université de Montréal, Département d'informatique et de recherche opérationnelle
Professeur associé, McGill University, École d'informatique
Sujets de recherche
Apprentissage profond

Biographie

Hugo Larochelle est le directeur scientifique de Mila, l’un des plus importants centres de recherche en intelligence artificielle au monde. Avec une communauté de près de 2000 chercheurs et professionnels, l’institut s’est imposé comme un pilier de l'écosystème canadien de l’IA, avec un rayonnement qui s’étend bien au-delà des frontières du pays.

En tant que chercheur pionnier et leader industriel, il possède une connaissance unique tant sur les grands laboratoires de recherche en entreprise que sur l’excellence de la communauté universitaire canadienne en IA. Il a bâti ses assises académiques auprès de deux des « pères fondateurs » de l'intelligence artificielle : Yoshua Bengio et Geoffrey Hinton.

Au fil des ans, ses recherches ont mené à plusieurs découvertes que l'on retrouve dans les systèmes d'IA modernes. Ses travaux sur les auto-encodeurs débruiteurs (DAE) ont établi la reconstruction de données pures à partir de versions corrompues comme un paradigme évolutif pour l'apprentissage de représentations significatives à partir de grandes quantités de données non étiquetées. Grâce à des modèles tels que l'Estimateur de distribution autorégressif neuronal (NADE) et l'Auto-encodeur masqué pour l'estimation de distribution (MADE), il a contribué à populariser la modélisation autorégressive neuronales, désormais omniprésent dans l'IA générative. De plus, ses travaux sur l'apprentissage de nouvelles tâches sans données (Zero-Data Learning of New Tasks) ont introduit le concept, aujourd'hui courant, de l'apprentissage « zero-shot ».

Il a jeté un pont entre le milieu académique et l’industrie en cofondant la startup Whetlab, acquise par Twitter en 2015. Après un passage chez Twitter Cortex, il a été recruté pour diriger le laboratoire de recherche en IA de Google à Montréal (Google Brain), aujourd’hui intégré à Google DeepMind. Il demeure professeur associé à l'Université de Montréal et l’Université McGill, et est titulaire d'une chaire en IA Canada-CIFAR, formant ainsi la relève scientifique.

Parallèlement à ses fonctions de directeur scientifique chez Mila, il est également responsable scientifique chez Adaption Labs et conseille les start-ups Tiptree Systems et Prizmal.

Père de quatre enfants, Hugo Larochelle et sa conjointe, Angèle St-Pierre, ont également fait de multiples dons à l'Université de Montréal et à l'Université de Sherbrooke, particulièrement dans le domaine de l'IA pour l’environnement. Il a également fondé la conférence Techaide, mobilisant la communauté technologique de Montréal afin de collecter des fonds pour Centraide dans sa mission de lutte contre la pauvreté et l'exclusion sociale.

Étudiants actuels

Doctorat - UdeM
Superviseur⋅e principal⋅e :
Maîtrise professionnelle - McGill
Collaborateur·rice alumni - UdeM
Superviseur⋅e principal⋅e :
Visiteur de recherche indépendant - McGill
Postdoctorat - Polytechnique
Superviseur⋅e principal⋅e :

Publications

The Search for Squawk: Agile Modeling in Bioacoustics
Otilia Stretcu
Jenny Hamer
Lauren Harrell
Rob Laber
Amanda K. Navine
Patrick Hart
Ben Williams
Timothy A. C. Lamont
Tries B. Rasak
Mars Coral Restoration Team
Sheryn Brodie
Brendan Doohan
Philip Eichinski
Paul Roe
Lin Schwarzkopf
Tom Denton
Don't Flatten, Tokenize! Unlocking the Key to SoftMoE's Efficacy in Deep RL
Ghada Sokar
Johan Obando-Ceron
The use of deep neural networks in reinforcement learning (RL) often suffers from performance degradation as model size increases. While sof… (voir plus)t mixtures of experts (SoftMoEs) have recently shown promise in mitigating this issue for online RL, the reasons behind their effectiveness remain largely unknown. In this work we provide an in-depth analysis identifying the key factors driving this performance gain. We discover the surprising result that tokenizing the encoder output, rather than the use of multiple experts, is what is behind the efficacy of SoftMoEs. Indeed, we demonstrate that even with an appropriately scaled single expert, we are able to maintain the performance gains, largely thanks to tokenization.
Selective Unlearning via Representation Erasure Using Domain Adversarial Training
Eleni Triantafillou
James J. Clark
Daniel M. Roy
Assessing SAM for Tree Crown Instance Segmentation from Drone Imagery
Many-Shot In-Context Learning
Avi Singh
Lei M Zhang
Bernd Bohnet
Stephanie C.Y. Chan
Luis Rosias
Biao Zhang
Zaheer Abbas
Azade Nova
John D Co-Reyes
Eric Chu
Feryal Behbahani
Aleksandra Faust
Large language models (LLMs) excel at few-shot in-context learning (ICL) -- learning from a few examples provided in context at inference, w… (voir plus)ithout any weight updates. Newly expanded context windows allow us to investigate ICL with hundreds or thousands of examples – the many-shot regime. Going from few-shot to many-shot, we observe significant performance gains across a wide variety of generative and discriminative tasks. While promising, many-shot ICL can be bottlenecked by the available amount of human-generated outputs. To mitigate this limitation, we explore two new settings: (1) "Reinforced ICL" that uses model-generated chain-of-thought rationales in place of human rationales, and (2) "Unsupervised ICL" where we remove rationales from the prompt altogether, and prompts the model only with domain-specific inputs. We find that both Reinforced and Unsupervised ICL can be quite effective in the many-shot regime, particularly on complex reasoning tasks. We demonstrate that, unlike few-shot learning, many-shot learning is effective at overriding pretraining biases, can learn high-dimensional functions with numerical inputs, and performs comparably to supervised fine-tuning. Finally, we reveal the limitations of next-token prediction loss as an indicator of downstream ICL performance.
Optimisation of quantitative brain diffusion-relaxation MRI acquisition protocols with physics-informed machine learning.
Álvaro Planchuelo-Gómez
Maxime Descoteaux
Jana Hutter
Derek K. Jones
C. Tax
A density estimation perspective on learning from pairwise human preferences
Learning from human feedback (LHF) -- and in particular learning from pairwise preferences -- has recently become a crucial ingredient in tr… (voir plus)aining large language models (LLMs), and has been the subject of much research. Most recent works frame it as a reinforcement learning problem, where a reward function is learned from pairwise preference data and the LLM is treated as a policy which is adapted to maximize the rewards, often under additional regularization constraints. We propose an alternative interpretation which centers on the generative process for pairwise preferences and treats LHF as a density estimation problem. We provide theoretical and empirical results showing that for a family of generative processes defined via preference behavior distribution equations, training a reward function on pairwise preferences effectively models an annotator's implicit preference distribution. Finally, we discuss and present findings on"annotator misspecification"-- failure cases where wrong modeling assumptions are made about annotator behavior, resulting in poorly-adapted models -- suggesting that approaches that learn from pairwise human preferences could have trouble learning from a population of annotators with diverse viewpoints.
Consolidating Separate Degradations Model via Weights Fusion and Distillation
Dinesh Daultani
Real-world images prevalently contain different varieties of degradation, such as motion blur and luminance noise. Computer vision recogniti… (voir plus)on models trained on clean images perform poorly on degraded images. Previously, several works have explored how to perform image classification of degraded images while training a single model for each degradation. Nevertheless, it becomes challenging to host several degradation models for each degradation on limited hardware applications and to estimate degradation parameters correctly at the run-time. This work proposes a method for effectively combining several models trained separately on different degradations into a single model to classify images with different types of degradations. Our proposed method is four-fold: (1) train a base model on clean images, (2) fine-tune the base model in-dividually for all given image degradations, (3) perform a fusion of weights given the fine-tuned models for individual degradations, (4) perform fine-tuning on given task using distillation and cross-entropy loss. Our proposed method can outperform previous state-of-the-art methods of pretraining in out-of-distribution generalization based on degradations such as JPEG compression, salt-and-pepper noise, Gaussian blur, and additive white Gaussian noise by 2.5% on CIFAR-100 dataset and by 1.3% on CIFAR-10 dataset. Moreover, our proposed method can handle degra-dation used for training without any explicit information about degradation at the inference time. Code will be available at https://github.com/dineshdaultani/FusionDistill.
SatBird: Bird Species Distribution Modeling with Remote Sensing and Citizen Science Data
Mélisande Teng
Amna Elmustafa
Benjamin Akera
Hager Radi Abdelwahed
Neural Causal Structure Discovery from Interventions
Nan Rosemary Ke
Bernhard Schölkopf
Michael Curtis Mozer
Christopher Pal
Recent promising results have generated a surge of interest in continuous optimization methods for causal discovery from observational data.… (voir plus) However, there are theoretical limitations on the identifiability of underlying structures obtained solely from observational data. Interventional data, on the other hand, provides richer information about the underlying data-generating process. Nevertheless, extending and applying methods designed for observational data to include interventions is a challenging problem. To address this issue, we propose a general framework based on neural networks to develop models that incorporate both observational and interventional data. Notably, our method can handle the challenging and realistic scenario where the identity of the intervened upon variable is unknown. We evaluate our proposed approach in the context of graph recovery, both de novo and from a partially-known edge set. Our method achieves strong benchmark results on various structure learning tasks, including structure recovery of synthetic graphs as well as standard graphs from the Bayesian Network Repository.
Repository-Level Prompt Generation for Large Language Models of Code
Disha Shrivastava
Daniel Tarlow
With the success of large language models (LLMs) of code and their use as code assistants (e.g. Codex used in GitHub Copilot), techniques fo… (voir plus)r introducing domain-specific knowledge in the prompt design process become important. In this work, we propose a framework called Repo-Level Prompt Generator that learns to generate example-specific prompts using prompt proposals. The prompt proposals take context from the entire repository, thereby incorporating both the structure of the repository and the context from other relevant files (e.g. imports, parent class files). Our technique doesn't require any access to the weights of the LLM, making it applicable in cases where we only have black-box access to the LLM. We conduct experiments on the task of single-line code-autocompletion using code repositories taken from Google Code archives. We demonstrate that an oracle constructed from our prompt proposals gives a remarkably high relative improvement of 36% over Codex, showing the quality of these proposals. Further, we show that when we train a model to predict a prompt proposal, we can achieve significant performance gains over Codex and other baselines. We release our code, data, and trained checkpoints at: https://github.com/shrivastavadisha/repo_level_prompt_generation.
Bird Distribution Modelling using Remote Sensing and Citizen Science data
Mlisande Teng
Amna Elmustafa
Benjamin Akera