Yoshua Bengio

Biographie

*Pour toute demande média, veuillez écrire à medias@mila.quebec.

Pour plus d’information, contactez Cassidy MacNeil, adjointe principale et responsable des opérations cassidy.macneil@mila.quebec.

Reconnu comme une sommité mondiale en intelligence artificielle, Yoshua Bengio s’est surtout distingué par son rôle de pionnier en apprentissage profond, ce qui lui a valu le prix A. M. Turing 2018, le « prix Nobel de l’informatique », avec Geoffrey Hinton et Yann LeCun. Il est professeur titulaire à l’Université de Montréal, fondateur et conseiller scientifique de Mila – Institut québécois d’intelligence artificielle, et codirige en tant que senior fellow le programme Apprentissage automatique, apprentissage biologique de l'Institut canadien de recherches avancées (CIFAR). Il occupe également la fonction de conseiller spécial et directeur scientifique fondateur d’IVADO.

En 2018, il a été l’informaticien qui a recueilli le plus grand nombre de nouvelles citations au monde. En 2019, il s’est vu décerner le prestigieux prix Killam. Depuis 2022, il détient le plus grand facteur d’impact (h-index) en informatique à l’échelle mondiale. Il est fellow de la Royal Society de Londres et de la Société royale du Canada, et officier de l’Ordre du Canada.

Soucieux des répercussions sociales de l’IA et de l’objectif que l’IA bénéficie à tous, il a contribué activement à la Déclaration de Montréal pour un développement responsable de l’intelligence artificielle.

Étudiants actuels

Jamal Abou Haibeh

Collaborateur·rice alumni - McGill

Berkes Anaïs

Collaborateur·rice de recherche - Cambridge University

Rim Assouel

Doctorat - UdeM

Stefan Bauer

Visiteur de recherche indépendant

Shahana Chatterjee

Collaborateur·rice de recherche - N/A

Xiaoyin Chen

Doctorat - UdeM

Sanghyeok Choi

Collaborateur·rice de recherche - KAIST

Doctorat - UdeM

Collaborateur·rice alumni - UdeM

Desmond Elliott

Visiteur de recherche indépendant

Doctorat - UdeM

Doctorat - UdeM

Doctorat - UdeM

Doctorat

Doctorat - UdeM

Moksh Jain

Doctorat - UdeM

Doctorat - UdeM

Collaborateur·rice alumni - UdeM

Hyeonah Kim

Postdoctorat - UdeM

Postdoctorat - UdeM

Collaborateur·rice alumni

Collaborateur·rice alumni - UdeM

Cristian Dragos Manta

Doctorat - UdeM

Doctorat - UdeM

Visiteur de recherche indépendant - UdeM

Doctorat - UdeM

Ali Parviz

Collaborateur·rice de recherche - Ying Wu Coll of Computing

Lena Podina

Collaborateur·rice de recherche - University of Waterloo

Camille Rochefort-Boulanger

Nassim Rahaman

Collaborateur·rice alumni - Max-Planck-Institute for Intelligent Systems

Amine RAZIG

Collaborateur·rice de recherche - UdeM

Doctorat - UdeM

Postdoctorat - UdeM

Postdoctorat - UdeM

Doctorat - UdeM

Dragos Secrieru

Collaborateur·rice alumni - UdeM

Postdoctorat

Collaborateur·rice alumni - Polytechnique

Mélisande Astrid Crystal Teng

Doctorat - UdeM

Ivan Titov

Collaborateur·rice de recherche

Alex Tong

Collaborateur·rice alumni - UdeM

Collaborateur·rice alumni - UdeM

Doctorat - UdeM

Collaborateur·rice de recherche

Collaborateur·rice de recherche - UdeM

Doctorat - McGill

Doctorat - UdeM

Doctorat - UdeM

Harry Zhao

Collaborateur·rice alumni - McGill

Skipper : combiner l’abstraction spatiale et temporelle afin d’améliorer la généralisation

Billets de blogue

Generic thumbnail for Mila Blog articles.

22 février 2024

par

Mingde Harry Zhao

Safa Alver

Harm van Seijen

Romain Laroche

Doina Precup

Yoshua Bengio

Mise à l’échelle au service du raisonnement et de l’apprentissage automatique basé sur un modèle

Scaling in the service of reasoning & model-based ML

4 avril 2023

par

Yoshua Bengio

Edward J. Hu

Une collaboration entre Mila et Relation Therapeutics pour découvrir in vitro de nouvelles associations médicamenteuses synergiques

A collaboration between Mila and Relation Therapeutics to discover novel synergistic combinations of drugs in vitro

23 mars 2022

par

Paul Bertin

Jake P. Taylor-King

Yoshua Bengio

Les réseaux de flot génératifs

15 mars 2022

par

Yoshua Bengio

Publications

Bayesian Dynamic Causal Discovery

Alexander Tong

Lazar Atanackovic

Jason Hartford

Learning the causal structure of observable variables is a central focus for scientific discovery. Bayesian causal discovery methods tackle … (voir plus)this problem by learning a posterior over the set of admissible graphs that are equally likely given our priors and observations. Existing methods primarily consider observations from static systems and assume the underlying causal structure takes the form of a directed acyclic graph (DAG). In settings with dynamic feedback mechanisms that regulate the trajectories of individual variables, this acyclicity assumption fails unless we account for time. We treat causal discovery in the unrolled causal graph as a problem of sparse identification of a dynamical system. This imposes a natural temporal causal order between variables and captures cyclic feedback loops through time. Under this lens, we propose a new framework for Bayesian causal discovery for dynamical systems and present a novel generative flow network architecture (Dyn-GFN) tailored for this task. Dyn-GFN imposes an edge-wise sparse prior to sequentially build a k -sparse causal graph. Through evaluation on temporal data, our results show that the posterior learned with Dyn-GFN yields improved Bayes coverage of admissible causal structures relative to state of the art Bayesian causal discovery methods.

2022-11-29

NeurIPS.cc/2022/Workshop/CDS (poster)

Object-centric causal representation learning

Amin Mansouri

Jason Hartford

Kartik Ahuja

2022-11-06

NeurIPS.cc/2022/Workshop/NeurReps (poster)

Laurence Perreault-Levasseur

Posterior samples of source galaxies in strong gravitational lenses with score-based priors

Yashar Hezaveh

Inferring accurate posteriors for high-dimensional representations of the brightness of gravitationally-lensed sources is a major challenge,… (voir plus) in part due to the difficulties of accurately quantifying the priors. Here, we report the use of a score-based model to encode the prior for the inference of undistorted images of background galaxies. This model is trained on a set of high-resolution images of undistorted galaxies. By adding the likelihood score to the prior score and using a reverse-time stochastic differential equation solver, we obtain samples from the posterior. Our method produces independent posterior samples and models the data almost down to the noise level. We show how the balance between the likelihood and the prior meet our expectations in an experiment with out-of-distribution data.

2022-11-06

ArXiv (prépublication)

Bayesian learning of Causal Structure and Mechanisms with GFlowNets and Variational Bayes

Mizu Nishikawa-Toomey

Tristan Deleu

Jithendaraa Subramanian

Laurent Charlin

Bayesian causal structure learning aims to learn a posterior distribution over directed acyclic graphs (DAGs), and the mechanisms that defin… (voir plus)e the relationship between parent and child variables. By taking a Bayesian approach, it is possible to reason about the uncertainty of the causal model. The notion of modelling the uncertainty over models is particularly crucial for causal structure learning since the model could be unidentifiable when given only a finite amount of observational data. In this paper, we introduce a novel method to jointly learn the structure and mechanisms of the causal model using Variational Bayes, which we call Variational Bayes-DAG-GFlowNet (VBG). We extend the method of Bayesian causal structure learning using GFlowNets to learn not only the posterior distribution over the structure, but also the parameters of a linear-Gaussian model. Our results on simulated data suggest that VBG is competitive against several baselines in modelling the posterior over DAGs and mechanisms, while offering several advantages over existing methods, including the guarantee to sample acyclic graphs, and the flexibility to generalize to non-linear causal mechanisms.

2022-11-03

ArXiv (prépublication)

A General-Purpose Neural Architecture for Geospatial Systems

Nasim Rahaman

Martin Weiss

Frederik Träuble

Francesco Locatello

Alexandre Lacoste

Christopher Pal

Li Erran Li

Bernhard Schölkopf

2022-11-01

OpenReview.net/Anonymous_Preprint (inconnu)

Consistent Training via Energy-Based GFlowNets for Modeling Discrete Joint Distributions

Chanakya Ajit Ekbote

Moksh J. Jain

Payel Das

Generative Flow Networks (GFlowNets) have demonstrated significant performance improvements for generating diverse discrete objects …

2022-10-31

ArXiv (prépublication)

Discrete Factorial Representations as an Abstraction for Goal Conditioned Reinforcement Learning

Hongyu Zang

Xin Li

Romain Laroche

Remi Tachet des Combes

2022-10-31

ArXiv (prépublication)

Controlled Sparsity via Constrained Optimization or: How I Learned to Stop Tuning Penalties and Love Constraints

Jose Gallego-Posada

The performance of trained neural networks is robust to harsh levels of pruning. Coupled with the ever-growing size of deep learning models,… (voir plus) this observation has motivated extensive research on learning sparse models. In this work, we focus on the task of controlling the level of sparsity when performing sparse learning. Existing methods based on sparsity-inducing penalties involve expensive trial-and-error tuning of the penalty factor, thus lacking direct control of the resulting model sparsity. In response, we adopt a constrained formulation: using the gate mechanism proposed by Louizos et al. (2018), we formulate a constrained optimization problem where sparsification is guided by the training objective and the desired sparsity target in an end-to-end fashion. Experiments on CIFAR-{10, 100}, TinyImageNet, and ImageNet using WideResNet and ResNet{18, 50} models validate the effectiveness of our proposal and demonstrate that we can reliably achieve pre-determined sparsity targets without compromising on predictive performance.

2022-10-30

NeurIPS.cc/2022/Conference (accepté)

MAgNet: Mesh Agnostic Neural PDE Solver

The computational complexity of classical numerical methods for solving Partial Differential Equations (PDE) scales significantly as the res… (voir plus)olution increases. As an important example, climate predictions require fine spatio-temporal resolutions to resolve all turbulent scales in the fluid simulations. This makes the task of accurately resolving these scales computationally out of reach even with modern supercomputers. As a result, current numerical modelers solve PDEs on grids that are too coarse (3km to 200km on each side), which hinders the accuracy and usefulness of the predictions. In this paper, we leverage the recent advances in Implicit Neural Representations (INR) to design a novel architecture that predicts the spatially continuous solution of a PDE given a spatial position query. By augmenting coordinate-based architectures with Graph Neural Networks (GNN), we enable zero-shot generalization to new non-uniform meshes and long-term predictions up to 250 frames ahead that are physically consistent. Our Mesh Agnostic Neural PDE Solver (MAgNet) is able to make accurate predictions across a variety of PDE simulation datasets and compares favorably with existing baselines. Moreover, MAgNet generalizes well to different meshes and resolutions up to four times those trained on.

2022-10-30

NeurIPS.cc/2022/Conference (accepté)

Efficient Queries Transformer Neural Processes

Leo Feng

Hossein Hajimirsadeghi

Mohamed Osama Ahmed

2022-10-20

NeurIPS.cc/2022/Workshop/MetaLearn (poster)

FL Games: A federated learning framework for distribution shifts

Niladri Chatterjee

Federated learning aims to train predictive models for data that is distributed across clients, under the orchestration of a server. However… (voir plus), participating clients typically each hold data from a different distribution, whereby predictive models with strong in-distribution generalization can fail catastrophically on unseen domains. In this work, we argue that in order to generalize better across non-i.i.d. clients, it is imperative to only learn correlations that are stable and invariant across domains. We propose FL Games, a game-theoretic framework for federated learning for learning causal features that are invariant across clients. While training to achieve the Nash equilibrium, the traditional best response strategy suffers from high-frequency oscillations. We demonstrate that FL Games effectively resolves this challenge and exhibits smooth performance curves. Further, FL Games scales well in the number of clients, requires significantly fewer communication rounds, and is agnostic to device heterogeneity. Through empirical evaluation, we demonstrate that FL Games achieves high out-of-distribution performance on various benchmarks.

2022-10-19

NeurIPS.cc/2022/Workshop/Federated_Learning (présentation orale)

Inductive biases for deep learning of higher-level cognition

Anirudh Goyal

2022-10-11

Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences (publié)