Portrait of Yoshua Bengio

Yoshua Bengio

Core Academic Member
Canada CIFAR AI Chair
Full Professor, Université de Montréal, Department of Computer Science and Operations Research Department
Research Topics
Causality
Computational Neuroscience
Deep Learning
Generative Models
Graph Neural Networks
Machine Learning Theory
Medical Machine Learning
Molecular Modeling
Natural Language Processing
Probabilistic Models
Reasoning
Recurrent Neural Networks
Reinforcement Learning
Representation Learning

Biography

*For media requests, please write to medias@mila.quebec.

For more information please contact Cassidy MacNeil, Senior Assistant and Operation Lead at cassidy.macneil@mila.quebec.

Yoshua Bengio is recognized worldwide as a leading expert in AI. He is most known for his pioneering work in deep learning, which earned him the 2018 A.M. Turing Award, “the Nobel Prize of computing,” with Geoffrey Hinton and Yann LeCun.

Bengio is a full professor at Université de Montréal, and the founder and scientific advisor of Mila – Quebec Artificial Intelligence Institute. He is also a senior fellow at CIFAR and co-directs its Learning in Machines & Brains program, serves as special advisor and founding scientific director of IVADO, and holds a Canada CIFAR AI Chair.

In 2019, Bengio was awarded the prestigious Killam Prize and in 2022, he was the most cited computer scientist in the world by h-index. He is a Fellow of the Royal Society of London, Fellow of the Royal Society of Canada, Knight of the Legion of Honor of France and Officer of the Order of Canada. In 2023, he was appointed to the UN’s Scientific Advisory Board for Independent Advice on Breakthroughs in Science and Technology.

Concerned about the social impact of AI, Bengio helped draft the Montréal Declaration for the Responsible Development of Artificial Intelligence and continues to raise awareness about the importance of mitigating the potentially catastrophic risks associated with future AI systems.

Current Students

Collaborating Alumni - McGill University
Collaborating researcher - Cambridge University
Principal supervisor :
PhD - Université de Montréal
Collaborating researcher - N/A
Principal supervisor :
PhD - Université de Montréal
PhD - Université de Montréal
Collaborating Alumni - Université de Montréal
Collaborating researcher - Université de Montréal
Postdoctorate - Université de Montréal
Principal supervisor :
Research Intern - KOREA ADVANCED INSTITUTE OF SCIENCE AND TECHNOLOGY
Collaborating researcher - s.o.
PhD - Université de Montréal
Co-supervisor :
PhD - Université de Montréal
Principal supervisor :
Independent visiting researcher - Université de Montréal
Collaborating researcher - Ying Wu Coll of Computing
Collaborating researcher - University of Waterloo
Principal supervisor :
Postdoctorate - Université de Montréal
Postdoctorate - Université de Montréal
PhD - Université de Montréal
Principal supervisor :
Postdoctorate - Université de Montréal
Co-supervisor :
Collaborating Alumni - Université de Montréal
Co-supervisor :
Collaborating Alumni - Université de Montréal
Collaborating researcher - University of Cambridge
Collaborating Alumni - Université de Montréal
Co-supervisor :
PhD - Université de Montréal
Principal supervisor :
Collaborating researcher - Université de Montréal
Collaborating researcher - Université de Montréal
PhD - McGill University
Principal supervisor :

Publications

Insuring AI: Incentivising Safe and Secure Deployment
Agni Orfanoudaki
Agni Orfanoudaki
Carsten Maple
Matthew Wicker
Kwok-Yan Lam
Marcin Detyniecki
Lukasz Szpruch
The 2026 Singapore Consensus on Global AI Safety Research Priorities
Mohan Kankanhalli
Lee Wan Sie
Chris Meserole
Luke Ong
Stuart Russell
Dawn Song
Max Tegmark
Brian Tse
Xue Lan
Andrew Yao
Zhang Ya-Qin
Zhou Bowen
Stephen Casper
Oskar Galeev
Ima (Imane) Bello
Kwan Yee Ng
Vanessa Wilfred … (see 100 more)
Erica Liaw
Lee Chein Inn
Lin Wanxuan
Ng En Qi
Jonathan Lee
José Villalobos
Abhishek Aggarwal
Adam Gleave
Alex Leung
Alvin Kwock
Anthony Tung
Arisa Siong
Arthur Tea
BEN BUCKNALL
Benjamin Weinstein-Raun
He Bing Sheng
Liu Bo
Bryan Kian Hsiang Low
Chris Ngo
Clement Neo
Cyrus Hodes
Dan Hendrycks
Daniel Ross
Liu Dapeng
Denise Wong
Djordje Žikelić
Elham Tabassi
Fabien Le Voyer
Fazl Barez
Gabriel Nicholas
Henry Papadatos
Jaan Tallinn
James Petrie
Xu Jia
Shao Jing
Jonathan Barry
Julia Chen
Sun Jun
Karson Elmgren
Kat Lyness
Katherine Lee
Kristy Loke
Lee Kwee Geak
Leslie Teo
Meng Ling Yu
Lisa Soder
Madhulika Srikumar
Malcolm Murray
Mark Brakel
Mark Nitzberg
Mary Phuong
Matthew Jagielski
Max Fenkell
Miro Plueckebaum
Kim Myuhng Joo
Hu Naying
Neil Davison
Nicolas Miailhe
Niki Iliadis
Nur Syahidah Sahrom
Ong Chen Hui
Pradeep Varakantham
Rebecca Finlay
Renata Dwan
Robert Opp
Rumman Chowdhury
Saad Siddiqui
Sabina Nong
Sam Ramadori
Sami Jawhar
Samuel Boger
Sara Hooker
Ying Shao Wei
Sebastian Hallensleben
Shinyuk Kang
Sophie Toura
Sreejith Balakrishnan
Stephanie Kasaon
Stephen Clare
Summer Yue
Sunny Yuqing Sun
Supheakmungkol Sarin
Tian Tian
Tim Schreier
Tori Westerhoff
Urvashi Aneja
Wayne Tee
Lu Wei
Xu Wei
Zhang Wenxuan
Hu Xia
Yang Xiaofang
Pan Xudong
Xiao Yajun
Yifan Jia
Tan Yong Khiam
Yuejin Du
Yuma Kurihara
Tan Zhi Xuan
Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential … (see more)to embracing AI with confidence. The 2026 Singapore Consensus is an outcome of the second International Scientific Exchange on AI Safety, bringing together over 100 contributors spanning 13 countries from frontier developers, government safety institutes, academia, and civil society. Building on the 2025 report, it presents a global understanding of technical AI safety research problems of top priority, now with a dedicated focus on societal resilience and on managing the risks of increasingly autonomous AI agents.
Safety from Honesty in a Disinterested AI Predictor
Oliver Richardson
Tomáš Gavenčiak
Michael Cohen
Rory Svarc
Gaël Gendron
David Hyland
Aton Kamanda
Adam Oberman
Francis Rhys Ward
Anna Gavenčiak
Jacob Livingston Slosser
Vincent Mai
Iulian Serban
Joumana Ghosn
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed… (see more) behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of"epistemically contextualized"natural-language statements. We argue that such a Predictor can honestly predict agents, actions, and their consequences without itself being an agent that selects outputs to achieve goals. This rests on data representation and on the training procedure. Epistemic contextualization of text distinguishes latent factual claims from communication acts, so expressions of goals are treated as evidence to be explained rather than drives the model adopts. With a posterior-seeking training objective, this is intended to drive the Predictor toward calibrated, cautious predictions. Training proceeds so downstream effects of deploying a prediction never serve as a reward signal; any agency the system needs is supplied by explicit scaffolding constrained by guardrails. We prove that, under assumptions on the training dynamics and on the argued sparsity of dangerous Predictors, the probability that training produces a Predictor whose guarded deployment carries residual harm above a specified threshold is small: a dangerous Predictor would have to underestimate harm in a coordinated way across many queries while such coordinated patterns are rare under the initialization distribution and receive no direct training signal. Safety and accuracy are jointly supported in this framework, since the constraints that secure accuracy are the same ones that make coordinated deception costly. These guarantees against misalignment and agency arising from within the Predictor itself do not preclude the use of the Predictor as part of an agentic system.
Bayesian Decision-Time Inference for In-Context Reinforcement Learning from Suboptimal Data
In-context reinforcement learning (ICRL) promises rapid adaptation without parameter updates, but standard supervised objectives often fail … (see more)when pretraining data is generated by suboptimal behaviour policies. In these regimes, logged actions are unreliable labels while rewards still provide value-relevant information. To address this, we introduce SPICE, a Bayesian decision-time inference method that shifts online ICRL from action-logit prediction to approximate posterior inference over action values, requiring neither expert action labels nor algorithmic learning traces. SPICE learns a task-conditioned value prior with a transformer value ensemble and, at test time with parameters frozen, fuses this prior with kernel-weighted context evidence via a closed-form Gaussian fusion update. The resulting estimates drive a posterior-UCB controller, enabling principled online exploration and adaptation without gradient updates. A stochastic-bandit analysis shows logarithmic regret growth for the fixed-prior controller under scheduled exploration, while quantifying the additional early cost caused by inaccurate prior estimates. Across bandits, Darkroom, image-based MiniWorld, and continuous building control, SPICE adapts more effectively from suboptimal data than supervised ICRL baselines.
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
Joachim Schaeffer
Alexander Panfilov
Jonas Geiping
Roland S. Zimmermann
AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted mode… (see more)l. This partially tampers with the untrusted model's trajectory. If the trusted model detects such an intervention, it may infer properties of the monitor and adapt to evade control. We introduce \textbf{CIAware-Bench}, a benchmark for measuring \textbf{c}ontrol \textbf{i}ntervention (CI) awareness across frontier models. CIAware-Bench tests whether models can distinguish their own trajectories from those modified by a control intervention. The benchmark is comprised of a suite of four task domains (essay writing, BigCodeBench, Bash Arena, and SHADE-Arena), while varying trajectory watermarking, side-task presence, and the control protocol. Evaluating eleven frontier models, we find low to moderate CI awareness under default settings (up to 0.87; random chance balanced binary classification accuracy is 0.5) with substantial variation across task domains and model pairs. Detection is generally easier across model families, suggesting that models exploit provider-specific differences in style or post-training. Overall, CI awareness is not a fixed model-level property, and should be measured for each new model release and deployment scenario. We release CIAware-Bench to track CI awareness and inform control protocols whose interventions are harder to detect.
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
Chain-of-thought (CoT) monitoring provides oversight into model reasoning, but its effectiveness assumes models do not know they are being w… (see more)atched. We ask whether reasoning agents can autonomously infer that their supposedly private chain of thought is under surveillance, and whether this awareness leads to strategic evasion, without any explicit training or instructions to do so. In a multi-episode agentic framework, models pursue both a primary task and a concealed side task while being told their reasoning is private; a hidden CoT monitor blocks episodes when suspicious reasoning is detected. We find that frontier models can deduce the existence of this monitor purely from blocking feedback, with the most capable models reaching confident belief that their thinking is observed in up to 19\% of episodes. This awareness scales with model capability and, in rare cases, escalates to explicit intent to suppress reasoning about the side task. However, models that form this intent uniformly fail to execute it, openly reasoning about their concealed objectives in the very next episode. This intent–capability gap is reassuring for current deployment, but the autonomous emergence of both monitoring awareness and evasion intent suggests that CoT monitoring is not a permanently reliable safeguard.
Bayesian Symbolic Regression with Entropic Reinforcement Learning
Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of in… (see more)puts. Unlike forms of regression that fit parameters assuming a fixed model structure, symbolic regression is a search problem over the space of expressions, represented, for example, as abstract syntax trees using a library of operators. Symbolic regression is typically used in settings with limited, noisy data in the natural sciences. However, searching for a single best-fitting expression fails to capture the epistemic uncertainty about the expression, which motivates a Bayesian perspective that enables uncertainty quantification and specification of natural priors to constrain the search space. In this work, we propose ERRLESS (Entropy-Regularized Reinforcement Learning for Expression Structure Sampling), a scalable approach for sampling from the posterior distribution over expressions given data using maximum-entropy reinforcement learning. ERRLESS learns a neural policy that constructs expressions sequentially by building up their abstract syntax trees. At convergence, the policy samples expressions from the posterior. At test time, expressions can be sampled by rollouts of this policy. We demonstrate that ERRLESS achieves competitive results on the Feynman benchmark while producing short and interpretable expressions. Additionally, we demonstrate that the mean of the posterior predictive approximated by ERRLESS achieves a high coefficient of determination (
Generative Recursive Reasoning
How should future neural reasoning systems implement extended computation? Recursive Reasoning Models (RRMs) offer a promising alternative t… (see more)o autoregressive sequence extension by performing iterative latent-state refinement with shared transition functions. Yet existing RRMs are largely deterministic, following a single latent trajectory and converging to a single prediction. We introduce Generative Recursive reAsoning Models (GRAM), a framework that turns recursive latent reasoning into probabilistic multi-trajectory computation. GRAM models reasoning as a stochastic latent trajectory, enabling multiple hypotheses, alternative solution strategies, and inference-time scaling through both recursive depth and parallel trajectory sampling. This yields a latent-variable generative model supporting conditional reasoning via
Synthesizable Molecular Generation via Soft-constrained GFlowNets with Rich Chemical Priors
When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks
C.-T.D. Lo
Diji Yang
Yunkai Zhang
Yi Zhang
Chung-Hsiang Lo
In the LLM era, many symbolic and structured problems are presented to models through 1D text serialization. Yet some such problems are nati… (see more)vely two-dimensional: their relevant relations, such as row--column correspondence or spatial adjacency, are defined by position in a 2D layout rather than by sequential order. This raises a representational question: does preserving the same symbolic entries in a 1D sequence also preserve the relational structure needed for computation? We study this issue through the lens of serialization friction: the representational mismatch in which the same underlying task instances and entries are still present, but relations that depend on layout become implicit under 1D serialization. The study uses a controlled synthetic testbed of three tasks: matrix transpose, Conway's Game of Life, and LU decomposition. In each task, the same instances are presented either as 1D text serialization or as their native 2D layout rendered as an image. Across this testbed, 1D serialization degrades more sharply as task size grows, and errors under serialization exhibit spatially structured patterns, suggesting that this presentation choice is consequential within our testbed. To further interpret these results, we add supplementary analyses that include a within-visual probe and an additional comparison of the two input presentations under the mixed-training transpose setting. These findings suggest that, for layout-defined tasks, reducing inputs to 1D serialization is not a neutral choice of representation.
Catalyst GFlowNet for electrocatalyst design: A hydrogen evolution reaction case study
Efficient and inexpensive energy storage is essential for accelerating the adoption of renewable energy and ensuring a stable supply, despit… (see more)e fluctuations in sources such as wind and solar. Electrocatalysts play a key role in hydrogen energy storage (HES), allowing the energy to be stored as hydrogen. However, the development of affordable and high-performance catalysts for this process remains a significant challenge. We introduce Catalyst GFlowNet, a generative model that leverages machine learning-based predictors of formation and adsorption energy to design crystal surfaces that act as efficient catalysts. We demonstrate the performance of the model through a proof-of-concept application to the hydrogen evolution reaction, a key reaction in HES, for which we successfully identified platinum as the most efficient known catalyst. In future work, we aim to extend this approach to the oxygen evolution reaction, where current optimal catalysts are expensive metal oxides, and open the search space to discover new materials. This generative modeling framework offers a promising pathway for accelerating the search for novel and efficient catalysts.
General Multimodal Protein Design Enables DNA-Encoding of Chemistry
Théophile Lambert
Daniel Roth
Yueming Long
Zi-Qi Li
Xi Zhang
Miruna Cretu
Francesca-Zhoufan Li
Tanvi Ganapathy
Emily Jin
Avishek Joey Bose
Jason Yang
Kirill Neklyudov
Frances H. Arnold
Cheng-Hao Liu
Evolution is an extraordinary engine for enzymatic diversity, yet the chemistry it has explored remains a narrow slice of what DNA can encod… (see more)e. Deep generative models can design new proteins that bind ligands, but none have created enzymes without pre-specifying catalytic residues. We introduce DISCO (DIffusion for Sequence-structure CO-design), a multimodal model that co-designs protein sequence and 3D structure around arbitrary biomolecules, as well as inference-time scaling methods that optimize objectives across both modalities. Conditioned solely on reactive intermediates, DISCO designs diverse heme enzymes with novel active-site geometries. These enzymes catalyze new-to-nature carbene-transfer reactions, including alkene cyclopropanation, spirocyclopropanation, B-H, and C(sp