Offered by Mila and the Public Policy Forum, this program is designed to equip policy and decision makers with the tools to navigate the opportunities and risks of AI. The next cohort will be held in French on September 1-2, 2026, at Mila.
The Mila AI Policy Fellowship translates deep AI expertise into rigorous, public-interest policy. Read the newest publication Bridging the Expertise Gap: Knowledge Transfer Mechanisms for AI Regulation by Moritz von Knebel
This program supports AI startups at any time of the year. Benefit from cutting-edge resources and tailored support to accelerate your technology's development.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Publications
<scp>1D</scp> Pre‐Acquisition Navigator Correcting Respiratory‐Induced Field Fluctuations in Multi‐Echo Gradient‐Echo Imaging of the Thoracic Spinal Cord
Alicia E. Cronin
Alexandre D’Astous
Nathan Williams
Antoine Guénette
Aimee Salakhov
Seth Stubblefield
Colin D. Mcknight
Lipika Narisetti
Subramaniam Sriram
Seth A. Smith
Ryan K. Robison
Guillaume Gilbert
Julien Cohen‐Adad
Kristin P. O’Grady
PURPOSE: In the spinal cord (SC), multi-echo gradient echo (ME-GRE) increases gray (GM) and white matter (WM) contrast and improves sensitiv… (see more)ity to lesions in people with multiple sclerosis (pwMS). However, SC ME-GRE is susceptible to breathing-induced field fluctuations, causing ghosting artifacts and signal loss. Recent work introduced a 1D phase navigator following the last echo to measure field variations; however, susceptibility to phase wrapping increases at longer echo times. We propose a 1D phase navigator preceding the first echo, reducing phase accumulation and eliminating the need for respiratory monitoring. METHODS: ME-GRE data covering the lower (T9-T12 vertebrae) and upper (T4-T8 vertebrae) thoracic SC were acquired in 20 healthy volunteers and 3 pwMS at 3T. Standard and navigator-corrected images were acquired in the same acquisition. To evaluate image quality, WM and GM signal-to-noise ratio (SNR), WM/GM contrast-to-noise ratio (CNR), and background ghosting signals were measured and compared between the two reconstructions. Both were blindly assessed for artifacts, structural delineation, and diagnostic confidence in pwMS. RESULTS: Navigator correction significantly increased GM and WM SNR and CNR, reduced posterior ghosting across both thoracic regions, and significantly reduced artifacts while increasing structural delineation. Preliminary evaluation in three pwMS showed consistent improvements in artifact mitigation, structural delineation, and lesion conspicuity with navigator correction, providing proof-of-concept for potential clinical application. CONCLUSION: A 1D navigator prior to the first echo reduces ghosting and improves thoracic SC image quality without respiratory monitoring. This approach could improve the diagnostic value and enhance the reliability of thoracic SC ME-GRE.
Traditional recommendation systems rely on latent (dense) representations, making them difficult to interpret and control. We propose the Co… (see more)ntrollable and Content-Based Recommendations (CCBR) framework, which builds its recommendations from textual user profile representations. CCBR plugs into collaborative filtering models and introduces controllability via text bottlenecks. We show that CCBR enables text-based and multimodal interventions, allowing users to steer the model towards the directions they prefer. Different from existing controllable recommendation systems, CCBR infers the text summaries directly from item contents (images, audio or video). Across image-, audio-, and video-based datasets, we demonstrate that the proposed framework obtains competitive model performance with standard (latent-representation) models while providing controllable model summaries via text. The model also outperforms TEARS, a recent baseline for controllable recommendation systems. Through systematic interventions, we demonstrate the efficacy of the user steering mechanism.
OBJECTIVES: Prefrontal regions are implicated in explore-exploit decision-making during foraging. Older adults often show an exploitation bi… (see more)as, and this age period is also marked by deteriorating prefrontal myelination. To investigate whether these phenomena are linked, we examined whether lower magnetization transfer saturation (MTsat), a myelin-sensitive quantitative MRI (qMRI) measure, in these regions predicts greater exploitation bias during foraging, and whether cortical microstructure is a better predictor of bias than macrostructure (i.e., cortical thickness). METHODS: Cognitively healthy older adults with familial risk of Alzheimer's disease (AD) (N=118, 60-88 years) completed a foraging task indexing explore-exploit decision-making. qMRI was used to derive MTsat values for the frontopolar cortex (FPC), medial orbitofrontal cortex (OFC), rostral middle frontal gyrus (rMFG), dorsal anterior cingulate cortex (dACC), as well as the locus coeruleus (LC), a core subcortical region strongly implicated in explore-exploit decision-making. Secondary analyses examined associations between available AD risk markers and foraging. RESULTS: Lower MTsat in the FPC, OFC, rMFG, and LC was associated with an exploitation bias, with LC and FPC emerging as the strongest predictors. No relationship was observed for the dACC. MTsat remained a significant predictor of foraging after controlling for cortical thickness. Observed associations were largely unrelated to AD risk markers. DISCUSSION: Individual differences in cortical microstructural integrity within a well-defined explore-exploit circuit are associated with an exploitative decision-making bias in older adults. These findings highlight the value of qMRI microstructural integrity markers, beyond standard macrostructural assays, in characterizing the neural correlates of exploitation biases in later life.
2026-07-22
Journals of Gerontology Series B: Psychological Sciences and Social Sciences (published)
Protein language models have been increasingly successful on tasks ranging from fitness prediction to functional design, yet what biological… (see more) knowledge they acquire and where it is encoded within their internal representations remain underexplored. Through a high-resolution layer-by-layer interpretability analysis of 8 models from the ESM2 and AMPLIFY families on 22 concepts from human proteome annotations, we found that these models encode concepts of increasing levels of complexity along their depth: basic physicochemical properties and linear motifs are best captured by early-layer embeddings, secondary structure from subsequent layers, and domain-level semantics from middle layers. Principal component projections of these embeddings showed that they separate biologically meaningful protein groupings, and molecular-biology-inspired interventions demonstrated that pLM embeddings can discriminate phosphomimic-active from inactive mutants. Perhaps surprisingly, we observed that pretraining data and compute had a greater impact on the linear emergence of biological concepts than scaling up parameters. By revealing where biological knowledge is captured in pLMs and which choices shape its emergence, our work offers insights to develop more robust, biologically grounded protein language models.
Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement le… (see more)arning (RL), this trade-off is typically controlled through a fixed discount factor, which imposes a single exponentially discounted temporal horizon. However, biological agents exhibit flexible and adaptive temporal discounting, suggesting that effective planning requires multiple timescales. Here, we propose a multi-horizon approach that adaptively selects and combines temporal horizons, enabling robust adaptation to changes in reward structure without manual discount-factor tuning. This flexibility makes the method particularly suitable for continual learning scenarios involving task switches and varying environmental configurations. Empirically, we demonstrate that our approach identifies effective discount factors across a range of MiniGrid environments, including continual settings composed of three sequentially changing tasks. These results suggest that adaptive temporal discounting can improve parameter efficiency and enhance adaptability in both artificial and biologically inspired learning systems.
Temporally extended exploration via graph Laplacian-based options is a promising approach to sparse-reward reinforcement learning (RL), but … (see more)existing methods either do not explicitly target novelty or fail to scale to pixel-based domains under function approximation. Novel Exploration via Orthogonality (NEO) addresses the first issue by constructing options that navigate from highly visited regions toward less visited ones, yet prior results were limited to settings where exact eigenvectors can be computed. We present a scalable extension of NEO to pixel-based domains, built on three contributions. First, we use a novelty-weighted continuous Laplacian graph-drawing objective, which enables RL with continuous observations. Second, we embed the resulting eigen-potential options within a hierarchical reinforcement learning framework, enabling coherent temporally extended behavior. Third, we observe that learned eigen-potential rewards are directional but locally unreliable under online approximation; we therefore augment each option reward with a novelty bonus, a novel design idea that proves essential for stabilizing option learning while preserving novelty-directed exploration. Together, these contributions yield stronger and more persistent exploration, enabling longer option rollouts and better access to hard-to-reach novel states. Empirically, our method significantly outperforms both the prior scalable Laplacian-option baseline and a direct extension of NEO on sparse-reward benchmarks under a fixed budget. On Montezuma's Revenge, our best variant achieves approximately 1.8x higher return than both baselines. On Venture, both baselines yield returns near zero, whereas our method achieves a return of 1135. Across seven hard ProcGen games, our method achieves approximately 3.5x and 5.6x higher aggregate normalized return than the two baselines, respectively.
Training from offline data has allowed for substantial progress in domains such as robotics, leading to general-purpose policies that can be… (see more) easily applied zero-shot or efficiently finetuned for downstream tasks. However, training policies on offline data can lead to poor generalization, both due to the choice of modeling objective and from learning from a static dataset. In this work, we focus on the challenging task of zero-shot goal generalization, where a policy is evaluated on unseen tasks that require reusing its existing knowledge (compositional generalization). An avenue for improving a policy's generalization is through generating new experience through world models; however, such generation has proven difficult for longer horizons. Thus, to alleviate this issue, we propose TD-Aug, sampling from a geometric horizon model, which allows for directly imagining novel outcomes that can be achieved through composing existing knowledge. We demonstrate that training on these future outcomes as goals for goal-conditioned BC and offline RL policies improves generalization in stitching-based OGBench tasks.
Physical AI agents acting in the real world fail differently from language models: a misjudgment of trajectory, force, or contact can have i… (see more)mmediate and potentially irreversible consequences. As video world models are increasingly deployed as the perception and dynamics layer of such agents, understanding what physical structure they actually represent internally be- comes a precondition for trustworthy deployment. Yet today these models are largely studied as black boxes, evaluated only through behavioral benchmarks. We argue that this should change: world models can be opened up, and the latent variables they encode can be measured, interpreted, and eventually controlled. We take a first step in this direction across two state-of-the-art video encoders (V-JEPA 2 and VideoMAE-v2), using layerwise probing, subspace geometry, patch-level decoding, and targeted attention ablations. We identify a sharp intermediate-depth transition we call the Physics Emergence Zone, at which physical variables become accessible. Decomposing motion into explicit variables, we find that scalar quantities such as speed and acceleration are available from early layers onwards, whereas motion direction becomes accessible only at the Physics Emergence Zone, encoded as a high-dimensional circular population code that requires coordinated multi-feature intervention to steer. These findings argue against compact, reusable latent physics state and in favor of distributed, task-specific representations supported by a shared local-attention circuit. We discuss implications for safe deploy- ment of video world models in embodied physical AI, including why intermediate-layer features, multi-feature monitors, and the local-attention circuit at the Physics Emergence Zone are natural targets for runtime verification.
2026-07-19
SPAI @ International Joint Conference on Artificial Intelligence - European Conference on Artificial Intelligence (published)
We express an interest to Equivalent Numeric Approved Identity of Passport and Immigration Papers. [We are making a submission to the Journa… (see more)l of Banking and Financial Technology].
A point prediction that is well calibrated on average can still be systematically biased conditional on its own value, undermining its use i… (see more)n downstream decision-making. We consider two objectives for reliable uncertainty quantification: self-calibration, requiring a point prediction to be unbiased conditional on its own value, and prediction-conditional validity, requiring a prediction interval to attain nominal coverage conditional on the prediction. Self-Calibrating Conformal Prediction (SC-CP) attains both objectives exactly in finite samples, but requires refitting its calibrator for every candidate outcome, which is computationally prohibitive for continuous outcomes. We propose Isotonic Conformal Prediction (ICP), a framework that decouples calibration from prediction-set construction by fitting a single isotonic recalibration map and constructing prediction intervals within strata of similar recalibrated predictions. Within this framework we develop two procedures. Split Isotonic Conformal Prediction (SICP) attains prediction-conditional validity in finite samples and self-calibration asymptotically, at the computational cost of split conformal prediction. Transductive Isotonic Conformal Prediction (TICP) attains both objectives exactly in finite samples through a per-test-point inner loop that avoids refitting the isotonic calibrator. On synthetic heteroscedastic regression problems and a real-world healthcare-utilization dataset, both procedures match the coverage of SC-CP at substantially lower computational cost.