This program supports AI startups at any time of the year. Benefit from cutting-edge resources and tailored support to accelerate your technology's development.
Offered by Mila and the Public Policy Forum, this program is designed to equip policy and decision makers with the tools to navigate the opportunities and risks of AI. The next cohort will be held in French on September 1-2, 2026, at Mila.
Connect with a Mila academic advisor and current student-researchers to learn more about Mila's community and how to join us on August 19, 31 and September 11, 2026.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Publications
Silent Sabotage: Injecting Backdoors into AI Agents Through Fine-Tuning
The rise of AI agents that can use tools, browse the web and interact with computers on behalf of a user, has sparked strong interest in imp… (see more)roving these capabilities by explicitly fine-tuning the LLMs/VLMs that power these agents. Several researchers have proposed collecting data by letting the agents interact with their environment (e.g., a computer operating system, the web or a collection of APIs exposed as tools), and improve agent performance by fine tuning on this data. In this work, we show that such data collection can be manipulated by adversaries to insert poisoned traces. By modifying just 5% of collected traces, adversaries can embed stealthy bad behaviors into agents—like leaking confidential user information whenever the tool or webpage exposes a trigger. Our results raise important security concerns in the development of AI agents, and underscore the importance of careful scrutiny of all data collection processes used to improve agentic AI.
Foundation Models (FMs) have dramatically increased the potential and power of deep learning algorithms through general capacities over a va… (see more)riety of tasks. The performance increase they offer is obtained without elaborated specific trainings for domains such as natural language processing and computer vision. However, their application in specialized fields like biomedical imaging and fluorescence microscopy remains difficult due to distribution shifts and the scarcity of high-quality annotated datasets. The high cost of data acquisition and the requirement for in-domain expertise further exacerbate this challenge in microscopy. To address this we introduce STED-FM, a foundation model specifically designed for super-resolution STimulated Emission Depletion (STED) microscopy. STED-FM leverages a Vision Transformer architecture trained at scale with Masked Autoencoding on a new dataset of nearly one million STED images. STED-FM learns expressive latent representations without requiring extensive annotations, yielding robust performance across diverse downstream microscopy image analysis tasks. Unsupervised experiments demonstrate the discriminative structure of its learned latent space. These representations can be leveraged for multiple downstream applications, including fully supervised classification and segmentation with reduced annotation requirements. Moreover, STED-FM representations enhance the performance of deep learning–based image denoising and improve the quality of images generated by diffusion models, enabling latent attribute manipulation for the data-driven discovery of subtle nanostructures and phenotypes, as well as algorithmic super-resolution. Moreover, its powerful structure retrieval capabilities are integrated into automated STED microscopy acquisition pipelines, paving the way for smart microscopy. In sum, we demonstrate that STED-FM lays a robust foundation for state-of-the-art algorithms across a wide array of tasks, establishing it as a highly valuable and scalable resource for researchers in super-resolution microscopy.
Given the increasing prevalence of mental health problems among adolescents, early intervention and appropriate management are needed to dec… (see more)rease mortality and morbidity. Artificial intelligence’s (AI) potential contributions, although significant in the field of medicine, have not been adequately studied in the context of adolescents’ mental health.
This review aimed to identify AI interventions that have been tested, implemented, or both, for use in adolescents’ mental health care.
We used the Arksey and O’Malley framework, further refined by Levac et al, along with the Joanna Briggs Institute methodology, to guide this scoping review. We searched 5 electronic databases from the inception date through July 2024 (inclusive). Four independent reviewers screened the titles and abstracts, read the full texts, and extracted data using a validated data extraction form. Disagreements were resolved by consensus, and if this was not possible, the opinion of a fifth reviewer was sought. We evaluated the risk of bias (ROB) for prognosis and diagnosis-related studies using the Prediction Model Risk of Bias Assessment Tool. We followed the PRISMA-ScR (Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews) checklist for reporting.
Of the papers screened, 88 papers relevant to our eligibility criteria were identified. Among the included papers, AI was most commonly used for diagnosis (n=78), followed by monitoring and evaluation (n=19), treatment (n=10), and prognosis (n=6). As some studies addressed multiple applications, categories are not mutually exclusive. For diagnosis, studies primarily addressed suicidal behaviors (n=11) and autism spectrum disorder (n=7). Machine learning was the most frequently reported AI method across all application areas. The overall ROB for diagnostic and prognostic models was predominantly unclear (58%), while 20% of studies had a high ROB and 22% were assessed as low risk.
In our review, we found that AI is being applied across various areas of adolescent mental health care, spanning diagnosis, treatment planning, symptom monitoring, and prognosis. Interestingly, most studies to date have concentrated heavily on diagnostic tools, leaving other important aspects of care relatively underexplored. This presents a key opportunity for future research to broaden the scope of AI applications beyond diagnosis. Moreover, future studies should emphasize the meaningful and active involvement of end users in the design, development, and validation of AI interventions, alongside improved transparency in reporting AI models, data handling, and analytical processes to build trust and support safe clinical implementation.
Abstract Background Given the increasing prevalence of mental health problems among adolescents, early intervention and appropriate manageme… (see more)nt are needed to decrease mortality and morbidity. Artificial intelligence’s (AI) potential contributions, although significant in the field of medicine, have not been adequately studied in the context of adolescents’ mental health. Objective This review aimed to identify AI interventions that have been tested, implemented, or both, for use in adolescents’ mental health care. Methods We used the Arksey and O’Malley framework, further refined by Levac et al, along with the Joanna Briggs Institute methodology, to guide this scoping review. We searched 5 electronic databases from the inception date through July 2024 (inclusive). Four independent reviewers screened the titles and abstracts, read the full texts, and extracted data using a validated data extraction form. Disagreements were resolved by consensus, and if this was not possible, the opinion of a fifth reviewer was sought. We evaluated the risk of bias (ROB) for prognosis and diagnosis-related studies using the Prediction Model Risk of Bias Assessment Tool. We followed the PRISMA-ScR (Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews) checklist for reporting. Results Of the papers screened, 88 papers relevant to our eligibility criteria were identified. Among the included papers, AI was most commonly used for diagnosis (n=78), followed by monitoring and evaluation (n=19), treatment (n=10), and prognosis (n=6). As some studies addressed multiple applications, categories are not mutually exclusive. For diagnosis, studies primarily addressed suicidal behaviors (n=11) and autism spectrum disorder (n=7). Machine learning was the most frequently reported AI method across all application areas. The overall ROB for diagnostic and prognostic models was predominantly unclear (58%), while 20% of studies had a high ROB and 22% were assessed as low risk. Conclusions In our review, we found that AI is being applied across various areas of adolescent mental health care, spanning diagnosis, treatment planning, symptom monitoring, and prognosis. Interestingly, most studies to date have concentrated heavily on diagnostic tools, leaving other important aspects of care relatively underexplored. This presents a key opportunity for future research to broaden the scope of AI applications beyond diagnosis. Moreover, future studies should emphasize the meaningful and active involvement of end users in the design, development, and validation of AI interventions, alongside improved transparency in reporting AI models, data handling, and analytical processes to build trust and support safe clinical implementation.
DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models
Revant Teotia
Candace Ross
Karen Ullrich
Sumit Chopra
Adriana Romero
Melissa Hall
Matthew J. Muckley
Recent advances in text-to-image (T2I) models have achieved impressive quality and consistency. However, this has come at the cost of repres… (see more)entation diversity. While automatic evaluation methods exist for benchmarking model diversity, they either require reference image datasets or lack specificity about the kind of diversity measured, limiting their adaptability and interpretability. To address this gap, we introduce the Does-it/Can-it framework, DIM-CIM, a reference-free measurement of default-mode diversity ("Does" the model generate images with expected attributes?) and generalization capacity ("Can" the model generate diverse attributes for a particular concept?). We construct the COCO-DIMCIM benchmark, which is seeded with COCO concepts and captions and augmented by a large language model. With COCO-DIMCIM, we find that widely-used models improve in generalization at the cost of default-mode diversity when scaling from 1.5B to 8.1B parameters. DIMCIM also identifies fine-grained failure cases, such as attributes that are generated with generic prompts but are rarely generated when explicitly requested. Finally, we use DIMCIM to evaluate the training data of a T2I model and observe a correlation of 0.85 between diversity in training images and default-mode diversity. Our work provides a flexible and interpretable framework for assessing T2I model diversity and generalization, enabling a more comprehensive understanding of model performance.
Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention. Although widely adopted, trans… (see more)formers require scaling memory and compute linearly during inference. A recent stream of work linearized the softmax operation, resulting in powerful recurrent neural network (RNN) models with constant memory and compute costs such as DeltaNet, Mamba or xLSTM. These models can be unified by noting that their recurrent layer dynamics can all be derived from an in-context regression objective, approximately optimized through an online learning rule. Here, we join this line of work and introduce a numerically stable, chunkwise parallelizable version of the recently proposed Mesa layer (von Oswald et al., 2024), and study it in language modeling at the billion-parameter scale. This layer again stems from an in-context loss, but which is now minimized to optimality at every time point using a fast conjugate gradient solver. Through an extensive suite of experiments, we show that optimal test-time training enables reaching lower language modeling perplexity and higher downstream benchmark performance than previous RNNs, especially on tasks requiring long context understanding. This performance gain comes at the cost of additional flops spent during inference time. Our results are therefore intriguingly related to recent trends of increasing test-time compute to improve performance -- here by spending compute to solve sequential optimization problems within the neural network itself.
The digitalization of health records stands to improve decision-making at clinical, administrative, and policy level. Efforts follow various… (see more) paths and are closely intertwined with health system and organizational configurations. Problems persist in both uptake and use. This study explores the digitalization trajectories of academic health centers (AHCs) to understand tensions between organizational and government strategies and their impact on digital development.
AHCs play a leadership role within health systems in data-driven improvement. This retrospective case study draws on documentary, observational, and interview data to compare digitalization efforts over 3 decades in 4 AHCs in the province of Quebec (Canada).
At system level, strategy shifted from supporting multilayered development that encouraged bottom-up initiatives in the first decade of the 2000s, to harmonizing clinical information systems in a highly prescriptive manner after 2010. AHCs experienced the shift differently according to concurrent impacts of health system restructuring, and internal choices around electronic health record (EHR) systems and implementation priorities. Digital maturity remained low in all 4 AHCs.
Coordination between system strategies and organizational strategies in AHCs was neglected in early digital development in Québec and improved only after an intense period of prescription and resistance. Confrontation highlighted tensions around different objectives at AHC and system level, competing missions within AHCs, and trade-offs between relying on commercial EHRs and developing publicly owned systems, all of which ultimately influence EHR implementation.
The different experiences of focal organizations with digitalization underline the importance of adapting national strategies and providing support to implementers, building on acquired strengths, and arriving at the right balance of guidance from the top and autonomy to develop innovative capacities.
2025-06-04
Journal of the American Medical Informatics Association : JAMIA (published)
Evaluating machine translation (MT) quality for under-resourced African languages remains a significant challenge, as existing metrics often… (see more) suffer from limited language coverage and poor performance in low-resource settings. While recent efforts, such as AfriCOMET, have addressed some of the issues, they are still constrained by small evaluation sets, a lack of publicly available training data tailored to African languages, and inconsistent performance in extremely low-resource scenarios. In this work, we introduce SSA-MTE, a large-scale human-annotated MT evaluation (MTE) dataset covering 13 African language pairs from the News domain, with over 63,000 sentence-level annotations from a diverse set of MT systems. Based on this data, we develop SSA-COMET and SSA-COMET-QE, improved reference-based and reference-free evaluation metrics. We also benchmark prompting-based approaches using state-of-the-art LLMs like GPT-4o and Claude. Our experimental results show that SSA-COMET models significantly outperform AfriCOMET and are competitive with the strongest LLM (Gemini 2.5 Pro) evaluated in our study, particularly on low-resource languages such as Twi, Luo, and Yoruba. All resources are released under open licenses to support future research.
To thrive in complex environments, animals and artificial agents must learn to act adaptively to maximize fitness and rewards. Such adaptive… (see more) behavior can be learned through reinforcement learning1, a class of algorithms that has been successful at training artificial agents2–6 and at characterizing the firing of dopamine neurons in the midbrain7–9. In classical reinforcement learning, agents discount future rewards exponentially according to a single time scale, controlled by the discount factor. Here, we explore the presence of multiple timescales in biological reinforcement learning. We first show that reinforcement agents learning at a multitude of timescales possess distinct computational benefits. Next, we report that dopamine neurons in mice performing two behavioral tasks encode reward prediction error with a diversity of discount time constants. Our model explains the heterogeneity of temporal discounting in both cue-evoked transient responses and slower timescale fluctuations known as dopamine ramps. Crucially, the measured discount factor of individual neurons is correlated across the two tasks suggesting that it is a cell-specific property. Together, our results provide a new paradigm to understand functional heterogeneity in dopamine neurons, a mechanistic basis for the empirical observation that humans and animals use non-exponential discounts in many situations 10–14, and open new avenues for the design of more efficient reinforcement learning algorithms.
ABSTRACT Biotic interactions are expected to influence species' responses to global changes, but they are rarely considered across broad spa… (see more)tial extents. Abiotic factors are thought to operate at larger spatial scales, while biotic factors, such as species interactions, are considered more important at local scales within communities, in part because of the knowledge gap on species interactions at large spatial scales (i.e., the Eltonian shortfall). We assessed, at a continental scale, (i) the importance of biotic interactions, through food webs, on species distributions, and (ii) how biotic interactions under scenarios of climate and land‐use change may affect the distribution of the brown bear ( Ursus arctos ). We built a highly detailed, spatially dynamic, and empirically sampled food web based on the energy contribution of 276 brown bear food species from different taxa (plants, vertebrates, and invertebrates) and their ensemble habitat models at high resolution across Europe. Then, combining energy contribution and predicted habitat of food species, we modelled energy contribution across space and included these layers within Bayesian‐based models of the brown bear distribution in Europe. The inclusion of biotic interactions considerably improved our understanding of brown bear distribution at large (continental) scales compared with Bayesian models including only abiotic factors (climate and land use). Predicted future range shifts, which included changes in the distribution of food species, varied greatly when considering various scenarios of change in biotic factors, providing a warning that future indirect climate and land‐use change are likely to have strong but highly uncertain impacts on species biogeography. Our study confirmed that advancing our understanding of ecological networks of species interactions will improve future projections of biodiversity change, especially for modelling species distributions and their functional role under climate and land‐use change scenarios, which is key for effective conservation of biodiversity and ecosystem services.
Galaxy cluster characterization with machine learning techniques
M. Sadikov
J. Hlavacek-Larrondo
L. Perreault-Levasseur
C. L. Rhea
M. McDonald
M. Ntampaka
J. Zuhone
We present an analysis of the X-ray properties of the galaxy cluster population in the z=0 snapshot of the IllustrisTNG simulations, utilizi… (see more)ng machine learning techniques to perform clustering and regression tasks. We examine five properties of the hot gas (the central cooling time, the central electron density, the central entropy excess, the concentration parameter, and the cuspiness) which are commonly used as classification metrics to identify cool core (CC), weak cool core (WCC) and non cool core (NCC) clusters of galaxies. Using mock Chandra X-ray images as inputs, we first explore an unsupervised clustering scheme to see how the resulting groups correlate with the CC/WCC/NCC classification based on the different criteria. We observe that the groups replicate almost exactly the separation of the galaxy cluster images when classifying them based on the concentration parameter. We then move on to a regression task, utilizing a ResNet model to predict the value of all five properties. The network is able to achieve a mean percentage error of 1.8% for the central cooling time, and a balanced accuracy of 0.83 on the concentration parameter, making them the best-performing metrics. Finally, we use simulation-based inference (SBI) to extract posterior distributions for the network predictions. Our neural network simultaneously predicts all five classification metrics using only mock Chandra X-ray images.
This study demonstrates that machine learning is a viable approach for analyzing and classifying the large galaxy cluster datasets that will soon become available through current and upcoming X-ray surveys, such as eROSITA.