This program supports AI startups at any time of the year. Benefit from cutting-edge resources and tailored support to accelerate your technology's development.
Offered by Mila and the Public Policy Forum, this program is designed to equip policy and decision makers with the tools to navigate the opportunities and risks of AI. The next cohort will be held in French on September 1-2, 2026, at Mila.
Connect with a Mila academic advisor and current student-researchers to learn more about Mila's community and how to join us on August 19, 31 and September 11, 2026.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Publications
Imagining and building wise machines: The centrality of AI metacognition
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
Luke Marks
Alasdair Paren
David M. Krueger
Fazl Barez
Sparse Autoencoders (SAEs) have shown promise in improving the interpretability of neural network activations, but can learn features that a… (see more)re not features of the input, limiting their effectiveness. We propose \textsc{Mutual Feature Regularization} \textbf{(MFR)}, a regularization technique for improving feature learning by encouraging SAEs trained in parallel to learn similar features. We motivate \textsc{MFR} by showing that features learned by multiple SAEs are more likely to correlate with features of the input. By training on synthetic data with known features of the input, we show that \textsc{MFR} can help SAEs learn those features, as we can directly compare the features learned by the SAE with the input features for the synthetic data. We then scale \textsc{MFR} to SAEs that are trained to denoise electroencephalography (EEG) data and SAEs that are trained to reconstruct GPT-2 Small activations. We show that \textsc{MFR} can improve the reconstruction loss of SAEs by up to 21.21\% on GPT-2 Small, and 6.67\% on EEG data. Our results suggest that the similarity between features learned by different SAEs can be leveraged to improve SAE training, thereby enhancing performance and the usefulness of SAEs for model interpretability.
Predicting molecular impact on cellular function is a core challenge in therapeutic design.
Phenomic experiments, designed to capture cellu… (see more)lar morphology, utilize microscopy based techniques and demonstrate a high throughput solution for uncovering molecular impact on the cell.
In this work, we learn a joint latent space between molecular structures and microscopy phenomic experiments, aligning paired samples with contrastive learning.
Specifically, we study the problem of Contrastive PhenoMolecular Retrieval, which consists of zero-shot molecular structure identification conditioned on phenomic experiments.
We assess challenges in multi-modal learning of phenomics and molecular modalities such as experimental batch effect, inactive molecule perturbations, and encoding perturbation concentration.
We demonstrate improved multi-modal learner retrieval through (1) a uni-modal pre-trained phenomics model, (2) a novel inter sample similarity aware loss, and (3) models conditioned on a representation of molecular concentration.
Following this recipe, we propose MolPhenix, a molecular phenomics model.
MolPhenix leverages a pre-trained phenomics model to demonstrate significant performance gains across perturbation concentrations, molecular scaffolds, and activity thresholds.
In particular, we demonstrate an 8.1
Advancements in AI heavily rely on large-scale datasets meticulously curated and annotated for training. However, concerns persist regarding… (see more) the transparency and context of data collection methodologies, especially when sourced through crowdsourcing platforms. Crowdsourcing often employs low-wage workers with poor working conditions and lacks consideration for the representativeness of annotators, leading to algorithms that fail to represent diverse views and perpetuate biases against certain groups. To address these limitations, we propose a methodology involving a co-design model that actively engages stakeholders at key stages, integrating principles of Equity, Diversity, and Inclusion (EDI) to ensure diverse viewpoints. We apply this methodology to develop a dataset and AI model for evaluating public space quality using street view images, demonstrating its effectiveness in capturing diverse perspectives and fostering higher-quality data.
Over the past decade, hyperscanning has emerged as an important methodology to study neural processes underlying human interaction using fMR… (see more)I, EEG, fNIRS, and MEG. However, many methodological decisions regarding preprocessing and analysis of hyperscanning data have not yet been standardized in the hyperscanning community, yet may affect inter-brain estimates. Here we systematically investigate the effects common methodological choices can have on estimates of phase-based inter-brain synchronization (IBS) measures, using real and simulated hyperscanning (dual) EEG data. Notably, we introduce a new method to compute circular correlation (CCorr) coefficients in IBS studies, which performs more reliably in comparison to the standard approach, showing that the conventional CCorr implementation leads to large fluctuations in IBS estimates due to fluctuations in circular mean directions. Furthermore, we demonstrate how short epoch durations (of 1 second or less) can lead to inflated IBS estimates in scenarios with no strong underlying interaction. Finally, we show how signal-to-noise ratios and temporal factors may confound IBS estimates, particularly when comparing e.g., resting states with conditions involving motor actions. For each of these investigated effects, we provide recommendations for future research employing hyperscanning-EEG techniques, aimed at increasing validity and replicability of inter-brain synchronization studies.
Background/Objectives: Nutritional deficiencies have been proposed as possible etiological causes for autoimmune diseases, among which type … (see more)1 diabetes (T1D). Vitamin K (VK) has potentially positive effects on type 2 diabetes, but its role on T1D in humans remains largely unknown. We aimed to examine the presence of a causal association between VK and T1D using a Mendelian randomization (MR) approach. Methods: Genetic variants from a genome-wide association study (GWAS) for VK (N = 2138 Europeans) were used as instruments in our two-sample MR study to investigate whether circulating VK levels are causally associated with the risk of T1D in a large European T1D GWAS cohort (18,942 cases/520,580 controls). Through a multivariable MR (MVMR), the effects of both VK and specific gut microbiota on T1D were investigated given that the gut microbiome synthesizes VK. Results: We found that changes in levels of circulating VK did not affect T1D risk in our univariate two-sample MR, but this study had limited power to detect small effects of VK (OR for T1D of less than 0.8). However, our MVMR indicated a suggestive association of VK with the risk of T1D adjusting for two different gut microbiome populations. Conclusions: In conclusion, VK levels are unlikely to significantly affect the risk of T1D, but small effects cannot be excluded, and the role of gut microbiome in this association should be further investigated.
The CA1 region of the hippocampus is one of the most studied regions of the rodent brain, thought to play an important role in cognitive fun… (see more)ctions such as memory and spatial navigation. Despite a wealth of experimental data on its structure and function, it has been challenging to integrate information obtained from diverse experimental approaches. To address this challenge, we present a community-based, full-scale in silico model of the rat CA1 that integrates a broad range of experimental data, from synapse to network, including the reconstruction of its principal afferents, the Schaffer collaterals, and a model of the effects that acetylcholine has on the system. We tested and validated each model component and the final network model, and made input data, assumptions, and strategies explicit and transparent. The unique flexibility of the model allows scientists to potentially address a range of scientific questions. In this article, we describe the methods used to set up simulations to reproduce in vitro and in vivo experiments. Among several applications in the article, we focus on theta rhythm, a prominent hippocampal oscillation associated with various behavioral correlates and use our computer model to reproduce experimental findings. Finally, we make data, code, and model available through the hippocampushub.eu portal, which also provides an extensive set of analyses of the model and a user-friendly interface to facilitate adoption and usage. This community-based model represents a valuable tool for integrating diverse experimental data and provides a foundation for further research into the complex workings of the hippocampal CA1 region.
Deep clustering incorporates embedding into clustering in order to find a lower-dimensional space suitable for clustering task. Conventional… (see more) deep clustering methods aim to obtain a single global embedding subspace (aka latent space) for all the data clusters. In contrast, in this paper, we propose a deep multi-representation learning (DML) framework for data clustering whereby each difficult to cluster data group is associated with its own distinct optimized latent space, and all the easy to cluster data groups are associated to a general common latent space. Autoencoders are employed for generating the cluster-specific and general latent spaces. To specialize each autoencoder in its associated data cluster(s), we propose a novel and effective loss function which consists of weighted reconstruction and clustering losses of the data points, where higher weights are assigned to the samples more probable to belong to the corresponding cluster(s). Experimental results on benchmark datasets demonstrate that the proposed DML framework and loss function outperform state-of-the-art clustering approaches. In addition, the results show that the DML method significantly outperforms the SOTA on imbalanced datasets as a result of assigning an individual latent space to the difficult clusters. <br>
2024-10-31
IEEE Transactions on Neural Networks and Learning Systems (published)