This program supports AI startups at any time of the year. Benefit from cutting-edge resources and tailored support to accelerate your technology's development.
Offered by Mila and the Public Policy Forum, this program is designed to equip policy and decision makers with the tools to navigate the opportunities and risks of AI. The next cohort will be held in French on September 1-2, 2026, at Mila.
Connect with a Mila academic advisor and current student-researchers to learn more about Mila's community and how to join us on August 19, 31 and September 11, 2026.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Publications
Discovering Diverse Behaviors via Temporal Contrastive Learning
Effective exploration in reinforcement learning requires not only tracking where an agent has been, but also understanding how the agent per… (see more)ceives and represents the world. To learn powerful representations, an agent should actively explore states that contribute to its knowledge of the environment. Temporal representations can capture the information necessary to solve a wide range of potential tasks while avoiding the computational cost associated with full state reconstruction. In this paper, we propose an exploration method that leverages temporal contrastive representations to guide exploration, prioritizing states with unpredictable future outcomes. We demonstrate that such representations can enable the learning of complex exploratory behaviors in locomotion, manipulation, and embodied-AI tasks, revealing capabilities and behaviors that traditionally require extrinsic rewards. Unlike approaches that rely on explicit distance learning or episodic memory mechanisms (e.g., quasimetric-based methods), our method builds directly on temporal similarities, yielding a simpler yet effective strategy for exploration.
2025-12-31
International Conference on Learning Representations (Accept (Poster))
Model merging has emerged as an efficient and flexible paradigm for multi-task learning, with numerous methods being proposed in recent year… (see more)s. However, these state-of-the-art techniques are typically evaluated on benchmark suites that are highly favorable to model merging, and their robustness in more realistic settings remains largely unexplored. In this work, we first investigate the vulnerabilities of model-merging methods and pinpoint the source-model characteristics that critically underlie them. Specifically, we identify two factors that are particularly harmful to the merging process: (1) disparities in task vector norms, and (2) the low confidence of the source models. To address this issue, we propose **DisTaC** (**Dis**tillation for **Ta**sk vector **C**onditioning), a novel method that pre-conditions these problematic task vectors before the merge. DisTaC leverages knowledge distillation to adjust a task vector's norm and increase source-model confidence while preserving its essential task-specific knowledge. Our extensive experiments demonstrate that by pre-conditioning task vectors with DisTaC, state-of-the-art merging techniques can successfully integrate models that exhibit these harmful traits, where they would otherwise fail, and achieve significant performance gains.
2025-12-31
International Conference on Learning Representations (Accept (Poster))
DSL: Dual-Stage List Decoding for Polar Turbo Product Coded Massive MIMO Towards 6 G Extreme Connectivity
Jian Zheng
Yu Tian
Huayi Zhou
Xiaosi Tan
Yutai Sun
Warren J. Gross
Xiaohu You
Chuan Zhang
Targeting sixth generation (6 G) extreme connectivity, massive multiple-input multiple-output (mMIMO) systems face challenges in achieving b… (see more)oth high data rates and low latency. Polar turbo product codes (polar-TPCs) with near-optimal error correction performance under constrained latency are promising for 6 G mMIMO applications. In this letter, we propose a dual-stage list (DSL) decoding scheme for polar-TPC coded mMIMO systems, leveraging a turbo list decoder followed by a tree-search-based post-processing method to enhance receiver performance. Moreover, we introduce a two-phase codeword validity check strategy to address the complexity challenges. Empirical results demonstrate that the receiver with our proposed DSL decoding outperforms existing receivers, achieving a superior trade-off between performance and complexity.
2025-12-31
IEEE Transactions on Vehicular Technology (published)
Constrained optimization is a powerful framework for enforcing requirements on neural networks. These constrained deep learning problems are… (see more) typically solved using first-order methods on their min-max Lagrangian formulation, but such approaches often suffer from oscillations and can fail to find all local solutions. While the Augmented Lagrangian method (ALM) addresses these issues, practitioners often favor dual optimistic ascent schemes (PI control) on the standard Lagrangian, which perform well empirically but lack formal guarantees. In this paper, we establish a previously unknown equivalence between these approaches: dual optimistic ascent on the Lagrangian is equivalent to gradient descent-ascent on the Augmented Lagrangian. This finding allows us to transfer the robust theoretical guarantees of the ALM to the dual optimistic setting, proving it converges linearly to all local solutions. Furthermore, the equivalence provides principled guidance for tuning the optimism hyper-parameter. Our work closes a critical gap between the empirical success of dual optimistic methods and their theoretical foundation.
2025-12-31
International Conference on Learning Representations (Accept (Poster))
In recent years, many industries have utilized machine learning models (ML) in their systems. Ideally, machine learning models should be tra… (see more)ined on and applied to data from the same distributions. However, the data evolves over time in many application areas, leading to data and concept drift, which in turn causes the performance of the ML models to degrade over time. Therefore, maintaining up to date ML models plays a critical role in the MLOps pipeline. Existing ML model maintenance approaches are often computationally resource intensive, costly, time consuming, and model dependent. Thus, we propose an improved MLOps pipeline, a new model maintenance approach and a Similarity Based Model Reuse (SimReuse) tool to address the challenges of ML model maintenance. We identify seasonal and recurrent distribution patterns in time series datasets throughout a preliminary study. Recurrent distribution patterns enable us to reuse previously trained models for similar distributions in the future, thus avoiding frequent retraining. Then, we integrated the model reuse approach into the MLOps pipeline and proposed our improved MLOps pipeline. Furthermore, we develop SimReuse, a tool to implement the new components of our MLOps pipeline to store models and reuse them for inference of data segments with similar data distributions in the future. Our evaluation results on four time series datasets demonstrate that our model reuse approach can maintain the performance of models while significantly reducing maintenance time and costs. Our model reuse approach achieves ML performance comparable to the best baseline, while being 15 times more efficient in terms of computation time and costs. Therefore, industries and practitioners can benefit from our approach and use our tool to maintain the performance of their ML models in the deployment phase to reduce their maintenance costs.
Envisioning digital health ecosystem transformation in Canada «a conceptual foundation en Neuf Etapes.»
Nitika Pant Pai
Samira Abbasgholizadeh Rahimi
Juhi Tulsi
Susan Bartlett
Steven Grover
Ervin Sejdic
Canada’s journey towards digital health transformation is in a phase that precedes widespread catalytic change, trailing peer nations. In … (see more)this perspective piece, we discuss the nine steps that are key to catalyzing digital health transformation. We highlight the importance of foundational investments for health systems redesign. These investments in interoperability, unified digital core and ID, scalable health data systems, will enable precision-focused clinical care and prevention-focused public health with Smart Care Everywhere models. Our conceptual foundation highlights the importance of an agile health system with a unified digital core, capable of integrating multiple AI-enhanced digital tools, and managing the data deluge of multimodal data. We hereby advocate for a Smart, Scalable, Digitized, “Care Everywhere” model that can expand health care access to all of its populations: the served and the underserved. An essential component of the foundation is the creation of agile health systems and business models that prevent provider burnout while promoting collaborative, connected care that reaches served and under-served populations, with caring, compassion, enabling an improved engagement and connection. We also call for an investment in the continuous training of healthcare professionals, data professionals, and for an ethical, efficient implementation of AI/digital solutions everywhere from hospitals to community care settings. We also highlight the necessity of data governance policies to safeguard patient autonomy, promote data ownership, to ensure health data privacy, security, and confidentiality. This nine-step approach offers a framework for a unified, connected, patient-centred health ecosystem operationalized/made efficient with digital/AI solutions for patient communities, enabled by connectivity, caring, and compassion. Together, these nine steps serve as a conceptual foundation to enable a sustainable health system that advances access, equity, and efficiency in caring in health care nationwide.
Reinforcement learning from human feedback (RLHF) with proximal policy optimization (PPO) is widely used but often yields less diverse outpu… (see more)ts than supervised fine-tuning, suggesting an effect in which the policy’s support contracts during on-policy optimization. We formalize this “policy contraction” with the Support Retention Ratio (SRR)—the share of SFT completions that retain non-negligible probability under the RL policy—and additionally track token-entropy, Kullback–Leibler (KL) divergence to the reference, and repetition. We propose Contraction-Aware PPO (CaPPO), a minimum-norm multi-gradient update that co-optimizes reward, entropy, and KL, paired with a controller that steers exploration toward a target token entropy. On HH-RLHF, Summarize-from-Feedback, and UltraFeedback with Qwen2-7B, Qwen2.5-14B, Mistral-7B-Instruct, and Llama-3-8B-Instruct, CaPPO increases win rate by 2 to 4 points over PPO and improves diversity, gaining 0.2 to 0.3 higher SRR. The gains persist under decoding sweeps and are robust to reward scaling and critic variance. Treating reward, diversity, and stability as first-class objectives, CaPPO mitigates contraction without sacrificing alignment performance.
2025-12-31
International Conference on Learning Representations (Accept (Poster))
Protein-protein interactions (PPIs) are mediated at the residue level. Most sequence-based PPI models consider residue-residue interactions … (see more)across two proteins, which can yield accurate interaction scores but are too slow to scale. At proteome scale, identifying candidate PPIs requires evaluating nearly *all possible protein pairs*. For
2025-12-31
International Conference on Learning Representations (Accept (Poster))
Fast Sphere Decoding of Short Systematic Polar-like Codes
Huayi Zhou
Y. Liu
Xiaosi Tan
Chen Ji
Warren J. Gross
Chuan Zhang
Short polar-like codes are competitive for low latency requirements in future communications. Systematic polar codes have not been shown to … (see more)offer substantial benefits for decoders beyond improving the bit error rate. In this paper, we demonstrate that the sparsity of the equivalent generator matrix of systematic polar codes significantly reduces calculation complexity when using sphere decoding (SD). We propose a fast SD (Fast-SD) for systematic polar codes. Numerical results indicate that the proposed Fast-SD reduces calculation complexity by up to 33.25% compared to SD on short high-rate codes while maintaining maximum likelihood performance.
2025-12-31
IEEE Transactions on Vehicular Technology (published)