Publications

Discovering Diverse Behaviors via Temporal Contrastive Learning
Catherine Ji
Benjamin Eysenbach
Effective exploration in reinforcement learning requires not only tracking where an agent has been, but also understanding how the agent per… (voir plus)ceives and represents the world. To learn powerful representations, an agent should actively explore states that contribute to its knowledge of the environment. Temporal representations can capture the information necessary to solve a wide range of potential tasks while avoiding the computational cost associated with full state reconstruction. In this paper, we propose an exploration method that leverages temporal contrastive representations to guide exploration, prioritizing states with unpredictable future outcomes. We demonstrate that such representations can enable the learning of complex exploratory behaviors in locomotion, manipulation, and embodied-AI tasks, revealing capabilities and behaviors that traditionally require extrinsic rewards. Unlike approaches that rely on explicit distance learning or episodic memory mechanisms (e.g., quasimetric-based methods), our method builds directly on temporal similarities, yielding a simpler yet effective strategy for exploration.
DisTaC: Conditioning Task Vectors via Distillation for Robust Model Merging
Kotaro Yoshida
Yuji Naraki
Takafumi Horie
Ryotaro Shimizu
Model merging has emerged as an efficient and flexible paradigm for multi-task learning, with numerous methods being proposed in recent year… (voir plus)s. However, these state-of-the-art techniques are typically evaluated on benchmark suites that are highly favorable to model merging, and their robustness in more realistic settings remains largely unexplored. In this work, we first investigate the vulnerabilities of model-merging methods and pinpoint the source-model characteristics that critically underlie them. Specifically, we identify two factors that are particularly harmful to the merging process: (1) disparities in task vector norms, and (2) the low confidence of the source models. To address this issue, we propose **DisTaC** (**Dis**tillation for **Ta**sk vector **C**onditioning), a novel method that pre-conditions these problematic task vectors before the merge. DisTaC leverages knowledge distillation to adjust a task vector's norm and increase source-model confidence while preserving its essential task-specific knowledge. Our extensive experiments demonstrate that by pre-conditioning task vectors with DisTaC, state-of-the-art merging techniques can successfully integrate models that exhibit these harmful traits, where they would otherwise fail, and achieve significant performance gains.
DSL: Dual-Stage List Decoding for Polar Turbo Product Coded Massive MIMO Towards 6 G Extreme Connectivity
Jian Zheng
Yu Tian
Huayi Zhou
Xiaosi Tan
Yutai Sun
Warren J. Gross
Xiaohu You
Chuan Zhang
Targeting sixth generation (6 G) extreme connectivity, massive multiple-input multiple-output (mMIMO) systems face challenges in achieving b… (voir plus)oth high data rates and low latency. Polar turbo product codes (polar-TPCs) with near-optimal error correction performance under constrained latency are promising for 6 G mMIMO applications. In this letter, we propose a dual-stage list (DSL) decoding scheme for polar-TPC coded mMIMO systems, leveraging a turbo list decoder followed by a tree-search-based post-processing method to enhance receiver performance. Moreover, we introduce a two-phase codeword validity check strategy to address the complexity challenges. Empirical results demonstrate that the receiver with our proposed DSL decoding outperforms existing receivers, achieving a superior trade-off between performance and complexity.
Dual Optimistic Ascent (PI Control) is the Augmented Lagrangian Method in Disguise
Constrained optimization is a powerful framework for enforcing requirements on neural networks. These constrained deep learning problems are… (voir plus) typically solved using first-order methods on their min-max Lagrangian formulation, but such approaches often suffer from oscillations and can fail to find all local solutions. While the Augmented Lagrangian method (ALM) addresses these issues, practitioners often favor dual optimistic ascent schemes (PI control) on the standard Lagrangian, which perform well empirically but lack formal guarantees. In this paper, we establish a previously unknown equivalence between these approaches: dual optimistic ascent on the Lagrangian is equivalent to gradient descent-ascent on the Augmented Lagrangian. This finding allows us to transfer the robust theoretical guarantees of the ALM to the dual optimistic setting, proving it converges linearly to all local solutions. Furthermore, the equivalence provides principled guidance for tuning the optimism hyper-parameter. Our work closes a critical gap between the empirical success of dual optimistic methods and their theoretical foundation.
Efficiency at What Cost? Safety and Fairness in Parameter-Efficient Fine-Tuning of LLMs
An Efficient Model Maintenance Approach for MLOps
Heng Li
Amin Nikanjam
In recent years, many industries have utilized machine learning models (ML) in their systems. Ideally, machine learning models should be tra… (voir plus)ined on and applied to data from the same distributions. However, the data evolves over time in many application areas, leading to data and concept drift, which in turn causes the performance of the ML models to degrade over time. Therefore, maintaining up to date ML models plays a critical role in the MLOps pipeline. Existing ML model maintenance approaches are often computationally resource intensive, costly, time consuming, and model dependent. Thus, we propose an improved MLOps pipeline, a new model maintenance approach and a Similarity Based Model Reuse (SimReuse) tool to address the challenges of ML model maintenance. We identify seasonal and recurrent distribution patterns in time series datasets throughout a preliminary study. Recurrent distribution patterns enable us to reuse previously trained models for similar distributions in the future, thus avoiding frequent retraining. Then, we integrated the model reuse approach into the MLOps pipeline and proposed our improved MLOps pipeline. Furthermore, we develop SimReuse, a tool to implement the new components of our MLOps pipeline to store models and reuse them for inference of data segments with similar data distributions in the future. Our evaluation results on four time series datasets demonstrate that our model reuse approach can maintain the performance of models while significantly reducing maintenance time and costs. Our model reuse approach achieves ML performance comparable to the best baseline, while being 15 times more efficient in terms of computation time and costs. Therefore, industries and practitioners can benefit from our approach and use our tool to maintain the performance of their ML models in the deployment phase to reduce their maintenance costs.
Envisioning digital health ecosystem transformation in Canada «a conceptual foundation en Neuf Etapes.»
Nitika Pant Pai
Samira Abbasgholizadeh Rahimi
Juhi Tulsi
Susan Bartlett
Steven Grover
Ervin Sejdic
Canada’s journey towards digital health transformation is in a phase that precedes widespread catalytic change, trailing peer nations. In … (voir plus)this perspective piece, we discuss the nine steps that are key to catalyzing digital health transformation. We highlight the importance of foundational investments for health systems redesign. These investments in interoperability, unified digital core and ID, scalable health data systems, will enable precision-focused clinical care and prevention-focused public health with Smart Care Everywhere models. Our conceptual foundation highlights the importance of an agile health system with a unified digital core, capable of integrating multiple AI-enhanced digital tools, and managing the data deluge of multimodal data. We hereby advocate for a Smart, Scalable, Digitized, “Care Everywhere” model that can expand health care access to all of its populations: the served and the underserved. An essential component of the foundation is the creation of agile health systems and business models that prevent provider burnout while promoting collaborative, connected care that reaches served and under-served populations, with caring, compassion, enabling an improved engagement and connection. We also call for an investment in the continuous training of healthcare professionals, data professionals, and for an ethical, efficient implementation of AI/digital solutions everywhere from hospitals to community care settings. We also highlight the necessity of data governance policies to safeguard patient autonomy, promote data ownership, to ensure health data privacy, security, and confidentiality. This nine-step approach offers a framework for a unified, connected, patient-centred health ecosystem operationalized/made efficient with digital/AI solutions for patient communities, enabled by connectivity, caring, and compassion. Together, these nine steps serve as a conceptual foundation to enable a sustainable health system that advances access, equity, and efficiency in caring in health care nationwide.
AI Epistemic Risks: Emerging Mechanisms & Evidence
Mick Yang
Stephen Casper
Jonathan Stray
Jasmine Li
Cameron Jones
Anna Gausen
Natasha Jacques
Brian Christian
Bálint Gyevnár
Hannah Rose Kirk
Zhonghao He
Dan Zhao (285025)
Siao Si Looi
J. Levy
Kobi Hackenburg
Elizabeth Seger
Matt Kowal
Michelle Malonza
Luke Hewitt
Hause Lin … (voir 10 de plus)
Maarten Sap
Dylan Hadfield-Menell
Thomas Costello
David Rand
Atoosa Kasirzadeh
Gordon Pennycook
Escaping Policy Contraction: Contraction-Aware PPO (CaPPO) for Stable Language Model Fine-Tuning
Xue Liu
Reinforcement learning from human feedback (RLHF) with proximal policy optimization (PPO) is widely used but often yields less diverse outpu… (voir plus)ts than supervised fine-tuning, suggesting an effect in which the policy’s support contracts during on-policy optimization. We formalize this “policy contraction” with the Support Retention Ratio (SRR)—the share of SFT completions that retain non-negligible probability under the RL policy—and additionally track token-entropy, Kullback–Leibler (KL) divergence to the reference, and repetition. We propose Contraction-Aware PPO (CaPPO), a minimum-norm multi-gradient update that co-optimizes reward, entropy, and KL, paired with a controller that steers exploration toward a target token entropy. On HH-RLHF, Summarize-from-Feedback, and UltraFeedback with Qwen2-7B, Qwen2.5-14B, Mistral-7B-Instruct, and Llama-3-8B-Instruct, CaPPO increases win rate by 2 to 4 points over PPO and improves diversity, gaining 0.2 to 0.3 higher SRR. The gains persist under decoding sweeps and are robust to reward scaling and critic variance. Treating reward, diversity, and stability as first-class objectives, CaPPO mitigates contraction without sacrificing alignment performance.
Fast Proteome-Scale Protein Interaction Retrieval via Residue-Level Factorization
Narendra Chaudhary
Qian Cong
Jian Zhou
Sanchit Misra
Protein-protein interactions (PPIs) are mediated at the residue level. Most sequence-based PPI models consider residue-residue interactions … (voir plus)across two proteins, which can yield accurate interaction scores but are too slow to scale. At proteome scale, identifying candidate PPIs requires evaluating nearly *all possible protein pairs*. For
Fast Sphere Decoding of Short Systematic Polar-like Codes
Huayi Zhou
Y. Liu
Xiaosi Tan
Chen Ji
Warren J. Gross
Chuan Zhang
Short polar-like codes are competitive for low latency requirements in future communications. Systematic polar codes have not been shown to … (voir plus)offer substantial benefits for decoders beyond improving the bit error rate. In this paper, we demonstrate that the sparsity of the equivalent generator matrix of systematic polar codes significantly reduces calculation complexity when using sphere decoding (SD). We propose a fast SD (Fast-SD) for systematic polar codes. Numerical results indicate that the proposed Fast-SD reduces calculation complexity by up to 33.25% compared to SD on short high-rate codes while maintaining maximum likelihood performance.
FIN: Boosting binary code embedding by normalizing function inlinings
Mohammadhossein Amouei
Benjamin C. M. Fung
Philippe Charland