Publications

Loss Smoothing for Continual Adaptation
Neural networks are often adapted in nonstationary data distributions settings where the objective is to optimize performance on the current… (see more) task, and preserving accuracy on previous tasks is not required. As a result, existing methods primarily focus on improving plasticity, while stability is largely studied in the context of continual learning. In this work, we examine whether preserving stability can also be beneficial in model adaptation settings where past-task performance is irrelevant. We propose a simple loss smoothing approach that encourages selective adaptation by preserving task-shared features while modifying task-inconsistent ones. We evaluate our method on continual supervised model adaptation benchmarks and reinforcement learning benchmarks, and show that promoting representational stability during adaptation can improve performance across settings.
Value Drifts: Tracing Value Alignment During LLM Post-Training
As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to dra… (see more)w on their general knowledge but also to align with certain human value systems. Therefore, studying the alignment of LLMs with human values has become a crucial field of inquiry. Prior work, however, mostly focuses on evaluating the alignment of fully trained models, overlooking the training dynamics by which models learn to express human values. In this work, we investigate how and at which stage value alignment arises during the course of a model's post-training. Our analysis disentangles the effects of post-training algorithms and datasets, measuring both the magnitude and time of value drifts during training. Experimenting with Llama-3 and Qwen-3 models of different sizes and popular supervised fine-tuning (SFT) and preference optimization datasets and algorithms, we find that the SFT phase generally establishes a model's values, and subsequent preference optimization rarely re-aligns these values. Furthermore, using a synthetic preference dataset that enables controlled manipulation of values, we find that different preference optimization algorithms lead to different value alignment outcomes, even when preference data is held constant. Our findings provide actionable insights into how values are learned during post-training and help to inform data curation, as well as the selection of models and algorithms for preference optimization to improve model alignment to human values.
CIRCLE: A Framework for Evaluating AI from a Real-World Lens
Reva Schwartz
Carina E. I. Westling
Morgan Briggs
Marzieh Fadaee
Isar Nejadgholi
Matthew Holmes
Fariza Rashid
Maya Carlyle
Kyra Wilson
Peter Douglas
Theodora Skeadas
Gabriella Waters
Rumman Chowdhury
Thiago Lacerda
This paper proposes CIRCLE, a six-stage, lifecycle-based framework to bridge the reality gap between model-centric performance metrics and A… (see more)I's materialized outcomes in deployment. Current approaches such as MLOps frameworks and AI model benchmarks offer detailed insights into system stability and model capabilities, but they do not provide decision-makers outside the AI stack with systematic evidence of how these systems actually behave in real-world contexts or affect their organizations over time. CIRCLE operationalizes the Validation phase of TEVV (Test, Evaluation, Verification, and Validation) by formalizing the translation of stakeholder concerns outside the stack into measurable signals. Unlike participatory design, which often remains localized, or algorithmic audits, which are often retrospective, CIRCLE provides a structured, prospective protocol for linking context-sensitive qualitative insights to scalable quantitative metrics. By integrating methods such as field testing, red teaming, and longitudinal studies into a coordinated pipeline, CIRCLE produces systematic knowledge: evidence that is comparable across sites yet sensitive to local context. This, in turn, can enable governance based on materialized downstream effects rather than theoretical capabilities.
DYMAG: Rethinking Message Passing Using Dynamical-systems-based Waveforms
Dhananjay Bhaskar
Xingzhi Sun
Yanlei Zhang
Charles Xu
Oluwadamilola Fasina
Michael Perlmutter
We present DYMAG, a graph neural network based on a novel form of message aggregation. Standard message-passing neural networks, which often… (see more) aggregate local neighbors via mean-aggregation, can be regarded as convolving with a simple rectangular waveform which is non-zero only on 1-hop neighbors of every vertex. Here, we go beyond such local averaging. We will convolve the node features with more sophisticated waveforms generated using dynamics such as the heat equation, wave equation, and the Sprott model (an example of chaotic dynamics). Furthermore, we use snapshots of these dynamics at different time points to create waveforms at many effective scales. Theoretically, we show that these dynamic waveforms can capture salient information about the graph, including connected components, connectivity, and cycle structures. Empirically, we test DYMAG on both real and synthetic benchmarks to establish that DYMAG outperforms baseline models on recovery of graph persistence, generating parameters of random graphs, as well as property prediction for proteins, molecules and materials. Our code is available at https://github.com/KrishnaswamyLab/DYMAG.
Using virtual reality hypnosis during stem cell transplant for patients in hematology: A protocol for a feasibility randomized study
Audrey Laurin
Floriane Rousseaux
Isaiah Gitonga
Jean Roy
Mathieu Landry
Richard LeBlanc
Nadia Godin
Caroline Arbour
Philippe Richebé
Pierre Rainville
David Ogez
Valentyn Fournier
ClinicalTrials.gov NCT06817759.
Using virtual reality hypnosis during stem cell transplant for patients in hematology: A protocol for a feasibility randomized study
Audrey Laurin
Floriane Rousseaux
Isaiah Gitonga
Jean Roy
Mathieu Landry
Richard LeBlanc
Nadia Godin
Caroline Arbour
Philippe Richebé
Pierre Rainville
David Ogez
Valentyn Fournier
ClinicalTrials.gov NCT06817759.
CellPace: A temporal diffusion-forcing framework for simulation, interpolation and forecasting of single-cell dynamics
Abstract Single-cell omics technologies resolve cellular heterogeneity at high resolution but provide only static snapshots of continuous de… (see more)velopmental processes. This makes it difficult to recover coherent temporal dynamics when developmental stages are irregularly sampled or missing. While recent generative models can simulate observed cell states, they often treat timepoints as discrete categories, hindering interpolation across gaps and extrapolation to unobserved future stages. We present CellPace, a generative model that learns and generates developmental dynamics by leveraging a transformer-based temporal diffusion backbone conditioned on continuous, gap-aware temporal encodings. Across diverse mouse developmental lineages, CellPace achieves state-of-the-art performance in simulation, interpolation, and forecasting tasks. Beyond recovering global population statistics, generated cells preserve fine-grained biological structure, retaining dynamic gene regulatory programs and mapping accurately to anatomical regions in spatial transcriptomics data. Furthermore, CellPace extends naturally to multi-modal data, modeling joint RNA-chromatin dynamics even when temporal ordering is inferred from pseudotime. Together, these results position CellPace as a robust framework for modeling and generating continuous developmental dynamics from sparse, cross-sectional single-cell data.
InnerQ: Hardware-aware Tuning-free Quantization of KV Cache for Large Language Models
Mohammadreza Tayaranian
Amir Ardakani
Warren J. Gross
Reducing the hardware footprint of large language models (LLMs) during decoding is critical for efficient long-sequence generation. A key bo… (see more)ttleneck is the key-value (KV) cache, whose size scales with sequence length and easily dominates the memory footprint of the model. Previous work proposed quantization methods that are focused on compressing the KV cache while maintaining its information. We introduce InnerQ, a hardware-aware KV-cache quantization scheme that lowers decode latency without sacrificing accuracy. InnerQ applies group-wise quantization while grouping the cache matrices over their inner dimension. Unlike previous work that group over the outer dimension, InnerQ aligns dequantization with the vector-matrix multiplication and enables scale factor reuse across GPU compute units. This reduces memory accesses and accelerates dequantization, yielding up to
Takeuchi's Information Criteria as Generalization Measures for DNNs Close to NTK Regime
Taiji Suzuki
Rio Yokota
Masahiro Nomura
Kohta Ishikawa
Ikuro Sato
Generalization measures have been studied extensively in the machine learning community to better characterize generalization gaps. However,… (see more) establishing a reliable generalization measure for statistically singular models such as deep neural networks (DNNs) is difficult due to their complex nature. This study focuses on Takeuchi's information criterion (TIC) to investigate the conditions under which this classical measure can effectively explain the generalization gaps of DNNs. Importantly, the developed theory indicates the applicability of TIC near the neural tangent kernel (NTK) regime. In a series of experiments, we trained more than 5,000 DNN models with 12 architectures, including large models (e.g., VGG-16), on four datasets, and estimated the corresponding TIC values to examine the relationship between the generalization gap and the TIC estimates. We applied several TIC approximation methods with feasible computational costs and assessed the accuracy trade-off. Our experimental results indicate that the estimated TIC values correlate well with the generalization gap under conditions close to the NTK regime. However, we show both theoretically and empirically that outside the NTK regime such correlation disappears. Finally, we demonstrate that TIC provides better trial pruning ability than existing methods for hyperparameter optimization.
Optimizing renal mass histology prediction with artificial intelligence using handcrafted radiomics
Teodora Boblea Podasca
Marc-Antoine Blais
Marie-Lou Gadoury-Campbell
Stéphanie Boulet
Patrick O. Richard
Soil microbiome prediction using traditional machine learning and deep learning models
Zahia Aouabed
Vincent Therrien
Mohamed Achraf Bouaoune
Mohammadreza Bakhtyari
Mohamed Hijri
The accuracy of macrobiological community predictions largely depends on the taxonomic scale considered. Nowadays, the applicability of such… (see more) predictions remains an important challenge when extended to microbial soil communities. This is not only due to the lack of reliable benchmark data, but also to a greater diversity of the soil microorganisms compared to other environments. In this study, we use six traditional machine learning regression models and one deep learning regressor to predict relative frequencies of bacterial and fungal communities within the soil microbiome based on environmental factors. We analyze the data from two publicly available soil microbiome datasets: (1) Data collected by Averill and co-authors and analyzed in a recent Nature Ecology and Evolution article, and (2) Data extracted from the NEON database, to estimate the composition of bacterial and fungal communities at the functional (i.e. functional group level) and taxonomic scales (i.e. phylum, class, order, family, and genus levels). Our findings suggest the presence of a general pattern across the observed taxonomic scales according to which the predictability of the soil microbiome increases with taxonomic scale. However, a notable exception occurs when machine learning models are applied to predict bacterial communities at the functional group level for Averill et al.’s data when all of them fail to provide accurate predictions results. The best overall results obtained include the value of the coefficient of determination
Soil microbiome prediction using traditional machine learning and deep learning models
Zahia Aouabed
Vincent Therrien
Mohamed Achraf Bouaoune
Mohammadreza Bakhtyari
Mohamed Hijri
The accuracy of macrobiological community predictions largely depends on the taxonomic scale considered. Nowadays, the applicability of such… (see more) predictions remains an important challenge when extended to microbial soil communities. This is not only due to the lack of reliable benchmark data, but also to a greater diversity of the soil microorganisms compared to other environments. In this study, we use six traditional machine learning regression models and one deep learning regressor to predict relative frequencies of bacterial and fungal communities within the soil microbiome based on environmental factors. We analyze the data from two publicly available soil microbiome datasets: (1) Data collected by Averill and co-authors and analyzed in a recent Nature Ecology and Evolution article, and (2) Data extracted from the NEON database, to estimate the composition of bacterial and fungal communities at the functional (i.e. functional group level) and taxonomic scales (i.e. phylum, class, order, family, and genus levels). Our findings suggest the presence of a general pattern across the observed taxonomic scales according to which the predictability of the soil microbiome increases with taxonomic scale. However, a notable exception occurs when machine learning models are applied to predict bacterial communities at the functional group level for Averill et al.’s data when all of them fail to provide accurate predictions results. The best overall results obtained include the value of the coefficient of determination