Publications

Memorandum, Notes on Visitations and Ongoing Projects in Canada (Version -1)
Nonvikan Karl-Augustt Alahassa
Bidossessi R.U. Alahassa
Julie Carrier
Nathalie Lacelle
Samuel Bassetto
Leonard Wantchekon
Mylène Bédard
Emmanuel Stip
Bruno Rémillard
Suljo Linic
Joel TOSSA
Bakary Manga
Jérôme Théau
Victor M. Panaretos
David Haziza
Daniel F. Nadeau
Marlène Frigon
Christiane Rousseau
Maciej Augustyniak … (see 6 more)
Dimitrios Koukoulopoulos
Élise Vandomme
Wilfrid Gangbo
Raphael R. Kelani
J. B. Chabi Orou
Cyriaque Atindogbé
We are on the Refugee Claimant list in Canada. We would like to express few notes of Memorandum.
Memorandum, Notes on Visitations and Ongoing Projects in Canada (Version -1)
Nonvikan Karl-Augustt Alahassa
Bidossessi R.U. Alahassa
Julie Carrier
Nathalie Lacelle
Samuel Bassetto
Leonard Wantchekon
Mylène Bédard
Emmanuel Stip
Bruno Rémillard
Suljo Linic
J. Tossa
Bakary Manga
Jérôme Théau
Victor M. Panaretos
David Haziza
Daniel F. Nadeau
Marlène Frigon
Christiane Rousseau
Maciej Augustyniak … (see 8 more)
Dimitrios Koukoulopoulos
Vandomme, Élise
Gangbo, Wilfrid
Kelani, Raphael
Chabi Orou, Jean Bio
ATINDOGBE, Comlan Cyriaque
Wilfrid Gangbo
Raphael Kelani
We are on the Refugee Claimant list in Canada. We would like to express few notes of Memorandum.
The Project for Open Source Email Writing [or] Public Open Letters with Regular Updates and Contributives Comments
Bidossessi R.U. Alahassa
Nonvikan Karl-Augustt Alahassa
Julie Carrier
Joel TOSSA
Wilfrid Gangbo
Cyriaque Atindogbé
Marlène Frigon
Christiane Rousseau
Maciej Augustyniak
Dimitrios Koukoulopoulos
Samuel Bassetto
J. B. Chabi Orou
Raphael R. Kelani
Jérôme Théau
Damien Échevin
Leonard Wantchekon
David Haziza
Emmanuel Stip
Mylène Bédard … (see 8 more)
Élise Vandomme
Mouloud-Beallah Belbahri
Bakary Manga
Daniel F. Nadeau
Suljo Linic
Bruno Rémillard
Victor M. Panaretos
Nathalie Lacelle
We express an interest to crowd collaborative writing projects for e-mails. E-mails are programs as well that deserve updates and releases.
Building2Building: A Large Scale Benchmark for Generalizable Real-World Reinforcement Learning
Reinforcement learning (RL) has achieved strong results in control, yet learned policies remain brittle to changes in dynamics, action space… (see more)s, observation spaces, or goals, a critical limitation for real-world deployment. Existing benchmarks offer limited diversity and complexity, making it difficult to rigorously study transfer, multi-task learning, and meta-learning in RL. We introduce Building2Building (B2B), a large-scale suite of realistic Heating, Ventilation, and Air Conditioning (HVAC) control environments built on EnergyPlus, a state-of-the-art building simulator. B2B is fully compatible with the Gymnasium interface and features a parametric building generator, enabling the systematic generation of diverse building configurations with heterogeneous observation and action spaces. Based on this suite, we define benchmark tasks targeting key open challenges in RL, including goal adaptation, dynamics adaptation, action-space shifts, and cross-domain transfer. By providing a large-scale, diverse, and physically grounded testbed with standardized evaluation protocols, B2B enables systematic investigation of generalization and transfer in continuous control. Beyond advancing research on generalization in RL, this new benchmark also carries significant societal implications by enabling improved HVAC control at scale, one of the most energy-intensive systems in buildings.
Across-Site MRI Prediction of Substantial Lymphovascular Space Invasion in Endometrial Cancer: Radiomics versus Deep Learning Features
Daniel A. Di Giovanni
Akiyo Tanaka
Takuro Horikoshi
Takahiro Tsuboyama
H Yokota
Rita Zakarian
Yuka Matsumoto
Caroline Reinhold
Purpose: To compare the cross-site generalization of radiomic features and deep learning embeddings for MRI prediction of substantial lympho… (see more)vascular space invasion (LVSI) in endometrial cancer. Materials and Methods: This retrospective two-center study included 206 women (mean age, 59.8 years) with endometrial cancer who underwent preoperative 3-T MRI from March 2016 to March 2023. Hospital A (n = 130) was used for development and Hospital B (n = 76) for strict external testing. T2-weighted, reduced field-of-view diffusion-weighted, and apparent diffusion coefficient images were manually segmented. Radiomic features and seed-pooled embeddings from 3D ResNet18, DenseNet121, and U-NEXtractor were modeled with elastic-net logistic regression or XGBoost. Out-of-fold Platt calibration and sensitivity-targeted thresholds were estimated using development data only. AUCs were summarized with 95% bootstrap confidence intervals. Results: External radiomics with elastic-net achieved an AUC of 0.609 (95% CI: 0.464, 0.740) and sensitivity of 0 of 12 (0%). DenseNet121 with elastic-net had the highest external AUC (0.685; 95% CI: 0.538, 0.822) but sensitivity of 3 of 12 (25%). U-NEXtractor with elastic-net detected 10 of 12 positive cases (83.3%) with specificity of 32 of 64 (50.0%) and balanced accuracy of 0.667. XGBoost showed higher apparent development performance but weaker external operating behavior. Conclusion: Under real-world cross-site MRI acquisition shift, DenseNet121 and U-NEXtractor embeddings showed better external generalization than handcrafted radiomic features for substantial LVSI prediction.
Freeze, Diffuse, Decode: Task-Aware Adaptation of Transformer Embeddings for Antimicrobial Peptide Design
Pankhil Gawade
Adam Izdebski
Kevin R. Moon
Jake Slater Rhodes
Ewa Szczurek
Pretrained transformers provide general-purpose molecular embeddings for downstream tasks. While these embeddings provide task-agnostic stru… (see more)ctural patterns, they lack task-specific alignment, limiting downstream performance. Here, we introduce Freeze, Diffuse, Decode (FDD), a diffusion-based framework that adapts pretrained transformer embeddings to downstream tasks by building a task supervised diffusion geometry over the frozen embeddings, without any backbone training. Applied to antimicrobial peptide design, FDD yields low-dimensional, predictive, and interpretable representations that support property prediction, retrieval, and latent-space interpolation.
Relevant and Irrelevant: A Renormalization Group Analysis of Transformer Attention
Parviz Haggi-Mani
Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the t… (see more)rained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator. We derive a fixed-point shift formula and obtain four testable predictions for the fixed-point geometry, effective rank profile, layer specificity, and perturbation decay spectrum. Testing these on synthetic Markov chain sequences with controlled correlation length, we find: (1) For large chains(long correlation), attention is strongly relevant: it closes a residual loss gap the MLP cannot bridge and drives a phase transition in representation space, with effective rank jumping above input dimensionality at layer 1 and stabilizing at a high-dimensional plateau. (2) For short chains(short correlation), attention is irrelevant: the Transformer converges to the same loss and fixed-point geometry as the MLP, though it contracts perturbations faster. (3) The transition is dominated by the first-layer head (L0H0), which accounts for more than 4 times the representational shift of any subsequent head, consistent with the prediction that the relevant operator acts before the MLP begins integrating out positional variation. (4) Perturbation decay experiments reveal a regime reversal: in the long correlation regime the Transformer selectively preserves slow Markov modes (5.4 times the dynamic range in decay length vs. 1.3 times for the MLP); in the short correlation regime it suppresses all modes faster than the MLP, with no spectral selectivity. Together, these results show that the relevance of attention is not a property of the architecture but of the spectral structure of the data-generating process, and that a first-order RG perturbation framework provides a predictive account of that difference.
Toward a general understanding of neural representations learned by deep neural networks on group multiplications
We study the neural representations learned by deep neural networks trained on alternating group multiplication, discovering that they are S… (see more)chreier coset graphs, a generalization of Cayley graphs. Previous works inspected neural representations learned from cyclic group multiplication and identified them as Cayley graphs. Since cyclic group multiplication is Abelian, all subgroups are normal, and Schreier coset graphs reduce to Cayley graphs in this case. This contribution makes a step toward finding a general theory of what neural representations are in networks learning group multiplications.
Towards Distilling the Representational Geometry of Language Models
Andrew J Steindl
Dhananjay Bhaskar
Ian Adelstein
Large neural networks are increasingly combined through weight-space merging and distilled into smaller models, but weights and activations … (see more)are not canonical coordinates: two networks can implement similar functions while differing by permutations, rescalings, or basis changes. We ask whether these operations can instead be guided by the geometry of a model's representation space. For each model, we summarize the relationships among probe examples using a diffusion operator built from internal activations. We then test a simple distillation procedure in which a student is trained to match a fixed diffusion-geometry target derived from a teacher. Across Pythia, BERT, and DeBERTa-v3, this target does not generally make the student inherit the teacher's specific representation geometry. Instead, distilled students usually remain closer to matched task-only controls, and the objective has little effect on held-out performance. These results suggest that representation-geometry objectives must be evaluated against the geometry a student would learn from ordinary training alone.
Discrete Diffusion Models: A Unified Framework from Tokenization to Generation
Weien Li
Rui Song
Zeyu Li
Haochen Liu
Xiangyu Kong
Zixuan Dong
Jiaxin Huang
Changjiang Han
Yonghan Yang
Zichen Zhao
Xiuyuan Hu
Yankai Chen
Fengran Mo
Jikun Kang
Bowei He … (see 2 more)
Philip S. Yu
Xue Liu
Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete da… (see more)ta, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed, DDMs are fundamentally shaped by how the discrete state space is constructed: the tokenization scheme, the vocabulary topology, and domain-specific structural alphabets. This work introduces a unified conceptual framework that views discrete diffusion models through the construction of the underlying discrete state space. Within this framework, existing formulations, including transition-matrix, masking/absorbing-state, and score/ratio-based approaches, emerge as different instantiations of a common design space. The framework further exposes common design trade-offs across training objectives, inference algorithms, scaling behavior, systems optimization, and evaluation protocols, suggesting several promising directions for future research.
Leveraging unlabelled data for generalizable neural population decoding
Nanda H. Krishna
Avery Hee-Woon Ryoo
Matthew G. Perich
Robust and accurate neural decoders are integral to neurotechnologies such as brain-computer interfaces and closed-loop experiments. Recent … (see more)work has shown that tokenizing neural data at the spike level facilitates multi-session pretraining and delivers state-of-the-art decoding performance. However, current spike-based models are restricted to supervised learning (SL), limiting training to datasets with paired behavioural labels. To address this limitation, we introduce MOJO (Masked autOencoder-based JOint training), a training framework for spike-tokenizing models that jointly leverages self-supervised learning (SSL) via masked autoencoding and SL objectives. We evaluate MOJO on three spiking datasets spanning monkey motor cortex during reaching tasks and multi-regional mouse recordings during vision and decision making tasks, demonstrating superior performance over purely SL-trained models. This improvement is especially pronounced when training with limited labelled data, particularly in few-shot finetuning, where only a small amount of labelled data from a new session is available. Incorporating SSL also yields more interpretable neuronal representations, improving performance on brain region classification and spike-statistics prediction without explicit optimization for these tasks. We further show that MOJO generalizes beyond spiking data to human electrocorticography during speech, where it continues to outperform purely SL-trained models and achieves performance comparable to neuro-foundation models (NFMs) designed specifically for continuous signals. Overall, augmenting spike-tokenizing models with SSL improves performance in label-impoverished settings and enables the use of unlabelled data across various tasks and species, while generalizing to other neural modalities. These results suggest a path towards more flexible and scalable data usage when training NFMs.
Plausible Deniability Guarantees for Whistleblowers
Leo Richter
Matt J. Kusner
Whistleblowers are a key safeguard against organizational wrongdoing, but the threat of retaliation deters reporting. Existing whistleblower… (see more)-protection proposals lack formal privacy guarantees, and existing differential privacy mechanisms do not directly target the natural threat model -- one in which the audited organization itself observes auditor selection decisions and uses them to identify reporters. We formalize protection against a strong-adversary threat model as per-report