Publications

The Project for Open Source Email Writing [or] Public Open Letters with Regular Updates and Contributives Comments
Bidossessi R.U. Alahassa
Nonvikan Karl-Augustt Alahassa
Julie Carrier
Joel TOSSA
Wilfrid Gangbo
Cyriaque Atindogbé
Marlène Frigon
Christiane Rousseau
Maciej Augustyniak
Dimitrios Koukoulopoulos
Samuel Bassetto
J. B. Chabi Orou
Raphael R. Kelani
Jérôme Théau
Damien Échevin
Leonard Wantchekon
David Haziza
Emmanuel Stip
Mylène Bédard … (see 8 more)
Élise Vandomme
Mouloud-Beallah Belbahri
Bakary Manga
Daniel F. Nadeau
Suljo Linic
Bruno Rémillard
Victor M. Panaretos
Nathalie Lacelle
We express an interest to crowd collaborative writing projects for e-mails. E-mails are programs as well that deserve updates and releases.
Across-Site MRI Prediction of Substantial Lymphovascular Space Invasion in Endometrial Cancer: Radiomics versus Deep Learning Features
Daniel A. Di Giovanni
Akiyo Tanaka
Takuro Horikoshi
Takahiro Tsuboyama
H Yokota
Rita Zakarian
Yuka Matsumoto
Caroline Reinhold
Purpose: To compare the cross-site generalization of radiomic features and deep learning embeddings for MRI prediction of substantial lympho… (see more)vascular space invasion (LVSI) in endometrial cancer. Materials and Methods: This retrospective two-center study included 206 women (mean age, 59.8 years) with endometrial cancer who underwent preoperative 3-T MRI from March 2016 to March 2023. Hospital A (n = 130) was used for development and Hospital B (n = 76) for strict external testing. T2-weighted, reduced field-of-view diffusion-weighted, and apparent diffusion coefficient images were manually segmented. Radiomic features and seed-pooled embeddings from 3D ResNet18, DenseNet121, and U-NEXtractor were modeled with elastic-net logistic regression or XGBoost. Out-of-fold Platt calibration and sensitivity-targeted thresholds were estimated using development data only. AUCs were summarized with 95% bootstrap confidence intervals. Results: External radiomics with elastic-net achieved an AUC of 0.609 (95% CI: 0.464, 0.740) and sensitivity of 0 of 12 (0%). DenseNet121 with elastic-net had the highest external AUC (0.685; 95% CI: 0.538, 0.822) but sensitivity of 3 of 12 (25%). U-NEXtractor with elastic-net detected 10 of 12 positive cases (83.3%) with specificity of 32 of 64 (50.0%) and balanced accuracy of 0.667. XGBoost showed higher apparent development performance but weaker external operating behavior. Conclusion: Under real-world cross-site MRI acquisition shift, DenseNet121 and U-NEXtractor embeddings showed better external generalization than handcrafted radiomic features for substantial LVSI prediction.
Relevant and Irrelevant: A Renormalization Group Analysis of Transformer Attention
Parviz Haggi-Mani
Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the t… (see more)rained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator. We derive a fixed-point shift formula and obtain four testable predictions for the fixed-point geometry, effective rank profile, layer specificity, and perturbation decay spectrum. Testing these on synthetic Markov chain sequences with controlled correlation length, we find: (1) For large chains(long correlation), attention is strongly relevant: it closes a residual loss gap the MLP cannot bridge and drives a phase transition in representation space, with effective rank jumping above input dimensionality at layer 1 and stabilizing at a high-dimensional plateau. (2) For short chains(short correlation), attention is irrelevant: the Transformer converges to the same loss and fixed-point geometry as the MLP, though it contracts perturbations faster. (3) The transition is dominated by the first-layer head (L0H0), which accounts for more than 4 times the representational shift of any subsequent head, consistent with the prediction that the relevant operator acts before the MLP begins integrating out positional variation. (4) Perturbation decay experiments reveal a regime reversal: in the long correlation regime the Transformer selectively preserves slow Markov modes (5.4 times the dynamic range in decay length vs. 1.3 times for the MLP); in the short correlation regime it suppresses all modes faster than the MLP, with no spectral selectivity. Together, these results show that the relevance of attention is not a property of the architecture but of the spectral structure of the data-generating process, and that a first-order RG perturbation framework provides a predictive account of that difference.
Leveraging unlabelled data for generalizable neural population decoding
Nanda H. Krishna
Avery Hee-Woon Ryoo
Matthew G. Perich
Robust and accurate neural decoders are integral to neurotechnologies such as brain-computer interfaces and closed-loop experiments. Recent … (see more)work has shown that tokenizing neural data at the spike level facilitates multi-session pretraining and delivers state-of-the-art decoding performance. However, current spike-based models are restricted to supervised learning (SL), limiting training to datasets with paired behavioural labels. To address this limitation, we introduce MOJO (Masked autOencoder-based JOint training), a training framework for spike-tokenizing models that jointly leverages self-supervised learning (SSL) via masked autoencoding and SL objectives. We evaluate MOJO on three spiking datasets spanning monkey motor cortex during reaching tasks and multi-regional mouse recordings during vision and decision making tasks, demonstrating superior performance over purely SL-trained models. This improvement is especially pronounced when training with limited labelled data, particularly in few-shot finetuning, where only a small amount of labelled data from a new session is available. Incorporating SSL also yields more interpretable neuronal representations, improving performance on brain region classification and spike-statistics prediction without explicit optimization for these tasks. We further show that MOJO generalizes beyond spiking data to human electrocorticography during speech, where it continues to outperform purely SL-trained models and achieves performance comparable to neuro-foundation models (NFMs) designed specifically for continuous signals. Overall, augmenting spike-tokenizing models with SSL improves performance in label-impoverished settings and enables the use of unlabelled data across various tasks and species, while generalizing to other neural modalities. These results suggest a path towards more flexible and scalable data usage when training NFMs.
The topology of adolescent mental health
Maria B. Jelen
Alexa Mousley
Kayson Fakhar
Estherina Trachtenberg
Yuankai He
Robert Kohler
Varun Warrier
Sarah W. Yip
Duncan E. Astle
Abstract The increased vulnerability to mental health problems in adolescence is frequently reported but poorly understood, hampered by a ri… (see more)gid diagnostic system which fails to capture intertwining symptoms and only loosely aligns with biological axes of variability. Here, we reconceptualised the mental health symptoms of young adolescents in the ABCD cohort (N=11862) as a latent topology of overlapping symptom dimensions, using an unsupervised machine learning algorithm to establish how transdiagnostic dimensions co-occur and overlap within individuals. Combining this with a novel classification approach, we delineated zones within this landscape, within which specific profiles of symptoms were robustly represented. These data-driven profiles were leveraged to establish associated resting-state functional connectivity and genetic characteristics. In doing so we recaptured the commonly reported p- factor axis as well as further symptom-subtype dimensions. Gene ontology analysis revealed that shared neurobiological and cellular mechanisms embedded in both the genome and transcriptome may confer risk for psychopathology.
Optimal energy trading in residential prosumer clusters via graphon mean field games
Copositive Characterizations of Convex Hull Pricing
Madhusudan Ghosh
Joshua Adam Taylor
Due to the nonconvex binary constraints of unit commitment (UC), no uniform linear pricing scheme supports the optimal dispatch. Convex hull… (see more) pricing (CHP) and copositive duality pricing (CDP) both address this problem. CHP derives the price from the subgradient of the value function of the convex hull relaxation of UC. CDP refers to several different pricing mechanisms that can be constructed from the dual multipliers of the completely positive programming reformulation. In this work, we define a centralized convex hull price over the joint feasible set of UC and prove that, under non-degeneracy, it coincides with the marginal copositive duality price. Numerical experiments on the Scarf example validate this equivalence and quantify the pricing gap introduced by the semidefinite restriction.
From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP
Satwik Bhattamishra
Michael Hahn
A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). … (see more)There is much work analyzing the expressivity of attention-based models. By proposing handcrafted weights or using computational complexity arguments, a large amount of past theoretical works have sought to characterize which tasks are and which are not in the hypothesis class of Transformer models. However, little work investigates the learnability of such solutions. In this work, we make progress towards this goal. Inspired by recent loss landscape analysis work, we propose preliminary sample complexity bounds for learning C-RASP constructions with Transformers.
Special issue dedicated to the International Symposium on Mathematical Programming (ISMP), Montréal, 2024
Youssef Diouane
Franklin Djeumou Fomeni
Alain Hertz
Dominique Orban
Density Evolution of Soft-Decision Collapsed Projection-Aggregation Decoding for Reed-Muller Codes over the BIAWGN Channel
Jiajie Li
Marvin Rübenacke
Warren J. Gross
Reed-Muller (RM) codes have been shown to achieve capacity over a range of channels, and recently proposed projection-aggregation (PA) decod… (see more)ing has been experimentally shown to achieve near-maximum-likelihood decoding performance. These recent achievements motivate theoretical research on PA decoding. In this work, we analyze the density function of the soft output from collapsed projection-aggregation (CPA) decoding for RM codes over the binary-input additive white Gaussian noise (BIAWGN) channel. We prove that soft-decision CPA decoding returns an exact marginal probability and is symmetric. Based on the analysis, we build a density evolution model for CPA decoding. To simplify the density evolution, we approximate the projection and the fast Hadamard transform decoding using hard-decision decoding. Simulation results over the BIAWGN channel show that our proposed density evolution model captures the fast reduction in the mean and the variance of the soft information returned from the CPA decoding, which qualitatively explains the decoding mechanism and the fast convergence speed of the CPA decoding. We perform an asymptotic analysis based on the proposed density evolution, and we show that CPA decoding can achieve a vanishing error probability for RM codes with a vanishing code rate.
Exploring Test-time Scaling via Prediction Merging on Large-Scale Recommendation
Fuyuan Lyu
Z Chen
Jingyan Jiang
Lingjie Li
Xing Tang
xiuqiang He
Xue Liu
Inspired by the success of language models (LM), scaling up deep learning recommendation systems (DLRS) has become a recent trend in the com… (see more)munity. All previous methods tend to scale up the model parameters during training time. However, how to efficiently utilize and scale up computational resources during test time remains underexplored, which can prove to be a scaling-efficient approach and bring orthogonal improvements in LM domains. The key point in applying test-time scaling to DLRS lies in effectively generating diverse yet meaningful outputs for the same instance. We propose two ways: One is to explore the heterogeneity of different model architectures. The other is to utilize the randomness of model initialization under a homogeneous architecture. The evaluation is conducted across eight models, including both classic and SOTA models, on three benchmarks. Sufficient evidence proves the effectiveness of both solutions. We further prove that under the same inference budget, test-time scaling can outperform parameter scaling. Our test-time scaling can also be seamlessly accelerated with the increase in parallel servers when deployed online, without affecting the inference time on the user side. Code is available here. https://github.com/aTitye/TTS4CTR.
Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning
Charles Edward Gagnon
Steven H. H. Ding
Philippe Charland
Benjamin C. M. Fung
We present a practical pipeline for recovering source code from stripped binary functions by combining reverse engineering, anchor-based sou… (see more)rce code retrieval, and large language model reasoning. Our binary-to-source-code retrieval method attempts to identify the source function from a source code database, rather than generating approximate decompiled pseudocode. It extracts anchors such as strings, constants, external calls, and available function names using Ghidra, retrieves candidate files via an inverted-index search database, narrows candidates to likely function snippets, and re-ranks them with a large language model (LLM) based on disassembly, decompiled code, and source metadata. Confident matches can also serve as anchors in later passes. In an evaluation backed by our high-fidelity source code database on a stripped, optimized tcpdump binary, our proposed binary-to-source matching method achieves 95.2% assembly instruction coverage. Experiments on a GitHub-based retrieval database showed lower performance with 35.5% instruction coverage on average, mainly due to retrieval misses. These results show that source-level binary recovery excels with high-quality databases and remains a useful tool in noisy environments.