Publications

Cell type transcriptomic modules reveal shared molecular mechanisms in Alzheimer’s and Parkinson’s disease
Edward A. Fon
Alain Dagher
Yasser Iturria-Medina
Jo Anne Stratton
L. M. Hodgson
David A Bennett
Historically, Alzheimer's disease (AD) and Parkinson's disease (PD) have been investigated as two distinct disorders of the brain. However, … (voir plus)a few similarities in neuropathology and clinical symptoms have been documented over the years. Traditional single-gene centric studies, such as differential gene expression analyses, have struggled to unravel the molecular basis for the observed pathological links between AD and PD. To address this, we tailor a latent factor framework to analyze synchronous gene co-expression at sub-cell-type resolution. Utilizing large, single-nucleus transcriptomics datasets in AD (70,634 nuclei) and PD (340,902 nuclei) from postmortem human brains, we systematically extract and juxtapose disease-critical molecular signatures in the brain. Our transcriptomic analysis reveals shared molecular programs between AD and PD that systematically localize to specific glial and neuronal cell types. In neurons, convergent gene groups in AD and PD relate to cytoskeletal dynamics and mitochondrial stress mechanisms. Similarly, overlapping gene groups in microglia modules implicate T cell activation mechanisms and synapse pruning pathways. In parallel, AD- and PD-associated genes in astrocytes are involved in heavy metal processing; oligodendrocytes highlight convergent dysregulation in myelin synthesis. In addition, our analysis reveals APOE, an AD GWAS gene, has disease predictive roles in PD-associated gene modules. Conversely, SNCA, a PD GWAS gene, emerges within AD associated gene modules. Our multi-module sub-cell-type approach offers unique insights into the molecular basis of shared neuropathology in AD and PD.
Hydra: Towards Transferable Multi-Task Learning on Temporal Graphs
Kiarash Shamsi
Tran Gia Bao Ngo
Baris Coskunuzer
Michael M. Bronstein
Cuneyt Gurcan Akcora
Real-world evolving networks are naturally modeled as temporal graphs (TGs), where capturing temporal dynamics is essential for predicting f… (voir plus)uture graph properties that support downstream decision-making. Existing temporal graph methods have been developed primarily for single-task prediction, and little is known about their generalization across tasks or transfer to unseen networks. This leaves the challenge of multi-task graph property prediction in TGs largely open. We address this challenge by introducing Hydra, a novel architecture that integrates local connectivity features from temporal GNNs with a spectral learning module that captures global connectivity patterns. This design enables joint learning of local and global information under a multi-task objective. In multi-task classification, Hydra achieves an 8.9% relative gain in AUC over the strongest competitor. In multi-task regression, Hydra achieves competitive results in all three tasks, while obtaining the best results in two tasks with a 8.2% relative gain in MAE compared to the strongest baseline. Moreover, Hydra delivers these gains with a 22× reduction in training time compared to temporal transfer models. These results provide the first systematic evidence that multi-task transferable learning on temporal graphs is effective. By delivering consistent top-ranked performance, Hydra highlights multi-task training on temporal graphs as a promising direction toward adaptable foundation models for temporal graphs.
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
Flow matching with …
Sex-specific hormone-sensitive regulatory architecture in adolescence as a scaffold for depression vulnerability
Gladi Thng
Michel Garcia-Miranda
Kailu Song
Anjali Chawla
Reine Khoury
Minh Nguyen
Gabriella Frosi
Matthew Suderman
David Liao
Natalina Salmaso
Tie Yuan Zhang
Pan Wong Tak
Yashar Zeighami
Corina Nagy
External validation of cough-based algorithms for pulmonary tuberculosis screening from the CODA TB DREAM challenge using cough data from Peru
Alexandra J. Zimmer
Patricia Espinoza-Lopez
Vijay Ravi
Solveig K. Sieberts
Samira Abbasgholizadeh Rahimi
Madhukar Pai
Cesar Ugarte-Gil
Simon Grandjean Lapierre
The COugh Diagnostic Algorithm for Tuberculosis (CODA TB) DREAM Challenge recently evaluated the performance of artificial intelligence (AI)… (voir plus) algorithms for tuberculosis (TB) screening using cough sounds. Eleven AI models were developed using a dataset of 733,756 cough sounds collected from 2143 adults from seven countries. This study evaluates the CODA Challenge AI models with an external independent cough dataset from Peru. Cough recordings from 303 coughing adults were collected from health facilities in Lima, Peru. The AUCs of the models ranged from 0.480 to 0.615, showing a decrease in performance compared to their performance when internally validated using the CODA Challenge, which ranged from 0.689 to 0.743. The best performing model in the CODA Challenge was also the best performing model in this external validation. Sub-group analyses revealed that models performed better in older (≥ 35 years) populations and among people with prior TB. The external validation revealed limitations in the generalizability of the CODA Challenge models to other settings. While some models showed promise, the overall performance decline highlights the need for continued model validation on external datasets. It also underscores the importance of developing context-specific models to account for population-specific factors that influence cough characteristics and TB prevalence.
Mem-$π$: Adaptive Memory through Learning When and What to Generate
Chao Wang
Christopher Pal
Alexandre Lacoste
We present Mem-…
Model Stealing Through the Lens of Model Multiplicity
Model stealing attacks, where adversaries create high-fidelity surrogate models, are a significant threat to the intellectual property of ma… (voir plus)chine learning services. Conventional wisdom suggests these surrogates could provide adversaries with economic leverage comparable to the original service providers. This paper challenges this assumption by evaluating model stealing attacks beyond mere fidelity to the target model. Because query-based extraction provides only partial supervision of the target's input-output behavior, the surrogate is not uniquely identified: many near-optimal surrogates can achieve comparable fidelity while differing in deployment-relevant properties. Instead of performing a classic learning-based model stealing attack, we compute the Rashomon Set (i.e., the set of almost-equally-accurate models) of surrogate models, and evaluate its diversity using multiplicity metrics (ambiguity, discrepancy and rashomon capcity) and group fairness metrics. Our experiments on real-world datasets reveal that despite exhibiting similar fidelity to the target model, surrogate models can display significant variances in other critical performance metrics. These findings cast doubt on the presumed equivalence between high-fidelity surrogates and the target model in practical deployment scenarios.
Representations in vision and language converge in a shared, multidimensional space of perceived similarities
Katerina M. Simkova
Adrien Doerig
Clayton Hickey
Humans can effortlessly describe what they see, yet establishing a shared representational format between vision and language remains a sign… (voir plus)ificant challenge. Emerging evidence suggests that human brain representations in both vision and language are well predicted by semantic feature spaces obtained from large language models (LLMs). This raises the possibility that sensory systems converge in their inherent ability to transform their inputs onto shared, embedding-like representational space. However, it remains unclear how such a space manifests in human behavior. To investigate this, 63 participants performed behavioral similarity judgments separately on 100 natural scene images and 100 corresponding sentence captions from the Natural Scenes Dataset. We found that visual and linguistic similarity judgments not only converge at the behavioral level but also predict a remarkably similar network of functional magnetic resonance imaging brain responses evoked by viewing the natural scene images. Furthermore, computational models trained to map images onto LLM-embeddings outperformed both category-trained and AlexNet controls in predicting the behavioral similarity structure. These findings demonstrate that human visual and linguistic similarity judgments are grounded in a shared, modality-agnostic representational structure that mirrors how the visual system encodes experience. The convergence between sensory and artificial systems observed here suggests a common capacity of how conceptual representations are formed-not as arbitrary products of first order, modality-specific input, but as structured representations that reflect the stable, relational properties of the external world.
To Select or not to Select, that is the Question: Distilling Robot Skill Prediction into a Small Ensemble
Simon Roy
Euhid Aman
As robot fleets become more heterogeneous, including humanoids, rovers, quadrupeds, and drones, selecting the right robot for a task becomes… (voir plus) a core systems problem. We study robot skill prediction: mapping a natural-language task description to the physical capabilities required to execute it, such as fly, wheels, legs, surface water, under water and hands. Since labelled data that maps natural-language task descriptions to robot's physical capabilities does not exist, we construct a synthetic task-to-skill dataset using LLM-assisted generation and targeted label auditing. Trained on this data, a ~133M-parameter ensemble of two fine-tuned sentence encoders (mpnet + MiniLM) reaches 83.5% task-to-skill matching on a stratified 200 task dataset, outperforming Kimi K2 (1T MoE) at 72.0%, GPT-OSS-120B at 71.5%, and Llama-4-Scout-17B at 69.0% under the same zero-shot prompt. These results suggest that, for fixed robot skill taxonomies, small specialized models trained on synthetic data can outperform much larger general-purpose LLMs for fleet-level task routing.
Widespread use of invalid statistical tests in biomedical machine learning
Tianchu Zeng
Hui Li
Shaoshi Zhang
Yan Quan Tan
Fang Tian
Csaba Orbán
Lijun An
Wanyu Che
Jingwen Cheng
Joanna Su Xian Chong
Niousha Dehestani
Zijian Dong
Xin Li
Zhizhou Li
Mervyn Jun Rui Lim
Yi Lin
Qinrui Ling
Zijie Ling
Xi Zhi Low
Sina Mansour L. … (voir 24 de plus)
Kwun Kei Ng
Thuan Tinh Nguyen
Leon Qi Rong Ooi
Shreya Pande
Xing Qian
Jingxuan Ruan
Z WANG
Yapei Xie
Chen Zhang
Yichi Zhang
K Patil
Linden Parkes
Elvisha Dhamala
Sidhant Chopra
Andrew Zalesky
Avram Holmes
S Eickhoff
Juan Helen Zhou
Olivier Renaud
Nico Dosenbach
Konrad P. Kording
Thomas Nichols
B T Thomas Yeo
Abstract Machine learning is accelerating biomedical research. Cross-validation is widely used to compare predictive performance – not onl… (voir plus)y to benchmark algorithms, but also to inform scientific applications, such as ranking biomarkers. However, prediction performance estimates across cross-validation folds are not independent. Standard tests for comparing prediction performance (e.g., paired t-test) assume independence and can therefore inflate false positive rates. In a PRISMA-guided meta-analysis of 210 studies (impact factor ≥15, 1 June 2020 – 1 June 2025), we find that 97% ignored fold dependence when comparing prediction performance. This problem is ubiquitous across scientific fields and unaffected by impact factor, rigor-promoting policies, or open science practices. Simulations across 420 scenarios spanning four diverse datasets show that ignoring fold dependence leads to invalid false positive control in most settings. Repeated cross-validation further compounds this problem, with false positive rates rising toward 100% as the number of repetitions grows. Existing fold-dependence-aware tests rely on strong assumptions because the variance of fold-level statistics and the between-fold correlation cannot be disentangled under standard cross-validation. We therefore propose the SHARP (Split-HAlf RePeated) test, a simple modification to standard cross-validation that enables direct estimation of variance and correlation. Benchmarked against 12 tests, SHARP provides the best overall balance of false-positive control, statistical power, and confidence-interval calibration across simulation schemes. We conclude by providing best practices and reporting guidelines for valid model comparison inference in biomedical machine learning and beyond.
Characterization of limb representation in the pig’s motor cortex
David Bergeron
Hugo Delivet-Mongrain
Marina Martinez
Due to its large gyrencephalic brain, the pig is increasingly used for neuroscience research, especially for the preclinical testing of nove… (voir plus)l neuroprostheses. However, our understanding of the pig’s motor system remains limited compared to the common species used for neuroscience research. Here, we aimed to characterize the forelimb and hindlimb representation of the pig motor cortex using intracortical microstimulation (ICMS). Three domestic pigs ( Sus scrofa) were placed in a modified stereotactic frame and maintained under intravenous propofol sedation. We mapped the motor cortex using ICMS, applied at varying cortical coordinates and depths. For each site, we recorded the electrode depth eliciting the maximal limb response and determined the motor threshold. Responses were assessed visually and via electromyographic recordings. ICMS uncovered a large forelimb representation, with stereotypical contralateral responses. Conversely, the hindlimb representation was smaller and located within the interhemispheric fissure. The mean threshold of the five most responsive forelimb sites was 75 ± 25 μA, compared to 280 ± 45 μA for hindlimb sites (p<0.01). A summation of stimulations in the hindlimb representation of the motor cortex unilaterally triggered bilateral alternating hindlimb movements. These results suggest that while the porcine cortex can directly command forelimb movements via the corticospinal pathway, cortical control of hindlimb likely relies on polysynaptic pathways through the brainstem, such as the cortico-reticulospinal pathway.
Generative Recursive Reasoning
How should future neural reasoning systems implement extended computation? Recursive Reasoning Models (RRMs) offer a promising alternative t… (voir plus)o autoregressive sequence extension by performing iterative latent-state refinement with shared transition functions. Yet existing RRMs are largely deterministic, following a single latent trajectory and converging to a single prediction. We introduce Generative Recursive reAsoning Models (GRAM), a framework that turns recursive latent reasoning into probabilistic multi-trajectory computation. GRAM models reasoning as a stochastic latent trajectory, enabling multiple hypotheses, alternative solution strategies, and inference-time scaling through both recursive depth and parallel trajectory sampling. This yields a latent-variable generative model supporting conditional reasoning via