Publications

Sex-specific hormone-sensitive regulatory architecture in adolescence as a scaffold for depression vulnerability
Gladi Thng
Michel Garcia-Miranda
Kailu Song
Anjali Chawla
Reine Khoury
Minh Nguyen
Gabriella Frosi
Matthew Suderman
David Liao
Natalina Salmaso
Tie Yuan Zhang
Pan Wong Tak
Yashar Zeighami
Corina Nagy
External validation of cough-based algorithms for pulmonary tuberculosis screening from the CODA TB DREAM challenge using cough data from Peru
Alexandra J. Zimmer
Patricia Espinoza-Lopez
Vijay Ravi
Solveig K. Sieberts
Samira Abbasgholizadeh Rahimi
Madhukar Pai
César Ugarte-Gil
Simon Grandjean Lapierre
The COugh Diagnostic Algorithm for Tuberculosis (CODA TB) DREAM Challenge recently evaluated the performance of artificial intelligence (AI)… (see more) algorithms for tuberculosis (TB) screening using cough sounds. Eleven AI models were developed using a dataset of 733,756 cough sounds collected from 2143 adults from seven countries. This study evaluates the CODA Challenge AI models with an external independent cough dataset from Peru. Cough recordings from 303 coughing adults were collected from health facilities in Lima, Peru. The AUCs of the models ranged from 0.480 to 0.615, showing a decrease in performance compared to their performance when internally validated using the CODA Challenge, which ranged from 0.689 to 0.743. The best performing model in the CODA Challenge was also the best performing model in this external validation. Sub-group analyses revealed that models performed better in older (≥ 35 years) populations and among people with prior TB. The external validation revealed limitations in the generalizability of the CODA Challenge models to other settings. While some models showed promise, the overall performance decline highlights the need for continued model validation on external datasets. It also underscores the importance of developing context-specific models to account for population-specific factors that influence cough characteristics and TB prevalence.
Mem-$π$: Adaptive Memory through Learning When and What to Generate
Chao Wang
Christopher Pal
Alexandre Lacoste
We present Mem-…
Model Stealing Through the Lens of Model Multiplicity
Model stealing attacks, where adversaries create high-fidelity surrogate models, are a significant threat to the intellectual property of ma… (see more)chine learning services. Conventional wisdom suggests these surrogates could provide adversaries with economic leverage comparable to the original service providers. This paper challenges this assumption by evaluating model stealing attacks beyond mere fidelity to the target model. Because query-based extraction provides only partial supervision of the target's input-output behavior, the surrogate is not uniquely identified: many near-optimal surrogates can achieve comparable fidelity while differing in deployment-relevant properties. Instead of performing a classic learning-based model stealing attack, we compute the Rashomon Set (i.e., the set of almost-equally-accurate models) of surrogate models, and evaluate its diversity using multiplicity metrics (ambiguity, discrepancy and rashomon capcity) and group fairness metrics. Our experiments on real-world datasets reveal that despite exhibiting similar fidelity to the target model, surrogate models can display significant variances in other critical performance metrics. These findings cast doubt on the presumed equivalence between high-fidelity surrogates and the target model in practical deployment scenarios.
Representations in vision and language converge in a shared, multidimensional space of perceived similarities
Katerina M. Simkova
Adrien Doerig
Clayton Hickey
Humans can effortlessly describe what they see, yet establishing a shared representational format between vision and language remains a sign… (see more)ificant challenge. Emerging evidence suggests that human brain representations in both vision and language are well predicted by semantic feature spaces obtained from large language models (LLMs). This raises the possibility that sensory systems converge in their inherent ability to transform their inputs onto shared, embedding-like representational space. However, it remains unclear how such a space manifests in human behavior. To investigate this, 63 participants performed behavioral similarity judgments separately on 100 natural scene images and 100 corresponding sentence captions from the Natural Scenes Dataset. We found that visual and linguistic similarity judgments not only converge at the behavioral level but also predict a remarkably similar network of functional magnetic resonance imaging brain responses evoked by viewing the natural scene images. Furthermore, computational models trained to map images onto LLM-embeddings outperformed both category-trained and AlexNet controls in predicting the behavioral similarity structure. These findings demonstrate that human visual and linguistic similarity judgments are grounded in a shared, modality-agnostic representational structure that mirrors how the visual system encodes experience. The convergence between sensory and artificial systems observed here suggests a common capacity of how conceptual representations are formed-not as arbitrary products of first order, modality-specific input, but as structured representations that reflect the stable, relational properties of the external world.
To Select or not to Select, that is the Question: Distilling Robot Skill Prediction into a Small Ensemble
Simon Roy
Euhid Aman
As robot fleets become more heterogeneous, including humanoids, rovers, quadrupeds, and drones, selecting the right robot for a task becomes… (see more) a core systems problem. We study robot skill prediction: mapping a natural-language task description to the physical capabilities required to execute it, such as fly, wheels, legs, surface water, under water and hands. Since labelled data that maps natural-language task descriptions to robot's physical capabilities does not exist, we construct a synthetic task-to-skill dataset using LLM-assisted generation and targeted label auditing. Trained on this data, a ~133M-parameter ensemble of two fine-tuned sentence encoders (mpnet + MiniLM) reaches 83.5% task-to-skill matching on a stratified 200 task dataset, outperforming Kimi K2 (1T MoE) at 72.0%, GPT-OSS-120B at 71.5%, and Llama-4-Scout-17B at 69.0% under the same zero-shot prompt. These results suggest that, for fixed robot skill taxonomies, small specialized models trained on synthetic data can outperform much larger general-purpose LLMs for fleet-level task routing.
Widespread use of invalid statistical tests in biomedical machine learning
Tianchu Zeng
Hui Li
Shaoshi Zhang
Yan Quan Tan
Fang Tian
Csaba Orbán
Lijun An
Wanyu Che
Jingwen Cheng
Joanna Su Xian Chong
Niousha Dehestani
Zijian Dong
Xin Li
Zhizhou Li
Mervyn Jun Rui Lim
Yi Lin
Qinrui Ling
Zijie Ling
Xi Zhi Low
Sina Mansour L. … (see 24 more)
Kwun Kei Ng
Thuan Tinh Nguyen
Leon Qi Rong Ooi
Shreya Pande
Xing Qian
Jingxuan Ruan
Z WANG
Yapei Xie
Chen Zhang
Yichi Zhang
K Patil
Linden Parkes
Elvisha Dhamala
Sidhant Chopra
Andrew Zalesky
Avram Holmes
S Eickhoff
Juan Helen Zhou
Olivier Renaud
Nico Dosenbach
Konrad P. Kording
Thomas Nichols
B T Thomas Yeo
Abstract Machine learning is accelerating biomedical research. Cross-validation is widely used to compare predictive performance – not onl… (see more)y to benchmark algorithms, but also to inform scientific applications, such as ranking biomarkers. However, prediction performance estimates across cross-validation folds are not independent. Standard tests for comparing prediction performance (e.g., paired t-test) assume independence and can therefore inflate false positive rates. In a PRISMA-guided meta-analysis of 210 studies (impact factor ≥15, 1 June 2020 – 1 June 2025), we find that 97% ignored fold dependence when comparing prediction performance. This problem is ubiquitous across scientific fields and unaffected by impact factor, rigor-promoting policies, or open science practices. Simulations across 420 scenarios spanning four diverse datasets show that ignoring fold dependence leads to invalid false positive control in most settings. Repeated cross-validation further compounds this problem, with false positive rates rising toward 100% as the number of repetitions grows. Existing fold-dependence-aware tests rely on strong assumptions because the variance of fold-level statistics and the between-fold correlation cannot be disentangled under standard cross-validation. We therefore propose the SHARP (Split-HAlf RePeated) test, a simple modification to standard cross-validation that enables direct estimation of variance and correlation. Benchmarked against 12 tests, SHARP provides the best overall balance of false-positive control, statistical power, and confidence-interval calibration across simulation schemes. We conclude by providing best practices and reporting guidelines for valid model comparison inference in biomedical machine learning and beyond.
Characterization of limb representation in the pig’s motor cortex
David Bergeron
Hugo Delivet-Mongrain
Marina Martinez
Due to its large gyrencephalic brain, the pig is increasingly used for neuroscience research, especially for the preclinical testing of nove… (see more)l neuroprostheses. However, our understanding of the pig’s motor system remains limited compared to the common species used for neuroscience research. Here, we aimed to characterize the forelimb and hindlimb representation of the pig motor cortex using intracortical microstimulation (ICMS). Three domestic pigs ( Sus scrofa) were placed in a modified stereotactic frame and maintained under intravenous propofol sedation. We mapped the motor cortex using ICMS, applied at varying cortical coordinates and depths. For each site, we recorded the electrode depth eliciting the maximal limb response and determined the motor threshold. Responses were assessed visually and via electromyographic recordings. ICMS uncovered a large forelimb representation, with stereotypical contralateral responses. Conversely, the hindlimb representation was smaller and located within the interhemispheric fissure. The mean threshold of the five most responsive forelimb sites was 75 ± 25 μA, compared to 280 ± 45 μA for hindlimb sites (p<0.01). A summation of stimulations in the hindlimb representation of the motor cortex unilaterally triggered bilateral alternating hindlimb movements. These results suggest that while the porcine cortex can directly command forelimb movements via the corticospinal pathway, cortical control of hindlimb likely relies on polysynaptic pathways through the brainstem, such as the cortico-reticulospinal pathway.
Improved Ising Model Formulation for Polar Codes
Ryan Seah
Warren J. Gross
This paper presents an improved Ising model framework for polar codes, termed POLARIS, which reduces the number of binary variables by incor… (see more)porating rate-1 node structures and embedding elements of successive-cancellation decoding into the Ising formulation. The decoder scales efficiently to block lengths up to N = 64, doubling prior Ising-based limits. POLARIS achieves near-successive-cancellation list performance within 0.4 dB while reducing QUBO dimensionality from 192 to 126 variables. These advancements bring Ising-based polar decoding closer to practical realization, offering improved efficiency for implementation on both quantum and hybrid CMOS-classical annealing hardware.
LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series
Can language-pretrained transformers become effective time-series forecasters, and why? In this paper, we show that cross-modal transfer ari… (see more)ses because language pretraining preconditions time series training with a reusable manifold. A linear probe on frozen LLM states decodes realistic time-series trajectories without paired supervision, and retrieval in this projected space yields competitive forecasts, showing that structure and dynamics exist before finetuning. Pretrained initialization also improves optimization, producing coherent gradients and a highly anisotropic loss landscape unlike random initialization. Finetuning then acts as low-dimensional alignment, reusing existing directions rather than learning temporal primitives from scratch, as evidenced by low-rank updates, subspace alignment, and shared features for periodicity, trend, and repetition. Together, these results support a geometric account of LLM-to-time-series transfer: language pretraining builds the manifold, and finetuning projects numerical dynamics onto task-relevant directions.
Matérn Noise for Triangulation-Agnostic Flow Matching on Meshes
Arman Maesumi
Daniel Ritchie
This paper tackles the task of learning to generate signals over triangle meshes in a triangulation-agnostic manner, meaning the trained mod… (see more)el can be applied to different meshes and triangulations effectively. Practically, the paper adapts the flow matching (FM) paradigm to a mesh-based, triangulation-agnostic setting. Theoretically, it proposes a specific noise distribution which is triangulation agnostic, to be used inside the FM model's denoising process. While noise distributions are usually trivial to devise for, e.g., images, devising a triangulation-agnostic distribution proves to be a much more difficult task. We formulate a mathematical definition of triangulation agnosticism of distributions, via their spectrum. We then show that a discretization of a specific Gaussian random field called a Matérn process holds these desired properties, and provides a simple and efficient sampling algorithm. We use it as our noise model, and adapt FM to the triangulation-agnostic setting by using a state-of-the-art approach for learning signals on meshes in the gradient domain -- PoissonNet -- as the denoiser. We conduct experiments on elaborate tasks such as sampling elastic rest states, and generating poses of humanoids. Our method is shown to be capable of producing highly realistic results for meshes of over one million triangles, significantly exceeding the state-of-the-art in quality and diversity.
RFGWRK: a hybrid downscaling framework for high-resolution precipitation mapping in geohazard-prone mountainous regions
Simin Zhang
Zeshuang Zheng
Shengbing Yang
Yuan Zeng