Publications

Securing automotive data flow: A survey of telematics security across intra-vehicle, V2X, and cloud layers
Junjie Wu
Benjamin C. M. Fung
Natalia Stakhanova
Faiyaz Khan
Hanbo Yu
Spinal cord imaging for multiple sclerosis: Advances, priorities, and opportunities
Cornelia Laule
Atlee A Witt
Gabriele C De Luca
Cristina Granziera
B Mark Keegan
Anne Kerbrat
Eric C Klawiter
Shannon Kolind
Kristin P O’Grady
Jiwon Oh
Kurt G Schilling
Dinesh K Sivakolundu
Seth A Smith
Ceren Tozlu
Irene M Vavasour
Francesca Bagnato
Susan A Gauthier
Caterina Mainero
Eva Alonso-Ortiz … (voir 9 de plus)
Rohit Bakshi
Erin S Beck
Matthew R Brier
Christopher C Hemond
Stephen Krieger
David KB Li
Russell T Shinohara
Roland G Henry
North American Imaging in Multiple Sclerosis (NAIMS) Cooperative
The spinal cord plays a central role in the pathophysiology and clinical manifestations of multiple sclerosis (MS), yet remains under-studie… (voir plus)d compared with the brain. This review summarizes key insights from the 2025 North American Imaging in MS Spinal Cord Imaging Workshop, highlighting recent advances, ongoing challenges, and future opportunities in MS spinal cord imaging. We review pathological studies and outline the clinical relevance of spinal cord lesions and atrophy for diagnosis, prognosis, and disease monitoring, highlighting emerging biomarkers of progression independent of relapse activity. Correlations between magnetic resonance imaging, histopathology, and clinical outcomes support the validation and translational potential of advanced spinal cord imaging techniques. Finally, we discuss spinal cord–specific processing pipelines and reproducibility challenges. Collectively, these insights underscore the need to integrate advanced and quantitative spinal cord imaging into clinical trials, research studies, and—when feasible—clinical care, to fully capture the extent of MS pathology, and ultimately improve patient outcomes.
STING dampens the unfolded protein response to enable the presentation of self-antigens on MHC-I during inflammation
Ahmed M. Fahmy
Ali Ahmadi
Joël Lanoix
Tyler Cannon
moustafa Nouh Badr Elemeery
Camberly Hernandez Paredes
Benoit Barrette
Éric Bonneil
Yong Zhong Xu
Maha Ibrahim
Guillermo Arango-Duque
Éric Audemard
Éric Chevet
Erwin Schurr
P. Pierre
Samantha Gruenheid
Pierre Thibault
Heidi M. McBride
Michel Desjardins
Summary A growing body of evidence supports the contribution of the long-lasting adaptive immune system in Parkinson’s disease (PD). We sh… (voir plus)owed that the PD-associated protein PINK1 negatively regulates the presentation of mitochondrial antigens (MitAP) on MHC-I molecules. In vivo evidence indicated that MitAP activation in mice, in the absence of PINK1, led to cytotoxic CD8 + T cell stimulation and severe motor impairments, reversible by L-DOPA. We show here that following TLR4 activation, MitAP is engaged through a pathway involving cGAS-STING, which acts as a rheostat to dampen the unfolded protein response (UPR). Without STING, the stress response is amplified, leading to a translational attenuation that inhibits the expression of XBP1s, a transcription factor required for MitAP. STING activity also regulates the repertoire of peptides displayed at the cell surface during inflammation, highlighting a potential role in immunosurveillance. These findings establish STING and the UPR as key immune regulators targetable for therapeutic intervention during autoimmune diseases and PD.
The schema spectrum: Emergent structures and levels of abstraction in AI and the brain
Blake A. Richards
Bayesian Symbolic Regression with Entropic Reinforcement Learning
Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of in… (voir plus)puts. Unlike forms of regression that fit parameters assuming a fixed model structure, symbolic regression is a search problem over the space of expressions, represented, for example, as abstract syntax trees using a library of operators. Symbolic regression is typically used in settings with limited, noisy data in the natural sciences. However, searching for a single best-fitting expression fails to capture the epistemic uncertainty about the expression, which motivates a Bayesian perspective that enables uncertainty quantification and specification of natural priors to constrain the search space. In this work, we propose ERRLESS (Entropy-Regularized Reinforcement Learning for Expression Structure Sampling), a scalable approach for sampling from the posterior distribution over expressions given data using maximum-entropy reinforcement learning. ERRLESS learns a neural policy that constructs expressions sequentially by building up their abstract syntax trees. At convergence, the policy samples expressions from the posterior. At test time, expressions can be sampled by rollouts of this policy. We demonstrate that ERRLESS achieves competitive results on the Feynman benchmark while producing short and interpretable expressions. Additionally, we demonstrate that the mean of the posterior predictive approximated by ERRLESS achieves a high coefficient of determination (
Causal Discovery with Metadata-Informed Latent Types
Causal discovery seeks to recover causal structure from data but the underlying graph is typically identifiable only up to its Markov equiva… (voir plus)lence class. Yet real-world systems often exhibit redundancy, where groups of variables share similar causal roles. We introduce a Bayesian causal discovery framework that leverages variable-level metadata to infer latent types and constrain causal interactions across variables. We model causal graphs as type-consistent DAGs and propose t-DiBS, a fully differentiable method that jointly learns variable types, graph structure, and metadata representations. Our approach enables principled uncertainty quantification and integrates expressive neural models for metadata. We provide theoretical results showing that, under structured assumptions, metadata combined with typing can improve identifiability beyond classical limits. Empirically, we demonstrate improved performance over standard causal discovery methods on synthetic and pseudo-real datasets, with detailed analysis demonstrating the benefit of joint type and structure learning. These results establish metadata-driven typing as a principled approach to identifiable causal discovery.
Plausibility Is Not Prediction: Contrastive Evidence for LLM-Based Cellular Perturbation Reasoning
Yashi Zhang
Hongyu Guo
Perturbation experiments are central to understanding cellular mechanisms, but remain costly and sparse, motivating prediction of gene expre… (voir plus)ssion responses for unobserved conditions. A promising recent direction leverages large language models (LLMs) as"virtual cell"simulators-using stepwise, knowledge-grounded mechanistic reasoning to infer differential expression-pointing toward an interpretable, knowledge-driven paradigm that transcends purely data-driven approaches. However, we find that plausibility is not prediction: despite producing biologically plausible explanations, these methods fail to capture perturbation-specific effects: systematically overestimating differential expression, often underperforming a simple gene-frequency baseline in aggregate evaluations, and collapsing to chance-level performance at the per-gene level. This reveals a reliance on intrinsic gene response tendencies rather than true perturbation reasoning. We trace this failure to how evidence is presented: existing methods evaluate perturbation-gene pairs in isolation, without exposing how related perturbations differ in their effects on the same gene. To address this limitation, we introduce CORE (Contrastive Organization of Relational Evidence), which reframes prediction as a comparison task by organizing evidence into positive and negative outcomes from related perturbations. Using a biomedical knowledge graph for evidence retrieval, CORE improves calibration and substantially boosts perturbation-specific prediction in both LLM-based and non-LLM settings: for example, on drug-perturbation data, CORE-Reasoning improves Qwen3.5-9B aggregate metrics by up to 28.6%, while on generic perturbation data, CORE-Voting raises macro-per-gene AUROC from chance to 0.703 in average across four cell lines. This highlights contrastive evidence organization as essential to reliable LLM-based perturbation reasoning
Stop Probing, Start Coding: Why Linear Probes and Sparse Autoencoders Fail at Compositional Generalization
Vitória Barin-Pacela
Isabela Camacho
David Klindt
Foundational to interpreting pretrained representations of deep generative models, the linear representation hypothesis states that neural n… (voir plus)etwork activations encode high-level concepts as linear mixtures. However, linear representation does not imply linear accessibility of such concepts: under superposition, when the number of concepts exceeds the activation dimension, recovering the underlying latent factors requires sparse nonlinear inference, making methods such as linear probes insufficient. Sparse autoencoders (SAEs) perform nonlinear inference but amortize it into a fixed encoder, introducing a systematic amortization gap. We show this gap dominates all other error sources and persists as the number of training samples is increased, causing SAEs to fail under out-of-distribution (OOD) compositional shifts. In contrast, classical sparse coding with per-sample iterative inference leverages compressed sensing guarantees to recover latent factors robustly, maintaining near-zero gaps in the accuracy between in and out of distribution. Our results demonstrate that the recent OOD failures of SAEs can be attributed to amortization failures: per-sample inference at test time substantially improves OOD performance, even when using a dictionary learned by an SAE. This is observed along a spectrum of hybrid approaches that progressively undo amortization and recover OOD performance.
TECCI: Tricky Edits of Collected and Curated Images
Roy Hirsch
Yasumasa Onoe
Sherry Ben
Jason Baldridge
Despite tremendous recent progress, current text-guided image editing methods still struggle with many aspects of editing involving instruct… (voir plus)ion following, minimally editing the source image, and ensuring high visual quality. These problems are especially apparent when the requested edit is challenging, such as those that involve position, motion, viewpoint, scale and creative edits. To systematically test generative image editors, we propose a novel image editing benchmark -- TECCI: Tricky Edits of Collected and Curated Images. TECCI consists of a completely new set of images we are releasing. The images in TECCI span 7 image categories. The images and these categories were curated intentionally to target weaknesses of existing methods. The edit instructions in TECCI are automatically generated by Gemini, covering 5 edit types per source image. We also curated a set of 530 images for which we created challenging manually written edit instructions. Overall, TECCI contains 7550 pairs of images and edit instructions. We conduct human evaluations of five leading image editing models on TECCI. Humans judge outputs along three dimensions: 1) instruction following, 2) minimality of the edits, and 3) visual quality. To scale-up the evaluation, we also build an auto-rater using Gemini that achieves 74.7% accuracy in matching human evaluations. Our evaluations reveal that: 1) none of the models exceed a 22% overall success rate, demonstrating the challenging nature of TECCI, 2) Nano Banana Pro is the best performing model overall, 3) models perform significantly better at instruction following compared to minimal edits and visual quality, 4) models struggle with editing architecture and nature images which require strong understanding of spatial layout and intricate visual details. 5) reasoning and creative edits are the most difficult, whereas color and appearance edits are the easiest.
TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages
Senyu Li
Wassim Hamidouche
Waqas Zamir
Inbal Becker-Reshef
Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly Afric… (voir plus)an ones, critically underexplored. We introduce TUKABENCH, a jailbreak benchmark for seven African languages that extends JailbreakBench (JBB) beyond direct translation through four settings: human translation of JBB prompts, English adaptation to African contexts followed by human translation, human-curated prompts validated through interactions with GPT-5.2, and code-switched prompts combining English and African languages, isolating the effect of language, cultural grounding, and prompt evasiveness on model safety. Across closed and open models, prompting in African languages reduces refusal relative to English, with culturally adapted prompts leading to least refusal. The evaluation also surfaces two structural limitations: model comprehension failures and reduced LLM-as-a-judge reliability in LRLs. To capture the first, we introduce Deflection alongside Refused and Jailbroken; to assess the second, we validate outputs with human annotations, showing that judge-human agreement drops in lower-resource languages and less commonly supported scripts.
Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations
Identifiability in representation learning is commonly evaluated using standard metrics (e.g., MCC, DCI, R^2) on synthetic benchmarks with k… (voir plus)nown ground-truth factors. These metrics are assumed to reflect recovery up to the equivalence class guaranteed by identifiability theory. We show that this assumption holds only under specific structural conditions: each metric implicitly encodes assumptions about both the data-generating process (DGP) and the encoder. When these assumptions are violated, metrics become misspecified and can produce systematic false positives and false negatives. Such failures occur both within classical identifiability regimes and in post-hoc settings where identifiability is most needed. We introduce a taxonomy separating DGP assumptions from encoder geometry, use it to characterise the validity domains of existing metrics, and release an evaluation suite for reproducible stress testing and comparison.
Bayesian Last Layer for Neural Force Fields
Reliable uncertainty quantification is essential for deploying Machine Learning Interatomic Potentials (MLIPs), also known as Neural Force F… (voir plus)ields, especially when molecular dynamics or materials simulations encounter configurations outside the training distribution. Deep ensembles remain the strongest practical baseline for MLIP uncertainty, but training and storing several copies of a modern pretrained model is often prohibitively expensive. We show that Bayesian Linear Last Layers (BLLs) provide a scalable alternative for MLIPs: a single pretrained backbone supplies atomic features, while exact Bayesian inference over the final force-prediction layer gives predictive uncertainties. BLL is known to underestimate the uncertainties. We provide an in-depth analysis that shows two sources of miscalibration and introduce a simple post-hoc recalibration to address the issue. On MPtrj and rMD17 benchmarks, including both in-distribution tests and increasingly out-of-distribution regimes, BLLs that are recalibrated on in-distribution examples produce uncertainty estimates competitive with ensembles, while using only one base model.