Publications

Optimizing Dec-POMDP Agent-State Policies via Risk-Seeking Utility
Matthieu Geist
Solving decentralized decision-making problems modeled as Dec-POMDPs is notoriously NEXP-complete, as optimal solutions require policies con… (see more)ditioned on an agent's entire action-observation history. To maintain tractability, it is common to restrict agents to finite-memory models, known as agent-state policies. Although this constrained policy class may not contain the globally optimal solution, finding the highest-performing agent-state policy remains a critical objective for practical applications. Addressing the challenge of planning under bounded memory, we introduce an iterated best-response algorithm that converges monotonically to a local optimum in polynomial runtime relative to the Dec-POMDP model size. To discover superior policies within this restricted memory space, we employ a novel objective that pairs a risk-seeking incentive with conservative policy updates. Our experiments on standard Dec-POMDP benchmarks demonstrate that this approach is competitive with state-of-the-art methods, delivering near-optimal results despite the limited memory.
Representing Time Series as Structured Programs for LLM Reasoning
Jaeho Kim
Changhun Oh
Seokhyun Lee
Changhee Lee
Large language models (LLMs) have demonstrated strong reasoning and instruction-following capabilities, making them potentially powerful too… (see more)ls for time-series analysis. However, time series lie outside their native textual modality, raising a fundamental question: how should time series be represented so that LLMs can reason about them effectively? Existing work typically serializes raw numerical sequences or fine-tunes pre-trained LLMs on time-series data. These approaches place the burden of extracting temporal structure directly on the LLM, creating a modality mismatch that often degrades performance on long sequences and introduces substantial computational overhead. In this work, we introduce Time-Series-to-Structured-Program representation (T2SP), a deterministic, training-free method that represents a time series as a structured symbolic program. T2SP decomposes time series into trends, periods, and salient events, expressing them in a program-friendly format aligned with the textual and code-like modalities on which LLMs are natively trained. By shifting temporal-structure extraction from the model to the representation itself, T2SP enables off-the-shelf LLMs to leverage their existing reasoning capabilities for time-series understanding. We evaluate T2SP on three reasoning tasks -- editing, captioning, and question answering -- where it consistently improves performance, reduces reasoning time, and lowers failure rates compared with raw-string representations. Our results demonstrate that T2SP provides an effective interface between time series and LLMs.
SLowRL: Safe Low-Rank Adaptation for Bridging the Sim-to-Real Gap in Legged Locomotion
Shafeef Omar
Majid Khadiv
A simulator is, at best, a coarse low-fidelity model of the real world the agent eventually has to act in. Closing this residual gap on hard… (see more)ware is a canonical instance of operating in a big world: the real environment exposes contact dynamics, latencies, and disturbances that the agent was never given the capacity (parameters or data) to model during pretraining. Naive on-hardware fine-tuning is risky --- the policy can damage the robot before it improves --- and full-parameter updates require prohibitive interaction time. We propose SLowRL, a continual fine-tuning framework that confronts this big-world adaptation problem with two complementary forms of capacity limitation: (i) a rank-1 LoRA adapter applied per layer to both actor and critic, restricting each layer's update to a single direction in its image space (
The blueprint of human functional architecture shifts from cognition to anatomy during perturbations of consciousness
Andrea I. Luppi
Dragana Manasova
Justine Y. Hansen
Zhen-Qi Liu
Asa Farahani
Yonatan Sanz Perl
Jakub Vohryzek
Daniel Golkowski
Andreas Ranft
R. Ilg
Denis Jordan
Vincent Bonhomme
Audrey Vanhaudenhuyse
Athéna Demertzi
Océane Jaquet
Mohamed Ali Bahri
Naji Alnagger
Paolo Cardone
Lorina Naci
Adrian M. Owen … (see 9 more)
John Pickard
Guy Williams
Judith Allanson
Enrico Amico
Jacobo Sitt
David Menon
Emmanuel A. Stamatakis
Bratislav Misic
Consciousness and cognition arise from the ongoing interactions between brain regions. Synchronous fluctuations of fMRI signals may indicate… (see more) that two brain regions perform similar cognitive functions, but neural interactions are also constrained by anatomical connectivity and regions' molecular, cytoarchitectonic, and metabolic profiles. Here we disentangle the respective contributions of ongoing cognition and multimodal neurobiological constraints in shaping functional connectivity. We jointly contextualise haemodynamic FC against eight distinct multimodal representations of the human connectome: (i) structural connectivity from diffusion tractography; (ii) spatial embedding; (iii) similarity of transcriptional profiles from gene expression; (iv) similarity of receptor profiles from Positron Emission Tomography; (v) laminar profile similarity from histology; (vi) correlated electrophysiological activity from magnetoencephalography; (vii) correlated metabolic activity from PET glucose uptake; (viii) coordinated activation across 123 cognitive operations from the NeuroSynth meta-analytic engine. We demonstrate that cognitive co-activation is the dominant predictor of inter-regional fMRI synchrony in the awake human brain, even when quantified using intracranial electrical stimulation. Crucially, this predominance of cognitive co-activation for shaping functional connectivity is systematically obliterated across five datasets of pharmacological and pathological perturbations of consciousness (chronic disorders of consciousness; anaesthesia with sevoflurane, propofol, or ketamine) when cognition is disconnected from the environment or altogether abolished. Altogether, we show that multimodal predictors of functional architecture shift away from cognitive co-activation and toward anatomical-molecular constraints during pharmacological and pathological perturbations of consciousness.
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
Joachim Schaeffer
Alexander Panfilov
Jonas Geiping
Roland S. Zimmermann
AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted mode… (see more)l. This partially tampers with the untrusted model's trajectory. If the trusted model detects such an intervention, it may infer properties of the monitor and adapt to evade control. We introduce \textbf{CIAware-Bench}, a benchmark for measuring \textbf{c}ontrol \textbf{i}ntervention (CI) awareness across frontier models. CIAware-Bench tests whether models can distinguish their own trajectories from those modified by a control intervention. The benchmark is comprised of a suite of four task domains (essay writing, BigCodeBench, Bash Arena, and SHADE-Arena), while varying trajectory watermarking, side-task presence, and the control protocol. Evaluating eleven frontier models, we find low to moderate CI awareness under default settings (up to 0.87; random chance balanced binary classification accuracy is 0.5) with substantial variation across task domains and model pairs. Detection is generally easier across model families, suggesting that models exploit provider-specific differences in style or post-training. Overall, CI awareness is not a fixed model-level property, and should be measured for each new model release and deployment scenario. We release CIAware-Bench to track CI awareness and inform control protocols whose interventions are harder to detect.
Critical dynamics in spontaneous EEG predict perturbational complexity in disorders of consciousness with measurable evoked responses
Derek Newman
Charlotte Maschke
Jordan O‘Byrne
Michele Colombo
Angela Comanducci
Silvia Casarotto
Giuseppe Citerio
M Rosanova
Marcello Massimini
Stefanie Blain‐Moraes
Abstract Identifying which severely brain-injured patients retain the capacity for consciousness remains a major challenge in neurocritical … (see more)care. The perturbational complexity index (PCI) provides a reliable assessment of consciousness capacity, but its reliance on transcranial magnetic stimulation and EEG (TMS-EEG) limits bedside scalability. PCI and brain criticality capture complementary dimensions of brain dynamics: PCI quantifies the complexity of the brain’s evoked response to perturbation, whereas criticality characterizes the intrinsic organization of spontaneous activity. Here, we tested whether resting-state EEG signatures of criticality predict PCI max in disorders of consciousness, extending prior findings from anesthesia to severe brain injury. In 26 patients with vascular, traumatic, or anoxic brain injury, multivariate criticality related features did not generalize PCI max prediction across the full heterogeneous cohort. However, criticality features predicted PCI max when analyses were restricted to non-anoxic patients and when restricting analyses to patients with non-zero PCI max values. These findings suggest that spontaneous criticality measures index the brain’s intrinsic dynamical regime that supports complex perturbational responses, while their correspondence with PCI max depends on whether the injured brain retains sufficient capacity to sustain large-scale evoked responses. Together, our results extend the relationship between resting-state criticality and evoked perturbational complexity to disorders of consciousness and support the development of stratified EEG measures in severe brain injury.
Human learning of noninvasive brain–computer interfaces via manifold geometry
Erica L. Busch
E. Chandra Fincke
Nicholas B. Turk‐Browne
Overcoming Rank Collapse in Feedback Alignment
Gauthier Boeshertz
Backpropagation (BP) is widely viewed as biologically implausible, in part because it requires feedback weights to be the transpose of forwa… (see more)rd weights for error propagation. Interestingly, when training a network with fixed random feedback weights to circumvent this issue, learning aligns the forward weights with the feedback weights, leading the backpropagated error signal to become an approximation of the standard gradient used by BP. This process, called Feedback Alignment (FA), occurs in MLPs and very shallow CNNs but does not scale well to deeper architectures. In this work, we first investigated differences between BP and FA models, trained on CIFAR10, specifically focusing on the effective rank of the signal. We found that the FA error has a considerably lower rank and hence is constrained to a lower-dimensional subspace compared to BP, limiting exploration of the parameter space. Motivated by this observation, we evaluated two mechanisms for increasing the effective dimensionality of FA: Muon, an optimiser that orthogonalises weight updates; and hidden activity normalisation, which promotes activation orthogonality. Across larger architectures and benchmarks, we find that these methods consistently improve over FA baselines, for example, on CIFAR100 with a Resnet-18, accuracy increases by 9 percentage points. Our results identify low-dimensional gradient dynamics as a key obstacle to scaling FA and suggest that inducing higher-dimensional update geometry is a promising route toward scaling alternatives to backpropagation.
Rank Collapse, Fixed Points, and the Renormalization Group Structure of MLP Residual Networks
Parviz Haggi-Mani
The analogy between deep neural network forward passes and renormalization group (RG) flows has been repeatedly noted in the literature, but… (see more) existing treatments remain qualitative: depth is described as a coarse-graining scale, attention is likened to a partition function, and representations are said to flow toward fixed points. No existing work has defined a measurable RG order parameter, tested it under controlled variation of the input distribution, or made quantitative predictions that are empirically verified. We study the simplest architecture for which the analogy is tractable: a pure MLP residual stack trained on masked token prediction over synthetic Markov chain sequences with known spectral properties. We report three findings. (i) The effective rank of the residual stream decreases monotonically with depth after training, consistent with progressive integration of irrelevant degrees of freedom. (ii) This rank collapse is selective: it occurs for chains with short correlation length approximately 1 but is absent for chains with long correlation length approximately 7, measured at the position level to control for mean-pooling artifacts. The network preserves exactly the degrees of freedom relevant to the prediction task, the content of the RG relevance criterion. (iii) Inter-layer kernel drift is concentrated at one or two specific transitions, with the remainder of the network near a fixed point, consistent with a discrete fixed-point plateau. Together these findings constitute the first quantitative, position-level evidence that MLP residual networks implement a selective coarse-graining procedure governed by the spectral structure of the input distribution.
Unifying Local Communications and Local Updates for LLM Pretraining
Communication-efficient pre-training of LLMs is increasingly important as training draws on compute distributed across clusters, data center… (see more)s, and lower-bandwidth links. Many practical methods reduce communication frequency but still rely on synchronous All-Reduce operations that maintain identical model states and tie progress to global collectives. This can become a bottleneck when bandwidth or worker speed is heterogeneous. We introduce GASLoC, a novel decentralized pre-training algorithm that generalizes the notion of communication acceleration to the recently popular"outer optimizer"to allow a practical gossip-based training framework that is compatible with adaptive optimizers, allows for local optimizer steps, and can utilize sparse randomized peer communication. Empirically, on a number of standard LLM training tasks, we demonstrate that GASLoC outperforms state-of-the-art decentralized algorithms in single step per communication setting for a number of topologies and, unlike existing decentralized methods in the LLM setting, it allows to obtain performance competitive with DiLoCo when utilizing multiple local steps. In the heterogeneous bandwidth setting we demonstrate the advantage of GASLoC showing that it can significantly outperform DiLoCo.
Charting Cervical Spinal Cord Morphometry Across the Lifespan
Kurt Schilling
Michael E Kim
Matthew Amandola
Chenyu Gao
Karthik Ramadass
Praitayini Kanakaraj
Sam Bogdanov
G Rudravaram
Nancy R. Newlin
Derek B. Archer
Timothy J Hohman
Angela L Jefferson
Victoria L Morgan
Alexandra Roche
Dario J Englot
Murat Bilgel
Lori L Beason-Held
Luigi Ferrucci
Laurie Cutting
Laura A Barquero … (see 21 more)
Micah D’Archangel
Tin Q Nguyen
Kathryn L Humphreys
Yanbin Niu
Sophia Vinci-Booher
Carissa J. Cascio
Zhiyuan Li
Daniel Moyer
Simon Vandekar
Panpan Zhang
Samuelle St-Onge
Benjamin De Leener
John C Gore
Seth Smith
B A Landman
John C. Gore
Seth Smith
Bennett A. Landman
Abstract Spinal cord morphometry provides essential biomarkers of neurological health, but clinical interpretations are confounded by inter-… (see more)subject variability and a lack of normative references across the full human lifespan. We address this gap by generating the first comprehensive lifespan charts for cervical spinal cord morphometry. We leveraged 30 population-based brain MRI datasets, aggregating 78,269 scans from 41,042 individuals (ages 0–100) whose imaging protocols included cervical cord coverage. To overcome contrast variability, we employed a state-of-the-art contrast-agnostic deep learning segmentation method, extracting cross-sectional area (CSA), anteroposterior (AP) and right–left (RL/transverse), and shape indices (compression ratio, eccentricity, and solidity) from C1 to C7. Normative trajectories were modeled using Generalized Additive Models for Location, Scale, and Shape (GAMLSS). The resulting charts reveal distinct non-linear lifespan changes: rapid growth through childhood and adolescence, peak maturation occurring in early-to-mid adulthood (e.g., mid-30s for CSA), followed by gradual decreases. Significant regional variations along the cervical cord and consistent sex differences (males > females for size metrics) were quantified. Spinal cord trajectories showed strong temporal coupling with brain white matter and brainstem volumes, suggesting integrated CNS development and aging. These lifespan charts provide a robust normative framework, enabling age- and sex-specific centile scoring of individual spinal cord morphometry. This resource offers a critical tool for differentiating typical variation from pathological changes, enhancing the clinical utility of spinal cord MRI in studies of development and neurodegeneration.
Difference-Aware Retrieval Policies for Imitation Learning
Quinn Pfeifer
Ethan Pronovost
Paarth Shah
Siddhartha Srinivasa
Abhishek Gupta
Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding erro… (see more)rs during deployment. We show that reusing the training data during inference via a semi-parametric retrieval-based imitation learning approach can alleviate this challenge. We present Difference-Aware Retrieval Policies for Imitation Learning (DARP), a semi-parametric retrieval-based imitation learning approach that addresses this limitation by reparameterizing the imitation learning problem in terms of local neighborhood structure rather than direct state-to-action mappings. Instead of learning a global policy, DARP trains a model to predict actions based on