Publications

Exploiting weight-space symmetries for approximating curvature
Artem Artemev
Rui Xia
Benjamin M. Boyd
Youjing Yu
Guillaume Hennequin
Alberto Bernacchia
Many machine learning techniques rely on approximating a loss function's curvature, but this is notoriously hard to do at the scale of moder… (voir plus)n deep networks. Surprisingly, no previous work has exploited the curvature constraints that arise from well known weight-space symmetries in loss landscapes. By analytically averaging over group actions that leave the loss invariant, we construct structured Hessian approximations from single gradients that can be tractably estimated, stored, and inverted. The choice of user-specified symmetry group directly governs the trade-off between approximation accuracy and computational cost. Moreover, our framework provides a unifying theoretical lens for viewing existing methods; in particular, a specific choice of symmetry group recovers Shampoo/Muon-like curvature estimates. We validate our method on a range of network architectures, and deploy it to second-order optimization benchmarks, including a small language model. Our curvature estimation framework might find applications in other machine learning problems such as uncertainty estimation, continual learning, compression/pruning, training data attribution, and more.
Microlensing Detection and Inference via Learned Bayes Factors
We present a unified framework for gravitational microlensing event detection and parameter inference. Traditional pipelines use determinist… (voir plus)ic hard cuts on photometric statistics, systematically missing low-magnification events in the finite-source regime. We instead frame detection as Bayesian model comparison using Evidence Networks, which learn calibrated Bayes factors from binary-labeled simulations, and combine this with Neural Posterior Estimation (NPE) for amortized parameter inference. Both share a transformer encoder that handles irregularly-sampled time series without imputation. On simulated Roman Space Telescope data, our Evidence Network achieves
One More Time: Revisiting Neural Quantum States from a Reinforcement Learning Perspective
Juan Agustin Duque
Sergio García-Heredia
Vinicius Hernandes
Eliška Greplová
Thomas Spriggs
Anna Dawid
Neural quantum states (NQS) provides a flexible and scalable framework for approximating quantum many-body wavefunctions. Among NQS paramete… (voir plus)rizations, autoregressive models are especially attractive because they enable exact, independent sampling from the Born distribution, avoiding the autocorrelation and mixing issues of Markov chain methods. Yet their optimization remains comparatively underexplored: Adam is a scalable method but ignores function space geometry, while stochastic reconfiguration is principled but costly and numerically fragile in large models. To address this gap, we show that variational energy minimization can be viewed as an advantage policy-gradient problem over the Born distribution, motivating trust-region optimization for NQS training. We introduce \emph{Proximal Wavefunction Optimization} (PWO), a trust-region algorithm that clips probability-ratio changes in the amplitude channel and wrapped phase increments in the phase channel. PWO avoids explicit matrix inversion, reuses samples across inner updates, and preserves the scalability of first-order optimization. Across Ising, Heisenberg, and frustrated
Strong Gravitational Lensing Posterior Sampling in Pixel-Space Using Diffusion Models and Recurrent Inference Machines
Modeling galaxy-galaxy strong gravitational lenses to infer the brightness of the source galaxy and the mass distribution of the foreground … (voir plus)galaxy is computationally challenging, particularly for high-resolution, high signal-to-noise observations. In this regime, high-dimensional representations of both the source and the foreground mass distribution are necessary to model the data down to the noise level. This inference problem has been challenging for both traditional and machine learning-based methods because of its high dimensionality and its non-linearity in the foreground mass distribution. We present a method to generate joint posterior samples of the source galaxy and foreground mass distribution as pixelated images conditioned on observations. The method combines diffusion-based generative modeling and recurrent inference machines. It can model realistic gravitational lensing simulations with background and foreground galaxies drawn from cosmological hydrodynamical simulations down to the noise level.
DeSQ: Decomposition-based SPARQL Query Generation
Papa Abdou Karim Karou Diallo
Dominant approaches to Knowledge Base Question Answering (KBQA) fall into two categories. First is the generation of a formal query that suf… (voir plus)fers from brittleness and limited explainability, and the second is direct answer retrieval through KB exploration that is computationally costly and prone to hallucination. To combine the strengths of both paradigms while mitigating their respective weaknesses, we introduce DeSQ (Decomposition-based SPARQL Query Generation), a KB-agnostic framework that operates in three stages. First, it decomposes complex questions into Atomic Constraints (ACs) that mirror the relational structure of the underlying KB. Second, it generates a two-part structured output: (a) Mapping of each AC to its corresponding SPARQL Fragment, using standardized variable and URIs placeholders, and (b) URIs Grounding block describing each placeholder. Third, it assembles these fragments into a complete SPARQL query. DeSQ surpasses state-of-the-art approaches on four out of five major benchmarks and demonstrates superior robustness to lexical variation. Beyond performance gains, our framework greatly simplifies evaluation by eliminating the need for a live KB endpoint, and its structured output enables fine-grained error analysis, allowing more targeted interventions for improvement.
Drift Q-Learning
Offline reinforcement learning requires improving a policy from fixed data while avoiding out-of-distribution actions with unreliable value … (voir plus)estimates. Diffusion and flow policies handle this trade-off by modeling the behavior distribution to regularize the RL objective, but they require iterative denoising, solver integrations, and in more efficient variants, distillation or other approximations at inference. We propose DriftQL, which combines a drift-based behavioral regularizer with critic-driven policy improvement. The value signal biases the policy toward high-value regions of the data support, while attraction and repulsion together keep generated actions near the data and prevent collapse onto a single mode. DriftQL is implemented as a single network with a unified training objective and generates actions in a single forward pass. On D4RL and OGBench, DriftQL consistently outperforms diffusion and flow methods, advancing the state of the art. Under degraded data quality, where the baselines visibly struggle, DriftQL remains close to its clean-data performance, positioning it as a promising alternative to diffusion and flow-based methods while maintaining the simplicity and efficiency of deterministic approaches. Project page: https://driftql.github.io/
Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks
Biological and neuromorphic recurrent neural networks (RNNs) are subject to spatial and temporal locality constraints on the information tha… (voir plus)t can plausibly be used during learning. A common strategy to satisfy these constraints is to modify gradient descent by neglecting non-local terms to varying degrees, as in random feedback local online (RFLO) learning and truncated backpropagation through time (tBPTT). However, the learning dynamics of these algorithms, and how they compare with BPTT, remain poorly understood. We apply dynamical systems theory to data-aligned linear RNNs -- whose dynamics can be separated into orthogonal modes -- to compare stationary solutions, stability properties, and convergence rates, finding qualitatively distinct behaviour for RFLO versus BPTT and one-step tBPTT. We further observe that the solutions learned by RFLO are restricted to low-rank perturbations of initial parameters, a result which holds beyond the data-aligned setting. Our work provides analytical insight into how locality constraints shape learning dynamics, with implications for neuroscientific models of learning and alternative optimization approaches for RNNs.
LSD Reconfigures Cortical Dynamics Through Faster Brain Rhythms and Increased Fractal Dimension
Venkatesh Subramani
Annalisa Pascarella
Jérémy Brunel
Yorguin José Mantilla Ramos
Yann Harel
Suresh Muthukumaraswamy
Robin Carhart-Harris
Giulia Lioi
Nicolas Farrugia
Lysergic acid diethylamide (LSD) profoundly alters conscious experience, yet the electrophysiological mechanisms by which it reshapes neural… (voir plus) dynamics remain incompletely understood. A hallmark of psychedelic states is widespread cortical desynchronization, typically inferred from reductions in spectral power, but whether such effects reflect genuine weakening of neural oscillations or are confounded by shifts in oscillatory peak frequencies remains unresolved. Here, we address this gap by combining source-resolved magnetoencephalography (MEG), spectral parameterization, temporal complexity metrics, and interpretable machine learning in an LSD versus placebo design, with and without music. We show that LSD induces robust, spatially structured increases in alpha and beta peak frequencies alongside genuine attenuation of oscillatory power, with these effects displaying partly dissociable cortical patterns. Beyond rhythmic activity, LSD is associated with flattening of the aperiodic 1/f spectral slope and increased neural signal fractality and complexity, preferentially affecting sensory, language, emotion, and imagery-related networks while sparing motor cortex. Machine-learning analyses further identify peak-frequency shifts, aperiodic parameters, and complexity measures as key discriminators of the psychedelic state. Music does not robustly amplify these neural signatures and instead shows a trend toward attenuation. Together, these findings provide a comprehensive electrophysiological account of how LSD reorganizes large-scale human brain dynamics and highlight features that may differentiate its neural signature from that of other psychedelics.
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
Alexander Gurung
Issam H. Laradji
Rafael Pardinas
Deep research agents increasingly combine private local documents with external tools like web retrieval, creating a privacy risk: an agent'… (voir plus)s external queries may leak sensitive information from its local context. This risk is amplified by the mosaic effect, where individual queries may appear harmless but become revealing in aggregate. We introduce MosaicLeaks, a benchmark of 1,001 multi-hop deep research tasks that chain private enterprise documents and a public web corpus, forcing agents to make external queries that depend on local information. We evaluate leakage with an adversary LLM that observes only the agent's external queries and attempts to infer private information at three levels: the agent's research intent, answers to specific private questions and verifiable claims about the enterprise documents. We find that models across families and sizes frequently leak at all three levels, that zero-shot privacy prompting reduces but does not eliminate leakage and that reinforcement learning for task performance alone worsens leakage. To address this, we propose Privacy-Aware Deep Research (PA-DR), an RL framework that combines situational rewards for task success with a learned privacy classifier to provide dense credit assignment over both per-query and mosaic-level leakage. Training Qwen3-4B-Instruct with PA-DR improves accuracy from 48.7% to 58.7% and reduces answer and full-information leakage from 34.0% to 9.9%.
Quantitative Equational Logic
Giorgio Bacci
Radu Mardare
Gordon Plotkin
We develop a quantitative analogue of equational reasoning, which we call quantitative equational logic. The quantitative equations use, ins… (voir plus)tead of classical equality, quantitative equalities, which are equalities indexed with nonnegative reals. Thus, s = ε t means that “ s and t are points in a metric space and their distance is less than ε”. Quantitative equalities will be used to encode behavioural distances, with ε being an upper bound on the measure of dissimilarity between two terms. We develop the metatheory of this subject. We define a notion of quantitative algebra, which is the quantitative analogue of universal algebra. We prove a completeness theorem for quantitative equational logic, and we show that we obtain monads on suitable categories of metric spaces. We present a set of examples where the free algebra of a quantitative equational theory corresponds to some well-known structure. These examples are: Hausdorff metrics from quantitative semilattices; p -Wasserstein metrics (hence also the Kantorovich metric), and the total variation metric.
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
Reinforcement learning problems typically define the goal as maximizing the expected value of a scalar reward function. But, pairwise prefer… (voir plus)ences are often easier to specify than scalar rewards, and they express certain goals that scalar rewards cannot. Methods for reinforcement learning with pairwise preferences have thus received growing interest. Unfortunately, these methods are inefficient in problems with long time horizons, and they lack guarantees on the performance of Markov policies relative to history-dependent policies, which bridge the theory and practice of reinforcement learning. We therefore propose the \textit{Markov decision contest} as a new problem model for reinforcement learning with pairwise preferences. We prove that stationary Markov policies are optimal among all history-dependent policies, that solving a Markov decision contest exactly is in P, and that a simple iterative algorithm converges to an optimal policy at a sublinear rate. Lastly, in a set of high-dimensional decision problems with long time horizons, we show that our approximate algorithm is significantly more learning-efficient than prior work.
Reusable Low-Rank Subspaces Explain Why Cross-Modal Transfer Adapts with Tiny Updates
Parameter-efficient finetuning methods such as LoRA routinely adapt massive pretrained transformers to new tasks using only tiny low-rank up… (voir plus)dates, but the representational geometry that makes this possible remains unclear. We use cross-modal transfer from a language-pretrained transformer to time-series forecasting as a controlled probe of low-rank adaptation, asking why so few directions are sufficient. Across adaptation regimes, LoRA recovers most of the transfer benefit of full finetuning; effective-rank analyses show that pretrained representations already concentrate on a low-rank subspace that finetuning \emph{redistributes} rather than rebuilds; and a single linear projection over frozen hidden states aligns with realistic time-series trajectories without paired supervision. Randomly initialized models, by contrast, first construct a compressed representation through a uniform layer-wise collapse before they can specialize. These results support a view of cross-modal adaptation as low-rank \emph{direction selection} within reusable pretrained subspaces.