Publications

Invariant Causal Set Covering Machine
Baptiste Bauvin
Pascal Germain
Rule-based models, such as decision trees, appeal to practitioners due to their interpretable nature. However, the learning algorithms that … (see more)produce such models are often vulnerable to spurious associations, and thus, they are not guaranteed to extract causally relevant insights. This limitation reduces their utility in gaining mechanistic insights into a phenomenon of interest. In this work, we build on ideas from the invariant causal prediction literature to propose Invariant Causal Set Covering Machines, an extension of the classical Set Covering Machine (SCM) algorithm for conjunctions/disjunctions of binary-valued rules that provably avoids spurious associations. The proposed method leverages structural assumptions about the functional form of such models, enabling an algorithm that identifies the causal parents of a variable of interest in polynomial time. We demonstrate the validity and efficiency of our approach through a simulation study and highlight its favorable performance compared to SCM in uncovering causal variables across real-world datasets.
CrediBench: Building Web-Scale Network Datasets for Information Integrity
Online misinformation poses an escalating threat, amplified by the Internet's open nature and increasingly capable LLMs that generate persua… (see more)sive yet deceptive content. Existing misinformation detection methods typically focus on either textual content or network structure in isolation, failing to leverage the rich, dynamic interplay between website content and hyperlink relationships that characterizes real-world misinformation ecosystems. We introduce CrediBench: a large-scale data processing pipeline for constructing temporal web graphs that jointly model textual content and hyperlink structure for misinformation detection. Unlike prior work, our approach captures the dynamic evolution of general misinformation domains, including changes in both content and inter-site references over time. Our processed one-month snapshot extracted from the Common Crawl archive in December 2024 contains 45 million nodes and 1 billion edges, representing the largest web graph dataset made publicly available for misinformation research to date. From our experiments on this graph snapshot, we demonstrate the strength of both structural and webpage content signals for learning credibility scores, which measure source reliability. The pipeline and experimentation code are all available here, and the dataset is in this folder.
Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control
Qi Zhao
Guozheng Ma
Yilun Kong
Haoyu Wang
Zilin Wang
Tiantian Zhang
Yuxing Wang
Jian Sha
Yongzhe Chang
Xueqian Wang
Dacheng Tao
Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL s… (see more)ystem design to jointly account for many tightly coupled factors. Despite advances in individual algorithmic components, their functional interdependencies remain underexplored: do they exhibit mutual synergy or counterproductive interference? To bridge this gap, we conduct a systematic investigation and find that the efficacy of different components exhibits significant task-dependency, and naively stacking state-of-the-art techniques does not necessarily yield performance gains; instead, it often triggers emergent challenges, such as compounded non-stationarity. Building upon these findings, we distill a suite of actionable insights into the principled coordination of these components. Guided by these insights, we propose ROSER, an RL framework that coordinates three critical dimensions: Model-based Representation, Optimization Stability, and Experience Replay. Across diverse continuous-control benchmarks, ROSER consistently outperforms vanilla baselines and achieves 17.60% gains over naive stack. Our findings underscore the necessity of a holistic perspective in RL system design and paves the way for developing sample-efficient agents.
What Survives KV-Cache Compression for Reasoning? A Dual-Channel View of Structural and Associative Memory
Md Rifat Arefin
Long reasoning traces make Transformer key–value (KV) caches grow linearly with generation length. Existing compression methods adopt diff… (see more)erent memory topologies, but it is unclear how those choices affect the information that survives eviction. We observe a task–fidelity mismatch: dense summaries achieve low KV reconstruction error yet lose information required for structured reasoning. We explain this through a dual-channel model of compressed memory consisting of a temporally organized scaffold and sparse associative bindings, suggesting that structure and association should be stored separately. We instantiate this idea as TT–Delta, which combines a Tensor-Train structural state with a Delta-rule associative state. On a frozen Qwen2.5-0.5B task model, TT–Delta compresses an evicted prefix by 25.3× (46,680 persistent scalars) and improves exact-answer generation on sequential linear-system reasoning from 0.302±0.113 to 0.839±0.171. In contrast, multi-head Delta is strongest on associative recall and branching graph search. These results suggest that no single compressed-memory topology is universally optimal; the appropriate topology depends on the structure of the reasoning state.
A Passivity-Based Analysis of First-Order Momentum-Based Methods
Sepehr Moalemi
This paper presents a discrete-time passivity-based analysis of first-order momentum-based methods for a class of functions whose gradient h… (see more)as lower and upper sector bounds of
VIDP: Variable Impedance Diffusion Policy for Compliant Robot Manipulation from Diverse Demonstrations
Hisham Khalil
Neil Fernandes
Thomas M. Kwok
Contact-rich manipulation requires precise tracking and mechanical compliance, where variable impedance control can improve robustness in ta… (see more)sk success, whereas static compliance cannot adapt to varying contact constraints. Variable impedance skills can be learned from demonstrations, avoiding complex modeling, but compliance is a hidden variable in force-agnostic kinematic data. While existing methods infer compliance from trajectory variations, these variations may reflect geometric adaptation and not intentional compliance when subject to changing spatial layouts. Therefore, this letter introduces Variable Impedance Diffusion Policy (VIDP), an imitation learning-based variable impedance control framework leveraging a Task-Parameterized Directionality-Aware Mixture Model (TP-DAMM) to extract physically consistent trajectory distributions from diverse demonstrations. By mapping distributions to stiffness profiles, VIDP jointly predicts pose actions and task compliance without force sensors. Real-world experiments show that VIDP significantly outperforms fixed-impedance baselines in task success rate while reducing interaction forces with respect to high stiffness controllers and tracking errors with respect to low stiffness baselines.
A Chain Is Only as Strong as Its Weakest Link: A Scoping Review of System Integration Audits in AI
Dominic Martin
As AI systems become increasingly integrated into diverse interfaces and applications, model-centric audits are insufficient to address risk… (see more)s arising from interactions among system components and deployment environments. System integration has long been central to software audits in safety-critical domains such as aerospace. However, its role in AI auditing remains underexplored. Scanning through 4,259 documents, we present a scoping review of AI audits that treat system integration as a core tenet of evaluation (n = 58). Using reflexive thematic analysis, we analyze their elements, actors, enablers, and constraints. We find that the corpus represents an emerging yet still fragmented form of AI auditing: few existing measures target integration-specific risks; large gaps remain in meeting traditional audit expectations; and access to necessary information and resources significantly influences audit design. Nonetheless, integration can be categorized across three sites (inter-component, system-environment, and multi-system), each serving the functions of risk exploration, risk determination, coordination, and procedural regularity. Deviating from other types of evaluations, these audits assess qualities specific to system integration, including compatibility, completeness, and oversight. This review calls on the AI community to prioritize system integration as a core strategy for addressing AI risk, and to develop audit practices capable of capturing failures across components, environments, and systems beyond the reach of component-level evaluation.
DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data
Elizabeth Kourbatski
Hegang Chen
Ziyang Song
Gilles Boire
Marie Hudson
Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constraine… (see more)d by scarce, heterogeneous, and temporal complexity. Developing effective ML pipelines for such data remains time-consuming and error-prone, while existing automated machine learning (AutoML) systems only partially address this challenge because they largely rely on brute-force search over predefined spaces and lack explicit reasoning and memory. We therefore reformulate AutoML for small clinical data from exhaustive search to reasoning-driven refinement. We propose DoctorAgents, an agentic AI framework that autonomously constructs and optimizes end-to-end ML pipelines through specialized large language model (LLM) agents for generation, validation, and refinement. DoctorAgents backpropagates natural-language feedback through textual gradient descent to perform targeted updates without exhaustive search. Experiments across diverse clinical tasks show that DoctorAgents consistently outperforms established AutoML baselines while producing more interpretable task-specific representations.
Generative Meta-Models are Adversarial Classifiers
We study adversarial prompt detection in LLMs through the geometry of intermediate activations. Our hypothesis is that token-level adversari… (see more)al attacks induce hidden states that lie outside the activation manifold associated with typical prompts. To test this, we use Generative Latent Priors (GLP), diffusion-based meta-models trained on LLM residual stream activations, as proxies for this manifold (Luo et al., 2026). We evaluate several training-free anomaly scores derived from GLP, namely Guard-GLP, including reconstruction error, diffusion time estimation, and a density-based estimate using Hutchinson trace estimation. Empirically, GLP reconstruction error separates benign prompts from adversarial prompts and provides a competitive classifier without supervised fine-tuning. We also study GLP as a regularizer for activation steering, showing that denoising edited activations can reduce attack success while limiting over-refusal and nonsensical generations. Overall, our results suggest that generative priors over LLM activations provide a useful interface for both adversarial prompt detection and safer activation-level interventions.
Inference for subgraph densities in noisy dynamic networks
Peter W. MacDonald
Eric D. Kolaczyk
In this work we develop statistical methodology to estimate and perform inference on subgraph densities using time-indexed, or dynamic netwo… (see more)rk sequences. These estimates explicitly adjust for observation errors for the network edges, and have good theoretical properties as the size of the network grows. By specifying a stochastically evolving hidden Markov network model, we address two important directions for further investigation identified by Chang et al. (2022): robustness to non-identical network replicates, and efficient aggregation of multiple available network snapshots. These new methods vastly expand the analysis of noisy networks to new data settings, as network replicates are commonly observed dynamically. The methodology is also extended to consider joint inference for subgraph densities at multiple time points, to facilitate formal statistical comparison of dynamic network snapshots.
On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity
Andrei Liviu Nicolicioiu
On-policy self-distillation achieves strong pass@1 accuracy by using a single model as both teacher and student, with the teacher conditione… (see more)d on a correct demonstration to provide dense token-level feedback. We show that this could come at a hidden cost: rollout diversity decreases and pass@k curves flatten (i.e., generating more rollouts fails to improve accuracy). We trace this to compounding biases in the design of self-distillation with sampled demonstrations. The teacher scores each student rollout while conditioned on a sampled correct rollout, channeling its feedback through the model's own biases. We theoretically analyze the optimal self-distillation policy and show that it tilts the base distribution by a pointwise conditional mutual information score between the student's rollout and the correct rollout used as context. Unlike the ideal optimal on-policy reinforcement learning (RL), which preserves probability ratios among equally correct rollouts, self-distillation can amplify existing probability gaps, concentrating mass on already-dominant modes. On a controlled graph path-finding task and science question-answering benchmarks, self-distilled models match or exceed RL on average performance but exhibit substantially lower \textit{functional} and \textit{semantic} diversity, failing on out-of-distribution settings that require diverse strategies.
Dynamically Allocating Evaluation Effort for Model Ranking
Vilém Zouhar
Alon Lavie
Tom Kocmi
Matt Post
Ondrej Bojar
Mrinmaya Sachan
While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-… (see more)performing models, typical evaluation protocols waste effort by exhaustively evaluating all models on the entire benchmark, a safe but inefficient approach. In this work, we formalize multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model. By sampling adaptively based on the intermediate model rankings obtained on the samples so far, we can focus the annotation budget on the most competitive models. We prove the optimality of the proposed algorithms and show that it improves discrimination between top-performing models. This makes evaluations faster, cheaper and more aligned with large-scale competition evaluation goals.