Publications

Intermediate Bilevel Optimization: Modeling Endogenous Follower Tie-Breaking Behavior
Maria Bazotte
Thibaut Vidal
In bilevel optimization, optimistic and pessimistic follower behaviors are the most commonly used forms to define how the follower ties-brea… (see more)ks among multiple optimal solutions. In this work, we go beyond these extreme tie-breaking behaviors and investigate the intermediate bilevel optimization program (I-BO), where the follower's selected optimal response is a decision-dependent random event, with a probability measure influenced by the leader's decision. We formally introduce a class of such endogenous measures, including the special case of strong-weak decision-dependent I-BO. We reformulate the I-BO as a Transformed I-BO (T-I-BO) with exogenous uncertainty by defining inverse and Markov-chain transformations, which represent the follower's response as a function of the leader's decision and exogenous randomness. We handle the T-I-BO's uncertainty via sample-average approximation (SAA), and we propose tailored approaches for its SAA program according to the chosen transformation. Computationally, our methods solve reasonable-sized instances efficiently and outperform the deterministic equivalent when available. Furthermore, experiments stress the critical need to accurately model follower tie-breaking behavior, particularly depending on its alignment with the leader's objective, as misspecification leads to suboptimal leader decisions.
Perceived Regret: Evaluating Agents in Any World
Continual reinforcement learning considers an agent receiving and learning from a stream of experience. The agent only learns about its worl… (see more)d through these experiences and aims to maximize its accumulated reward. A fundamental problem is to evaluate this agent using only its stream of experience, without assuming a particular structure of the world. We outline a novel approach to this problem: Define an \textit{examiner} that observes the same stream of experience and, at every timestep, computes a \textit{perceived regret}, representing the agent's suboptimality gap in value from the examiner's perspective. We empirically demonstrate how perceived regret can be a useful performance measure, applicable to arbitrary streams of experience. Theoretically, we prove that there is no universal examiner, capable of accurately assessing all agents in any world but specific modeling biases can enable success in associated environments. Further experiments show the flexibility of this approach and explore its nuances.
QD-Learning for Continual Reinforcement Learning
Zijing Wu
Nishanth Anand
In continual reinforcement learning (CRL), an agent continually adapts to an unbroken stream of experience by balancing two conflicting obje… (see more)ctives: stability and plasticity. To balance these two objectives, we introduce QD-learning, a new class of algorithms rooted in KL-regularized RL. QD-learning algorithms combine \textit{an adaptive default policy} with \textit{a soft-Q policy} to determine the agent's behaviour at each time step. Similar to habit learning in the brain, the adaptive default policy is updated to mimic the agent's overall behaviour, while the soft-Q policy is updated using TD-error along with a regularizing term based on the KL divergence between the default and the agent's current policies. Our paradigm, QD-learning, encompasses capacity-constrained RL and existing KL-regularized RL approaches as special cases while remaining general. We present three instantiations of QD-learning in dynamic programming and sample-based settings and analyze their theoretical properties via policy-improvement theorems. Additionally, we empirically evaluate our algorithms on both single-task and continual RL settings using a tabular chain MDP. The results show that QD-learning clearly outperforms baselines, particularly in environments with a large number of actions.
Replay Buffer with Efficient Local Forgetting for Adaptation to Local Reward Changes in Deep Model-Based Reinforcement Learning
Ruis MacDonald
Dilith Jayakody
Sageev Oore
We study adaptation in the Local Change Adaptation (LoCA) setup, a problem closely related to continual RL, where an agent must adapt to a l… (see more)ocal reward change while retaining prior knowledge of unchanged regions, without an explicit change signal. This requires updating the learned environment model in the changed region while retaining still-valid knowledge elsewhere. Standard First-In-First-Out (FIFO) replay buffers can fail by mixing outdated and updated transitions in the changed region while discarding transitions from regions not recently visited. The Local Forgetting (LoFo) replay buffer addresses this interference-forgetting dilemma by discarding the oldest transition from the same local region as the newly observed transition. However, its current instantiation identifies this local region by comparing each new transition's learned state embedding with all stored transitions in the buffer, which is expensive. We propose LoFoV2, which uses hash-based assignment of samples to local regions and per-region FIFO buffers to preserve localized discarding without buffer-wide distance comparisons. On local reward-change tasks with Deep Dyna-Q in low- and high-dimensional environments, LoFoV2 improves computational efficiency while maintaining adaptive performance.
The Three Regimes of Offline-to-Online Reinforcement Learning
Offline-to-online reinforcement learning (RL) has emerged as a practical paradigm that leverages offline datasets for pretraining and online… (see more) interactions for fine-tuning. However, its empirical behavior is highly inconsistent: design choices of online-fine tuning that work well in one setting can fail completely in another. We propose a stability--plasticity principle that can explain this inconsistency: we should preserve the knowledge of pretrained policy or offline dataset during online fine-tuning, whichever is better, while maintaining sufficient plasticity. This perspective identifies three regimes of online fine-tuning, each requiring distinct stability properties. We validate this framework through a large-scale empirical study, finding that the results strongly align with its predictions in 45 of 63 cases. This work provides a principled framework for guiding design choices in offline-to-online RL based on the relative performance of the offline dataset and the pretrained policy.
Abstention-Aware Personalized Object Rearrangement via Uncertainty-Guided LLM Assistance
Sam Collin
Robotic assistance in household environments requires not only predicting where objects should be placed, but also reasoning about when obje… (see more)cts should not be placed at all. Existing approaches to personalized object rearrangement primarily focus on placement decisions under the assumption of clean observations and complete actionability, limiting their applicability in realistic, cluttered, and partially erroneous settings. In this paper, we introduce APOLLO, a hybrid framework for abstention-aware personalized object rearrangement that combines a lightweight, personalized embedding model (PEM) with selective large language model (LLM) assistance. PEM is trained for each user-environment pair using a small number of demonstrations, operates entirely on CPU, and produces uncertainty estimates, which are used to selectively invoke LLM-based reasoning only for ambiguous decisions, balancing efficiency, privacy, and reasoning capability. To evaluate this formulation beyond existing benchmarks, we introduce APOR, a synthetic, LLM-generated dataset that captures room-level, multi-furniture environments, diverse organizational profiles, explicit abstention behavior, and noisy partial scene context. Extensive experiments on both PARSEC and APOR provide initial evidence that APOLLO improves over prior LLM-based baselines in controlled benchmark settings while substantially reducing LLM usage. Code is available at https://github.com/PaInt-Lab/APOLLO.
A Functional Approach to Synthesizing Routable Programmable Accelerators for Neural Networks
Paul Teng
T V Paul
Producing optimized accelerators is tedious, as even modern HDLs (Hardware Description Languages) such as Chisel, require reasoning about lo… (see more)w-level concepts. Recent functional approaches, such as Aetherling and SHIR, treat hardware as composition of pure operators. This raises the abstraction level, allowing for systematic optimizations through rewriterules for FPGAs (Field Programmable Gate Arrays).
Rethinking Groups in Critic-Free RLVR
Yihong Wu
Lingfeng Xiao
Muzhi Li
Xinyu Wang
Yingxue Zhang
Jian-Yun Nie
Reinforcement learning (RL) has become a central paradigm for post-training large language models. Existing critic-free RL methods typically… (see more) generate a group of rollouts for the same question to estimate value baselines for advantage computation. However, this design suffers from data inefficiency, group synchronization barriers, and inflexibility with structured rollouts. In this work, we revisit the role of the ``group'' and show that its underlying function is not merely to estimate baselines but to prevent false penalties on negative samples. Building on this insight, we propose negative token filtering, a simple and effective strategy that enables stable single-rollout training. We apply it to two batch-level advantage methods, achieving comparable performance on reasoning tasks and stronger performance on agentic tasks relative to group-based RL techniques.
3D Scene Graphs: Open Challenges and Future Directions
Dennis Rotondi
Sebastian Koch
Nathan Hughes
Martin Buechner
Johanna Wald
Lukas Rosenberger Schmid
Daniele Nardi
Abhinav Valada
Federico Tombari
Luca Carlone
Kai O. Arras
3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relation… (see more)al abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including manipulation, navigation, task planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real-world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are built from raw sensory observations, discussing the most common terminologies, conventions, and techniques. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed content, accessible at https://3dscenegraphs.com/.
Sirop: A Small IR for HLS with Parallel Patterns
L.T. Hildebrand
Designers of custom streaming accelerators traditionally use HDLs (Hardware Description Languages), but this is time-consuming and requires … (see more)advanced hardware expertise. C-based HLS (High-Level Synthesis) offers a higher level of abstraction and faster design time, but still requires some hardware expertise and performance is often left on the table. A promising direction is to use HLS with high-level functional parallel patterns such as map and reduce. Prior works have shown that high performance is achievable this way. However, designing such compiler systems is challenging because the optimizer must handle a large number of language primitives and interactions between them.
VISTA: Scale-Aware Visual Navigation via Action History Conditioning
Vision Navigation Foundation Models (VNMs) promise end-to-end learned navigation policies capable of zero-shot deployment across diverse emb… (see more)odiments and environments. To maintain generality, many vision-based navigation models predict normalized actions. However, this normalization introduces a critical deployment vulnerability: applying different scaling factors to the same normalized trajectory alters its physical geometry, which degrades navigation performance and increases collision risks. We address this vulnerability by conditioning the model on normalized action histories alongside image observations, providing explicit context on the relationship between the model's predictions and the robot's actual physical displacement. Furthermore, current VNMs often struggle in visually repetitive environments that lack distinct features. To resolve this issue, we integrate a DINOv3 encoder, whose richer representations enable our model to capture both spatial and geometric dimensions between observations. VISTA generalizes robustly to out-of-distribution environments, achieving 100% goal prediction accuracy in zero-shot, real-world deployment in Outdoor, Forest and Office settings, and an average of 95% checkpoints crossed, demonstrating consistent path following in unseen environments.
On Defining Erasure Harms for NLP
Arnav Goel
Jackie Chi Kit Cheung
Ziang Xiao
The deployment of NLP systems has raised concerns about harms they might produce, including representational harms. Recent literature has be… (see more)gun to conceptualize and measure one such harm, the harm of erasure. Nevertheless, the field lacks a clear and cohesive conceptual foundation for identifying and measuring erasure. Existing conceptualizations of erasure are often broad -- making it difficult to identify what is needed to establish and measure erasure -- or else specific to particular settings -- facilitating measurement for those settings but potentially challenging to adapt to other settings. To address this gap, we develop and propose a structured definition of erasure that clarifies what components are necessary for establishing whether erasure has occurred, which practitioners need to explicitly articulate and operationalize in order to measure erasure.