The upcoming meeting, taking place on November 10 at Mila, will explore how we can collectively develop, govern, and deploy high-performing, reliable, and secure agentic systems by connecting academic researchers, industry experts, and practitioners.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
In bilevel optimization, optimistic and pessimistic follower behaviors are the most commonly used forms to define how the follower ties-brea… (see more)ks among multiple optimal solutions. In this work, we go beyond these extreme tie-breaking behaviors and investigate the intermediate bilevel optimization program (I-BO), where the follower's selected optimal response is a decision-dependent random event, with a probability measure influenced by the leader's decision. We formally introduce a class of such endogenous measures, including the special case of strong-weak decision-dependent I-BO. We reformulate the I-BO as a Transformed I-BO (T-I-BO) with exogenous uncertainty by defining inverse and Markov-chain transformations, which represent the follower's response as a function of the leader's decision and exogenous randomness. We handle the T-I-BO's uncertainty via sample-average approximation (SAA), and we propose tailored approaches for its SAA program according to the chosen transformation. Computationally, our methods solve reasonable-sized instances efficiently and outperform the deterministic equivalent when available. Furthermore, experiments stress the critical need to accurately model follower tie-breaking behavior, particularly depending on its alignment with the leader's objective, as misspecification leads to suboptimal leader decisions.
Continual reinforcement learning considers an agent receiving and learning from a stream of experience. The agent only learns about its worl… (see more)d through these experiences and aims to maximize its accumulated reward. A fundamental problem is to evaluate this agent using only its stream of experience, without assuming a particular structure of the world.
We outline a novel approach to this problem: Define an \textit{examiner} that observes the same stream of experience and, at every timestep, computes a \textit{perceived regret}, representing the agent's suboptimality gap in value from the examiner's perspective.
We empirically demonstrate how perceived regret can be a useful performance measure, applicable to arbitrary streams of experience.
Theoretically, we prove that there is no universal examiner, capable of accurately assessing all agents in any world but specific modeling biases can enable success in associated environments.
Further experiments show the flexibility of this approach and explore its nuances.
In continual reinforcement learning (CRL), an agent continually adapts to an unbroken stream of experience by balancing two conflicting obje… (see more)ctives: stability and plasticity. To balance these two objectives, we introduce QD-learning, a new class of algorithms rooted in KL-regularized RL. QD-learning algorithms combine \textit{an adaptive default policy} with \textit{a soft-Q policy} to determine the agent's behaviour at each time step. Similar to habit learning in the brain, the adaptive default policy is updated to mimic the agent's overall behaviour, while the soft-Q policy is updated using TD-error along with a regularizing term based on the KL divergence between the default and the agent's current policies. Our paradigm, QD-learning, encompasses capacity-constrained RL and existing KL-regularized RL approaches as special cases while remaining general. We present three instantiations of QD-learning in dynamic programming and sample-based settings and analyze their theoretical properties via policy-improvement theorems. Additionally, we empirically evaluate our algorithms on both single-task and continual RL settings using a tabular chain MDP. The results show that QD-learning clearly outperforms baselines, particularly in environments with a large number of actions.
We study adaptation in the Local Change Adaptation (LoCA) setup, a problem closely related to continual RL, where an agent must adapt to a l… (see more)ocal reward change while retaining prior knowledge of unchanged regions, without an explicit change signal. This requires updating the learned environment model in the changed region while retaining still-valid knowledge elsewhere. Standard First-In-First-Out (FIFO) replay buffers can fail by mixing outdated and updated transitions in the changed region while discarding transitions from regions not recently visited. The Local Forgetting (LoFo) replay buffer addresses this interference-forgetting dilemma by discarding the oldest transition from the same local region as the newly observed transition. However, its current instantiation identifies this local region by comparing each new transition's learned state embedding with all stored transitions in the buffer, which is expensive. We propose LoFoV2, which uses hash-based assignment of samples to local regions and per-region FIFO buffers to preserve localized discarding without buffer-wide distance comparisons. On local reward-change tasks with Deep Dyna-Q in low- and high-dimensional environments, LoFoV2 improves computational efficiency while maintaining adaptive performance.
Offline-to-online reinforcement learning (RL) has emerged as a practical paradigm that leverages offline datasets for pretraining and online… (see more) interactions for fine-tuning. However, its empirical behavior is highly inconsistent: design choices of online-fine tuning that work well in one setting can fail completely in another. We propose a stability--plasticity principle that can explain this inconsistency: we should preserve the knowledge of pretrained policy or offline dataset during online fine-tuning, whichever is better, while maintaining sufficient plasticity. This perspective identifies three regimes of online fine-tuning, each requiring distinct stability properties. We validate this framework through a large-scale empirical study, finding that the results strongly align with its predictions in 45 of 63 cases. This work provides a principled framework for guiding design choices in offline-to-online RL based on the relative performance of the offline dataset and the pretrained policy.
Robotic assistance in household environments requires not only predicting where objects should be placed, but also reasoning about when obje… (see more)cts should not be placed at all. Existing approaches to personalized object rearrangement primarily focus on placement decisions under the assumption of clean observations and complete actionability, limiting their applicability in realistic, cluttered, and partially erroneous settings. In this paper, we introduce APOLLO, a hybrid framework for abstention-aware personalized object rearrangement that combines a lightweight, personalized embedding model (PEM) with selective large language model (LLM) assistance. PEM is trained for each user-environment pair using a small number of demonstrations, operates entirely on CPU, and produces uncertainty estimates, which are used to selectively invoke LLM-based reasoning only for ambiguous decisions, balancing efficiency, privacy, and reasoning capability. To evaluate this formulation beyond existing benchmarks, we introduce APOR, a synthetic, LLM-generated dataset that captures room-level, multi-furniture environments, diverse organizational profiles, explicit abstention behavior, and noisy partial scene context. Extensive experiments on both PARSEC and APOR provide initial evidence that APOLLO improves over prior LLM-based baselines in controlled benchmark settings while substantially reducing LLM usage. Code is available at https://github.com/PaInt-Lab/APOLLO.
Producing optimized accelerators is tedious, as even modern HDLs (Hardware Description Languages) such as Chisel, require reasoning about lo… (see more)w-level concepts. Recent functional approaches, such as Aetherling and SHIR, treat hardware as composition of pure operators. This raises the abstraction level, allowing for systematic optimizations through rewriterules for FPGAs (Field Programmable Gate Arrays).
2026-06-14
Proceedings of the 27th ACM SIGPLAN/SIGBED International Conference on Languages, Compilers, and Tools for Embedded Systems (published)
Reinforcement learning (RL) has become a central paradigm for post-training large language models. Existing critic-free RL methods typically… (see more) generate a group of rollouts for the same question to estimate value baselines for advantage computation. However, this design suffers from data inefficiency, group synchronization barriers, and inflexibility with structured rollouts. In this work, we revisit the role of the ``group'' and show that its underlying function is not merely to estimate baselines but to prevent false penalties on negative samples. Building on this insight, we propose negative token filtering, a simple and effective strategy that enables stable single-rollout training. We apply it to two batch-level advantage methods, achieving comparable performance on reasoning tasks and stronger performance on agentic tasks relative to group-based RL techniques.
3D Scene Graphs (3DSGs) have emerged as a powerful representation for spatial AI by combining geometric grounding with semantic and relation… (see more)al abstractions of the environment. Their expressiveness has made them relevant to a broad range of problems in robotics and computer vision, including manipulation, navigation, task planning, scene understanding, and many others. However, the field remains fragmented: different communities adopt distinct formulations, construction pipelines, and evaluation protocols, making it difficult to compare methods, identify common assumptions, and assess remaining challenges for robust real-world deployment. This survey provides a unified and critical review of 3DSGs, with particular emphasis on open challenges and future directions. We first formalize 3DSGs under a common definition and analyze the principal modeling choices that characterize existing formulations, including node and edge attributes, hierarchical structure, dynamic scene representations, and affordance-aware extensions. We then review how 3DSGs are built from raw sensory observations, discussing the most common terminologies, conventions, and techniques. Finally, we examine downstream applications and evaluation strategies, from intrinsic graph quality to task-level performance. To support the community, we also provide a dedicated website that organizes and extends the surveyed content, accessible at https://3dscenegraphs.com/.
Designers of custom streaming accelerators traditionally use HDLs (Hardware Description Languages), but this is time-consuming and requires … (see more)advanced hardware expertise. C-based HLS (High-Level Synthesis) offers a higher level of abstraction and faster design time, but still requires some hardware expertise and performance is often left on the table. A promising direction is to use HLS with high-level functional parallel patterns such as map and reduce. Prior works have shown that high performance is achievable this way. However, designing such compiler systems is challenging because the optimizer must handle a large number of language primitives and interactions between them.
2026-06-14
Proceedings of the 27th ACM SIGPLAN/SIGBED International Conference on Languages, Compilers, and Tools for Embedded Systems (published)
Vision Navigation Foundation Models (VNMs) promise end-to-end learned navigation policies capable of zero-shot deployment across diverse emb… (see more)odiments and environments. To maintain generality, many vision-based navigation models predict normalized actions. However, this normalization introduces a critical deployment vulnerability: applying different scaling factors to the same normalized trajectory alters its physical geometry, which degrades navigation performance and increases collision risks. We address this vulnerability by conditioning the model on normalized action histories alongside image observations, providing explicit context on the relationship between the model's predictions and the robot's actual physical displacement. Furthermore, current VNMs often struggle in visually repetitive environments that lack distinct features. To resolve this issue, we integrate a DINOv3 encoder, whose richer representations enable our model to capture both spatial and geometric dimensions between observations. VISTA generalizes robustly to out-of-distribution environments, achieving 100% goal prediction accuracy in zero-shot, real-world deployment in Outdoor, Forest and Office settings, and an average of 95% checkpoints crossed, demonstrating consistent path following in unseen environments.
The deployment of NLP systems has raised concerns about harms they might produce, including representational harms. Recent literature has be… (see more)gun to conceptualize and measure one such harm, the harm of erasure. Nevertheless, the field lacks a clear and cohesive conceptual foundation for identifying and measuring erasure. Existing conceptualizations of erasure are often broad -- making it difficult to identify what is needed to establish and measure erasure -- or else specific to particular settings -- facilitating measurement for those settings but potentially challenging to adapt to other settings. To address this gap, we develop and propose a structured definition of erasure that clarifies what components are necessary for establishing whether erasure has occurred, which practitioners need to explicitly articulate and operationalize in order to measure erasure.