The upcoming meeting, taking place on November 10 at Mila, will explore how we can collectively develop, govern, and deploy high-performing, reliable, and secure agentic systems by connecting academic researchers, industry experts, and practitioners.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Publications
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under … (see more)asymmetric information and strategic deception due to conflicting or hidden objectives. In these settings, misalignment with collective goals becomes a central concern. We propose a novel framework for evaluating objective misalignment using the social deduction game Werewolf, modifying the objective of a single agent while preserving its assigned role. Across LLMs from four different model families and sizes, four player roles, and three objective formulations, we introduce a dual analysis of the agents' internal reasoning and their public cheap-talk behavior (i.e costless, non-binding communication that does not directly affect the agents' utilities), complemented by an analysis of game outcomes. Our results show that objective misalignment undermines outcomes in inherently adversarial environments, an effect exacerbated by asymmetric information and specialized roles. While compromised agents consistently develop distinct objective-dependent reasoning strategies, these adaptations remain largely invisible in their public behavior. More broadly, our findings suggest that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.
Abstract Introduction T cells play a central role in adaptive immune responses, enabling vertebrates to fight infections and eliminate cance… (see more)r cells. Cancer immunotherapies, and in particular adoptive T-cell therapies, thus aim to harness their tumor-destroying capabilities to treat, or even cure, cancer patients. Despite recent advances in the development of these therapies, a major limitation remains our inability to predict, a priori, which tumor-infiltrating T cells will be best equipped to carry out anti-tumor immunity. Indeed, this multifactorial problem requires a model that integrates information on T cell receptor (TCR) specificity, tumor antigen abundance, and T cell activation history within the complex tumor microenvironment, yet such considerations are largely absent from current T cell selection strategies. Methods To address this gap, we are using a combination approach of high-throughput robotic multiplexing with computational and machine learning techniques to identify signatures of T cell phenotype that best predict response to tumor antigens. Specifically, we aim to study the roles of antigen presentation, TCR/antigen affinity, and inflammatory milieu on shaping T cell phenotype at the single-cell level, and training machine learning models to re-derive the activation history of T cells both in vivo and ex vivo. Results Preliminary results from in vitro co-culture data of T cells with cognate antigen-bearing splenocytes suggests that T cell antigen strength can indeed be back-calculated from single cell phenotype as measured by spectral flow cytometry. Conclusion With a validated model of tumor antigenicity in T cells, this research project aims to better inform T cell selection and optimal preparation strategies for adoptive T-cell therapies. Funding Source n/a Topic Categories Computational and Systems Immunology (COMP)
We are studying the Human Body as a System with its Nutrition Requirement as energy pre-requisites & necessities. We find out that it is pos… (see more)sible to submit our daily nutrition requirements to some restaurants & Jounals : this practice is called Dietary Requests with Accomodations.
Post-translational modifications (PTMs) expand protein function by encoding context-dependent regulatory states, and their dysregulation con… (see more)tributes to cancer, neurodegeneration and metabolic disease. However, existing methods treat PTMs as independent residue labels, limiting their ability to distinguish contextually permissible sites, model crosstalk and infer functional consequences. Here we introduce ProtSyntax, a PTM-aware foundation protein language model combining protein-aware positional encoding, bidirectional state-space propagation, geometry-constrained attention and adaptive multi-objective learning. This design integrates residue chemistry, motif order, long-range context and three-dimensional microenvironments while coupling PTM recognition to enzyme function. Across 40 PTM-site benchmarks, ProtSyntax exceeded the strongest baselines in mean MCC and AP by 12.66% and 10.67%. ProtSyntax also recovered masked PTM types and sites, rejected structural decoys, generalized to data-scarce modifications, reconstructed crosstalk and linked PTM perturbations to enzyme kinetics. Applications to pathogenic variants, biomolecular condensates and disease-associated PTM landscapes demonstrate its potential to decode the regulatory language of the modified proteome.
Much like humans benefit from guidance while learning, reinforcement learning algorithms may benefit from additional supervision beyond rewa… (see more)rds. Leveraging additional information during training to learn better representations and behaviors has been the focus of asymmetric reinforcement learning. This learning paradigm has proven effective under partial observability when additional state information is available, but also under full observability when more refined state information is available. Focusing on model-based reinforcement learning, we study the effect of asymmetric learning on observation representations and on privileged information representations. First, we identify a limitation in the privileged information representations learned by an asymmetric model-based algorithm known as the Informed Dreamer. Then, we propose a novel asymmetric representation learning objective using latent guidance, resulting in a new algorithm called the Reinformed Dreamer. Experiments across several benchmarks show a more consistent improvement over Dreamer than previous asymmetric approaches.
Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large l… (see more)anguage models (LLMs) and their users. In this paper, we evaluate the ability of LLMs to recognize unspoken beliefs made through implicatures and to understand their updates through implicature cancellation: the pragmatic phenomenon whereby an utterance's implied meaning is weakened or negated. We create the first expert-annotated implicature cancellation dataset, ImplicatureX, crowdsourced for human judgements of implicatures and their corresponding cancellations. We find that LLM belief update understanding lags behind that of humans, especially in more naturally-occurring scenarios. Additional control experiments suggest that successes in LLM belief updates may stem in part from a reliance on prior beliefs, and that failures in belief updates may depend on their type and on their form. Overall, our study suggests that current LLMs have not yet reached human-level understanding of unspoken beliefs and belief updates. Code and data are available at https://github.com/cesare-spinoso/ImplicatureX.
This work proposes an adaptation of the attention mechanism for triangle meshes. The core observation is that endowing the attention mechani… (see more)sm with critical properties for learning over meshes -- intrinsicality and triangulation-agnosticism -- enables it to attain state-of-the-art results over several learning-based tasks in geometry-processing. The above is achieved by modifying the attention mechanism from the bottom up based on simple principles from geometry-processing. Namely, the quantities used within attention -- queries, keys and values -- are created by an intrinsic, triangulation-agnostic network, and treated as discretizations of continuous functions. From that, we devise an appropriate attention mechanism that operates over triangle meshes through standard FEM discretization of the resulting integrals of the above functions. Surprisingly, as far as we know, this straightforward approach has not been utilized for learning over meshes. Experiments show our method exceeds current state of the art, including both mesh-based architectures as well as point cloud transformers. Namely, we show significant improvements on several common benchmarks and tasks -- predicting canonical high-frequency signals; predicting deformations; computing dense correspondences, both between full shapes and partial ones; and predicting feature descriptors.
Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstrom's mod… (see more)el of moral hazard in teams, the Dialogue Moral Hazard Game instantiates this hidden-action structure as a textual environment for language agents. An agent chooses between keeping an immediate local reward and paying a query cost to reveal a hidden safety fact that helps another agent's downstream decision. We evaluate fourteen open-weight and four frontier models using measures of information acquisition, communication, downstream use, and team success. In matched 3,015-decision-per-model experiments, GPT-5.6 Sol, Claude Opus 4.8, and Nemotron-3 Ultra track the derived private-share boundary across nine query costs, with mean absolute errors of 0.013, 0.030, and 0.024. Muse Spark 1.1 responds directionally, whereas Fable 5 remains query-saturated. SFT, RLOO, SFT+RLOO, and GEPA produce heterogeneous mechanism changes. GEPA raises Muse team success from 22.2% to 100.0% while reducing query use from 51.1% to 0.3%. Frozen-prompt interventions show that this success depends on a learned rank-label mapping rather than direct revelation: changing the mapping reduces team success from 100.0% to 12.5% and then 0.0%. We introduce CREDIT (Counterfactual Replay for Evidence-Driven Information Transfer), a mechanism-aligned multi-agent prompt-optimization algorithm that uses matched hidden-state twins and total-action replay to reward robust causal contribution rather than query frequency. Across five models and multiple seeds, CREDIT preserves query-mediated behavior while revealing model-specific acquisition and downstream-use bottlenecks. Optimization can reach the same aggregate outcome through direct revelation or a learned effective information structure, motivating mechanism-level evaluation and optimization rather than team success alone.
Testing quantum programs on NISQ (Noisy Intermediate-Scale Quantum) backends is challenging because the noise disturbs outcome distributions… (see more) and can affect pass/fail decisions. We present Q-BRIDGE, a graph learning-based approach that converts noisy observations into denoised distributions suitable for oracle-based verification. Q-BRIDGE uses a graph transformer architecture to encode a transpiled quantum circuit, capturing the characteristics of its gates and their connectivity; the physical backend information is encoded together with the logical structure of the circuit. An additional conditioning layer, based on FiLM (Feature-Wise Linear Modulation), takes the encoding as input and integrates noisy observations to produce denoised outcomes. We evaluate Q-BRIDGE on 23 IBM noise backends and 6 circuit families representative of practical workloads. In the first setting, we train a separate Q-BRIDGE model for each backend; in the second setting, we train a single general model shared across all backends. Across both settings, Q-BRIDGE outperforms the state-of-the-art baseline in noise mitigation by a large margin. In testing scenarios with noisy executions, Q-BRIDGE achieves 93.97%-94.90% precision and 82.50%-83.51% recall in detecting bug-induced test failures, significantly outperforming the state-of-the-art baseline. These results indicate that considering the graph structure of the transpiled circuits and the physical characteristics of specific quantum backends is a practical route to more reliable noise-aware quantum program testing.
Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning. E… (see more)xisting approaches either focus on restrictive classes of risk measures or rely on access to a simulator, limiting their applicability in fully online settings. In this work, we propose computationally efficient online learning algorithms for policy evaluation in Markov decision processes (MDPs) with dynamic utility-based shortfall risk (UBSR) measures under linear function approximation. Specifically, we introduce the UBSR-TD algorithm, establish conditions under which it converges almost surely, and develop several variants designed to accelerate convergence. Our formulation shows that existing policy evaluation algorithms for risk-neutral MDPs can be readily adapted to dynamic UBSR settings by incorporating a loss function into the temporal-difference error. Numerical experiments support our theoretical findings, and an application to a perishable inventory management problem with shelf-life uncertainty demonstrates the practical effectiveness of the proposed methods.
Abstract Purpose To evaluate how segmentation architecture and dataset-adaptive configuration influence uterine MRI segmentation across hete… (see more)rogeneous benign and malignant tasks. Methods U-Net, Swin-UNETR, and MedNeXt were compared with nnU-Net as a self-configuring reference across T2-weighted MRI datasets: public multiclass UMD anatomy/fibroid segmentation (n=300), institutional endometrial cancer tumor segmentation (n=206), and institutional uterine mass lesion segmentation (n=234). A relabeled external UMD-style cohort (n=12) assessed domain shift. Models used fixed partitions, fold ensembling, Dice, HD95, ASSD, volume error, and paired bootstrap comparisons with Holm correction. Results MedNeXt was the strongest manually controlled architecture. nnU-Net achieved the highest performance on all internal datasets and external testing. Macro-Dice reached 0.761, 0.746, and 0.814 for nnU-Net on UMD, endometrial cancer, and uterine mass datasets, respectively, versus 0.722, 0.726, and 0.789 for MedNeXt. The nnU-Net-MedNeXt gap was largest for multiclass UMD segmentation and smaller in binary tasks. External testing degraded all models; nnU-Net remained highest (0.542), followed by MedNeXt (0.490), U-Net (0.396), and Swin-UNETR (0.287). Conclusions Uterine MRI segmentation performance depended on task, architecture, and evaluation domain. MedNeXt supported modern convolutional design as a strong manual baseline, but nnU-Net remained the most robust overall, emphasizing the importance of dataset-adaptive configuration and external validation.