The upcoming meeting, taking place on November 10 at Mila, will explore how we can collectively develop, govern, and deploy high-performing, reliable, and secure agentic systems by connecting academic researchers, industry experts, and practitioners.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Publications
A game theory for foundation models shows new paths to rational cooperation through similarity inference
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles… (see more) governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,'where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the `embedded Bayesian agent,'a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms. We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner. We formalize this mechanism of similarity inference through the `embedded equilibrium,'a novel solution concept replacing the Nash equilibrium to provide a foundational game theory for the social behavior of modern AI agents.
Large language models (LLMs) are trained for downstream tasks by updating their parameters (e.g., via RL). However, updating parameters forc… (see more)es them to absorb task-specific information, which can result in catastrophic forgetting and loss of plasticity. In contrast, in-context learning with fixed LLM parameters can cheaply and rapidly adapt to task-specific requirements (e.g., prompt optimization), but cannot by itself typically match the performance gains available through updating LLM parameters. There is no good reason for restricting learning to being in-context or in-weights. Moreover, humans also likely learn at different time scales (e.g., System 1 vs 2). To this end, we introduce a fast-slow learning framework for LLMs, with model parameters as "slow" weights and optimized context as "fast" weights. These fast "weights" can learn from textual feedback to absorb the task-specific information, while allowing slow weights to stay closer to the base model and persist general reasoning behaviors. Fast-Slow Training (FST) is up to 3x more sample-efficient than only slow learning (RL) across reasoning tasks, while consistently reaching a higher performance asymptote. Moreover, FST-trained models remain closer to the base LLM (up to 70% less KL divergence), resulting in less catastrophic forgetting than RL-training. This reduced drift also preserves plasticity: after training on one task, FST trained models adapt more effectively to a subsequent task than parameter-only trained models. In continual learning scenarios, where task domains change on the fly, FST continues to acquire each new task while parameter-only RL stalls.
Model merging provides a way of cheaply combining individual models to produce a model that inherits each individual's capabilities.
While s… (see more)ome merging methods can approach the performance of multitask training, they are often heuristically motivated and lack theoretical justification.
A principled alternative is to pose model merging as a layer-wise optimization problem that directly minimizes interference between tasks.
However, this formulation requires estimating per-layer covariance matrices from data, which may not be available when performing merging.
In contrast, many of the heuristically-motivated methods do not require auxiliary data, making them practically advantageous.
In this work, we revisit the interference minimization framework and show that, under certain conditions, covariance matrices can be estimated directly from difference matrices, eliminating the need for data while also reducing computational costs.
We validate our approach across vision and language benchmarks on models ranging from
Research in artificial intelligence is undergoing a paradigm shift from prioritizing model innovations and benchmark scores towards emphasiz… (see more)ing problem definition and rigorous real-world evaluation. As the field enters the "second half," the central challenge becomes real utility in long-horizon, dynamic, and user-dependent settings such as agentic coding, deep research, and computer use,
where LLM-based agents face context explosion beyond fixed context windows and must continuously accumulate, manage, and selectively reuse large volumes of information across extended interactions. Memory, with hundreds of papers released in 2025, therefore emerges as the critical solution to fill the utility gap. Beyond serving as passive storage, memory is increasingly the substrate
through which agents self-evolve: short-term memory gates which experiences are perceived, selected, and abstracted during execution, while long-term memory accumulates and consolidates them into reusable knowledge and skills, forming the loop through which agents improve from their own experience and sustain continual learning. In this survey, we provide a unified view of foundation
agent memory along three dimensions: memory substrate (internal parametric state and external retrieval-augmented stores), cognitive mechanism (sensory, working, episodic, semantic, and procedural), and memory subject (user-centric personalization and agent-centric experience). We then analyze how memory is operated under single- and multi-agent topologies and highlight learning
policies over memory operations, showing how memory management itself is becoming a trainable, self-evolving capability that spans reinforcement-learned context curation, experience consolidation at decision time, and the emerging ecosystem of explicit, portable, and shareable agent skills surfaced through agent harnesses, context engineering, and standardized tool-mediation protocols. Finally, we review evaluation benchmarks and metrics for assessing memory utility, and outline various open challenges and future directions.
2026-08-03
Transactions on Machine Learning Research (accepted)
Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictive uncertainty. Propagat… (see more)ing full state distributions analytically offers a principled way to do this, but has traditionally required restrictive policy or reward structures to remain tractable. Consequently, modern deep reinforcement learning has largely retreated to either stochastic sampling, which introduces significant target variance, or deterministic point estimates that ignore predictive covariance entirely. We investigate whether distribution-aware planning is possible without these constraints. Using a quadratic action-value parameterization, we first reduce the Bellman backup to an expectation over the state-value function alone; the key idea is then a compatibility principle between the predictive transition distribution and the value function class, under which this expectation is analytic in the distribution's moments. We instantiate this principle with a Gaussian transition model paired with a radial-basis value function, yielding a closed-form backup that propagates both predictive mean and covariance. Empirically, our approach reduces target variance and yields well-calibrated predictive uncertainty under stochastic observations in continuous control, providing a principled framework for planning with learned distribution models.
Superforecasters are people that make forecasts that are statistically better than the average person; they offer an interesting testbed for… (see more) persona modelling.
In this preliminary work, we investigate whether large language models (LLMs) can faithfully model superforecaster-like personas making similar forecasts and generating similar rationales behind the forecasts.
When we compare LLM-generated responses to the ground truth responses of superforecasters, we find that few-shot prompting generates numerically closer forecasts than instruction-prompting baselines.
We also measure the similarity of generated and real rationales, finding that the rationales generated based on examples are judged more similar to those of superforecasters than instruction-prompting baselines.
However, an analysis of the reasoning patterns in the rationales shows significant differences between human and LLM forecasters, pointing to a gap between stylistic imitation and deep reasoning similarity.
These preliminary results raise several interesting questions about how getting superforecaster-like behaviour from LLMs actually works, and open new avenues to explore for improving the forecasting behaviour of LLMs.
2026-08-02
Social_Sim @ Conference on Language Modeling (poster)
Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive … (see more)reasoning, which limits systematic search, refinement, and backtracking. While recursive models such as Hierarchical Reasoning Model (HRM) and Tiny Recursive Model (TRM) address this limitation through iterative latent-state refinement, they are typically task-specific and do not leverage pretrained language priors. We propose R-Qwen, a recursive reasoning framework built upon a pretrained Qwen backbone. R-Qwen repeatedly refines a candidate solution through programmatic self-recursion and deep supervision, combining the structured iterative computation of recursive models with the linguistic and reasoning priors of pretrained LLMs. We further adapt Hierarchical Supervision Weighting (HSW) to autoregressive models by exponentially weighting losses across recursive steps. HSW reduces gradient variance by at least 50\%, improves the signal-to-noise ratio of stochastic gradients, and accelerates convergence. Across eight challenging benchmarks, R-Qwen consistently outperforms prior recursive reasoning models and substantially larger LLMs while using a comparable number of trainable parameters. Notably, on ARC-AGI dataset, our model achieves a 27.6\% improvement over the baseline, highlighting the effectiveness of recursive refinement for general symbolic reasoning. These results suggest that recursive reasoning mechanisms and pretrained language model priors are complementary approaches for improving symbolic puzzle-solving. Code and models will be released after acceptance.
While signed-directed graphs have been studied using linear Laplacians in the design of graph neural networks, relatively little research ha… (see more)s focused on developing non-linear Laplacian operators for such networks. We introduce a non-linear Laplacian operator specific to signed and directed networks (NLSD). This non-linear operator extends the concepts of the signed Laplacian for signed graphs and the Laplacian for directed graphs. The NLSD calculates node-specific potentials based on features More precisely, if the potential discrepancy is not aligned with the edge direction, we ignore it (and vice versa) leveraging message-passing techniques only across edges where potential discrepancies align with the edge's direction. Utilizing this novel operator, we propose an efficient spectral GNN framework (NLSD-GNN). We conducted comprehensive evaluations focusing on node classification and link prediction, examining scenarios involving signed, directional, or both types of information. Our findings reveal that this spectral GNN framework not only integrates signed and directional data effectively but also achieves superior performance across diverse datasets.
De novo L-(+)-tartaric acid biosynthesis in multi-modular engineered yeasts
Xuan Zhou
Jiaheng Hou
Zikai Wang
Zhendong Li
Yang Li
Xitong Li
Xianhao Xu
Yanfeng Liu
Jianghua Li
Guocheng Du
Dacheng Ma
J. Tang
Jian Chen
Xueqin Lv
Long Liu
L-(+)-tartaric acid (L-TA) is a high-value chiral organic acid essential for food and pharmaceuticals. Despite its industrial importance, su… (see more)stainable green production is constrained by the lack of a fully defined biosynthetic pathway. Here, we report the de novo biosynthesis of L-TA in Saccharomyces cerevisiae through reaction-guided enzyme mining, experimental validation, and Enzyme Commission-specific Catalytic Hybrid Optimizer (ECHO)-assisted enzyme prioritization. We first elucidate the elusive two-step conversion from precursor 5-keto-D-gluconic acid (5-KGA) to L-TA, catalyzed by transketolase (TK) and succinate semialdehyde dehydrogenase (SSDH). To optimize this critical step, we develop the ECHO. This multimodal framework integrates sequence, substrate, and pocket-aware structural information to identify high-performance TK-SSDH pairs. By integrating this pathway with de novo precursor synthesis, cofactor engineering, and semi-rational protein engineering, a final L-TA titer of 6.59 mg L−1 was achieved in a 5-L bioreactor. By connecting computational mining and metabolic assembly through a multi-module engineering strategy, our study establishes a green platform for L-TA production and demonstrates an effective workflow for synthetic pathway design. L-(+)-tartaric acid (L-TA) is a high-value chiral organic acid for food and pharmaceuticals. Here the authors produce L-TA in S. cerevisiae through reaction-guided enzyme mining and Enzyme Commission-specific Catalytic Hybrid Optimizer (ECHO)-assisted enzyme prioritization.
Preparation of MgAl2O4-reinforced magnesium composite refractories from high silicon magnesite tailings: α-Al2O3 /AlN reaction mechanism and thermal shock resistance