Publications

A game theory for foundation models shows new paths to rational cooperation through similarity inference
Alexander Meulemans
Maciej Wołczyk
Marissa A. Weis
Rajai Nasser
Roberta Rocca
Seijin Kobayashi
Angelika Steger
Marcus Hutter
James Manyika
Rif A. Saurous
João Sacramento
Blaise Agüera y Arcas
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles… (see more) governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,'where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the `embedded Bayesian agent,'a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms. We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner. We formalize this mechanism of similarity inference through the `embedded equilibrium,'a novel solution concept replacing the Nash equilibrium to provide a foundational game theory for the social behavior of modern AI agents.
Learning, Fast and Slow: Towards LLMs That Adapt Continually
Rishabh Tiwari
Lakshya A Agrawal
Joseph E. Gonzalez
Matei Zaharia
Kurt Keutzer
Inderjit S Dhillon
Devvrit Khatri
Large language models (LLMs) are trained for downstream tasks by updating their parameters (e.g., via RL). However, updating parameters forc… (see more)es them to absorb task-specific information, which can result in catastrophic forgetting and loss of plasticity. In contrast, in-context learning with fixed LLM parameters can cheaply and rapidly adapt to task-specific requirements (e.g., prompt optimization), but cannot by itself typically match the performance gains available through updating LLM parameters. There is no good reason for restricting learning to being in-context or in-weights. Moreover, humans also likely learn at different time scales (e.g., System 1 vs 2). To this end, we introduce a fast-slow learning framework for LLMs, with model parameters as "slow" weights and optimized context as "fast" weights. These fast "weights" can learn from textual feedback to absorb the task-specific information, while allowing slow weights to stay closer to the base model and persist general reasoning behaviors. Fast-Slow Training (FST) is up to 3x more sample-efficient than only slow learning (RL) across reasoning tasks, while consistently reaching a higher performance asymptote. Moreover, FST-trained models remain closer to the base LLM (up to 70% less KL divergence), resulting in less catastrophic forgetting than RL-training. This reduced drift also preserves plasticity: after training on one task, FST trained models adapt more effectively to a subsequent task than parameter-only trained models. In continual learning scenarios, where task domains change on the fly, FST continues to acquire each new task while parameter-only RL stalls.
Merging Adapted Models via Data-Free Covariance Estimation
Derek Tam
Pascal Junior Tikeng Notsawo
Colin Raffel
Model merging provides a way of cheaply combining individual models to produce a model that inherits each individual's capabilities. While s… (see more)ome merging methods can approach the performance of multitask training, they are often heuristically motivated and lack theoretical justification. A principled alternative is to pose model merging as a layer-wise optimization problem that directly minimizes interference between tasks. However, this formulation requires estimating per-layer covariance matrices from data, which may not be available when performing merging. In contrast, many of the heuristically-motivated methods do not require auxiliary data, making them practically advantageous. In this work, we revisit the interference minimization framework and show that, under certain conditions, covariance matrices can be estimated directly from difference matrices, eliminating the need for data while also reducing computational costs. We validate our approach across vision and language benchmarks on models ranging from
A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
Wei-Chieh Huang
Weizhi Zhang
Yueqing Liang
Yuanchen Bei
Yankai Chen
Tao Feng
xinyu Pan
Zhen Tan
Yu Wang
Tianxin Wei
Shanglin Wu
Ruiyao Xu
Liangwei Yang
Rui Yang
Wooseong Yang
Chin-Yuan Yeh
Hanrong Zhang
Haozhen Zhang
Siqi Zhu
Henry Peng Zou … (see 40 more)
Wanjia Zhao
Song Wang
Wujiang Xu
Zixuan Ke
Zheng Hui
Dawei Li
Yaozu Wu
Langzhou He
Chen Wang
Xiongxiao Xu
Baixiang Huang
Juntao Tan
Shelby Heinecke
Huan Wang
Caiming Xiong
Ahmed Metwally
Jun Yan
Chen-Yu Lee
Hanqing Zeng
Yinglong Xia
Xiaokai Wei
Ali Payani
Yu Wang
Haitong Ma
Wenya Wang
Chenguang Wang
Yu Zhang
Xin Eric Wang
Yongfeng Zhang
Jiaxuan You
Hanghang Tong
Xiao Luo
Xue Liu
Yizhou Sun
Wei Wang
Julian McAuley
James Zou
Jiawei Han
Philip S. Yu
Kai Shu
Research in artificial intelligence is undergoing a paradigm shift from prioritizing model innovations and benchmark scores towards emphasiz… (see more)ing problem definition and rigorous real-world evaluation. As the field enters the "second half," the central challenge becomes real utility in long-horizon, dynamic, and user-dependent settings such as agentic coding, deep research, and computer use, where LLM-based agents face context explosion beyond fixed context windows and must continuously accumulate, manage, and selectively reuse large volumes of information across extended interactions. Memory, with hundreds of papers released in 2025, therefore emerges as the critical solution to fill the utility gap. Beyond serving as passive storage, memory is increasingly the substrate through which agents self-evolve: short-term memory gates which experiences are perceived, selected, and abstracted during execution, while long-term memory accumulates and consolidates them into reusable knowledge and skills, forming the loop through which agents improve from their own experience and sustain continual learning. In this survey, we provide a unified view of foundation agent memory along three dimensions: memory substrate (internal parametric state and external retrieval-augmented stores), cognitive mechanism (sensory, working, episodic, semantic, and procedural), and memory subject (user-centric personalization and agent-centric experience). We then analyze how memory is operated under single- and multi-agent topologies and highlight learning policies over memory operations, showing how memory management itself is becoming a trainable, self-evolving capability that spans reinforcement-learned context curation, experience consolidation at decision time, and the emerging ecosystem of explicit, portable, and shareable agent skills surfaced through agent harnesses, context engineering, and standardized tool-mediation protocols. Finally, we review evaluation benchmarks and metrics for assessing memory utility, and outline various open challenges and future directions.
Analytic Planning under Uncertainty with Moment Closure
Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictive uncertainty. Propagat… (see more)ing full state distributions analytically offers a principled way to do this, but has traditionally required restrictive policy or reward structures to remain tractable. Consequently, modern deep reinforcement learning has largely retreated to either stochastic sampling, which introduces significant target variance, or deterministic point estimates that ignore predictive covariance entirely. We investigate whether distribution-aware planning is possible without these constraints. Using a quadratic action-value parameterization, we first reduce the Bellman backup to an expectation over the state-value function alone; the key idea is then a compatibility principle between the predictive transition distribution and the value function class, under which this expectation is analytic in the distribution's moments. We instantiate this principle with a Gaussian transition model paired with a radial-basis value function, yielding a closed-form backup that propagates both predictive mean and covariance. Empirically, our approach reduces target variance and yields well-calibrated predictive uncertainty under stochastic observations in continuous control, providing a principled framework for planning with learned distribution models.
Can LLMs Faithfully Enact Superforecaster Personas?
Andrew Robert Williams
Evan Jiang
Nasim Rahaman
Superforecasters are people that make forecasts that are statistically better than the average person; they offer an interesting testbed for… (see more) persona modelling. In this preliminary work, we investigate whether large language models (LLMs) can faithfully model superforecaster-like personas making similar forecasts and generating similar rationales behind the forecasts. When we compare LLM-generated responses to the ground truth responses of superforecasters, we find that few-shot prompting generates numerically closer forecasts than instruction-prompting baselines. We also measure the similarity of generated and real rationales, finding that the rationales generated based on examples are judged more similar to those of superforecasters than instruction-prompting baselines. However, an analysis of the reasoning patterns in the rationales shows significant differences between human and LLM forecasters, pointing to a gap between stylistic imitation and deep reasoning similarity. These preliminary results raise several interesting questions about how getting superforecaster-like behaviour from LLMs actually works, and open new avenues to explore for improving the forecasting behaviour of LLMs.
Reasoning with Preference Constraints: A Benchmark for Language Models in Many-to-One Matching Markets
Recursive Vision Language Models for General Symbolic Reasoning
Omid Nejati Manzari
Hassan Rivaz
Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive … (see more)reasoning, which limits systematic search, refinement, and backtracking. While recursive models such as Hierarchical Reasoning Model (HRM) and Tiny Recursive Model (TRM) address this limitation through iterative latent-state refinement, they are typically task-specific and do not leverage pretrained language priors. We propose R-Qwen, a recursive reasoning framework built upon a pretrained Qwen backbone. R-Qwen repeatedly refines a candidate solution through programmatic self-recursion and deep supervision, combining the structured iterative computation of recursive models with the linguistic and reasoning priors of pretrained LLMs. We further adapt Hierarchical Supervision Weighting (HSW) to autoregressive models by exponentially weighting losses across recursive steps. HSW reduces gradient variance by at least 50\%, improves the signal-to-noise ratio of stochastic gradients, and accelerates convergence. Across eight challenging benchmarks, R-Qwen consistently outperforms prior recursive reasoning models and substantially larger LLMs while using a comparable number of trainable parameters. Notably, on ARC-AGI dataset, our model achieves a 27.6\% improvement over the baseline, highlighting the effectiveness of recursive refinement for general symbolic reasoning. These results suggest that recursive reasoning mechanisms and pretrained language model priors are complementary approaches for improving symbolic puzzle-solving. Code and models will be released after acceptance.
Inferring brain-wide interactions using data-constrained recurrent neural network models
Matthew G. Perich
Charlotte Arlt
Sofia Soares
Siyan Zhou
Manuel Beiran
Aaron S. Andalman
Tyler Benster
Megan E. Young
Clayton P. Mosher
Juri Minxha
Eugene Carter
Ueli Rutishauser
Peter H. Rudebeck
Christopher D. Harvey
Karl Deisseroth
Kanaka Rajan
Nonlinear Laplacians Improve Signed-Directed Graph Learning
Yuichi Yoshida
While signed-directed graphs have been studied using linear Laplacians in the design of graph neural networks, relatively little research ha… (see more)s focused on developing non-linear Laplacian operators for such networks. We introduce a non-linear Laplacian operator specific to signed and directed networks (NLSD). This non-linear operator extends the concepts of the signed Laplacian for signed graphs and the Laplacian for directed graphs. The NLSD calculates node-specific potentials based on features More precisely, if the potential discrepancy is not aligned with the edge direction, we ignore it (and vice versa) leveraging message-passing techniques only across edges where potential discrepancies align with the edge's direction. Utilizing this novel operator, we propose an efficient spectral GNN framework (NLSD-GNN). We conducted comprehensive evaluations focusing on node classification and link prediction, examining scenarios involving signed, directional, or both types of information. Our findings reveal that this spectral GNN framework not only integrates signed and directional data effectively but also achieves superior performance across diverse datasets.
De novo L-(+)-tartaric acid biosynthesis in multi-modular engineered yeasts
Xuan Zhou
Jiaheng Hou
Zikai Wang
Zhendong Li
Yang Li
Xitong Li
Xianhao Xu
Yanfeng Liu
Jianghua Li
Guocheng Du
Dacheng Ma
J. Tang
Jian Chen
Xueqin Lv
Long Liu
L-(+)-tartaric acid (L-TA) is a high-value chiral organic acid essential for food and pharmaceuticals. Despite its industrial importance, su… (see more)stainable green production is constrained by the lack of a fully defined biosynthetic pathway. Here, we report the de novo biosynthesis of L-TA in Saccharomyces cerevisiae through reaction-guided enzyme mining, experimental validation, and Enzyme Commission-specific Catalytic Hybrid Optimizer (ECHO)-assisted enzyme prioritization. We first elucidate the elusive two-step conversion from precursor 5-keto-D-gluconic acid (5-KGA) to L-TA, catalyzed by transketolase (TK) and succinate semialdehyde dehydrogenase (SSDH). To optimize this critical step, we develop the ECHO. This multimodal framework integrates sequence, substrate, and pocket-aware structural information to identify high-performance TK-SSDH pairs. By integrating this pathway with de novo precursor synthesis, cofactor engineering, and semi-rational protein engineering, a final L-TA titer of 6.59 mg L−1 was achieved in a 5-L bioreactor. By connecting computational mining and metabolic assembly through a multi-module engineering strategy, our study establishes a green platform for L-TA production and demonstrates an effective workflow for synthetic pathway design. L-(+)-tartaric acid (L-TA) is a high-value chiral organic acid for food and pharmaceuticals. Here the authors produce L-TA in S. cerevisiae through reaction-guided enzyme mining and Enzyme Commission-specific Catalytic Hybrid Optimizer (ECHO)-assisted enzyme prioritization.
Preparation of MgAl2O4-reinforced magnesium composite refractories from high silicon magnesite tailings: α-Al2O3 /AlN reaction mechanism and thermal shock resistance
Sheng Wu
Qingdong Hou
Xudong Luo
Cairan Wang
Jinfan Xu