The upcoming meeting, taking place on November 10 at Mila, will explore how we can collectively develop, govern, and deploy high-performing, reliable, and secure agentic systems by connecting academic researchers, industry experts, and practitioners.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Publications
Supp Fig6+ Legend from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
<p>Supplementary Fig6. Gating strategy for phenotyping of immune cells by flow cytomtery. EM: Effector memory, TEMRA: Effector memory … (see more)Re-expression RA, CM: Central Memory</p>
2026-06-16
American Association for Cancer Research (accepted)
Supp Fig7+ Legend from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
<p>Supplementary Fig7. Characterizing peripheral cellular immune environment in patients receiving PGV001. Multicolor flowcytometry pe… (see more)rformed to phenotype circulating immune cells in patients and healthy donors (HD). PID-017 was excluded from analysis. a) Tabulated view of age and sex of the seven healthy donors whose PBMCs are used in the study. b) Frequency of lymphoid immune cell subsets over the course of treatment. Each dot represents a subject. # p value indicates un-paired two-sided Student’s T-test comparing HD with patient cohort # <0.05, ## <0.01. c) Frequency of myeloid immune cell subsets over the course of treatment. Each dot represents a subject. d) Depicting Fold change from “Pre” in listed immune cell subsets over the course of treatment. e) Pie charts depicting CD4+ and CD8+ T cell states in patient blood and healthy donors. f) �4+ T cells expressing TIGIT and CTLA4 shown as a fold change from baseline in patient blood. *p value indicates paired two-sided Student’s T-test comparing post treatment samples with “pre”. * <0.05. Data in pie charts depicts median. Patient samples, N=12. Healthy donor samples, N=7. Data in graphs shown as mean with bar graphs showing +/- SEM.</p>
2026-06-16
American Association for Cancer Research (accepted)
Supplementary Table1. from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
<p>Supplementary Table1. Staging at time of enrollment, adjuvant treatment following curative intent treatment until the end of vaccin… (see more)ation and vaccination timing</p>
2026-06-16
American Association for Cancer Research (accepted)
Supplementary Table2 from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
Supplementary Table3 from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
Supplementary Table4 from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
Supplementary Table5 from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement le… (see more)arning (RL), this trade-off is typically controlled through a fixed discount factor, which imposes a single exponentially discounted temporal horizon. However, biological agents exhibit flexible and adaptive temporal discounting, suggesting that effective planning requires multiple timescales. Here, we propose a multi-horizon approach that adaptively selects and combines temporal horizons, enabling robust adaptation to changes in reward structure without manual discount-factor tuning. This flexibility makes the method particularly suitable for continual learning scenarios involving task switches and varying environmental configurations. Empirically, we demonstrate that our approach identifies effective discount factors across a range of MiniGrid environments, including continual settings composed of three sequentially changing tasks. These results suggest that adaptive temporal discounting can improve parameter efficiency and enhance adaptability in both artificial and biologically inspired learning systems.
In-context reinforcement learning (ICRL) promises rapid adaptation without parameter updates, but standard supervised objectives often fail … (see more)when pretraining data is generated by suboptimal behaviour policies. In these regimes, logged actions are unreliable labels while rewards still provide value-relevant information. To address this, we introduce SPICE, a Bayesian decision-time inference method that shifts online ICRL from action-logit prediction to approximate posterior inference over action values, requiring neither expert action labels nor algorithmic learning traces. SPICE learns a task-conditioned value prior with a transformer value ensemble and, at test time with parameters frozen, fuses this prior with kernel-weighted context evidence via a closed-form Gaussian fusion update. The resulting estimates drive a posterior-UCB controller, enabling principled online exploration and adaptation without gradient updates. A stochastic-bandit analysis shows logarithmic regret growth for the fixed-prior controller under scheduled exploration, while quantifying the additional early cost caused by inaccurate prior estimates. Across bandits, Darkroom, image-based MiniWorld, and continuous building control, SPICE adapts more effectively from suboptimal data than supervised ICRL baselines.
In continual reinforcement learning, carefully managing the stability-plasticity tradeoff remains a core challenge. Recent work by Abel et a… (see more)l. (2025) formalized this dilemma by defining plasticity as the generalized directed information from an agent’s observations to its actions, and empowerment as the generalized directed information from its actions to its observations. This formulation successfully reframes the traditional stability-plasticity tradeoff as an empowerment-plasticity tradeoff. However, while extensive literature exists on optimizing for empowerment, there is currently no research addressing the optimization of plasticity under this new definition. This paper presents preliminary work toward optimizing plasticity within Markov decision processes. We show that there exists a Bellman optimality equation for optimizing plasticity similar to previous work for empowerment.
Recent advances in mobile agents are dominated by the GUI paradigm, in which agents perceive UI information and emit screen interactions. Ho… (see more)wever, mobile platforms also expose a command-line interface (CLI) that provides direct access to device services and data. We argue CLI deserves first-class consideration alongside GUI. We evaluate three coding agents (Claude Code, Terminus-2, mini-swe-agent) across four model APIs on AndroidWorld and MobileWorld without any mobile-specific post-training, comparing against three reproducible GUI baselines (GUI-Owl-1.5-32B, MAI-UI, Qwen3-VL-32B). Claude Code (Opus 4.7) reaches 71.8\% and 51.9\%, outperforming every reproducible GUI baseline (69.3/68.1/57.8\% on AndroidWorld; 43.2/26.3/13.3\% on MobileWorld), while every other CLI configuration remains competitive. To establish the paradigm's ceiling, we provide oracle CLI solutions that reach 88.8\% on AndroidWorld (103/116 tasks CLI-solvable) and 86.3\% on MobileWorld (101/117 tasks CLI-solvable), indicating substantial room for future improvement. To cover everyday user intents beyond the GUI scope, we introduce the \textbf{CLI-Advantage Task Suite}, comprising 45 templates across five categories: bulk operations, multi-condition filtering, aggregation, cross-app workflows, and hidden device state. Every CLI agent outperforms every GUI baseline in all five categories, with substantially fewer steps per task (10.7 vs.\ 18.6). To support future research on mobile CLI agents, we will open-source agent implementations, oracle solutions, the CLI-Advantage suite, and evaluation infrastructure.
Breaking down large tasks into smaller sub-tasks, either to accelerate learning or enable transfer across related environments, remains a ce… (see more)ntral challenge in reinforcement learning (RL). One way Hierarchical Reinforcement Learning addresses this is by introducing temporal abstractions, often instantiated as options: temporally extended action sequences directed toward sub-goals. While prior work largely focuses on algorithms that explicitly learn such options, we ask a different question: can temporal abstractions emerge naturally within general deep reinforcement learning agents? To this end, we introduce Decorrelate Cluster Temporal Activation (DCTA) Analysis, a tool for detecting temporal abstractions in agents that do not explicitly model them. We validate this approach on a custom Four-Room environment and Atari benchmarks. Most importantly, we show that DQN and PPO naturally develop internal representations with semi-Markov consistent temporal structure — the defining statistical property of temporal abstractions — without explicit option learning objective.