La prochaine rencontre, qui aura lieu le 10 novembre à Mila, permettra d'explorer comment pouvons-nous collectivement développer, encadrer et déployer des systèmes agentiques performants, fiables et sécuritaires en connectant chercheur·euse·s académiques, expert·e·s industriel·le·s et praticien·ne·s.
Ce programme à temps partiel offre aux scientifiques l'opportunité de tester leur intérêt pour l'entrepreneuriat. Vous avez jusqu'au 5 octobre pour postuler.
Avantage IA : productivité dans la fonction publique
Apprenez à tirer parti de l’IA générative pour soutenir et améliorer votre productivité au travail. La prochaine cohorte se déroulera en ligne les 6 et 8 octobre 2026, en anglais.
Nous utilisons des témoins pour analyser le trafic et l’utilisation de notre site web, afin de personnaliser votre expérience. Vous pouvez désactiver ces technologies à tout moment, mais cela peut restreindre certaines fonctionnalités du site. Consultez notre Politique de protection de la vie privée pour en savoir plus.
Paramètre des cookies
Vous pouvez activer et désactiver les types de cookies que vous souhaitez accepter. Cependant certains choix que vous ferez pourraient affecter les services proposés sur nos sites (ex : suggestions, annonces personnalisées, etc.).
Cookies essentiels
Ces cookies sont nécessaires au fonctionnement du site et ne peuvent être désactivés. (Toujours actif)
Cookies analyse
Acceptez-vous l'utilisation de cookies pour mesurer l'audience de nos sites ?
Lecteur Multimédia
Acceptez-vous l'utilisation de cookies pour afficher et vous permettre de regarder les contenus vidéo hébergés par nos partenaires (YouTube, etc.) ?
Publications
Supp Fig6+ Legend from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
Supp Fig7+ Legend from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
<p>Supplementary Fig7. Characterizing peripheral cellular immune environment in patients receiving PGV001. Multicolor flowcytometry pe… (voir plus)rformed to phenotype circulating immune cells in patients and healthy donors (HD). PID-017 was excluded from analysis. a) Tabulated view of age and sex of the seven healthy donors whose PBMCs are used in the study. b) Frequency of lymphoid immune cell subsets over the course of treatment. Each dot represents a subject. # p value indicates un-paired two-sided Student’s T-test comparing HD with patient cohort # <0.05, ## <0.01. c) Frequency of myeloid immune cell subsets over the course of treatment. Each dot represents a subject. d) Depicting Fold change from “Pre” in listed immune cell subsets over the course of treatment. e) Pie charts depicting CD4+ and CD8+ T cell states in patient blood and healthy donors. f) �4+ T cells expressing TIGIT and CTLA4 shown as a fold change from baseline in patient blood. *p value indicates paired two-sided Student’s T-test comparing post treatment samples with “pre”. * <0.05. Data in pie charts depicts median. Patient samples, N=12. Healthy donor samples, N=7. Data in graphs shown as mean with bar graphs showing +/- SEM.</p>
2026-06-16
American Association for Cancer Research (accepté)
Supplementary Table1. from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
<p>Supplementary Table1. Staging at time of enrollment, adjuvant treatment following curative intent treatment until the end of vaccin… (voir plus)ation and vaccination timing</p>
2026-06-16
American Association for Cancer Research (accepté)
Supplementary Table2 from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
Supplementary Table3 from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
Supplementary Table4 from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
Supplementary Table5 from PGV001, a Multi-Peptide Personalized Neoantigen Vaccine Platform: Phase I Study in Patients with Solid and Hematologic Malignancies in the Adjuvant Setting
Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement le… (voir plus)arning (RL), this trade-off is typically controlled through a fixed discount factor, which imposes a single exponentially discounted temporal horizon. However, biological agents exhibit flexible and adaptive temporal discounting, suggesting that effective planning requires multiple timescales. Here, we propose a multi-horizon approach that adaptively selects and combines temporal horizons, enabling robust adaptation to changes in reward structure without manual discount-factor tuning. This flexibility makes the method particularly suitable for continual learning scenarios involving task switches and varying environmental configurations. Empirically, we demonstrate that our approach identifies effective discount factors across a range of MiniGrid environments, including continual settings composed of three sequentially changing tasks. These results suggest that adaptive temporal discounting can improve parameter efficiency and enhance adaptability in both artificial and biologically inspired learning systems.
In-context reinforcement learning (ICRL) promises rapid adaptation without parameter updates, but standard supervised objectives often fail … (voir plus)when pretraining data is generated by suboptimal behaviour policies. In these regimes, logged actions are unreliable labels while rewards still provide value-relevant information. To address this, we introduce SPICE, a Bayesian decision-time inference method that shifts online ICRL from action-logit prediction to approximate posterior inference over action values, requiring neither expert action labels nor algorithmic learning traces. SPICE learns a task-conditioned value prior with a transformer value ensemble and, at test time with parameters frozen, fuses this prior with kernel-weighted context evidence via a closed-form Gaussian fusion update. The resulting estimates drive a posterior-UCB controller, enabling principled online exploration and adaptation without gradient updates. A stochastic-bandit analysis shows logarithmic regret growth for the fixed-prior controller under scheduled exploration, while quantifying the additional early cost caused by inaccurate prior estimates. Across bandits, Darkroom, image-based MiniWorld, and continuous building control, SPICE adapts more effectively from suboptimal data than supervised ICRL baselines.
In continual reinforcement learning, carefully managing the stability-plasticity tradeoff remains a core challenge. Recent work by Abel et a… (voir plus)l. (2025) formalized this dilemma by defining plasticity as the generalized directed information from an agent’s observations to its actions, and empowerment as the generalized directed information from its actions to its observations. This formulation successfully reframes the traditional stability-plasticity tradeoff as an empowerment-plasticity tradeoff. However, while extensive literature exists on optimizing for empowerment, there is currently no research addressing the optimization of plasticity under this new definition. This paper presents preliminary work toward optimizing plasticity within Markov decision processes. We show that there exists a Bellman optimality equation for optimizing plasticity similar to previous work for empowerment.
Recent advances in mobile agents are dominated by the GUI paradigm, in which agents perceive UI information and emit screen interactions. Ho… (voir plus)wever, mobile platforms also expose a command-line interface (CLI) that provides direct access to device services and data. We argue CLI deserves first-class consideration alongside GUI. We evaluate three coding agents (Claude Code, Terminus-2, mini-swe-agent) across four model APIs on AndroidWorld and MobileWorld without any mobile-specific post-training, comparing against three reproducible GUI baselines (GUI-Owl-1.5-32B, MAI-UI, Qwen3-VL-32B). Claude Code (Opus 4.7) reaches 71.8\% and 51.9\%, outperforming every reproducible GUI baseline (69.3/68.1/57.8\% on AndroidWorld; 43.2/26.3/13.3\% on MobileWorld), while every other CLI configuration remains competitive. To establish the paradigm's ceiling, we provide oracle CLI solutions that reach 88.8\% on AndroidWorld (103/116 tasks CLI-solvable) and 86.3\% on MobileWorld (101/117 tasks CLI-solvable), indicating substantial room for future improvement. To cover everyday user intents beyond the GUI scope, we introduce the \textbf{CLI-Advantage Task Suite}, comprising 45 templates across five categories: bulk operations, multi-condition filtering, aggregation, cross-app workflows, and hidden device state. Every CLI agent outperforms every GUI baseline in all five categories, with substantially fewer steps per task (10.7 vs.\ 18.6). To support future research on mobile CLI agents, we will open-source agent implementations, oracle solutions, the CLI-Advantage suite, and evaluation infrastructure.
Breaking down large tasks into smaller sub-tasks, either to accelerate learning or enable transfer across related environments, remains a ce… (voir plus)ntral challenge in reinforcement learning (RL). One way Hierarchical Reinforcement Learning addresses this is by introducing temporal abstractions, often instantiated as options: temporally extended action sequences directed toward sub-goals. While prior work largely focuses on algorithms that explicitly learn such options, we ask a different question: can temporal abstractions emerge naturally within general deep reinforcement learning agents? To this end, we introduce Decorrelate Cluster Temporal Activation (DCTA) Analysis, a tool for detecting temporal abstractions in agents that do not explicitly model them. We validate this approach on a custom Four-Room environment and Atari benchmarks. Most importantly, we show that DQN and PPO naturally develop internal representations with semi-Markov consistent temporal structure — the defining statistical property of temporal abstractions — without explicit option learning objective.