Publications

Long-Horizon Model-Based Offline Reinforcement Learning Without Conservatism
Do machine learning methods make better predictions than conventional ones in pharmacoepidemiology? A systematic review, meta-analysis, and network meta-analysis.
Ana Paula Bruno Pena-Gralle
Mireille E. Schnitzer
Sofia-Nada Boureguaa
Félix Morin
Caroline Sirois
Alice Dragomir
Lucie Blais
MANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasks
Lirong Che
Shuo Wen
Shan Huang
Chuang Wang
Yuzhe Yang
Xueqian Wang
Jian Su
Real-world robotic tasks are long-horizon and often span multiple floors, requiring complex spatial reasoning. Existing embodied benchmarks,… (see more) however, are largely confined to single-floor homes, failing to evaluate agents on realistic, building-scale tasks. We introduce MANSION, a language-driven framework for generating building-scale, multi-floor 3D environments for long-horizon tasks. Using this framework, we release MansionWorld, a large-scale dataset featuring over 1,000 diverse, non-residential buildings. These environments support cross-floor skills and long-horizon task generation on reusable building layouts. Experiments show that current methods degrade sharply on our multi-floor tasks, highlighting both the challenge and the value of this setting for advancing embodied AI.
Mapping of Subjective Accounts into Interpreted Clusters (MOSAIC): Topic Modelling and LLM applied to Stroboscopic Phenomenology
Romy Beauté
David J. Schwartzman
Jennifer Crook
Fiona Macpherson
Adam B. Barrett
Anil K. Seth
Abstract Stroboscopic light stimulation (SLS) on closed eyes typically induces simple visual hallucinations, characterized by vivid, geometr… (see more)ic, and colourful patterns. A dataset of 898 sentences, extracted from 407 open subjective reports, was recently compiled as part of the Dreamachine programme (https://dreamachine.world/) (Collective Act, 2022), an immersive multisensory experience that combines SLS and spatial sound in a collective setting. Although open reports extend the range of reportable phenomenology, their analysis presents significant challenges, particularly in systematically identifying patterns. To address this challenge, we implemented a data-driven approach leveraging large language models and topic modelling to uncover and interpret latent experiential topics directly from the Dreamachine’s text-based reports. Our analysis confirmed the presence of simple visual hallucinations typically documented in scientific studies of SLS, while also revealing experiences of altered states of consciousness and complex hallucinations. Building on these findings, our computational approach expands the systematic study of subjective experience by enabling data-driven analyses of open-ended phenomenological reports, capturing experiences not readily identified through standard questionnaires. By revealing rich and multifaceted aspects of experiences, our study broadens our understanding of stroboscopically induced phenomena while highlighting the potential of natural language processing and large language models in the field of computational phenomenology. More generally, this approach provides a practically applicable methodology for uncovering subtle hidden patterns of subjective experience across diverse research domains. Open-source implementation and an interactive web application are provided to facilitate application of this methodology.
MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings
Jean-Philippe Corbeil
Minseon Kim
Maxime Griot
Sheela Agarwal
Francois Beaulieu
Paul Vozila
Jean-Philippe Corbeil, Minseon Kim, Maxime Griot, Sheela Agarwal, Alessandro Sordoni, Francois Beaulieu, Paul Vozila. Proceedings of the 19t… (see more)h Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track). 2026.
MIRA: A Score for Conditional Distribution Accuracy and Model Comparison
We present Mira, a method for estimating the expected probability that samples from a candidate conditional distribution match the true, unk… (see more)nown conditional distribution, for which only data-label pairs are available. We derive theoretical bounds obtained when the candidate distribution matches the true one and when the conditional distributions are independent. This framework thus enables model comparison by quantifying the alignment between the conditional distribution of a candidate model and the data-label pairs of the true model. Consequently, Mira enables Bayesian model comparison through direct posterior validation, bypassing the challenging evidence computation. We demonstrate its effectiveness across several toy problems and Bayesian inference tasks.
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
As multilingual large language models become more widely used, ensuring their safety and fairness across diverse linguistic contexts present… (see more)s unique challenges. While existing research on machine unlearning has primarily focused on monolingual settings, typically English, multilingual environments introduce additional complexities due to cross-lingual knowledge transfer and biases embedded in both pretraining and fine-tuning data. In this work, we study multilingual unlearning using the Aya-Expanse 8B model under two settings: (1) data unlearning and (2) concept unlearning. We extend benchmarks for factual knowledge and stereotypes to ten languages through translation: English, French, Arabic, Japanese, Russian, Farsi, Korean, Hindi, Hebrew, and Indonesian. These languages span five language families and a wide range of resource levels. Our experiments show that unlearning in high-resource languages is generally more stable, with asymmetric transfer effects observed between typologically related languages. Furthermore, our analysis of linguistic distances indicates that syntactic similarity is the strongest predictor of cross-lingual unlearning behavior.
Neurosymbolic Large Neighbourhood Search
Arnaud Delage-Reid
Gilles Pesant
Recent advances in ai have spurred interest in NeSy architectures that integrate neural and symbolic methods. In particular, combining a Con… (see more)straint Programming (cp) model with a language model for constrained sequence generation tasks allows the neural component to capture domain knowledge while cp enforces structural constraints. In this paper we propose combining cp with a Masked Language Model (mlm) to perform Large Neighbourhood Search (lns). Unlike conventional left-to-right Large Language Models, mlm s can complete sequences with gaps in arbitrary positions, making them well-suited for this task. Meanwhile, lns provides a cp-based iterative framework to explore constrained subspaces whenever searching the whole space would be intractable. We evaluate NeSylns on tasks in constrained text generation and molecule discovery. Our experiments show that it can quickly generate many high-quality sentences and molecules, even for highly-constrained tasks.
Nonlinear Observer Design for Visual-Inertial Odometry
Mouaad Boughellaba
Abdelhamid Tayebi
Soulaimane Berkane
This paper addresses the problem of Visual-Inertial Odometry (VIO) for rigid body systems evolving in three-dimensional space. We introduce … (see more)a novel matrix Lie group structure, denoted SE_{3+n}(3), that unifies the pose, gravity, linear velocity, and landmark positions within a consistent geometric framework tailored to the VIO problem. Building upon this formulation, we design an almost globally asymptotically stable nonlinear geometric observer that tightly integrates data from an Inertial Measurement Unit (IMU) and visual sensors. Unlike conventional Extended Kalman Filter (EKF)-based estimators that rely on local linearization and thus ensure only local convergence, the proposed observer achieves almost global stability through the decoupling of the rotational and translational dynamics. A globally exponentially stable Riccati-based translational observer along with an almost global input-to-state stable attitude observer are designed such that the overall cascaded observer enjoys almost global asymptotic stability. This cascaded architecture guarantees robust and consistent estimation of the extended state, including orientation, position, velocity, gravity, and landmark positions, up to the VIO unobservable directions (i.e., a global translation and rotation about gravity). The effectiveness of the proposed scheme is demonstrated through numerical simulations as well as experimental validation on the EuRoC MAV dataset, highlighting its robustness and suitability for real-world VIO applications.
OBELiX: a curated dataset of crystal structures and experimentally measured ionic conductivities for lithium solid-state electrolytes
Rhiannon Hendley
Leah Wairimu Mungai
Sun Sun
Alain Tchagang
Jiang Su
Hongyu Guo
Homin Shin
OBELiX is a database of 599 synthesized solid electrolyte materials and their experimentally measured room temperature ionic conductivities … (see more)gathered from literature and curated by domain experts.
Online HD-tRNS over the Right Temporoparietal Junction Enhances Mentalizing during Social Interactions
Vincent Chamberland
Quentin Moreau
Lisane Moses
Gabriela Milanova
Operationalizing the Superficial Alignment Hypothesis via Task Complexity
The superficial alignment hypothesis (SAH) posits that large language models learn most of their knowledge during pre-training, and that pos… (see more)t-training merely surfaces this knowledge. The SAH, however, lacks a precise definition, which has led to (i) different and seemingly orthogonal arguments supporting it, and (ii) important critiques to it. We propose a new metric called **Task Complexity**: the length of the shortest program that achieves a target performance on a task. In this framework, the SAH claims that pre-trained models drastically reduce the task complexity of achieving high performance on many tasks. Our definition unifies prior arguments supporting the SAH, interpreting them as different strategies to find such short programs. Experimentally, we estimate task complexities of mathematical reasoning, machine translation, and instruction following tasks and show that their respective task complexities can be remarkably low when conditioned on a pre-trained model. Further, we find that pre-training enables access to strong performances on our tasks, but it can require programs of gigabytes of length to access them. Post-training, on the other hand, collapses the complexity of reaching this same performance by several orders of magnitude. Overall, our results highlight that task adaptation can require remarkably little information—often just a few kilobytes.