Publications

Tactile Modality Fusion for Vision-Language-Action Models
We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models. … (see more)While recent advances in VLA models have introduced robot policies that are both generalizable and semantically grounded, these models mainly rely on vision-based perception. Vision alone, however, cannot capture the complex interaction dynamics that occur during contact-rich manipulation, including contact forces, surface friction, compliance, and shear. While recent attempts to integrate tactile signals into VLA models often increase complexity through token concatenation or large-scale pretraining, the heavy computational demands of behavioural models necessitate more lightweight fusion strategies. To address these challenges, TacFiLM outlines a post-training finetuning approach that conditions intermediate visual features on pretrained tactile representations using feature-wise linear modulation (FiLM). Experimental results on insertion tasks demonstrate consistent improvements in success rate, direct insertion performance, completion time, and force stability across both in-distribution and out-of-distribution tasks. Together, these results support our method as an effective approach to integrating tactile signals into VLA models, improving contact-rich manipulation behaviours.
Toward Self-Driven Microscopy Exploration for the Characterization of Functional Materials
Claudia M. Bazán
Ramzi Zidani
Maxime Goulet
Jean-Nicolas Deraspe
Jeanine Looman
Delphine Bouilly
Audrey Laventure
BATIS: Bayesian Approaches for Targeted Improvement of Species Distribution Models
Benjamin Akera
Mélisande Teng
Species distribution models (SDMs), which aim to predict species occurrence based on environmental variables, are widely used to monitor and… (see more) respond to biodiversity change. Recent deep learning advances for SDMs have been shown to perform well on complex and heterogeneous datasets, but their effectiveness remains limited by spatial biases in the data. In this paper, we revisit deep SDMs from a Bayesian perspective and introduce BATIS, a novel and practical framework wherein prior predictions are updated iteratively using limited observational data. Models must appropriately capture both aleatoric and epistemic uncertainty to effectively combine fine-grained local insights with broader ecological patterns. We benchmark an extensive set of uncertainty quantification approaches on a novel dataset including citizen science observations from the eBird platform. Our empirical study shows how Bayesian deep learning approaches can greatly improve the reliability of SDMs in data-scarce locations, which can contribute to ecological understanding and conservation efforts.
Identifying and Analyzing Performance-Critical Tokens in Large Language Models
Heyan Huang
Sanxing Chen
Marc-Antoine Rondeau
Yang Gao
Jackie Chi Kit Cheung
In-context learning (ICL) has emerged as an effective solution for few-shot learning with large language models (LLMs). However, how LLMs le… (see more)verage demonstrations to specify a task and learn a corresponding computational function through ICL is underexplored. Drawing from the way humans learn from content-label mappings in demonstrations, we categorize the tokens in an ICL prompt into content, stopword, and template tokens. Our goal is to identify the types of tokens whose representations directly influence LLM's performance, a property we refer to as being performance-critical. By ablating representations from the attention of the test example, we find that the representations of informative content tokens have less influence on performance compared to template and stopword tokens, which contrasts with the human attention to informative words. We give evidence that the representations of performance-critical tokens aggregate information from the content tokens. Moreover, we demonstrate experimentally that lexical meaning, repetition, and structural cues are the main distinguishing characteristics of these tokens. Our work sheds light on how LLMs learn to perform tasks from demonstrations and deepens our understanding of the roles different types of tokens play in LLMs.
PlantTraitNet: An Uncertainty-Aware Multimodal Framework for Global-Scale Plant Trait Inference from Citizen Science Data
Ayushi Sharma
Johanna Trost
Daniel Lusk
Johannes Dollinger
Julian Schrader
Christian Rossi
Javier Lopatin
Simon Haberstroh
Jana Eichel
Daniel Mederer
Jose Miguel Cerda-Paredes
Shyam S. Phartyal
Lisa-Maricia Schwarz
Anja Linstädter
Maria Conceição Caldeira
Teja Kattenborn
Global plant maps of plant traits, such as leaf nitrogen or plant height, are essential for understanding ecosystem processes, including the… (see more) carbon and energy cycles of the Earth system. However, existing trait maps remain limited by the high cost and sparse geographic coverage of field-based measurements. Citizen science initiatives offer a largely untapped resource to overcome these limitations, with over 50 million geotagged plant photographs worldwide capturing valuable visual information on plant morphology and physiology. In this study, we introduce PlantTraitNet, a multi-modal, multi-task uncertainty-aware deep learning framework that predicts four key plant traits (plant height, leaf area, specific leaf area, and nitrogen content) from citizen science photos using weak supervision. By aggregating individual trait predictions across space, we generate global maps of trait distributions. We validate these maps against independent vegetation survey data (sPlotOpen) and benchmark them against leading global trait products. Our results show that PlantTraitNet consistently outperforms existing trait maps across all evaluated traits, demonstrating that citizen science imagery, when integrated with computer vision and geospatial AI, enables not only scalable but also more accurate global trait mapping. This approach offers a powerful new pathway for ecological research and Earth system modeling.
Profiling Pre-service Teachers’ Computational Thinking
Tanya Chichekian
Annie Savard
Yi-Mei Zhang
Computational thinking (CT) is a vital skill set for pre-service teachers who will need to foster computational literacy in K–12 classroom… (see more)s, yet the factors influencing their CT skills remain less understood than those for K–12 students or in-service teachers. This study leverages multimodal data to investigate how pre-service teachers (n=128) differ in CT skills, the predictive role of metacognitive strategies and prior coding experience, and variations in online behaviours. Using latent profile analysis, we identified three profiles based on digital literacy, problem-solving, and coding comfort (Novice, Developing, and Proficient), revealing heterogeneity in CT, and supporting non-linear skill acquisition. Linear discriminant analysis revealed that metacognitive strategies and prior coding experience significantly predict profile membership, validating the interplay of technical and cognitive factors in the development of CT skills. Behavioural data from an interactive problem-solving task showed that, compared to Novices and Developing learners, Proficient learners were more task efficient and perceived fewer challenges during task completion. Implications for designing a learning analytics dashboard to visualize profiles and behavioural metrics to support adaptive, equitable, and personalized teacher training are discussed, thereby enhancing pre-service teachers’ readiness to integrate CT into K–12 education.
Unbiased characterization of COVID-19 endotypes leads to prognostication of high-risk individuals using routine blood tests
Catherine Allard
Madeleine Durand
Karine Tremblay
Simon Rousseau
Unsupervised proteomic analysis identified biologically coherent endotypes that advance understanding of acute lung injury in COVID‑19 and… (see more) support improved diagnostic and prognostic strategies.
What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
Mengtao Zhou
Qi Sima
We investigate the capacity of Large Language Models (LLMs) for imaginative reasoning—the proactive construction, testing, and revision of… (see more) hypotheses in information-sparse environments. Existing benchmarks, often static or focused on social deduction, fail to capture the dynamic, exploratory nature of this reasoning process. To address this gap, we introduce a comprehensive research framework based on the classic "Turtle Soup" game, integrating a benchmark, an agent, and an evaluation protocol. We present TurtleSoup-Bench, the first large-scale, bilingual, interactive benchmark for imaginative reasoning, comprising 800 turtle soup stories sourced from both the Internet and expert authors. We also propose Mosaic-Agent, a novel agent designed to assess LLMs' performance in this setting. To evaluate reasoning quality, we develop a multi-dimensional protocol measuring logical consistency, detail completion, and conclusion alignment. Experiments with leading LLMs reveal clear capability limits, common failure patterns, and a significant performance gap compared to humans. Our work offers new insights into LLMs' imaginative reasoning and establishes a foundation for future research on exploratory agent behavior.
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
Shiva Krishna Reddy Malay
Jishnu Nair
Sagar Davasam
Aman Tiwari
Sathwik Tejaswi
Sridhar Krishna Nemala
Srinivas Sunkara
Sai Rajeswar
Large language models are shifting from passive information providers to active agents intended for complex workflows. However, their deploy… (see more)ment as reliable AI workers in enterprise is stalled by benchmarks that fail to capture the intricacies of professional environments, specifically, the need for long-horizon planning amidst persistent state changes and strict access protocols. In this work, we introduce EnterpriseOps-Gym, a benchmark designed to evaluate agentic planning in realistic enterprise settings. Specifically, EnterpriseOps-Gym features a containerized sandbox with 164 database tables and 512 functional tools to mimic real-world search friction. Within this environment, agents are evaluated on 1,150 expert-curated tasks across eight mission-critical verticals (including Customer Service, HR, and IT). Our evaluation of 14 frontier models reveals critical limitations in state-of-the-art models: the top-performing Claude Opus 4.5 achieves only 37.4% success. Further analysis shows that providing oracle human plans improves performance by 14-35 percentage points, pinpointing strategic reasoning as the primary bottleneck. Additionally, agents frequently fail to refuse infeasible tasks (best model achieves 53.9%), leading to unintended and potentially harmful side effects. Our findings underscore that current agents are not yet ready for autonomous enterprise deployment. More broadly, EnterpriseOps-Gym provides a concrete testbed to advance the robustness of agentic planning in professional workflows.
Machine learning–based prediction of Metabolic Syndrome risk in the Quebec population
Shayan Nejadshamsi
Stella S. Daskalopoulou
Samira Abbasgholizadeh Rahimi
Objective This study evaluates multiple machine learning approaches to predict metabolic syndrome (MetS) risk in the Quebec, Canada populati… (see more)on. We further perform explainability analysis to interpret model predictions and identify key features driving risk classification. Methods and analysis This study followed the Minimum Information about Clinical Artificial Intelligence Modeling (MI-CLAIM) guideline for reporting. We used cross-sectional data from the Canadian Community Health Survey (2015–2018) for the population living in the province of Quebec, which includes 42,279 participants. Partial sampling was used to obtain a balanced dataset for model development. We evaluated seven machine learning models for the defined classification task, including Logistic Regression, XGBoost, LightGBM, TabNet, NODE, 1D-CNN and Regularisation Cocktails. Performance was assessed using accuracy, precision, recall, F1-score, AUROC, and AUPRC, and interpretability was examined using SHAP to identify key predictors of MetS risk. Results After partial sampling, 7,866 participants (4,856 high-risk and 3,010 low-risk MetS cases) were included in the machine learning analysis. XGBoost and NODE showed the strongest performance. XGBoost achieved the highest accuracy (80.4%) and AUROC (84.1%), while NODE achieved the highest precision (80.1%) and AUPRC (86.0%). Explainability analysis identified age, perceived health, and sex as the most important features contributing to MetS risk predictions. Conclusion This study shows that machine learning can accurately predict MetS risk using self-reported health survey data from the Quebec population. Comparison of classical and deep learning approaches identified the optimal predictive model, and explainability analyses identified the most important features contributing to the risk predictions, which align with established clinical evidence. These results support a machine learning–driven initial screening framework for population-level early identification of high-risk individuals, enabling targeted interventions and efficient allocation of healthcare resources.
MANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasks
Lirong Che
Shuo Wen
Shan Huang
Chuang Wang
Yuzhe Yang
Xueqian Wang
Jian Su
Real-world robotic tasks are long-horizon and often span multiple floors, demanding rich spatial reasoning. However, existing embodied bench… (see more)marks are largely confined to single-floor in-house environments, failing to reflect the complexity of real-world tasks. We introduce MANSION, the first language-driven framework for generating building-scale, multi-floor 3D environments. Being aware of vertical structural constraints, MANSION generates realistic, navigable whole-building structures with diverse, human-friendly scenes, enabling the development and evaluation of cross-floor long-horizon tasks. Building on this framework, we release MansionWorld, a dataset of over 1,000 diverse buildings ranging from hospitals to offices, alongside a Task-Semantic Scene Editing Agent that customizes these environments using open-vocabulary commands to meet specific user needs. Benchmarking reveals that state-of-the-art agents degrade sharply in our settings, establishing MANSION as a critical testbed for the next generation of spatial reasoning and planning.
Theta Dual-Brain Stimulation of rTPJ Shapes Joint Agency
Yuto Kurihara
Ayaka Tsuchiya
Rieko Osu
Summary Joint agency, the shared feeling of “we are doing this together”, has been linked to inter-brain synchrony, but its causal role … (see more)in shaping this experience remains unclear. We applied dual transcranial alternating current stimulation (dual-tACS) over the right temporo-parietal junction (rTPJ) to 13 dyads performing an alternating tapping task (target ITI = 0.5 s; 180 deg. relative phase), manipulating in- and anti-phase coupling at theta (6 Hz), alpha (10 Hz), and beta (20 Hz). As a result, tapping in the theta anti-phase condition was significantly slower than the memorized reference tempo, whereas the other stimulation conditions did not influence the inter-tap interval. Meanwhile, the relative phase remained close to 180 deg. across all conditions. In the theta condition, anti-phase stimulation produced significantly lower joint agency than in-phase stimulation. Furthermore, mediation analysis suggested that the inter-tap interval may partially account for the effect of theta dual-brain stimulation on joint agency, although this indirect pathway did not reach statistical significance. These findings suggest that anti-phase theta stimulation over the rTPJ lowers joint agency, possibly by reducing coordination efficiency while preserving the overall 180 deg. alternation structure.