Publications

Why Open Source? A Game-Theoretic Analysis of the AI Race
In recent years, with the advancement of frontier AI, we have observed certain dynamics in open-sourcing and closed-sourcing decisions. We p… (voir plus)ropose a game-theoretic model to analyze these dynamics in the current landscape of the AI race. Our model builds on an R&D race framework under a winner-takes-all setting, and it accounts for the cases where the players' actions can be either discrete or continuous (i.e., partial open-sourcing, such as open weights). We show that determining the existence of a discrete pure non-trivial Nash equilibrium is NP-hard in general but that we can transform the discrete Nash existence computation into a MIP (Mixed-Integer Programming) problem, making it tractable for small instances using a standard MIP solver. Next, we show the existence and tractability of pure Nash equilibria in the continuous version of our problem, leveraging standard convex analysis results, and constructing an equivalent MIP formulation. Throughout this work, we leverage both our main technical results as well as surrounding technical analysis, to derive socially relevant insights that we believe can serve both to understand already existing decisions and dynamics and to potentially inform new policies.
The Golden Rule of Big Memory: Persistence Is Not Harmful
Yu Hua
Xue Liu
Ion Stoica
Seeking a transformative memory scheme that grows in performance and capacity as the infrastructure expands.
Neurovascular Coupling as Early, High-Sensitive Biomarker for Cognitive Decline and Vascular Pathology: Protocol for Systematic Review and Meta-Analysis
V. D. Abramova
Veronika Egovtseva
Ksenya Pronyaeva
Shamsa H. Alshamsi
Marta Estrada
Rustam Talybov
Taleb~M. Almansoori
Bassem Sadek
Mohammed Khogali
Mohammad~I.K. Hamad
Milos Ljubisavljevic
Yauhen Statsenko
RobotPan: A 360$^\circ$ Surround-View Robotic Vision System for Embodied Perception
Jiahao Ma
Qiang Zhang
Peiran Liu
Zeran Su
Pihai Sun
Gang Han
Wen Zhao
Wei Cui
Zhang Zhang
Zhiyuan Xu
Renjing Xu
Miaomiao Liu
Yijie Guo
Surround-view perception is increasingly important for robotic navigation and loco-manipulation, especially in human-in-the-loop settings su… (voir plus)ch as teleoperation, data collection, and emergency takeover. However, current robotic visual interfaces are often limited to narrow forward-facing views, or, when multiple on-board cameras are available, require cumbersome manual switching that interrupts the operator's workflow. Both configurations suffer from motion-induced jitter that causes simulator sickness in head-mounted displays. We introduce a surround-view robotic vision system that combines six cameras with LiDAR to provide full 360
AI Methods for Implementation Science (AIM-IS): developing a framework, toolkit, and reporting standard for the responsible use of AI in implementation practice and research
Guillaume Fontaine
Susan Michie
Rinad S. Beidas
Elvin Geng
Christine Fahim
Byron J. Powell
Vivian Welch
James Thomas
J. Chan
France Légaré
Janna Hastings
Sylvie D. Lambert
Justin Presseau
Sharon E. Straus
Ruopeng An
Ashrita Saran
Natalie Taylor
Open Science Framework, March 15, 2026: https://doi.org/10.17605/OSF.IO/BX35K.
Translating Brain Encoding Models to Clinical Cohorts: Challenges of Domain Adaptation
Marie St‐Laurent
Julie Boyle
Basile Pinsard
Elizabeth DuPré
We’ve optimized fMRI biomarkers to generalize across participants, not cognitive states. Brain encoding models might finally let us model … (voir plus)both and change what “good data” means. These are the slides of a presentation by Dr Lune Bellec at the AI4health workshop, ÉTS, Montréal, Feb 2026.
Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
Manuela González‐González
Soufiane Belharbi
Muhammad Zeeshan
Masoumeh Sharafi
Muhammad Haseeb Aslam
Lorenzo Sia
Nicolas Richet
Alessandro L. Koerich
Simon L Bacon
Eric Granger
Using behavioural science, health interventions focus on behaviour change by providing a framework to help patients acquire and maintain hea… (voir plus)lthy habits that improve medical outcomes. In-person interventions are costly and difficult to scale, especially in resource-limited regions. Digital health interventions offer a cost-effective approach, potentially supporting independent living and self-management. Automating such interventions, especially through machine learning, has gained considerable attention recently. Ambivalence and hesitancy (A/H) play a primary role for individuals to delay, avoid, or abandon health interventions. A/H are subtle and conflicting emotions that place a person in a state between positive and negative evaluations of a behaviour, or between acceptance and refusal to engage in it. They manifest as affective inconsistency across modalities or within a modality, such as language, facial, vocal expressions, and body language. While experts can be trained to recognize A/H, integrating them into digital health interventions is costly and less effective. Automatic A/H recognition is therefore critical for the personalization and cost-effectiveness of digital health interventions. Here, we explore the application of deep learning models for A/H recognition in videos, a multi-modal task by nature. In particular, this paper covers three learning setups: supervised learning, unsupervised domain adaptation for personalization, and zero-shot inference via large language models (LLMs). Our experiments are conducted on the unique and recently published BAH video dataset for A/H recognition. Our results show limited performance, suggesting that more adapted multi-modal models are required for accurate A/H recognition. Better methods for modeling spatio-temporal and multimodal fusion are necessary to leverage conflicts within/across modalities.
Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
Samuel Cahyawijaya
Peerat Limkonchotiwat
Tack Hwa Wong
Hitesh Laxmichand Patel
Amit Agarwal
Manuel Antonio Rufino
Carlos Rafael Catalan
Muhammad Reza Qorib
Vicky Feliren
Holy Lovenia
Aye Hninn Khine
Frederikus Hudi
David Anugraha
Alham Fikri Aji
Romrawin Chumpu
Viet-Thanh Pham
Minghan Wang
Mohamed Fazli Imam
Ruochen Zhang
Joseph Marvin Imperial … (voir 28 de plus)
Khumaisa Nur'aini
Do Xuan Long
Musa Izzanardi Wijanarko
Joel Ruben Antony Moniz
Patrick Amadeus Irawan
Hanif Muhammad Zhafran
Isaiah Flores
Salsabila Zahirah Pranida
Jun Kevin
Jostin Jerico Rosal
Patricia Nicole Monderin
Kun Kerdthaisong
Ahmad Mustafid
My Chiffon Nguyen
Natchapon Jongwiriyanurak
Siva Worajitwannakul
Haochen Li
Adrian Xuan Wei Lim
Bin Wang
Muhammad Ravi Shulthan Habibi
Lynnette Hui Xian Ng
Mithil Bangera
Yeshil Bangera
Priyaranjan Pattnayak
Dun Li Chan
Sherissa Caren Djuniwar
Cho Chan Myei Oo
Hee Ming Shan
While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple langua… (voir plus)ges and domains, there is still no dedicated framework for assessing human-centric alignment in vision-language systems. We offer two contributions to address this gap. First, we introduce Anthropogenic Regional Adaptation: a novel paradigm that aims to optimize model relevance to specific regional contexts while ensuring the retention of global generalization capabilities. Second, we present a simple, but effective adaptation method named Geographical-generalization-made-easy (GG-EZ), which utilizes regional data filtering and model merging. Through comprehensive experiments on 3 VL architectures: large vision-language models, text-to-image diffusion models, and vision-language embedding models, and a case study in Southeast Asia (SEA) regional adaptation, we demonstrate the importance of Anthropogenic Regional Adaptation and the effectiveness of GG-EZ, showing 5-15% gains in cultural relevance metrics across SEA while maintaining over 98% of global performance and even occasionally surpassing it. Our findings establish Anthropogenic Regional Alignment as a foundational paradigm towards applicability of multimodal vision-language models in diverse regions and demonstrate a simple-yet-effective baseline method that optimizes regional value alignment while preserving global generalization.
DASB - Discrete Audio and Speech Benchmark
Jarod Duret
Darius Petermann
Anastasia Kuznetsova
Discrete audio tokens have recently gained considerable attention for their potential to bridge audio and language processing, enabling mult… (voir plus)imodal language models that can both generate and understand audio. However, preserving key information such as phonetic content, speaker identity, and paralinguistic cues remains a major challenge. Identifying the optimal tokenizer and configuration is further complicated by inconsistent evaluation settings across existing studies. To address this, we introduce the Discrete Audio and Speech Benchmark (DASB), a comprehensive framework for benchmarking discrete audio tokens across speech, general audio, and music domains on a range of discriminative and generative tasks. Our results show that discrete representations are less robust than continuous ones and require careful tuning of factors such as model architecture, data size, learning rate, and capacity. Semantic tokens generally outperform acoustic tokens, but a gap remains between discrete tokens and continuous features, highlighting the need for further research. DASB codes, evaluation setup, and leaderboards are publicly available at https://poonehmousavi.github.io/DASB-website/.
Early detection of common reed ( <i>Phragmites australis</i> ) using unoccupied aerial vehicles and deep learning
EXPRESS: Climate Communications in IPOs: Unpacking the Influence of Climate Disclosure Volume, Sender, and Message Characteristics
Alok R. Saboo
Ritesh Adhyapak
Climate disclosures have emerged as a prominent communication tool for firms facing growing pressure to address climate challenges, yet thei… (voir plus)r impact on firm performance remains unclear. This study proposes a nonlinear (U-shaped) relationship between climate disclosure volume and IPO firm performance, grounded in a damage-limitation logic. At low to moderate levels, disclosures amplify risk salience and proprietary costs, damaging valuations. At higher levels, offsetting benefits related to information, stewardship, and climate-friendly reputation outweigh these costs. Using multi-sourced data from 1,586 IPO firms, a BERT-based large language model to identify climate-related text in prospectuses, and econometric methods that address endogeneity, the authors find support for the proposed U-shaped relationship. The research further demonstrates that sender characteristics (underwriter reputation, customer concentration, and market orientation) and message characteristics (discretionary disclosure and message clarity) moderate the nonlinear relationship. Post-hoc analyses decomposing disclosure content reveal that climate risk disclosures damage valuations. In contrast, climate risk-management disclosures (governance, strategy, and metrics/targets) generate positive effects, suggesting that disclosure effectiveness depends on both volume and content composition. These effects persist in the long-term performance of firms. The findings provide actionable insights for firms developing disclosure strategies and policymakers encouraging climate-related communication.
A Mechanistic Analysis of Looped Reasoning Language Models
Hugh Blayney
Álvaro Arroyo
Johan Obando-Ceron
Michael M. Bronstein
Xiaowen Dong
Reasoning has become a central capability in large language models. Recent research has shown that reasoning performance can be improved by … (voir plus)looping an LLM's layers in the latent dimension, resulting in looped reasoning language models. Despite promising results, few works have investigated how their internal dynamics differ from those of standard feedforward models. In this paper, we conduct a mechanistic analysis of the latent states in looped language models, focusing in particular on how the stages of inference observed in feedforward models compare to those observed in looped ones. To this end, we analyze cyclic recurrence and show that for many of the studied models each layer in the cycle converges to a distinct fixed point; consequently, the recurrent block follows a consistent cyclic trajectory in the latent space. We provide evidence that as these fixed points are reached, attention-head behavior stabilizes, leading to constant behavior across recurrences. Empirically, we discover that recurrent blocks learn stages of inference that closely mirror those of feedforward models, repeating these stages in depth with each iteration. We study how recurrent block size, input injection, and normalization influence the emergence and stability of these cyclic fixed points. We believe these findings help translate mechanistic insights into practical guidance for architectural design.