Publications

From Visual Widgets to UI Code: Efficient Tool-Grounded Generation
Houston H. Zhang
Linfeng Ye
Yuanhao Yu
Xinxin Zuo
Zhixiang Chi
Existing screenshot-to-code systems face a trade-off between flexibility and controllability. Direct multimodal generation can hallucinate v… (see more)isible details, whereas structured pipelines reduce such errors through component-wise decomposition, predefined templates, and customized intermediate representations. These structures, however, introduce additional generative orchestration and restrict outputs to designs covered by the representation. We investigate whether selective tool grounding can improve the fidelity--efficiency trade-off of direct widget-to-code generation. We introduce \textbf{WidgetGen}, a lightweight tool-grounded framework that extracts observable text and color evidence, performs high-level layout and optional chart reasoning, and directly generates executable JavaScript XML (\emph{JSX}). This design reduces reliance on component-wise generation while avoiding a fixed UI schema. Across six multimodal models and \(1{,}000\) held-out widgets, WidgetGen outperforms direct prompting and the structured Widget2Code pipeline on most visual reconstruction metrics, with consistent gains in area, legibility, and style. Finally, reconstruction-derived image-code pairs improve six Qwen-family open-weight models across every reported metric through supervised fine-tuning. These results establish WidgetGen as a strong lightweight baseline and show that selective evidence grounding offers an effective alternative to extensive representation constraints.
From Complex Behavior to Intelligent Human-Centered AI: The 11th ABAW Workshop & Competition
Dimitrios Kollias
Stefanos Zafeiriou
Irene Kotsia
Eric Granger
Alessandro Lameiras Koerich
Simon L Bacon
Oya Celiktutan
Soufiane Belharbi
Muhammad Osama Zeeshan
Muhammad Haseeb Aslam
Chunchang Shao
Guanyu Hu
The 11th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop and Competition, held in conjunction with ECCV 2026, continues to advance… (see more) research on the modelling, analysis and understanding of human affect and behavior in unconstrained real-world environments. The workshop maintains its dual structure, comprising both a competition and a paper track. The ABAW Competition features two challenging benchmarks targeting key problems in affective and behavioral understanding: Multi-Task Learning (MTL), requiring the joint estimation of valence-arousal, facial expressions and action units from a static version of the Aff-Wild2 database, and Ambivalence/Hesitancy (AH) Video Recognition, focusing on the recognition of complex behavioral states from multimodal videos using the BAH dataset. In parallel, the paper track presents recent advances spanning multimodal AI and deep learning, affective computing, behavior understanding, datasets and benchmarks, human-centered AI, human-robot interaction and social robotics, robustness, explainability and real-world applications. Overall, the 11th ABAW Workshop and Competition continues to serve as a leading forum for benchmarking, collaboration and innovation, fostering the development of the next generation of multimodal, trustworthy and human-centered AI systems.
Complexity-driven feature selection for enhancing tuberculosis detection
Sana Ben Mahjouba
Johannes C. Ayena
Youssef Ouakrim
Simon Grandjean Lapierre
Mihaja Raberahona
Neila Mezghani
Most existing machine learning approaches for tuberculosis (TB) screening typically utilize large, high-dimensional acoustic feature sets wi… (see more)thout examining their intrinsic discriminative power. To address this gap, we introduce a complexity-based feature selection approach that evaluates temporal and spectral descriptors using Fisher score (F1), class-distribution overlap (F2), and Shannon entropy (F4). Applied to the CODA-TB dataset (9772 audio recordings from 1105 participants), the proposed method identified 7 highly informative features from the original 26 features, primarily consisting of mel-frequency cepstral coefficients (MFCC) derivatives and spectral-shape measures. The resulting model achieved performance comparable to full-feature baselines while reducing feature dimensionality by 73% and computational cost by up to 14×. Comparative evaluation against four established feature selection techniques, supported by ablation and statistical analyses, confirmed the efficiency and robustness of the complexity-driven strategy, with no statistically significant loss in performance. These findings highlight the potential of lightweight, interpretable, and computationally efficient models for TB cough-based screening in resource-constrained environments.
Robust inference and correlates from genetic associations with personality
Ted Schwaba
Margaret L. Clapp Sullivan
Wonuola A. Akingbuwa
Kerli Ilves
Peter T. Tanksley
Camille M. Williams
Yavor Dragostinov
Travis T. Mallard
Justin D. Tubbs
Wangjingyi Liao
Lindsay S. Ackerman
Josephine C. M. Fealy
Gibran Hemani
Javier de la Fuente
George Davey Smith
Priya Gupta
Murray B. Stein
Joel Gelernter
Daniel F. Levey
Urmo Võsa … (see 121 more)
Liisi Ausmees
Anu Realo
Tõnu Esko
Mariliis Vaht
Jüri Allik
Tõnu Esko
René Mõttus
Uku Vainik
Gudrun A. Jonsdottir
Gudmar Thorleifsson
Árni Freyr Gunnarsson
Gyda Bjornsdottir
Thorgeir E. Thorgeirsson
Hreinn Stefansson
Kari Stefansson
Rosa Cheesman
Qi Qin
Elizabeth C. Corfield
Helga Ask
Fartein Ask Torvik
Eivind Ystrom
Martin Tesli
Dorret I. Boomsma
Eco J. C. de Geus
Jouke-Jan Hottenga
Dener Cardoso Melo
Harold Snieder
Catharina A. Hartman
Charley Xia
Archie Campbell
Michelle Luciano
Ian J. Deary
W. David Hill
Seon-Kyeong Jang
Scott I. Vrieze
Gonçalo Abecasis
Michelle K. Lupton
Brittany L. Mitchell
Petra V. Viher
Lucía Colodro-Conde
Nicholas G. Martin
Sarah E. Medland
Eske M. Derks
Briar Wormington
Jaakko Kaprio
Karri Silventoinen
Teemu Palviainen
Agnieszka Musial
Kaili Rimfeld
Robert Plomin
Margherita Malanchini
Danielle M. Dick
Fazil Aliev
COGA Collaborators
The Spit for Science Working Group
Laura W. Wesseldijk
Fredrik Ullén
Miriam A. Mosing
Henry R. Kranzler
Yaira Nunez
Sarah Beck
Renato Polimanti
Tobias Edwards
Alexandros Giannelis
Emily A. Willoughby
James J. Lee
Matt McGue
Antonio Terracciano
Michele Marongiu
Edoardo Fiorillo
Francesco Cucca
Angelina R. Sutin
Peter J. van der Most
Albertine J. Oldehinkel
Tina Kretschmer
Andrey A. Shabalin
Anna R. Docherty
Robert F. Krueger
Colin D. Freilich
Binisha H. Mishra
Terho Lehtimäki
Olli T. Raitakari
Mika Kähönen
Aino Saarinen
Henrik Dobewall
Liisa Keltikangas-Järvinen
Klaus Berger
Marisol Herrera-Rivero
Fabian Streit
Swapnil Awasthi
Stephanie H. Witt
Johanna Tuhkanen
Katri Räikkönen
Johan G. Eriksson
Jari Lahti
Gail Davies
Paul Redmond
Adele Taylor
Janie Corley
Tom C. Russ
Marina Ciullo
Teresa Nutile
Yong Qian
Toshiko Tanaka
Luigi Ferrucci
Lea Zillich
Lea Sirignano
K. Paige Harden
Erhan Genç
Patrick D. Gajewski
Stephan Getzmann
Christoph Fraenz
Javier E. Schneider Peñate
Stefanie Lis
Alisha S. M. Hall
Christian Schmahl
Sabine C. Herpertz
Abdel Abdellaoui
Michel G. Nivard
Elliot M. Tucker-Drob
Personality traits describe stable differences in how people think, feel and behave, and how they interact with and experience their social … (see more)and physical environments1,2. Many questions remain unanswered about associations between DNA and personality traits, such as their robustness, their generalizability and the biological and social pathways through which they act. Here we meta-analyse data across 46 cohorts comprising 611,037 to 1.14 million participants with European-like and African-like genomes for genome-wide association studies (GWAS) of the Big Five personality traits (extraversion, agreeableness, conscientiousness, neuroticism and openness to experience), and data from up to 50,725 participants for within-family GWAS. We identify 1,260 lead genetic variants associated with personality, including 824 novel variants3. Common genetic variants explain a moderate 4.8-9.3% of the variance in measures of each trait, and 9.3-13.3% among instruments with typical measurement reliability. Genetic associations with personality are highly consistent but not identical across geography, reporter (self versus close other), age group and measurement instrument, and we find minimal spousal assortment for personality in recent history. In contrast to many other social and behavioural traits4,5, within-family GWAS and polygenic index analyses indicate that genetic associations with personality are minimally confounded by the shared family environment. Polygenic prediction, genetic correlation and Mendelian randomization analyses indicate that personality traits have widespread, potentially causal associations with consequential behaviours and life outcomes. Overall, we find that the genetic architecture of personality is robustly generalizable, minimally confounded and widely relevant to human experience.
Do Tabular Foundation Models Know Physics? Contamination, Units, and the Deterministic Limit
Laurence Perreault Levasseur
Tabular foundation models (TFMs) learn to fill in tables the way language models fill in text, and tables are arguably the format in which m… (see more)ost physical measurement arrives. Did they learn any physics in the process? They are Bayesian by construction, so the question is what their prior contains. We probe it directly, evaluating four of them (TabPFN-3, TabICLv2, TabDPT and Real-TabPFN-2.5) against six baselines on datasets sampled from 316 physical equations, in and out of domain. TFMs dominate, out of the box and after tuning. But we show that their prior can represent neither a noiseless mechanism nor physical units, which is why they interpolate physics without yet being able to act as physical models.
Beyond Hard Writes and Rigid Preservation: Soft Recursive Least-Squares for Lifelong LLM Editing
Xinyu Wang
Yu Gu
Peng Lu
Yufei Cui
Xiao-Wen Chang
Model editing updates a pre-trained LLM with new facts or rules without retraining while preserving unrelated behavior. In real deployment, … (see more)edits arrive as long streams, creating a plasticity-stability dilemma: repeated locate-then-edit "hard writes" can accumulate interference over time, while rigid preservation constraints may protect only explicitly constrained directions, allowing past edits or unconstrained behaviors to deviate. We propose RLSEdit, a recursive least-squares editor for long sequential editing. RLSEdit formulates editing as an online quadratic optimization with soft constraints, minimizing a cumulative key-value fitting objective together with two regularizers that control deviation from the pre-trained weights and from a designated anchor mapping. This objective admits an efficient Woodbury-based online recursion, with per-edit cost independent of history length and scaling only with the current edit size. We further provide deviation bounds and an asymptotic characterization of the adherence-preservation trade-off in the many-edits regime. Experiments on CounterFact and ZsRE across multiple model families show stable scaling to 10K edits, outperforming strong baselines in both edit success and holistic stability, while retaining early edits and preserving general capabilities on GLUE and held-out reasoning/code benchmarks.
CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA
Hanseok Oh
Hyunji Lee
Paul Liang
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants … (see more)by requiring external information beyond a provided image to answer questions. KI-VQA involves multiple sub-problems -referring expression understanding, visual grounding, object recognition, knowledge retrieval, and reasoning-yet existing benchmarks typically report only end-task accuracy, obscuring where failures arise. To analyze the full KI-VQA pipeline, we introduce CRAG-MM-Diagnostics, a diagnostic benchmark with stage-wise data annotations that isolate 1) language-based visual grounding, 2) object identification, and 3) knowledge retrieval and reasoning. We evaluate fully parametric and retrieval-augmented VLMs, providing fine-grained analyses using newly collected metadata, such as target ROIs, entity names, and visual complexity scores. Our results point to knowledge retrieval and reasoning as the primary bottleneck, but also highlight issues in the other parts of the KI-VQA pipeline, such as the fact that VLMs struggle with target object identification or that image retrievers struggle to integrate textual cues. These findings expose fundamental limitations in current KI-VQA systems and motivate stage-aware evaluation. We, lastly, leverage these findings to propose a grounded bimodal RAG pipeline that integrates a visual grounding module to crop targets before image retrieval, boosting GPT-5 and Qwen's respective accuracies by 13.3 and 8.5 percentage points.
Creating reference data for early hospital outbreak detection algorithms from experts’ ratings
Brice Leclère
Didier Lepelletier
David L. Buckeridge
BACKGROUND Early outbreak detection algorithms can be useful tools to prevent healthcare-acquired infections in hospitals. However, their de… (see more)velopment and evaluation are hindered by the lack of available labelled data. AIM The aim of this study was to build a reference dataset for outbreak detection in hospitals using two different consensus approaches used to build this dataset. METHODS 25 Canadian and French experts were asked to review one-year time series of weekly incidence from different types of microorganisms, based on the data of a French university hospital. For each time series, experts also add access to additional surveillance data (locations and investigations). Each time series was submitted to three experts, whose role was to identify potential outbreak periods and rank their probability on a web platform. These rankings were summarized using two approaches: a majority vote wherein the most prevalent ranking was used for each week, and a hidden Markov model (HMM) in which the answers of the experts were used as observable variables that related to latent epidemiological states. FINDINGS three experts reviewed a total of 36 times series i.e., 1899 surveillance weeks for 14 different types of microorganisms. Overall, the concordance between the two approaches was the highest for identifying high-ranking weeks. All rankings considered, 89 potential events were identified by majority vote and 96 by the HMM. CONCLUSION We constructed a reliable reference standard data set for the development, evaluation and comparison of algorithms for nosocomial outbreak detection within hospitals.
Resting-state neural oscillations predict individual differences in verbal learning and encoding strategy use
Victor Oswald
Mathieu Landry
Sarah Lippé
Philippe Robaey
Individuals adopt different encoding strategies to facilitate learning, yet few studies have examined the neurophysiological basis of these … (see more)strategies across individuals. The present work addresses this gap by extending our previous findings on the direct relationship between cortical spectral power, measured via resting-state magnetoencephalography, and standard cognitive performance, to test whether resting-state neural features predict individual differences in encoding strategy preferences. Our results highlight the complex interactions between endogenous brain oscillations, learning, and verbal encoding strategies assessed by the California Verbal Learning Test-Second Edition (CVLT-2). First, resting-state theta oscillations were significantly associated with verbal learning and subjective clustering strategies. Second, semantic clustering was facilitated by oscillatory patterns in the left sensory-motor regions. Finally, serial and semantic clustering strategies showed opposite regression patterns, indicating a competitive interaction. Together, these findings provide insights into resting-state neural markers associated with diverse encoding strategies in verbal learning.
SHINIER: An open-source Python package for controlling low-level image properties
Mathias Salvas-Hébert
Nicolas Dupuis-Roy
Catherine Landry
Frédéric Gosselin
CURATE: Automatic Curriculum Learning for Reinforcement Learning Agents through Competence-Based Curriculum Policy Search in Structured Task Spaces
Nan Rosemary Ke
Sarvesh Patil
Annya Dahmani
Eunice Yiu
Alison Gopnik
Oliver Kroemer
Due to fundamental exploration challenges without informed priors or specialized algorithms, agents may be unable to consistently receive in… (see more)formative rewards, leading to inefficient or intractable learning. To address these challenges, we introduce CURATE, an automatic curriculum learning algorithm for reinforcement learning agents in structured task spaces of monotonic difficulty. Through "exploration by exploitation," CURATE dynamically scales the task difficulty to match the agent's current competence. By exploiting its current capabilities that were learned in easier tasks, the agent improves its exploration in more difficult tasks. Our key insight is that the learning improvement in tasks that are close to those used for training is inversely proportional to their difficulty, and an agent that chooses a nearby distribution of the easiest unsolved tasks at any given time can automatically induce an easiest-to-hardest curriculum in these task spaces. To achieve this, CURATE conducts policy search in the task space to learn the best task distribution for training. As the agent's mastery grows, the learned curriculum adapts in an approximately easiest-to-hardest and task-directed fashion, efficiently culminating in a performant agent. Our experiments across three diverse domains (MiniGrid, Procgen, BipedalWalker) demonstrate that CURATE learns effective curricula for sample efficiency and proficiency with the potential for yielding broadly capable agents, matching or exceeding prior curriculum methods that do not require informed initialization or predefined schedules.
One Policy Is Enough: Single-Agent Reinforcement Learning Outperforms Tree Search for Chemistry Tool Learning
Armin Dariani
Entao Yang
Chemistry questions often demand exact computation and database lookups that a language model cannot supply from its parameters, so it must … (see more)reach for external tools. Tool use here is a three-part problem: select the right tool from a large pool, fill it with correctly typed arguments, and chain calls so that each consumes the outputs of the last. CheMatAgent, a previously published system, addresses this with hierarchical evolutionary MCTS: separate policy and execution models searching tool-call trees under two learned critics, one regressed partly onto GPT-assigned scores. We show that a single policy suffices. Our model interleaves reasoning, tool calls, and returns in one left-to-right generation, trained by a supervised warm-up and then outcome-level reinforcement learning against a programmatic reward read directly off the gold call chain, which leaves no learned critic and no judge in the training loop. On ChemToolBench multiple-tool comprehensive chemistry, on both backbones CheMatAgent use, we improve Tool F1 by 5.5% and Return F1 by 9.6% on Qwen-2.5-7B, and by 3.7% and 3.9% on Llama-3.1-8B, compared with their strongest search configuration, at one model invocation per question, against a search whose cost grows with the tree; we also lead answer Pass Rate on Qwen-2.5-7B.