Publications

AQY volume 100 issue 411 Cover and Front matter
Robin Skeates
Lindsey Elstub
Max Storey
G. Barker
Amy Bogaard
Robin Coningham
B. Cunliffe
A. Harding
Carl Heron
M. Millett
Nicky Milner
Cameron Petrie
Chris Gosden
Sue Hamilton
Mike Pitts
Marie Louise
Stig Sørensen Peter
Bellwood
Xingcan Chen
C. Higham … (see 149 more)
S. Houston
S. McIntosh
I. Kuijt
K. Lane
Akira Matsuda
Peter Mitchell
T. Pauketat
Michael D. Petraglia
Susan Pollock
Natalia Shishlina
Benjamin Smith
Monica Smith
Miriam Stark
Sarah Tarlow
Bruce Anthony
Tumbaco Vega
María Soledad
Solórzano Venegas
Bryan G. Valencia Castillo
Sergio Alarcón Robledo
Ricardo E. Basso Rial
Gabriel García Atiénzar
Yolanda Carrión Marco
Paul La
Sierra Pareja
Virginia Barciela González
Mauro S. Hernández Pérez
Wenhui Cheng
Zhangqiaochu Yang
Caixia Bai
Qiyue Wang
Yongsheng Zhao
Sophia Adams
J. Armstrong
A. Bayliss
Tom Moore
Emily Williams
Johannes Eber
S. Gur-Arieh
Robert C. Power
Maxime Rageot
Philipp Wolfgang
Stockhammer
C. Klesner
Rosie R. Crawford
Jasmine Vieri
M. Martinón-Torres
Geir Grønnesby
Hanne Bryn
Lars Forseth
Bente Philippsen
Knut Paasche
Christian Løchsen
Arne Abel Stamnes
A. Krol
A. Khrustalev
Arkady Savinetsky
A. Chirkova
Christina T. Halperin
Carmen Ramos Hernandez
L. Gauthier
J. Flexner
Grzegorz Kiarszys
Marek Lemiesz
Mary Lewis
Rebecca Pitt
Solange Bohling
Stefano Costanzo
Nevio Danelon
Martina Ciavardini
Francesca Balossi
Giana María
Di Nocera
K. Olsen
Bruhns. 2024. Ancient South
Marion Uckelmann
Mohammad Masoumian
N. Hariri
Tim Boaz
Bruun Skuldbøl
S. Renette
Samran Asiabani
Mojgan Fateh Zarefar
Seif Panahi
Hamzeh Mohammadpour
Amar Tazik
Ali Behnia
Michael Brown
F. J. Kreppner
Ergül Kodaş
Charlotte Labedan Kodaş
Bahattin I pek
Laura Dietrich
C. V. Tycowicz
Michael Brandl
J. Mayer
Lohengrin Baunack
Iris Schmidt
Wulf Hein
Marina Eguíluz Valentini
Simone Meinecke
W. Carleton
Dan Lawrence
D. Brotherson
Claire E. Ebert
José Lobo
Scott G. Ortman
Michael E. Smith
Thon Tho
I. Romanowska
Sarah Klassen
Patrick Roberts
Svitlana Ivanova
Alexey G. Nikitin
S. Radchenko
Dmytro Kiosak
Anke Hein
Andrew Womack
Katherine Brunson
Jada Ko
Michel Lee
A. Vahdati
K. Mohammadkhani
Zeinab Mahjoub
John E. Parkington
S. Milton-Dean
Ashley Christowitz
Stephen Wessels
C. Poggenpoel
Liora Kolska
Horwitz
Emma Messinger
B. Hanks
M. Bermann
Jaime J. Awe
M. Malekzadeh
R. Naseri
Elena Fausti
A. Cesaretti
Roberto Dan
Andrzej Pydyn
Mateusz Popek
Konrad Lewek
Andrzej Kowalczyk
Patricia Ayipey
Alexa Höhn
Dela Kuma
J. Beneš
Automated diagnosis of usual interstitial pneumonia on chest CT via the mean curvature of isophotes
Peter Savadjiev
Morteza Rezanejad
Sahir Bhatnagar
David Camirand
Claude Kauffmann
Ronald J. Dandurand
Patrick Bourgouin
Carl Chartrand-Lefebvre
Alexandre Semionov
To test whether the mean curvature of isophotes (MCI), a geometric image transformation, can be used to improve automatic detection on chest… (see more) CT of Usual Interstitial Pneumonia (UIP), a determining radiological pattern in the diagnosis of Interstitial Lung Diseases (ILD). This retrospective study included chest CT scans from 234 patients (123 female,111 male; mean age: 61.6 years; age range: 18-90 years) obtained at two independent institutions between 2007 and 2024. Three different classification models were trained on the original CT images and separately on MCI-transformed CT images: (1) a previously published deep learning model for classifying fibrotic lung disease on chest CT, (2) a classification pipeline based on the EfficientNet-V2 convolutional neural network architecture, and (3) a non-deep-learning model based on the functional principal component analysis (FPCA) of density functions of voxel intensity. All models were trained on data from the first institution and evaluated on data from the second institution with the recall-macro, precision-macro and F1-macro scores. Performance difference between classifier pairs was tested with the Stuart-Maxwell marginal homogeneity test. For a fixed model architecture and training algorithm, MCI-transformed images yield comparable or better classification performance than the original CT images. The best performance improvement achieved with MCI compared to CT was: recall-macro 0.83 vs 0.57, precision-macro 0.81 vs 0.50, F1-macro 0.80 vs 0.49, p=4.2e-5. MCI may be a valuable addition to existing AI systems for screening for UIP on chest CT. Machine learning methods for identifying usual interstitial pneumonia on chest CT perform better when the input CT images are transformed via the mean curvature of isophotes (MCI), a geometric transformation method known from classical computer vision. Three machine learning models were trained on a dataset of 158 patients from one institution and tested on another dataset of 76 patients from an independent institution to discriminate for usual interstitial pneumonia (UIP) on chest CT in a 3-group classification task. When keeping the network architecture and parameters fixed, changing the input image domain from the original CT to MCI-transformed images improved classification performance (Stuart-Maxwell test, p < 5e-3) MCI may be a valuable addition to existing machine learning systems for screening for UIP on chest CT, whether based on deep learning or on simpler shallow classifiers.
Automated robust segmentation of the spinal canal on MRI
Abel Salmona
Maxime Bouthillier
Gergely David
Maryam Seif
Armin Curt
Nikolai Pfender
Markus Hupp
Patrick Freund
Tomáš Horák
Petr Kudlička
Josef Bednařík
Fauziyya Muhammad
Zachary A. Smith
ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning
Baicheng Peng
Ziyang Song
Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively … (see more)model structured longitudinal electronic health records (EHRs). In contrast, EHR foundation models can learn predictive patient representations, yet lack interpretable language-based reasoning. To bridge this gap, we propose ChatHealthAI, a multimodal reasoning framework that aligns structured EHR representations from a pretrained EHR foundation model with the semantic space of a frozen LLM through a task-aware resampler. By integrating longitudinal patient representations with refined clinical event descriptions, ChatHealthAI enables clinically grounded natural-language reasoning while maintaining accurate patient prediction. We evaluated ChatHealthAI on three clinical predictive tasks from the EHRSHOT benchmark. Results show that ChatHealthAI improves reasoning quality and interpretability while preserving competitive predictive performance. These findings highlight the potential of integrating EHR foundation models with pretrained LLMs for interpretable clinical prediction.
Danilo Bzdok
Gamer in the scanner : Event-related analysis of fMRI activity during retro videogame play guided by automated annotations of game content
Yann Harel
Basile Pinsard
Julie A. Boyle
Valentina Borghesani
Paul-Henri Mignot
André Cyr
Abstract In recent years, videogames have gathered interest in cognitive neuroscience for their potential to study cognition in dynamical an… (see more)d naturalistic contexts. Yet, the complexity of game environments often challenges traditional modeling approaches, and current annotation methods—typically manual or based on modified games—remain labor-intensive and limited in scope. Here, we introduce a flexible and scalable framework using the gym-retro Python library to emulate a classic action-platformer, Shinobi III: Return of the Ninja Master (Sega, 1993), and automatically annotate gameplay events directly from the game’s memory states. This setup enables the identification of both player actions (e.g., jumping, hitting) and feedback events (e.g., killing an enemy, being hit), without modifying the game. Four individuals played the videogame for a combined total of 32 hours (&gt;7 hours each) while undergoing functional magnetic resonance imaging (fMRI). Resulting activation maps revealed distributed engagement of visual, motor, executive, and limbic systems, consistent with the cognitive demands of gameplay. Within-participant reproducibility of brain responses across sessions was robust across event types (r ≈ .25–.55), with some consistency observed even for rarer events like HealthLoss. Between-participant correlations were notably lower, reflecting participant-specific neural signatures. Multivoxel pattern analysis showed that brain responses to different in-game events were highly discriminable, with classification accuracy typically around or above 90%, though occasionally dropping to ~40% for less frequent events. These findings demonstrate that automated emulator-based annotations enable robust, interpretable, and scalable mapping of naturalistic cognitive processes using commercial videogames.
Key Issues and Future Directions in the Construction and Control of Geocentric Orbit Constellations for Gravitational Wave Detection
Yue LIU
Borui YAO
Meng LU
Yanchao HE
Ming LI
Lihua ZHANG
Jianying WANG
Mingying HUO
Lacuna: A Research Map for Machine Learning Problem Formulation
Alejandro Hernandez
Miles Q. Li
Christopher Pal
Nasim Rahaman
Research problem formulation is the cognitive task of turning a vague scientific idea into a testable hypothesis. \textbf{Lacuna} is a resea… (see more)rch map that supports this process for machine learning, using LLMs to turn papers and scholarly metadata into markdown summaries, concept elements, research directions, author and venue pages, and research proposals. Each item keeps links to the primary source records and papers that support it. We release the map with web, markdown, and MCP interfaces, plus scripts for reproducing the agent run. On a theorem-proving seed idea, the agent reaches a research proposal in 85.5 seconds and 7 tool calls, vs.\ 289.2 seconds and 27 tool calls for a sequential-PDF baseline. On ScholarQA-CS-ML, an ML/AI subset of the OpenScholar benchmark, Lacuna-GPT-4o scores 0.694 under the ScholarQABench rubric judge, compared with 0.672 for the OpenScholar-GPT-4o baseline on the same questions.
MobEvolve: An Agentic Self-Evolving Heuristic System for Interpretable Human Mobility Generation
Junlin He
Yihong Tang
Tong Nie
Ao Qu
Yuebing Liang
Hamzeh Alizadeh
Wei Ma
Lijun Sun
Human mobility generation aims to synthesize realistic trip chains for target populations based on individual features. Existing paradigms, … (see more)including deep generative models, LLM-based methods, and traditional heuristics, struggle to satisfy the complex demands of this task while simultaneously maintaining interpretability, behavioral plausibility, population-level distributional alignment, and inference efficiency. To bridge this gap, we introduce MobEvolve, the first agentic self-evolving heuristic framework for human mobility generation. MobEvolve initializes a behavior-inspired heuristic system and employs an LLM agent to iteratively evolve its internal logic. By diagnosing empirical misalignments and failure cases on a validation set, the agent proposes targeted updates and accumulates evolution memory for cumulative self-improvement. Extensive evaluations on the Singapore and Montreal benchmarks demonstrate that MobEvolve significantly outperforms state-of-the-art deep generative and LLM-based methods in individual trajectory fidelity, population-level distribution alignment, and behavioral plausibility, while preserving interpretability and high inference efficiency.
Molecular pathways of immune checkpoint inhibitor–induced hepatitis.
Erika Bushatsky
Natasha Ryan
Manuel Flores Molina
Steph A. Pang
Judith Lapierre
Madelyn Abraham
Sonia del Rincón
Marie Hudson
Wilson H. Miller
2573 Background: Immune checkpoint inhibitor (ICI) related hepatitis is a clinically significant immune-related adverse event (irAE) a… (see more)nd a common cause of treatment interruption. It occurs in roughly 5 to 10 percent of patients receiving anti PD-(L)1 monotherapy and in up to one third of those treated with combination ICI therapy. Despite increasing clinical recognition, the molecular mechanisms and predictive factors underlying ICI hepatitis remain poorly defined. The Montreal Immune-Related Adverse Events (MIRAE)-led hepatitis project aims to characterize the immune cell populations and underlying transcriptional programs associated with ICI-hepatitis pathogenesis. Methods: This translational study is conducted within the MIRAE biobank, a prospective multicenter cohort of ICI-treated patients with and without irAEs. The hepatitis cohort includes patients with longitudinal plasma samples collected at baseline, on treatment, and at irAE onset. Ongoing immune profiling efforts include plasma-based cytokine and chemokine analysis, high-throughput plasma proteomics, and single cell RNA sequencing of PBMCs. Preliminary analysis focused on plasma proteomics. Five patients with high-grade ICI-hepatitis and five ICI-treated controls without irAEs were selected and matched by age, sex, and primary tumor. Plasma samples were analyzed using the SomaScan 11K assay to identify differentially expressed proteins and enriched immune pathways. Results: ICI-related hepatitis was clinically severe, requiring systemic corticosteroids in all cases and additional immunosuppressive therapies in most patients. ICI-hepatitis cases showed significantly higher plasma levels of liver injury markers, including ALT and AST, compared with matched controls. Widespread alterations were observed in the circulating proteome, with strong upregulation of liver-enriched proteins and inflammatory mediators. Gene set enrichment analyses revealed enrichment of liver-associated pathways including xenobiotic and bile acid metabolism, as well as IL-12 signaling, interferon-α and γ, neutrophil-associated pathways, and liver-resident macrophage signatures. Pathway analysis of single cell data revealed enhanced cytotoxic activity of CD8 T cells during ICI hepatitis, as exemplified by upregulation of the CTL and IL-6 pathways. Conclusions: ICI-hepatitis was associated with circulating immune signature characterized by liver injury markers, inflammatory mediators, and enrichment of innate immune pathways. These findings provide molecular insight into the immunopathogenesis of ICI hepatitis and inform future biomarker discovery, druggable pathways, and risk stratification.
A qualitative study on XAI techniques for Software Defect Prediction
Saumendu Roy
Banani Roy
Chanchal K. Roy
Context: Machine learning (ML) models are increasingly used in Software Defect Prediction (SDP) to identify defect-prone software modules. H… (see more)owever, many ML models operate as black boxes, making their predictions difficult for developers to interpret and trust. Although Explainable Artificial Intelligence (XAI) techniques such as LIME, SHAP, and BreakDown are widely adopted to improve transparency, different explainers often produce inconsistent and conflicting explanations for the same prediction outcomes. Objective: This study investigates the use of XAI techniques in SDP, evaluates the consistency and reliability of commonly used explainers, examines practitioners’ challenges in interpreting explanations, and explores strategies for improving explanation trustworthiness and usability. Method: We conducted a mixed-methods study consisting of four phases: (1) a systematic literature review of 93 studies on XAI in SDP; (2) a controlled comparative evaluation of multiple explainability techniques, including LIME, SHAP, PyExplainer, PDP, ICE, and BreakDown; (3) a scenario-based analysis of explanation behavior using JIRA-based defect datasets and Random Forest prediction models; and (4) a practitioner survey involving 71 participants to investigate interpretability, usability, and trust-related concerns. Explanation consistency was evaluated using Feature Agreement (FA), Rank Agreement (RA), and Sign Agreement (SA), supported by statistical validation using Friedman ranking and Kendall Tau correlation analysis. Results: The results reveal substantial disagreement among explainers in terms of feature importance, feature ranking, and feature influence direction. Statistical analysis further confirms significant variation in explanation consistency across different techniques, with SHAP demonstrating comparatively stronger agreement behavior. The practitioner study also identified several practical challenges, including low trust, contradictory explanations, and limited integration support in software engineering workflows. Conclusion: Current XAI techniques for SDP still face important reliability and interpretability challenges that limit their practical adoption. Our findings highlight the need for agreement-aware evaluation, explanation stability analysis, and developer-centered explainability support to improve the trustworthiness and usability of XAI systems in software engineering.
Retrieval-augmented generation for natural language processing: a survey
Shangyu Wu
Ying Xiong
Yufei Cui
Can Chen
Lianming Huang
Xue Liu
Tei-Wei Kuo
Nan Guan
Chun Jason Xue
Large language models (LLMs) have achieved strong empirical performance in various fields, benefiting from their huge amount of parameters t… (see more)hat store knowledge. However, LLMs still suffer from several key issues, such as hallucination problems, knowledge update issues, and lacking domain-specific expertise. The appearance of retrieval-augmented generation (RAG), which leverages an external knowledge base to augment LLMs, mitigates these limitations. This paper presents a systematic review of RAG techniques for natural language processing (NLP), with a focus on retrievers and retrieval fusions. We introduce a novel taxonomy of retrieval fusions, such as query-based, logits-based, latent, and parametric fusion, and provide structured comparisons across accessibility, efficiency, and use cases. The paper further examines RAG applications across diverse NLP tasks, discusses evaluation methodologies and benchmark limitations, and analyzes training paradigms with and without knowledge base updates. Finally, we explore industrial deployment considerations and identify emerging challenges and future directions, including security, efficiency, and graph-based retrieval.