Publications

Genetic connectivity of Uroteuthis sibogae (Cephalopoda: Loliginidae) in the Sulu Sea with notes on morphology and statolith microchemistry
Jessica M. Legaspi
Lorenzo C. Halasan
Kris Angeli S. Sanchez-Delos Reyes
Tomoyo Okumura
Hsiu‐Chin Lin
MinatoLoader: Accelerating Machine Learning Training Through Efficient Data Preprocessing
Stella Bitchebe
Ricardo Macedo
Machine learning (ML) frameworks, such as PyTorch and TensorFlow, rely on data loaders to preprocess data before feeding it to accelerators.… (see more) When preprocessing is inefficiently pipelined, GPUs can remain idle over long periods of time, leading to substantial training delays. For example, PyTorch's default data loaders can cause up to 76% GPU idleness. A key bottleneck is the variability in preprocessing time across samples within the same dataset. Existing data loaders are oblivious to this variability, training all samples uniformly. In this case, a single slow sample can stall the entire batch, causing head-of-line blocking.
PostLearn: Towards A Learned Index For PostgreSQL
Abrar Fuad
Bettina Kemme
Assessing Computational Thinking Skills in K–12 Education: A Systematic Review
Yimei Zhang
Yajie Song
Catalyst GFlowNet for electrocatalyst design: A hydrogen evolution reaction case study
Efficient and inexpensive energy storage is essential for accelerating the adoption of renewable energy and ensuring a stable supply, despit… (see more)e fluctuations in sources such as wind and solar. Electrocatalysts play a key role in hydrogen energy storage (HES), allowing the energy to be stored as hydrogen. However, the development of affordable and high-performance catalysts for this process remains a significant challenge. We introduce Catalyst GFlowNet, a generative model that leverages machine learning-based predictors of formation and adsorption energy to design crystal surfaces that act as efficient catalysts. We demonstrate the performance of the model through a proof-of-concept application to the hydrogen evolution reaction, a key reaction in HES, for which we successfully identified platinum as the most efficient known catalyst. In future work, we aim to extend this approach to the oxygen evolution reaction, where current optimal catalysts are expensive metal oxides, and open the search space to discover new materials. This generative modeling framework offers a promising pathway for accelerating the search for novel and efficient catalysts.
Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding
Fanghua Ye
Qiang Gao
Jian Li
Yuxing Tian
Sijing Duan
Nan Du
Xiaolong Li
Xue Liu
Large language models (LLMs) often produce content that contradicts or overlooks information provided in the input context, a phenomenon kno… (see more)wn as faithfulness hallucination. In this paper, we propose Context-Fidelity Boosting (CFB), a lightweight and general decoding-time framework that reduces such hallucinations by increasing the generation probability of source-supported tokens. Motivated by logit-shaping principles from watermarking techniques, CFB applies additive token-level logit adjustments based on a token's degree of support from the input context. Specifically, we develop three boosting strategies: static boosting, which applies a fixed bias to source-supported tokens; context-aware boosting, which scales this bias using the divergence between next-token distributions with and without context; and token-aware boosting, which further redistributes the adaptive bias according to local relevance estimated from source-position attention and source-scoped semantic similarity. CFB requires no retraining or architectural changes, making it compatible with a wide range of LLMs. Experiments on summarization and question answering tasks across multiple open-source LLMs show that CFB consistently improves faithfulness metrics with minimal generation overhead. Our implementation is fully open-sourced.
Generalizable spinal cord multiple sclerosis lesion segmentation across MRI contrasts, protocols, and centers
Pierre‐Louis Benveniste
Laurent Létourneau‐Guillon
David Araújo
Lydia Chougar
Dumitru Fetco
Masaaki Hori
Kouhei Kamiya
Steven Messina
Charidimos Tsagkas
Bertrand Audoin
Rohit Bakshi
Élise Bannier
Daniel Blezek
Jean‐Christophe Brisset
Virginie Callot
Erik Charlson
Michelle Chen
Olga Ciccarelli
Sarah Demortière
Gilles Edan … (see 36 more)
M Filippi
Tobias Granberg
Cristina Granziera
Christopher C. Hemond
B. Mark Keegan
Anne Kerbrat
J Kirschke
Petr Kudlička
Pierre Labauge
Lisa Eunyoung Lee
Yaou Liu
Caterina Mainero
Julian McGinnis
Mark Mühlau
Govind Nair
Kristin P. O’Grady
Jiwon Oh
Russell Ouellette
Alexandre Prat
Daniel S. Reich
Maria A. Rocca
Timothy M. Shepherd
Seth A. Smith
Leszek Stawiarz
Jason Talbott
Roger Tam
Shahamat Tauhid
Anthony Traboulsee
Constantina A. Treaba
Paola Valsasina
Zachary Vavasour
Marios Yiannakas
Shannon Kolind
The proposed model can achieve accurate and reliable spinal cord MS lesion segmentation across heterogeneous MRI data, addressing a key barr… (see more)ier to clinical translation. The model is available in the Spinal Cord Toolbox v7.2 and higher.Code repository: https://github.com/ivadomed/seg-sc-ms-lesion-multicontrast.
How Supply Chain Dependencies Complicate Bias Measurement and Accountability Attribution in AI Hiring Applications
The increasing adoption of AI systems in hiring has raised concerns about algorithmic bias and accountability, prompting regulatory response… (see more)s including the EU AI Act, NYC Local Law 144, and Colorado's AI Act. While existing research examines bias through technical or regulatory lenses, both perspectives overlook a fundamental challenge: modern AI hiring systems operate within complex supply chains where responsibility fragments across data vendors, model developers, platform providers, and deploying organizations. This paper investigates how these dependency chains complicate bias evaluation and accountability attribution. Drawing on literature review and regulatory analysis, we demonstrate that fragmented responsibilities create two critical problems. First, bias emerges from component interactions rather than isolated elements, yet proprietary configurations prevent integrated evaluation. A resume parser may function without bias independently but contribute to discrimination when integrated with specific ranking algorithms and filtering thresholds. Second, information asymmetries mean deploying organizations bear legal responsibility without technical visibility into vendor-supplied algorithms, while vendors control implementations without meaningful disclosure requirements. Each stakeholder may believe they are compliant; nevertheless, the integrated system may produce biased outcomes. Analysis of implementation ambiguities reveals these challenges in practice. We propose multi-layered interventions including system-level audits, vendor guidelines, continuous monitoring mechanisms, and documentation across dependency chains. Our findings reveal that effective governance requires coordinated action across technical, organizational, and regulatory domains to establish meaningful accountability in distributed development environments.
Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
Changjiang Han
Yuxing Tian
Jikun Kang
Xue Liu
Large Language Models (LLMs) exhibit strong implicit personalization ability, yet most existing approaches treat this behavior as a black bo… (see more)x, relying on prompt engineering or fine tuning on user data. In this work, we adopt a mechanistic interpretability perspective and hypothesize the existence of a sparse set of Preference Heads, attention heads that encode user specific stylistic and topical preferences and exert a causal influence on generation. We introduce Differential Preference Steering (DPS), a training free framework that (1) identifies Preference Heads through causal masking analysis and (2) leverages them for controllable and interpretable personalization at inference time. DPS computes a Preference Contribution Score (PCS) for each attention head, directly measuring its causal impact on user aligned outputs. During decoding, we contrast model predictions with and without Preference Heads, amplifying the difference between personalized and generic logits to selectively strengthen preference aligned continuations. Experiments on widely used personalization benchmarks across multiple LLMs demonstrate consistent gains in personalization fidelity while preserving content coherence and low computational overhead. Beyond empirical improvements, DPS provides a mechanistic explanation of where and how personalization emerges within transformer architectures. Our implementation is publicly available.
Assessing the impact of dimensionality reduction on clustering performance -- a systematic study
Ousmane Assani Amate
Mohammadreza Bakhtyari
Émilie Roy
Dimensionality reduction is a critical preprocessing step for clustering high-dimensional data, yet comprehensive evaluation of its impact a… (see more)cross diverse methods and data types remains limited. In this study, we systematically assess the influence of five dimensionality reduction techniques - Principal Component Analysis (PCA), Kernel Principal Component Analysis (Kernel PCA), Variational Autoencoder (VAE), Isometric Mapping (Isomap), and Multidimensional Scaling (MDS) - on the performance of four popular clustering algorithms - k-means, Agglomerative Hierarchical Clustering (AHC), Gaussian Mixture Models (GMM), and Ordering Points to Identify the Clustering Structure (OPTICS). We evaluate clustering quality using the Adjusted Rand Index (ARI), comparing results without and with dimensionality reduction at different reduction levels recommended in the literature (i.e., k-1, where k is the number of clusters, and 25% and 50% of the original number of dimensions). Our findings underscore the importance of a careful selection of the dimensionality reduction technique and the dimensionality reduction level that should be tailored to intrinsic data geometry and clustering algorithms under consideration.
Geometry-aware graph attention networks to explain single-cell chromatin states and gene expression with SEAGALL
Patrick Hanel
Anna Danese
Maria Colomé-Tatché
High-throughput single-cell sequencing is widely used to study cell identity. We present SEAGALL (Single-cell Explainable Geometry-Aware Gra… (see more)ph Attention Learning pipeLine), a deep learning method to quantify the impact of molecular features on cellular phenotype, based on geometry-regularised autoencoders (GRAE) and explainable graph attention networks (X-GAT). The GRAE embeds the data into a latent space to build a reliable cell-cell graph. The GAT is trained to learn the annotations and XAI is used to explain the predictions, unravelling the features driving cell identity. SEAGALL extracts specific and stable signatures from multiple omics experiments, going beyond differential marker genes.
Papillae growth and molecular responses of juvenile sea cucumbers (Apostichopus japonicus) exposed to different light intensities
Weiyan Li
Ziyu Liu
Yajie Deng
Jiaqi Liu
Jinge Yu
Haoran Xiao
Fenglin Tian
Lingshu Han
Chong Zhao