Publications

Graphylo Var: Predicting the impact of non-coding variants using a multi-species sequence model
Dongjoon Lim
MOTIVATION: Understanding the functional impact of genetic variants is a key problem for precision medicine. Tools like CADD, PhyloP, and Ph… (see more)astCons are useful, but they often look at each position in the genome in isolation. This means they can miss important information from the evolutionary history that connects different species. In this paper, we extend our previous model, Graphylo, to predict the effects of variants. Our new model, GraphyloVar, is built to directly utilize the phylogenetic tree that relates the species. RESULTS: GraphyloVar is a deep learning model that considers both DNA sequence and evolutionary patterns from many species. It uses two main components: Graph Convolutional Networks (GCNs) to process the phylogenetic tree, and Transformer encoders to extract features from the DNA sequences. Pre-trained to predict population-level allele frequencies on the TOPMed whole-genome sequencing cohort, GraphyloVar achieves an AUROC of 0.6246 zero-shot on ∼149M held-out variants, and an ensemble with CADD reaches 0.6442 (+0.020, p 10-15). Fine-tuned GraphyloVar achieves the highest AUROC across all 13 MPRA benchmark datasets. By integrating deep learning with explicit phylogenetic input, GraphyloVar offers a powerful and complementary approach to variant effect prediction that utilizes the full evolutionary history from many species to better identify and prioritize important non-coding variants. AVAILABILITY AND IMPLEMENTATION: Code and datasets are available at https://github.com/DongjoonLim/GraphyloVar under DOI : 10.5281/zenodo.20616818.
Patterns of Muscle Health in Single- and Multi-Site Chronic Pain: A UK Biobank Normative Modeling Study
Merve Kaptan
Yiyu Wang
Augustijn de Boer
Ananya Goyal
Skylar Holmes
Kerem Ozkan
Teresa Indriolo
Christine S W Law
Dario Pfyffer
Joel Fundaun
Estifanos Berhe
Garry E. Gold
Akshay Chaudhari
Anoosha Pai S
Anthony A. Gatti
Feliks Kogan
Brian A. Hargreaves
Scott L. Delp
John Ratliff … (see 21 more)
Serena Hu
Anand Veeravagu
Atman Desai
Suzanne Tharin
Todd Alamin
Andrew C. Smith
Marnee J. McKay
Brian Kim
Robert Walsh
Alec Schielke
Dean Dennis
Johannes Decker
Benjamin De Leener
Zachary A. Smith
Fauziyya Muhammad
James M. Elliott
Andre F. Marquand
Sean Mackey
Evert Onno Wesselink
Kenneth A. Weber II
Abstract Background Chronic pain is associated with impaired muscle health, but whether these changes reflect site-specific factors, broader… (see more) systemic factors, or both remains unclear. The purpose of this study is to determine whether normative markers of muscle health derived from MRI show site-specific patterns in chronic pain. Methods UK Biobank participants who underwent whole-body MRI from 2006 to 2010 were included in this retrospective cross-sectional study. The MuscleMap Toolbox quantified volume and intramuscular fat (IMF) in 42 muscles of the abdomen, pelvis, and thigh. Normative models trained on a no pain group generated muscle-specific deviations from normal (i.e., Z-scores) for single- and multi-site chronic and acute pain. Results Of 17,843 participants, the primary site-specific analysis included 9,704 no pain, 885 single-site chronic back pain (CBP), 438 single-site chronic hip pain (CHP), and 1,315 single-site chronic knee pain (CKP) participants (n=12,342; mean age 63.7±7.5 years; 52.7% female). Additional analyses included single-site chronic neck/shoulder pain, acute pain, and multi-site chronic pain groups. In CBP, deviations were localized to abdominal muscles, with decreased volume in 6/8 and increased IMF in 6/8. In CHP, deviations were broad, with decreased volume in 3/8 of the abdominal and 14/26 of the thigh muscles, and increased IMF in 6/8 of the abdominal, 5/8 of the pelvic, and 4/26 of the thigh muscles. In CKP, deviations were localized to thigh muscles, with decreased volume in 8/26 and increased IMF in 6/26. Acute pain groups showed no significant differences except for decreased volume in one thigh muscle in acute knee pain. With each additional chronic pain site, volume decreased (β=−.078;IQR:−0.100−0.051), and IMF increased (β=.085;IQR:0.066−0.101). Combined Z-scores classified chronic pain groups better than chance (accuracy: 48.6%;p<.001), but not acute pain groups (accuracy: 39.0%;p=.20). Conclusions Whole-body MRI combined with AI-driven muscle segmentation and normative modeling revealed site-specific patterns of muscle health in single-site chronic pain.
Foundation Models for Epileptogenic Zone Identification in Drug-Resistant Epilepsy
Thi Kieu Khanh Ho
Petr Klimes
Jan Cimbálník
Martin Pail
Milan Brázdil
Birgit Frauscher
Accurate identification of the epileptogenic zone (EZ) is essential for seizure freedom after resective surgery in drug-resistant epilepsy, … (see more)yet seizure freedom rates remain below 50%. We developed EpiiSLM, a dual foundation model system for EZ identification with stereo-electroencephalography (sEEG), by training a signal foundation model on 104,990 minutes of sEEG recordings from the Montreal Neurological Institute & Hospital, while leveraging all recordings regardless of surgical outcome and anchoring EZ biomarker extraction on non-epileptic signals. A language foundation model then integrates sEEG-derived outputs with multimodal clinical information to produce interpretable predictions. Under leave-one-patient-out evaluation, EpiiSLM achieved 0.978 contact-level positive predictive value (PPV), outperforming the seizure onset zone(SOZ)-as-EZ baseline by 15.1% (p < 0.05), and 100% region-level accuracy; on an external dataset, EpiiSLM achieved 0.857 contact-level PPV. EpiiSLM requires only one night of interictal sleep data, suggesting potential to reduce invasive sEEG monitoring duration and improve surgical outcomes.
Drowning in Routine: Signal Dilution in Multi-Turn Agent Training
Vi Retault
Multi-turn agents interleave consequential decisions with routine execution: some actions change the downstream return distribution, while o… (see more)thers are necessary but reward-equivalent. The cost of trajectory-level credit assignment, often attributed to long horizons, is in fact governed by decision density
An empirical study on logging evolution on stack overflow: trends, topics, and challenges
Patrick Loic Foalem
Andre Nguimbous
Heng Li
Ettore Merlo
Wagg: Cost-aware Aggregation of Windowing Operators in Stream Processing
Pritish Mishra
Ruoyu Deng
Alexandre da Silva Veith
Eyal de Lara
Accelerated and Stable Convergence with Anchored Optimistic Method
We study first-order methods for solving monotone variational inequalities arising in min-max optimization. Classical approaches such as the… (see more) extragradient method rely on two gradient queries per iteration, which limits their analysis and applicability in the online and stochastic settings. We propose a family of Generalized Optimistic Methods with Anchoring (GOMA), which combine two-time-scale optimistic updates with an anchoring term inspired by Halpern iteration. In the deterministic setting, GOMA achieves the optimal accelerated last-iterate rate
Geometric Path Following for Autonomous Dynamic Soaring
Zihao Zhuo
Meyer Nahon
Autonomous dynamic soaring can be used to increase the endurance and range of unmanned aerial vehicles by harvesting energy from the vertica… (see more)l gradient of the horizontal wind. This study aims to develop a guidance and control strategy that allows precise following of an optimal dynamic-soaring path for a glider vehicle. The proposed control architecture combines a geometric path-following guidance law with an SO(3)-based attitude control law. High-fidelity six-degree-of-freedom simulation shows that the proposed method can achieve a position accuracy of 0.1 m for a glider with a wingspan of 2 m while adhering to the constraints present on a glider airframe. The high tracking accuracy makes it possible to conduct autonomous dynamic-soaring operations with patterns that were considered impossible in previous studies, such as a travel pattern mimicking the albatrosses’ dynamic-soaring pattern in close proximity to the ocean surface.
PrivacyAlign: Contextual Privacy Alignment for LLM Agents
Manveer Singh Tamber
Marc-Etienne Brunet
Jimmy Lin
AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decisions must align with wh… (see more)at they actually want. Privacy is an important alignment problem for agents: every message, post, or tool call an agent makes is a contextual judgment about what is appropriate to share, with whom, and under which conditions. Because such judgments depend on social expectations and norms, human judgment does not merely label privacy violations but also helps define them. While existing work relies on unreliable proxies for both training and evaluation, we place human judgment at the center of agentic privacy alignment. We introduce PrivacyAlign, a dataset of 1,350 samples with 3,516 detailed annotations from 599 unique annotators across diverse scenarios where current LLMs actually leak, and use it to ground both alignment training and automated evaluation in human privacy norms. Building on these annotations, we first show that conditioning LLM judges on human annotations and explanations for reference responses to the same prompt makes their judgments more reliable. We then introduce annotation-conditioned reward modeling, which uses these annotations to score new responses during RL, and show that small open-weight agents trained with this reward better align with human privacy norms, with strong gains on PrivacyAlign and existing privacy benchmarks for agents.
The digital heartbeat: a qualitative descriptive study on women's views on preventing cardiovascular disease in primary care
Ilhem Chaima Bousbiat
Samira Abbasgholizadeh Rahimi
Roland Grad
Charo Rodriguez
BACKGROUND: This empirical study aims to explore women's perspectives on cardiovascular disease and the use of digital health interventions … (see more)(DHIs) for their primary prevention and to gather insights on essential features for developing artificial intelligent-enabled technologies. METHODS: Adopting a qualitative descriptive research design, we conducted 15 semi-structured, in-depth interviews via Zoom with women at higher risk for cardiovascular disease. Participants were women over 40 years old, residing in Quebec, with at least one cardiovascular disease risk factor, and proficient in English. Recruitment was from a McGill University-affiliated clinic. An inductive thematic analysis approach was used for data analysis. RESULTS: Five major themes were identified: (i) understanding cardiovascular disease in a variety of ways, (ii) barriers and challenges to preventing cardiovascular disease in women, (iii) women taking charge of their cardiovascular well-being, (iv) mixed perspectives regarding artificial intelligent-enabled technologies for cardiovascular disease prevention such as Xi-Care, and (v) range of suggestions for the format and design of a prospective artificial intelligent-enabled technologies. CONCLUSIONS: Despite the prevalence of cardiovascular disease, there is a significant knowledge gap among women regarding the chronic nature and manifestations of these diseases. Artificial intelligent-enabled technologies like Xi-Care, with the potential for customization and interactive engagement, could enhance the primary prevention of cardiovascular disease in women, providing valuable insights for the subsequent phases of the project leading to Xi-Care's development.
FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching
Miguel Saavedra-Ruiz
Daniele Nardi
Joint spatial and temporal understanding of 3D scenes is a crucial requirement for robots deployed in everyday household environments. Such … (see more)agents must not only comprehend and navigate spatial layouts, but also reason about how these spaces evolve over time. In particular, humans interact with objects daily, causing them to change position throughout the environment and making it difficult for robots to reliably associate current observations with previously seen objects. However, these interactions are not random: human habits and routines induce spatio-temporally consistent patterns in object locations, which robotic agents can potentially learn and then exploit for downstream tasks such as navigation. To this end, we introduce FlowMaps, a latent flow matching model for estimating multimodal distributions over the future locations of dynamic objects in a continuous 3D space. By learning the implicit dependencies among objects and their temporal evolution, FlowMaps predicts likely changes in object locations conditioned on past human interactions, while supporting generalization across previously unseen environments that share similar object routines. To demonstrate the utility of this method, we deploy FlowMaps in a downstream dynamic Object Navigation task in both simulated and real-world environments. Across more than 600 episodes, FlowMaps outperforms state-of-the-art approaches, showing that modeling object dynamics through continuous, multimodal spatio-temporal distributions improves robotic search and navigation in changing household environments. Code and additional material is available at https://fra-tsuna.github.io/flowmaps/.
Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior
Ali Khalil
Aly M. Kassem
Mohamed Abdelrazek
Santu Rana
We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled… (see more) into reusable jailbreak attacks. Using an emergent-misalignment organism and a refusal-ablated jailbroken organism, we transplant harmful CoTs into