Portrait of Jun Ding

Jun Ding

Affiliate Member
Assistant professor, McGill University, Department of Medicine
Research Topics
Computational Biology
Medical Machine Learning
Representation Learning

Biography

Jun Ding is an assistant professor in the Department of Medicine of the Faculty of Medicine and Health Sciences at McGill University.

Alongside his team, he is dedicated to employing machine learning techniques to decipher the complex dynamics of cells in various diseases, such as developmental disorders, pulmonary diseases and cancers. The diverse and intricate nature of these conditions necessitates innovative approaches, prompting the use of state-of-the-art single-cell technologies to meticulously profile individual cell states. The result is a rich source of data for our machine learning models.

These technologies present unprecedented opportunities to advance understanding, particularly in fields like developmental and cancer biology. However, the challenge is to develop computational models capable of linking this intricate biomedical data to potential discoveries.

Ding’s primary focus lies in the development and refinement of machine learning methodologies, especially probabilistic graphical models, to effectively analyze, model and visualize both single-cell and bulk omics data, often featuring longitudinal or spatial dimensions. The goal is to harness these advanced machine learning techniques to deepen the comprehension of cellular dynamics, and so develop groundbreaking diagnostic and therapeutic strategies that can significantly benefit public health.

Current Students

PhD - McGill University
Principal supervisor :

Publications

DOLPHIN advances single-cell transcriptomics beyond gene level by leveraging exon and junction reads
Kailu Song
Yumin Zheng
Bowen Zhao
David H. Eidelman
The advent of single-cell sequencing has revolutionized the study of cellular dynamics, providing unprecedented resolution into the molecula… (see more)r states and heterogeneity of individual cells. However, the rich potential of exon-level information and junction reads within single cells remains underutilized. Conventional gene-count methods overlook critical exon and junction data, limiting the quality of cell representation and downstream analyses such as subpopulation identification and alternative splicing detection. We introduce DOLPHIN, a deep learning method that integrates exon-level and junction read data, representing genes as graph structures. These graphs are processed by a variational graph autoencoder to improve cell embeddings. DOLPHIN not only demonstrates superior performance in cell clustering, biomarker discovery, and alternative splicing detection but also provides a distinct capability to detect subtle transcriptomic differences at the exon level that are often masked in gene-level analyses. By examining cellular dynamics with enhanced resolution, DOLPHIN provides new insights into disease mechanisms and potential therapeutic targets.
A deep generative model for deciphering cellular dynamics and in silico drug discovery in complex diseases
Yumin Zheng
Jonas C. Schupp
Taylor Adams
Geremy Clair
Aurelien Justet
Farida Ahangari
Xiting Yan
Paul Hansen
Marianne Carlon
Emanuela Cortesi
Marie Vermant
Robin Vos
Laurens J. De Sadeleer
Iván O. Rosas
Ricardo Pineda
John Sembrat
Melanie Königshoff
John E. McDonough
Bart M. Vanaudenaerde
Wim A. Wuyts … (see 2 more)
Naftali Kaminski
Human diseases are characterized by intricate cellular dynamics. Single-cell transcriptomics provides critical insights, yet a persistent ga… (see more)p remains in computational tools for detailed disease progression analysis and targeted in silico drug interventions. Here we introduce UNAGI, a deep generative neural network tailored to analyse time-series single-cell transcriptomic data. This tool captures the complex cellular dynamics underlying disease progression, enhancing drug perturbation modelling and screening. When applied to a dataset from patients with idiopathic pulmonary fibrosis, UNAGI learns disease-informed cell embeddings that sharpen our understanding of disease progression, leading to the identification of potential therapeutic drug candidates. Validation using proteomics reveals the accuracy of UNAGI’s cellular dynamics analysis, and the use of the fibrotic cocktail-treated human precision-cut lung slices confirms UNAGI’s predictions that nifedipine, an antihypertensive drug, may have anti-fibrotic effects on human tissues. UNAGI’s versatility extends to other diseases, including COVID, demonstrating adaptability and confirming its broader applicability in decoding complex cellular dynamics beyond idiopathic pulmonary fibrosis, amplifying its use in the quest for therapeutic solutions across diverse pathological landscapes.
Alveolar epithelial cell plasticity and injury memory in human pulmonary fibrosis
Taylor Adams
Jonas C. Schupp
Agshin Balayev
Johad Khoury
A. Justet
Fadi Nikola
Laurens De Sadeleer
Juan Cala-García
Marta Zapata‐Ortega
Panayiotis V. Benos
John E. McDonough
Farida Ahangari
Melanie Königshoff
Robert Homer
Iván O. Rosas
Xiting Yan
Bart Vanaudenaerde
Wim Wuyts
Naftali Kaminski
Acute and repetitive lung epithelial injury can lead to irreversible and even progressive pulmonary fibrosis; Idiopathic pulmonary fibrosis … (see more)(IPF) is a fatal disease and quintessential example of this phenomenon. The composition of epithelial cells in human pulmonary fibrosis – irrespective of disease etiology – is marked by the presence of Aberrant Basaloid cells: an abnormal cell phenotype with pro-fibrotic and senescent features, localized to the surface of fibrotic lesions. Despite their relevance to human pulmonary fibrosis, the exotic molecular profile of Aberrant Basaloid cells has obscured their etiology, preventing insights into how or why these cells emerge with fibrosis. Here we identify cellular intermediaries between Aberrant Basaloid and normal alveolar epithelial cells in human IPF tissue. We track the emergence of Aberrant Basaloid cells from alveolar epithelial cells ex vivo and uncover a role for similar cells in epithelial regeneration under normal conditions. Lastly, we characterize the epigenetic changes that distinguish Aberrant Basaloid cells from their progenitors and identify hallmarks of AP-1 injury memory retention. This study elucidates the phenomenon of maladaptive epithelial plasticity and regeneration in pulmonary fibrosis and re-contextualizes therapeutic strategies for epithelial dysfunction.
Advancing global antifungal development to combat invasive fungal infection
Xiu-Li Wang
Koon Ho Wong
Chen Ding
Chang-Bin Chen
Wen-Juan Wu
Ningning Liu
Harnessing agent-based frameworks in CellAgentChat to unravel cell–cell interactions from single-cell and spatial transcriptomics
Understanding cell–cell interactions (CCIs) is essential yet challenging owing to the inherent intricacy and diversity of cellular dynamic… (see more)s. Existing approaches often analyze global patterns of CCIs using statistical frameworks, missing the nuances of individual cell behavior owing to their focus on aggregate data. This makes them insensitive in complex environments where the detailed dynamics of cell interactions matter. We introduce CellAgentChat, an agent-based model (ABM) designed to decipher CCIs from single-cell RNA sequencing and spatial transcriptomics data. This approach models biological systems as collections of autonomous agents governed by biologically inspired principles and rules. Validated across eight diverse single-cell data sets, CellAgentChat demonstrates its effectiveness in detecting intricate signaling events across different cell populations. Moreover, CellAgentChat offers the ability to generate animated visualizations of single-cell interactions and provides flexibility in modifying agent behavior rules, facilitating thorough exploration of both close and distant cellular communications. Furthermore, CellAgentChat leverages ABM features to enable intuitive in silico perturbations via agent rule modifications, facilitating the development of novel intervention strategies. This ABM method unlocks an in-depth understanding of cellular signaling interactions across various biological contexts, thereby enhancing in silico studies for cellular communication–based therapies.
DTractor enhances cell type deconvolution in spatial transcriptomics by integrating deep neural networks, transfer learning, and matrix factorization
Yong Jin Kweon
Chenyu Liu
Gregory Fonseca
Spatial transcriptomics (ST) captures gene expression with spatial context but lacks single-cell resolution. Single-cell RNA sequencing (scR… (see more)NA-seq) offers high-resolution profiles without spatial information. Accurate spot-level decomposition requires effective integration of both. We present DTractor, a deep learning-based framework that improves cell-type deconvolution in ST data through spatial constraints and transfer learning. DTractor achieves dual utilization of scRNA-seq reference data by incorporating both a cell-type-specific gene expression matrix and learned latent embeddings into a unified matrix factorization model. This joint modeling enables accurate estimation of cell-type proportions and cell-type-resolved gene expression within each spatial spot, while preserving biological and spatial coherence. DTractor further applies spatial regularization to maintain local tissue structure. Across multiple ST platforms and tissue types, DTractor demonstrates improved decomposition accuracy, robustness, and interpretability compared to existing methods. The results from DTractor support downstream applications such as spatial domain analysis and the study of spatially organized cellular behaviors.
Efficient and scalable construction of clinical variable networks for complex diseases with RAMEN.
Yiwei Xiong
Jingtao Wang
Tingting Chen
Douglas D. Fraser
Gregory Fonseca
Simon Rousseau
scCobra allows contrastive cell embedding learning with domain adaptation for single cell data integration and harmonization
Bowen Zhao
Kailu Song
Dong-Qing Wei
Yi Xiong
Single-Cell Multi-Omics Profiling of Immune Cells Isolated from Atherosclerotic Plaques in Male ApoE Knockout Mice Exposed to Arsenic
Kiran Makhani
Xiuhui Yang
France Dierick
Nivetha Subramaniam
Natascha Gagnon
Talin Ebrahimian
Stephanie Lehoux
Hao Wu
Koren K. Mann
Millions worldwide are exposed to elevated levels of arsenic that significantly increase their risk of developing atherosclerosis, a patholo… (see more)gy primarily driven by immune cells. While the impact of arsenic on immune cell populations in atherosclerotic plaques has been broadly characterized, cellular heterogeneity is a substantial barrier to in-depth examinations of the cellular dynamics for varying immune cell populations. This study aimed to conduct single-cell multi-omics profiling of atherosclerotic plaques in apolipoprotein E knockout (ApoE–/–) mice to elucidate transcriptomic and epigenetic changes in immune cells induced by arsenic exposure. The ApoE–/– mice were fed a high-fat diet and were exposed to either 200 ppb arsenic in drinking water or a tap water control, and single-cell multi-omics profiling was performed on atherosclerotic plaque-resident immune cells. Transcriptomic and epigenetic changes in immune cells were analyzed within the same cell to understand the effects of arsenic exposure. Our data revealed that the transcriptional profile of macrophages from arsenic-exposed mice were significantly different from that of control mice and that differences were subtype specific and associated with cell–cell interaction and cell fates. Additionally, our data suggest that differences in arsenic-mediated changes in chromosome accessibility in arsenic-exposed mice were statistically more likely to be due to factors other than random variation compared to their effects on the transcriptome, revealing markers of arsenic exposure and potential targets for intervention. These findings in mice provide insights into how arsenic exposure impacts immune cell types in atherosclerosis, highlighting the importance of considering cellular heterogeneity in studying such effects. The identification of subtype-specific differences and potential intervention targets underscores the significance of understanding the molecular mechanisms underlying arsenic-induced atherosclerosis. Further research is warranted to validate these findings and explore therapeutic interventions targeting immune cell dysfunction in arsenic-exposed individuals. https://doi.org/10.1289/EHP14285
DTPSP: A Deep Learning Framework for Optimized Time Point Selection in Time-Series Single-Cell Studies
Michel Hijazin
Pumeng Shi
Jingtao Wang
Time-series studies are critical for uncovering dynamic biological processes, but achieving comprehensive profiling and resolution across mu… (see more)ltiple time points and modalities (multi-omics) remains challenging due to cost and scalability constraints. Current methods for studying temporal dynamics, whether at the bulk or single-cell level, often require extensive sampling, making it impractical to deeply profile all time points and modalities. To overcome these limitations, we present DTPSP, a deep learning framework designed to identify the most informative time points in any time-series study, enabling resource-efficient and targeted analyses. DTPSP models temporal gene expression patterns using readily obtainable data, such as bulk RNA-seq, to select time points that capture key system dynamics. It also integrates a deep generative module to infer data for non-sampled time points based on the selected time points, reconstructing the full temporal trajectory. This dual capability enables DTPSP to prioritize key time points for in-depth profiling, such as single-cell sequencing or multi-omics analyses, while filling gaps in the temporal landscape with high fidelity. We apply DTPSP to developmental and disease-associated time courses, demonstrating its ability to optimize experimental designs across bulk and single-cell studies. By reducing costs, enabling strategic multi-omics profiling, and enhancing biological insights, DTPSP provides a scalable and generalized solution for investigating dynamic systems.
MATES: a deep learning-based model for locus-specific quantification of transposable elements in single cell
Ruohan Wang
Yumin Zheng
Zijian Zhang
Kailu Song
Erxi Wu
Xiaopeng Zhu
Tao P. Wu
Transposable elements (TEs) are crucial for genetic diversity and gene regulation. Current single-cell quantification methods often align mu… (see more)lti-mapping reads to either ‘best-mapped’ or ‘random-mapped’ locations and categorize them at the subfamily levels, overlooking the biological necessity for accurate, locus-specific TE quantification. Moreover, these existing methods are primarily designed for and focused on transcriptomics data, which restricts their adaptability to single-cell data of other modalities. To address these challenges, here we introduce MATES, a deep-learning approach that accurately allocates multi-mapping reads to specific loci of TEs, utilizing context from adjacent read alignments flanking the TE locus. When applied to diverse single-cell omics datasets, MATES shows improved performance over existing methods, enhancing the accuracy of TE quantification and aiding in the identification of marker TEs for identified cell populations. This development facilitates the exploration of single-cell heterogeneity and gene regulation through the lens of TEs, offering an effective transposon quantification tool for the single-cell genomics community.
scCobra: Contrastive cell embedding learning with domain-adaptation for single-cell data integration and harmonization
Bowen Zhao
Dong-Qing Wei
Yi Xiong
The rapid development of single-cell technologies has underscored the need for more effective methods in the integration and harmonization o… (see more)f single-cell sequencing data. The prevalent challenge of batch effects, resulting from technical and biological variations across studies, demands accurate and reliable solutions for data integration. Traditional tools often have limitations, both due to reliance on gene expression distribution assumptions and the common issue of over-correction, particularly in methods based on anchor alignments. Here we introduce scCobra, a deep neural network tool designed specifically to address these challenges. By leveraging a deep generative model that combines a contrastive neural network with domain adaptation, scCobra effectively mitigates batch effects and minimizes over-correction without depending on gene expression distribution assumptions. Additionally, scCobra enables online label transfer across datasets with batch effects, facilitating the continuous integration of new data without retraining, and offers features for batch effect simulation and advanced multi-omic batch integration. These capabilities make scCobra a versatile data integration and harmonization tool for achieving accurate and insightful biological interpretations from complex datasets.