Portrait of Guy Wolf

Guy Wolf

Core Academic Member
Canada CIFAR AI Chair
Full Professor, Université de Montréal, Department of Mathematics and Statistics
Concordia University
CHUM - Montreal University Hospital Center
Research Topics
Data Mining
Deep Learning
Dynamical Systems
Graph Neural Networks
Information Retrieval
Learning on Graphs
Machine Learning Theory
Medical Machine Learning
Molecular Modeling
Multimodal Learning
Representation Learning
Spectral Learning

Biography

Guy Wolf is a Full Professor in the Department of Mathematics and Statistics (DMS) at the Université de Montréal (UdeM), a Canada CIFAR AI Chair & Core Academic Member at Mila (the Quebec AI institute), an Associate Researcher with CRCHUM (the Montreal university hospital research center), and a participating PI in the Helmholtz International Lab for Causal Cell Dynamics.

In 2024 he has been awarded a Humboldt Experienced Research Fellowship, as part of which he was a visiting professor at Heidelberg University (2024) and Helmholtz Munich (2024-2026) in Germany. Prior to joining UdeM and Mila, he was a Gibbs Assistant Professor (2015-2018) in the Applied Math Program and an Associate Research Scientist in the Department of Genetics (2018) at Yale University (CT, USA). Previously, he was a Postdoctoral Researcher (2013-2015) in the Department of Computer Science at École Normale Supérieure in Paris (France). He holds a Ph.D. in Computer Science from Tel Aviv University (Israel), and has five years of prior experience in IT software design & development for data analysis in military settings.

His current research focuses on guided representation learning for data exploration, including methods that leverage manifold learning and geometric deep learning for dimensionality reduction, visualization, denoising, data augmentation, and coarse graining. While relevant for a wide range of applications, he is particularly interested in the intersection of AI & health, including tools supporting exploratory analysis of biomedical data, e.g., in single-cell multiomics, drug discovery, and neuroscience.

Current Students

PhD - Université de Montréal
Collaborating researcher - University of Tübingen
Master's Research - Université de Montréal
Co-supervisor :
Master's Research - Concordia University
Principal supervisor :
Collaborating Alumni - Université de Montréal
PhD - Concordia University
Principal supervisor :
PhD - Université de Montréal
Independent visiting researcher - Helmholtz Munich
PhD - Université de Montréal
Co-supervisor :
Master's Research - Concordia University
Principal supervisor :
PhD - Université de Montréal
PhD - Université de Montréal
Co-supervisor :
Postdoctorate - Concordia University
Principal supervisor :
PhD - Université de Montréal
PhD - Concordia University
Principal supervisor :
Collaborating researcher - BYU
Independent visiting researcher - University of Fribourg
PhD - Université de Montréal
Principal supervisor :
PhD - Concordia University
Principal supervisor :
PhD - Université de Montréal
Master's Research - Université de Montréal
Master's Research - Université de Montréal
Collaborating Alumni - Université de Montréal
Co-supervisor :

Publications

Hierarchical Graph Neural Nets Can Capture Long-Range Interactions
Graph neural networks (GNNs) based on message passing between neighboring nodes are known to be insufficient for capturing long-range intera… (see more)ctions in graphs. In this project we study hierarchical message passing models that leverage a multi-resolution representation of a given graph. This facilitates learning of features that span large receptive fields without loss of local information, an aspect not studied in preceding work on hierarchical GNNs. We introduce Hierarchical Graph Net (HGNet), which for any two connected nodes guarantees existence of message-passing paths of at most logarithmic length w.r.t. the input graph size. Yet, under mild assumptions, its internal hierarchy maintains asymptotic size equivalent to that of the input graph. We observe that our HGNet outperforms conventional stacking of GCN layers particularly in molecular property prediction benchmarks. Finally, we propose two benchmarking tasks designed to elucidate capability of GNNs to leverage long-range interactions in graphs.
MURAL: An Unsupervised Random Forest-Based Embedding for Electronic Health Record Data
Michal Gerasimiuk
Dennis Shung
Adrian Stanley
Michael Schultz
Jeffrey Ngu
Loren Laine
A major challenge in embedding or visualizing clinical patient data is the heterogeneity of variable types including continuous lab values, … (see more)categorical diagnostic codes, as well as missing or incomplete data. In particular, in EHR data, some variables are {\em missing not at random (MNAR)} but deliberately not collected and thus are a source of information. For example, lab tests may be deemed necessary for some patients on the basis of suspected diagnosis, but not for others. Here we present the MURAL forest -- an unsupervised random forest for representing data with disparate variable types (e.g., categorical, continuous, MNAR). MURAL forests consist of a set of decision trees where node-splitting variables are chosen at random, such that the marginal entropy of all other variables is minimized by the split. This allows us to also split on MNAR variables and discrete variables in a way that is consistent with the continuous variables. The end goal is to learn the MURAL embedding of patients using average tree distances between those patients. These distances can be fed to nonlinear dimensionality reduction method like PHATE to derive visualizable embeddings. While such methods are ubiquitous in continuous-valued datasets (like single cell RNA-sequencing) they have not been used extensively in mixed variable data. We showcase the use of our method on one artificial and two clinical datasets. We show that using our approach, we can visualize and classify data more accurately than competing approaches. Finally, we show that MURAL can also be used to compare cohorts of patients via the recently proposed tree-sliced Wasserstein distances.
Topological Analysis of Single-Cell Hierarchy Reveals Inflammatory Glial Landscape of Macular Degeneration
Manik Kuchroo
Marcello DiStasio
Eric Song
Eda Calapkulu
Maryam Ige
Amar H. Sheth
Madhvi Menon
Abhinav Godavarthi
Yu Xing
Scott Gigante
Holly Steach
Janhavi Narain
George Mourgkos
Rahul M. Dhodapkar
Matthew J. Hirn
Bastian Rieck … (see 3 more)
Brian P. Hafler
Topological Analysis of Single-Cell Hierarchy Reveals Inflammatory Glial Landscape of Macular Degeneration
Manik Kuchroo
Marcello DiStasio
Eric Song
Eda Calapkulu
Maryam Ige
Amar H. Sheth
Madhvi Menon
Abhinav Godavarthi
Yu Xing
Scott Gigante
Holly Steach
Janhavi Narain
George Mourgkos
Rahul M. Dhodapkar
Matthew J. Hirn
Bastian Rieck … (see 3 more)
Brian P. Hafler
Uncovering the Folding Landscape of RNA Secondary Structure Using Deep Graph Embeddings
Egbert Castro
Andrew Benz
Biomolecular graph analysis has recently gained much attention in the emerging field of geometric deep learning. Here we focus on organizing… (see more) biomolecular graphs in ways that expose meaningful relations and variations between them. We propose a geometric scattering autoencoder (GSAE) network for learning such graph embeddings. Our embedding network first extracts rich graph features using the recently proposed geometric scattering transform. Then, it leverages a semi-supervised variational autoencoder to extract a low-dimensional embedding that retains the information in these features that enable prediction of molecular properties as well as characterize graphs. We show that GSAE organizes RNA graphs both by structure and energy, accurately reflecting bistable RNA structures. Also, the model is generative and can sample new folding trajectories.
Scattering GCN: Overcoming Oversmoothness in Graph Convolutional Networks
Graph convolutional networks (GCNs) have shown promising results in processing graph data by extracting structure-aware features. This gave … (see more)rise to extensive work in geometric deep learning, focusing on designing network architectures that ensure neuron activations conform to regularity patterns within the input graph. However, in most cases the graph structure is only accounted for by considering the similarity of activations between adjacent nodes, which limits the capabilities of such methods to discriminate between nodes in a graph. Here, we propose to augment conventional GCNs with geometric scattering transforms and residual convolutions. The former enables band-pass filtering of graph signals, thus alleviating the so-called oversmoothing often encountered in GCNs, while the latter is introduced to clear the resulting features of high-frequency noise. We establish the advantages of the presented Scattering GCN with both theoretical results establishing the complementary benefits of scattering and GCN features, as well as experimental results showing the benefits of our method compared to leading graph neural networks for semi-supervised node classification, including the recently proposed GAT network that typically alleviates oversmoothing using graph attention mechanisms.
Uncovering the Topology of Time-Varying fMRI Data using Cubical Persistence
Bastian Rieck
Tristan Yates
Christian Bock
Karsten Borgwardt
Nicholas Turk-Browne
Functional magnetic resonance imaging (fMRI) is a crucial technology for gaining insights into cognitive processes in humans. Data amassed f… (see more)rom fMRI measurements result in volumetric data sets that vary over time. However, analysing such data presents a challenge due to the large degree of noise and person-to-person variation in how information is represented in the brain. To address this challenge, we present a novel topological approach that encodes each time point in an fMRI data set as a persistence diagram of topological features, i.e. high-dimensional voids present in the data. This representation naturally does not rely on voxel-by-voxel correspondence and is robust to noise. We show that these time-varying persistence diagrams can be clustered to find meaningful groupings between participants, and that they are also useful in studying within-subject brain state trajectories of subjects performing a particular task. Here, we apply both clustering and trajectory analysis techniques to a group of participants watching the movie 'Partly Cloudy'. We observe significant differences in both brain state trajectories and overall topological activity between adults and children watching the same movie.
Learning General Transformations of Data for Out-of-Sample Extensions
Matthew Amodio
David van Dijk
While generative models such as GANs have been successful at mapping from noise to specific distributions of data, or more generally from on… (see more)e distribution of data to another, they cannot isolate the transformation that is occurring and apply it to a new distribution not seen in training. Thus, they memorize the domain of the transformation, and cannot generalize the transformation out of sample. To address this, we propose a new neural network called a Neuron Transformation Network (NTNet) that isolates the signal representing the transformation itself from the other signals representing internal distribution variation. This signal can then be removed from a new dataset distributed differently from the original one trained on. We demonstrate the effectiveness of our NTNet on more than a dozen synthetic and biomedical single-cell RNA sequencing datasets, where the NTNet is able to learn the data transformation performed by genetic and drug perturbations on one sample of cells and successfully apply it to another sample of cells to predict treatment outcome.
Geometric Wavelet Scattering Networks on Compact Riemannian Manifolds
Michael Perlmutter
Feng Gao
Matthew Hirn
The Euclidean scattering transform was introduced nearly a decade ago to improve the mathematical understanding of convolutional neural netw… (see more)orks. Inspired by recent interest in geometric deep learning, which aims to generalize convolutional neural networks to manifold and graph-structured domains, we define a geometric scattering transform on manifolds. Similar to the Euclidean scattering transform, the geometric scattering transform is based on a cascade of wavelet filters and pointwise nonlinearities. It is invariant to local isometries and stable to certain types of diffeomorphisms. Empirical results demonstrate its utility on several geometric learning tasks. Our results generalize the deformation stability and local translation invariance of Euclidean scattering, and demonstrate the importance of linking the used filter structures to the underlying geometry of the data.
TrajectoryNet: A Dynamic Optimal Transport Network for Modeling Cellular Dynamics
It is increasingly common to encounter data from dynamic processes captured by static cross-sectional measurements over time, particularly i… (see more)n biomedical settings. Recent attempts to model individual trajectories from this data use optimal transport to create pairwise matchings between time points. However, these methods cannot model continuous dynamics and non-linear paths that entities can take in these systems. To address this issue, we establish a link between continuous normalizing flows and dynamic optimal transport, that allows us to model the expected paths of points over time. Continuous normalizing flows are generally under constrained, as they are allowed to take an arbitrary path from the source to the target distribution. We present TrajectoryNet, which controls the continuous paths taken between distributions to produce dynamic optimal transport. We show how this is particularly applicable for studying cellular dynamics in data from single-cell RNA sequencing (scRNA-seq) technologies, and that TrajectoryNet improves upon recently proposed static optimal transport-based models that can be used for interpolating cellular distributions.
Advantages of biologically-inspired adaptive neural activation in RNNs during learning
Dynamic adaptation in single-neuron response plays a fundamental role in neural coding in biological neural networks. Yet, most neural activ… (see more)ation functions used in artificial networks are fixed and mostly considered as an inconsequential architecture choice. In this paper, we investigate nonlinear activation function adaptation over the large time scale of learning, and outline its impact on sequential processing in recurrent neural networks. We introduce a novel parametric family of nonlinear activation functions, inspired by input-frequency response curves of biological neurons, which allows interpolation between well-known activation functions such as ReLU and sigmoid. Using simple numerical experiments and tools from dynamical systems and information theory, we study the role of neural activation features in learning dynamics. We find that activation adaptation provides distinct task-specific solutions and in some cases, improves both learning speed and performance. Importantly, we find that optimal activation features emerging from our parametric family are considerably different from typical functions used in the literature, suggesting that exploiting the gap between these usual configurations can help learning. Finally, we outline situations where neural activation adaptation alone may help mitigate changes in input statistics in a given task, suggesting mechanisms for transfer learning optimization.
Visualizing structure and transitions in high-dimensional biological data
Kevin R. Moon
David van Dijk
Zheng Wang
Scott Gigante
Daniel B. Burkhardt
William S. Chen
Kristina Yim
Antonia van den Elzen
Matthew Hirn
Ronald R. Coifman
Natalia Ivanova