Portrait of Smita Krishnaswamy

Smita Krishnaswamy

Affiliate Member
Associate Professor, Yale University
Université de Montréal
Yale
Research Topics
AI in Health
Brain-computer Interfaces
Cognitive Science
Computational Biology
Computational Neuroscience
Data Geometry
Data Science
Data Sparsity
Deep Learning
Dynamical Systems
Generative Models
Geometric Deep Learning
Graph Neural Networks
Information Theory
Manifold Learning
Molecular Modeling
Representation Learning
Spectral Learning

Biography

Our lab works on developing foundational mathematical machine learning and deep learning methods that incorporate graph-based learning, signal processing, information theory, data geometry and topology, optimal transport and dynamics modeling that are capable of exploratory analysis, scientific inference, interpretation and hypothesis generation big biomedical datasets ranging from single-cell data, to brain imaging, to molecular structural datasets arising from neuroscience, psychology, stem cell biology, cancer biology, healthcare, and biochemistry. Our works have been instrumental in dynamic trajectory learning from static snapshot data, data denoising, visualization, network inference, molecular structure modeling and more.

Current Students

Collaborating researcher - Yale University
Principal supervisor :

Publications

Diffusion Earth Mover's Distance and Distribution Embeddings
Kincaid MacDonald
Manik Kuchroo
Ronald R. Coifman
We propose a new fast method of measuring distances between large numbers of related high dimensional datasets called the Diffusion Earth Mo… (see more)ver's Distance (EMD). We model the datasets as distributions supported on common data graph that is derived from the affinity matrix computed on the combined data. In such cases where the graph is a discretization of an underlying Riemannian closed manifold, we prove that Diffusion EMD is topologically equivalent to the standard EMD with a geodesic ground distance. Diffusion EMD can be computed in
Topological analysis of single-cell data reveals shared glial landscape of macular degeneration and neurodegenerative diseases
Manik Kuchroo
Marcello DiStasio
Eda Calapkulu
Maryam Ige
Amar H. Sheth
Madhvi Menon
Yu Xing
Scott Gigante
Rahul M. Dhodapkar
Bastian Rieck
Brian P. Hafler
1 A novel topological machine learning approach applied to single-nucleus RNA sequencing from human retinas… (see more) with age-related macular degeneration identifies interacting disease phase-specific glial activation states shared with Alzheimer’s disease and multiple sclerosis. 2 Neurodegeneration occurs in a wide range of diseases, including age-related macular degeneration (AMD), Alzheimer’s disease (AD), and multiple sclerosis (MS), each with distinct inciting events. To determine whether glial transcriptional states are shared across phases of degeneration, we sequenced 50,498 nuclei from the retinas of seven AMD patients and six healthy controls, generating the first single-cell transcriptomic atlas of AMD. We identified groupings of cells implicated in disease pathogenesis by applying a novel topologically-inspired machine learning approach called ‘diffusion condensation.’ By calculating diffusion homology features and performing persistence analysis, diffusion condensation identified activated glial states enriched in the early phases of AMD, AD, and MS as well as an AMD-specific proangiogenic astrocyte state promoting pathogenic neovascularization in advanced AMD. Finally, by mapping the expression of disease-associated genes to glial states, we identified key signaling interactions creating hypotheses for therapeutic intervention. Our topological analysis identified an integrated disease-phase specific glial landscape that is shared across neurodegenerative conditions affecting the central nervous system.
Finding Archetypal Spaces Using Neural Networks
David van Dijk
Daniel B. Burkhardt
Matthew Amodio
Archetypal analysis is a data decomposition method that describes each observation in a dataset as a convex combination of "pure types" or a… (see more)rchetypes. These archetypes represent extrema of a data space in which there is a trade-off between features, such as in biology where different combinations of traits provide optimal fitness for different environments. Existing methods for archetypal analysis work well when a linear relationship exists between the feature space and the archetypal space. However, such methods are not applicable to systems where the feature space is generated non-linearly from the combination of archetypes, such as in biological systems or image transformations. Here, we propose a reformulation of the problem such that the goal is to learn a non-linear transformation of the data into a latent archetypal space. To solve this problem, we introduce Archetypal Analysis network (AAnet), which is a deep neural network framework for learning and generating from a latent archetypal representation of data. We demonstrate state-of-the-art recovery of ground-truth archetypes in non-linear data domains, show AAnet can generate from data geometry rather than from data density, and use AAnet to identify biologically meaningful archetypes in single-cell gene expression data.
MURAL: An Unsupervised Random Forest-Based Embedding for Electronic Health Record Data
Michal Gerasimiuk
Dennis Shung
Adrian Stanley
Michael Schultz
Jeffrey Ngu
Loren Laine
A major challenge in embedding or visualizing clinical patient data is the heterogeneity of variable types including continuous lab values, … (see more)categorical diagnostic codes, as well as missing or incomplete data. In particular, in EHR data, some variables are {\em missing not at random (MNAR)} but deliberately not collected and thus are a source of information. For example, lab tests may be deemed necessary for some patients on the basis of suspected diagnosis, but not for others. Here we present the MURAL forest -- an unsupervised random forest for representing data with disparate variable types (e.g., categorical, continuous, MNAR). MURAL forests consist of a set of decision trees where node-splitting variables are chosen at random, such that the marginal entropy of all other variables is minimized by the split. This allows us to also split on MNAR variables and discrete variables in a way that is consistent with the continuous variables. The end goal is to learn the MURAL embedding of patients using average tree distances between those patients. These distances can be fed to nonlinear dimensionality reduction method like PHATE to derive visualizable embeddings. While such methods are ubiquitous in continuous-valued datasets (like single cell RNA-sequencing) they have not been used extensively in mixed variable data. We showcase the use of our method on one artificial and two clinical datasets. We show that using our approach, we can visualize and classify data more accurately than competing approaches. Finally, we show that MURAL can also be used to compare cohorts of patients via the recently proposed tree-sliced Wasserstein distances.
Topological Analysis of Single-Cell Hierarchy Reveals Inflammatory Glial Landscape of Macular Degeneration
Manik Kuchroo
Marcello DiStasio
Eric Song
Eda Calapkulu
Maryam Ige
Amar H. Sheth
Madhvi Menon
Abhinav Godavarthi
Yu Xing
Scott Gigante
Holly Steach
Janhavi Narain
George Mourgkos
Rahul M. Dhodapkar
Matthew J. Hirn
Bastian Rieck … (see 3 more)
Brian P. Hafler
Topological Analysis of Single-Cell Hierarchy Reveals Inflammatory Glial Landscape of Macular Degeneration
Manik Kuchroo
Marcello DiStasio
Eric Song
Eda Calapkulu
Maryam Ige
Amar H. Sheth
Madhvi Menon
Abhinav Godavarthi
Yu Xing
Scott Gigante
Holly Steach
Janhavi Narain
George Mourgkos
Rahul M. Dhodapkar
Matthew J. Hirn
Bastian Rieck … (see 3 more)
Brian P. Hafler
Uncovering the Folding Landscape of RNA Secondary Structure Using Deep Graph Embeddings
Egbert Castro
Andrew Benz
Biomolecular graph analysis has recently gained much attention in the emerging field of geometric deep learning. Here we focus on organizing… (see more) biomolecular graphs in ways that expose meaningful relations and variations between them. We propose a geometric scattering autoencoder (GSAE) network for learning such graph embeddings. Our embedding network first extracts rich graph features using the recently proposed geometric scattering transform. Then, it leverages a semi-supervised variational autoencoder to extract a low-dimensional embedding that retains the information in these features that enable prediction of molecular properties as well as characterize graphs. We show that GSAE organizes RNA graphs both by structure and energy, accurately reflecting bistable RNA structures. Also, the model is generative and can sample new folding trajectories.
Uncovering the Topology of Time-Varying fMRI Data using Cubical Persistence
Bastian Rieck
Tristan Yates
Christian Bock
Karsten Borgwardt
Nicholas Turk-Browne
Functional magnetic resonance imaging (fMRI) is a crucial technology for gaining insights into cognitive processes in humans. Data amassed f… (see more)rom fMRI measurements result in volumetric data sets that vary over time. However, analysing such data presents a challenge due to the large degree of noise and person-to-person variation in how information is represented in the brain. To address this challenge, we present a novel topological approach that encodes each time point in an fMRI data set as a persistence diagram of topological features, i.e. high-dimensional voids present in the data. This representation naturally does not rely on voxel-by-voxel correspondence and is robust to noise. We show that these time-varying persistence diagrams can be clustered to find meaningful groupings between participants, and that they are also useful in studying within-subject brain state trajectories of subjects performing a particular task. Here, we apply both clustering and trajectory analysis techniques to a group of participants watching the movie 'Partly Cloudy'. We observe significant differences in both brain state trajectories and overall topological activity between adults and children watching the same movie.
Learning General Transformations of Data for Out-of-Sample Extensions
Matthew Amodio
David van Dijk
While generative models such as GANs have been successful at mapping from noise to specific distributions of data, or more generally from on… (see more)e distribution of data to another, they cannot isolate the transformation that is occurring and apply it to a new distribution not seen in training. Thus, they memorize the domain of the transformation, and cannot generalize the transformation out of sample. To address this, we propose a new neural network called a Neuron Transformation Network (NTNet) that isolates the signal representing the transformation itself from the other signals representing internal distribution variation. This signal can then be removed from a new dataset distributed differently from the original one trained on. We demonstrate the effectiveness of our NTNet on more than a dozen synthetic and biomedical single-cell RNA sequencing datasets, where the NTNet is able to learn the data transformation performed by genetic and drug perturbations on one sample of cells and successfully apply it to another sample of cells to predict treatment outcome.
TrajectoryNet: A Dynamic Optimal Transport Network for Modeling Cellular Dynamics
It is increasingly common to encounter data from dynamic processes captured by static cross-sectional measurements over time, particularly i… (see more)n biomedical settings. Recent attempts to model individual trajectories from this data use optimal transport to create pairwise matchings between time points. However, these methods cannot model continuous dynamics and non-linear paths that entities can take in these systems. To address this issue, we establish a link between continuous normalizing flows and dynamic optimal transport, that allows us to model the expected paths of points over time. Continuous normalizing flows are generally under constrained, as they are allowed to take an arbitrary path from the source to the target distribution. We present TrajectoryNet, which controls the continuous paths taken between distributions to produce dynamic optimal transport. We show how this is particularly applicable for studying cellular dynamics in data from single-cell RNA sequencing (scRNA-seq) technologies, and that TrajectoryNet improves upon recently proposed static optimal transport-based models that can be used for interpolating cellular distributions.
Image-to-image Mapping with Many Domains by Sparse Attribute Transfer
Visualizing structure and transitions in high-dimensional biological data
Kevin R. Moon
David van Dijk
Zheng Wang
Scott Gigante
Daniel B. Burkhardt
William S. Chen
Kristina Yim
Antonia van den Elzen
Matthew Hirn
Ronald R. Coifman
Natalia Ivanova