Portrait de Mathieu Blanchette

Mathieu Blanchette

Membre académique associé
Directeur et professeur associé, McGill University, École d'informatique
Sujets de recherche
Apprentissage profond
Biologie computationnelle
Réseaux de neurones en graphes

Biographie

Mathieu Blanchette est professeur associé et directeur de l'École d'informatique de l'Université McGill.

Après avoir obtenu un doctorat (Université de Washington, 2002) et un postdoctorat (Université de Californie à Santa Cruz, 2003), il s'est joint à l'École d'informatique de l’Université McGill et a fondé le Laboratoire de génomique computationnelle. Les recherches effectuées par son équipe d’exception ont fait l'objet de plus de 70 publications. Récemment élu membre du Collège de nouveaux chercheurs et créateurs en art et science de la Société royale du Canada, il a été boursier Sloan (2009) et a reçu le prix Outstanding Young Computer Scientist Researcher de l'Association canadienne de l'informatique (2012) ainsi que le prix Chris Overton (2006). Il adore enseigner et superviser les étudiant·e·s, et a d’ailleurs reçu le prix Leo Yaffe pour l'enseignement (2008).

Étudiants actuels

Publications

TULIPS decorate the three-dimensional genome of PFA ependymoma
Michael J. Johnston
John J.Y. Lee
Bo Hu
Ana Nikolic
Elham Hasheminasabgorji
Audrey Baguette
Seungil Paik
Haifen Chen
Sachin Kumar
Carol C.L. Chen
Selin Jessa
Polina Balin
Vernon Fong
Melissa Zwaig
Kulandaimanuvel Antony Michealraj
Xun Chen
Yanlin Zhang
Srinidhi Varadharajan
Pierre Billon
Nikoleta Juretic … (voir 30 de plus)
Craig Daniels
Amulya Nageswara Rao
Caterina Giannini
Eric M. Thompson
Miklos Garami
Peter Hauser
Timea Pocza
Young Shin Ra
Byung-Kyu Cho
Seung-Ki Kim
Kyu-Chang Wang
Ji Yeoun Lee
Wieslawa Grajkowska
Marta Perek-Polnik
Sameer Agnihotri
Stephen Mack
Benjamin Ellezam
Alex Weil
Jeremy Rich
Guillaume Bourque
Jennifer A. Chan
V. Wee Yong
Mathieu Lupien
Jiannis Ragoussis
Claudia Kleinman
Jacek Majewski
Nada Jabado
Michael D. Taylor
Marco Gallo
Learning the Game: Decoding the Differences between Novice and Expert Players in a Citizen Science Game with Millions of Players.
Eddie Cai
Roman Sarrazin-Gendron
Renata Mutalova
Parham Ghasemloo Gheidari
Alexander Butyaev
Gabriel Richard
Sébastien Caisse
Rob Knight
Attila Szantner
Jérôme Waldispühl
In recent years, video games have surged in popularity, attracting millions of players across platforms. Citizen science games (CSGs) levera… (voir plus)ge the processing power of gamers to solve computational and scientific problems. Borderlands Science (BLS) is a mini-game within the mass market game Borderlands 3 that turns multiple sequence alignment (MSA) problems into puzzles. Parallel research demonstrated that BLS players outperformed classical approaches solving small sequence alignment tasks. This study aims to analyze the strategical differences in player solutions in BLS as they gain experience. Through the many collected player solutions from players of different experience level, we gained insights into players’ strategies, differences between expert and non-expert players, and how strategies evolve. We developed a Markov chain trained on solutions from players of different experience levels to understand their actions and outcomes. Results indicate that expert players utilize more gaps and achieve more matches, gradually improving and converging toward unique strategies. Our findings reveal distinct and evolving player strategies. For future citizen science projects, it will be important to consider the identification of player strategies and their evolution over time to improve the game design and data processing.
PhyloGFN: Phylogenetic Inference with Generative Flow Networks
Phylogenetics is a branch of computational biology that studies the evolutionary relationships among biological entities. Its long history a… (voir plus)nd numerous applications notwithstanding, inference of phylogenetic trees from sequence data remains challenging: the high complexity of tree space poses a significant obstacle for the current combinatorial and probabilistic techniques. In this paper, we adopt the framework of generative flow networks (GFlowNets) to tackle two core problems in phylogenetics: parsimony-based and Bayesian phylogenetic inference. Because GFlowNets are well-suited for sampling complex combinatorial structures, they are a natural choice for exploring and sampling from the multimodal posterior distribution over tree topologies and evolutionary distances. We demonstrate that our amortized posterior sampler, PhyloGFN, produces diverse and high-quality evolutionary hypotheses on real benchmark datasets. PhyloGFN is competitive with prior works in marginal likelihood estimation and achieves a closer fit to the target distribution than state-of-the-art variational inference methods. Our code is available at https://github.com/zmy1116/phylogfn.
Improving microbial phylogeny with citizen science within a mass-market video game
Roman Sarrazin-Gendron
Parham Ghasemloo Gheidari
Alexander Butyaev
Timothy Keding
Eddie Cai
Renata Mutalova
Julien Mounthanyvong
Yuxue Zhu
Elena Nazarova
Chrisostomos Drogaris
Kornél Erhart
David Michael Joshua Mathieu Vincent Steven Dan Jonathan Bélanger Bouffard Davidson Falaise Fiset Hebert He
David Michael Joshua Mathieu Vincent Steven Dan Jonathan Seung Jonathan David Steve Ludger Bélanger
David Bélanger
Michael Bouffard
Joshua Davidson
Mathieu Falaise
Vincent Fiset
Steven Hébert … (voir 16 de plus)
Dan Hewitt
Jonathan Huot
Seung Kim
Jonathan Moreau-Genest
David Najjab
Steve Prince
Ludger Saintélien
Amélie Brouillette
Gabriel Richard
Randy Pitchford
Sébastien Caisse
Daniel McDonald
Rob Knight
Attila Szantner
Jérôme Waldispühl
Citizen science video games are designed primarily for users already inclined to contribute to science, which severely limits their accessib… (voir plus)ility for an estimated community of 3 billion gamers worldwide. We created Borderlands Science (BLS), a citizen science activity that is seamlessly integrated within a popular commercial video game played by tens of millions of gamers. This integration is facilitated by a novel game-first design of citizen science games, in which the game design aspect has the highest priority, and a suitable task is then mapped to the game design. BLS crowdsources a multiple alignment task of 1 million 16S ribosomal RNA sequences obtained from human microbiome studies. Since its initial release on 7 April 2020, over 4 million players have solved more than 135 million science puzzles, a task unsolvable by a single individual. Leveraging these results, we show that our multiple sequence alignment simultaneously improves microbial phylogeny estimations and UniFrac effect sizes compared to state-of-the-art computational methods. This achievement demonstrates that hyper-gamified scientific tasks attract massive crowds of contributors and offers invaluable resources to the scientific community.
Posterior inference of Hi-C contact frequency through sampling
Yanlin Zhang
Christopher J. F. Cameron
Hi-C is one of the most widely used approaches to study three-dimensional genome conformations. Contacts captured by a Hi-C experiment are r… (voir plus)epresented in a contact frequency matrix. Due to the limited sequencing depth and other factors, Hi-C contact frequency matrices are only approximations of the true interaction frequencies and are further reported without any quantification of uncertainty. Hence, downstream analyses based on Hi-C contact maps (e.g., TAD and loop annotation) are themselves point estimations. Here, we present the Hi-C interaction frequency sampler (HiCSampler) that reliably infers the posterior distribution of the interaction frequency for a given Hi-C contact map by exploiting dependencies between neighboring loci. Posterior predictive checks demonstrate that HiCSampler can infer highly predictive chromosomal interaction frequency. Summary statistics calculated by HiCSampler provide a measurement of the uncertainty for Hi-C experiments, and samples inferred by HiCSampler are ready for use by most downstream analysis tools off the shelf and permit uncertainty measurements in these analyses without modifications.
Multi-ancestry polygenic risk scores using phylogenetic regularization
Accurately predicting phenotype using genotype across diverse ancestry groups remains a significant challenge in human genetics. Many state-… (voir plus)of-the-art polygenic risk score models are known to have difficulty generalizing to genetic ancestries that are not well represented in their training set. To address this issue, we present a novel machine learning method for fitting genetic effect sizes across multiple ancestry groups simultaneously, while leveraging prior knowledge of the evolutionary relationships among them. We introduce DendroPRS, a machine learning model where SNP effect sizes are allowed to evolve along the branches of the phylogenetic tree capturing the relationship among populations. DendroPRS outperforms existing approaches at two important genotype-to-phenotype prediction tasks: expression QTL analysis and polygenic risk scores. We also demonstrate that our method can be useful for multi-ancestry modelling, both by fitting population-specific effect sizes and by more accurately accounting for covariate effects across groups. We additionally find a subset of genes where there is strong evidence that an ancestry-specific approach improves eQTL modelling.
PERFUMES: pipeline to extract RNA functional motifs and exposed structures
Arnaud Chol
Roman Sarrazin-Gendron
Éric Lécuyer
Jérôme Waldispühl
Abstract Motivation Up to 75% of the human genome encodes RNAs. The function of many non-coding RNAs relies on their ability to fold into 3D… (voir plus) structures. Specifically, nucleotides inside secondary structure loops form non-canonical base pairs that help stabilize complex local 3D structures. These RNA 3D motifs can promote specific interactions with other molecules or serve as catalytic sites. Results We introduce PERFUMES, a computational pipeline to identify 3D motifs that can be associated with observable features. Given a set of RNA sequences with associated binary experimental measurements, PERFUMES searches for RNA 3D motifs using BayesPairing2 and extracts those that are over-represented in the set of positive sequences. It also conducts a thermodynamics analysis of the structural context that can support the interpretation of the predictions. We illustrate PERFUMES’ usage on the SNRPA protein binding site, for which the tool retrieved both previously known binder motifs and new ones. Availability and implementation PERFUMES is an open-source Python package (https://jwgitlab.cs.mcgill.ca/arnaud_chol/perfumes).
Graphylo: A deep learning approach for predicting regulatory DNA and RNA sites from whole-genome multiple alignments
Dongjoon Lim
Changhyun Baek
ARGV: 3D genome structure exploration using augmented reality
Chrisostomos Drogaris
Yanlin Zhang
Éric Zhang
Elena Nazarova
Roman Sarrazin-Gendron
Sélik Wilhelm-Landry
Yan Cyr
Jacek Majewski
Jérôme Waldispühl
Over the past two decades, scientists have increasingly realized the importance of the three-dimensional (3D) genome organization in regulat… (voir plus)ing cellular activity. Hi-C and related experiments yield 2D contact matrices that can be used to infer 3D models of chromosome structure. Visualizing and analyzing genomes in 3D space remains challenging. Here, we present ARGV, an augmented reality 3D Genome Viewer. ARGV contains more than 350 pre-computed and annotated genome structures inferred from Hi-C and imaging data. It offers interactive and collaborative visualization of genomes in 3D space, using standard mobile phones or tablets. A user study comparing ARGV to existing tools demonstrates its benefits.
H3K27me3 spreading organizes canonical PRC1 chromatin architecture to
regulate developmental programs
Brian Krug
Bo Hu
Haifen Chen
Adam Ptack
Xiao Chen
Kristjan H. Gretarsson
Shriya Deshmukh
Nisha Kabir
Augusto Faria Andrade
Elias Jabbour
Ashot S. Harutyunyan
John J. Y. Lee
Maud Hulswit
Damien Faury
Caterina Russo
Xinjing Xu
Michael J. Johnston
Audrey Baguette
Nathan A. Dahl
Alexander G. Weil … (voir 12 de plus)
Benjamin Ellezam
Rola Dali
Khadija Wilson
Benjamin A. Garcia
Rajesh Kumar Soni
Marco Gallo
Michael D. Taylor
Claudia L. Kleinman
Jacek Majewski
Nada Jabado
Chao Lu
Polycomb Repressive Complex 2 (PRC2)-mediated histone H3K27 tri-methylation (H3K27me3) recruits canonical PRC1 (cPRC1) to maintain heterochr… (voir plus)omatin. In early development, polycomb-regulated genes are connected through long-range 3D interactions which resolve upon differentiation. Here, we report that polycomb looping is controlled by H3K27me3 spreading and regulates target gene silencing and cell fate specification. Using glioma-derived H3 Lys-27-Met (H3K27M) mutations as tools to restrict H3K27me3 deposition, we show that H3K27me3 confinement concentrates the chromatin pool of cPRC1, resulting in heightened 3D interactions mirroring chromatin architecture of pluripotency, and stringent gene repression that maintains cells in progenitor states to facilitate tumor development. Conversely, H3K27me3 spread in pluripotent stem cells, following neural differentiation or loss of the H3K36 methyltransferase NSD1, dilutes cPRC1 concentration and dissolves polycomb loops. These results identify the regulatory principles and disease implications of polycomb looping and nominate histone modification-guided distribution of reader complexes as an important mechanism for nuclear compartment organization. The confinement of H3K27me3 at PRC2 nucleation sites without its spreading correlates with increased 3D chromatin interactions. The H3K27M oncohistone concentrates canonical PRC1 that anchors chromatin loop interactions in gliomas, silencing developmental programs. Stem and progenitor cells require factors promoting H3K27me3 confinement, including H3K36me2, to maintain cPRC1 loop architecture. The cPRC1-H3K27me3 interaction is a targetable driver of aberrant self-renewal in tumor cells.
Player-Guided AI outperforms standard AI in Sequence Alignment Puzzles.
Renata Mutalova
Roman Sarrazin-Gendron
Parham Ghasemloo Gheidari
Eddie Cai
Gabriel Richard
Sébastien Caisse
Rob Knight
Attila Szantner
Jérôme Waldispühl
Although Artificial Intelligence (AI) has gained widespread popularity across different fields, it is essential to recognize that AI systems… (voir plus), while impressive, do not consistently exhibit robust generalization, particularly for difficult problems such as the Multiple Sequence Alignment (MSA). In this study, we focus on bridging this performance gap by integrating human solutions into AI training. To illustrate these principles, we leverage data from Borderlands Science, a popular citizen science game in which small instances of the MSA problem are represented as puzzles. Our goal is to leverage the collective intelligence of human players to enhance the capabilities of AI agents. To achieve this, we have developed a Player-guided AI system that enables the AI model to learn from both standard training processes and the solutions provided by players. Our findings demonstrate that incorporating human-annotated information into the AI model improves its performance on puzzle tasks. Furthermore, the Player-guided AI model shows a decrease in noise compared to a pure AI model. This advancement allows for leveraging the model to align new sequences with improved accuracy and effectiveness. Moreover, this research brings attention to the potential of integrating AI and human expertise to address other challenges where the performance of AI models may be unsatisfactory.
Reference panel-guided super-resolution inference of Hi-C data
Yanlin Zhang
Abstract Motivation Accurately assessing contacts between DNA fragments inside the nucleus with Hi-C experiment is crucial for understanding… (voir plus) the role of 3D genome organization in gene regulation. This challenging task is due in part to the high sequencing depth of Hi-C libraries required to support high-resolution analyses. Most existing Hi-C data are collected with limited sequencing coverage, leading to poor chromatin interaction frequency estimation. Current computational approaches to enhance Hi-C signals focus on the analysis of individual Hi-C datasets of interest, without taking advantage of the facts that (i) several hundred Hi-C contact maps are publicly available and (ii) the vast majority of local spatial organizations are conserved across multiple cell types. Results Here, we present RefHiC-SR, an attention-based deep learning framework that uses a reference panel of Hi-C datasets to facilitate the enhancement of Hi-C data resolution of a given study sample. We compare RefHiC-SR against tools that do not use reference samples and find that RefHiC-SR outperforms other programs across different cell types, and sequencing depths. It also enables high-accuracy mapping of structures such as loops and topologically associating domains. Availability and implementation https://github.com/BlanchetteLab/RefHiC.