Publications

Performative Prediction in Time Series: A Case Study
Jennifer Jones
David Langelier
Anthony Reiman
Jonathan Greenland
Kristin Campbell
Reference panel guided topological structure annotation of Hi-C data
Yanlin Zhang
APOE alleles are associated with sex-specific structural differences in brain regions affected in Alzheimer's disease and related dementia
Sylvia Villeneuve
AmanPreet Badhwar
Kimia Shafighi
Chris Zajner
Vaibhav Sharma
Sarah A. Gagliano Taliun
Sali Farhan
Judes Poirier
Alzheimer’s disease is marked by intracellular tau aggregates in the medial temporal lobe (MTL) and extracellular amyloid aggregates in th… (see more)e default network (DN). Here, we examined codependent structural variations between the MTL’s most vulnerable structure, the hippocampus (HC), and the DN at subregion resolution in individuals with Alzheimer’s disease and related dementia (ADRD). By leveraging the power of the approximately 40,000 participants of the UK Biobank cohort, we assessed impacts from the protective APOE ɛ2 and the deleterious APOE ɛ4 Alzheimer’s disease alleles on these structural relationships. We demonstrate ɛ2 and ɛ4 genotype effects on the inter-individual expression of HC-DN co-variation structural patterns at the population level. Across these HC-DN signatures, recurrent deviations in the CA1, CA2/3, molecular layer, fornix’s fimbria, and their cortical partners related to ADRD risk. Analyses of the rich phenotypic profiles in the UK Biobank cohort further revealed male-specific HC-DN associations with air pollution and female-specific associations with cardiovascular traits. We also showed that APOE ɛ2/2 interacts preferentially with HC-DN co-variation patterns in estimating social lifestyle in males and physical activity in females. Our structural, genetic, and phenotypic analyses in this large epidemiological cohort reinvigorate the often-neglected interplay between APOE ɛ2 dosage and sex and link APOE alleles to inter-individual brain structural differences indicative of ADRD familial risk.
Autism incidence and spatial analysis in more than 7 million pupils in English schools: a retrospective, longitudinal, school registry study.
Andres Roman-Urrestarazu
Justin Christopher Yang
R. van Kessel
Varun Warrier
H. Jongsma
Gabriel Gatica-bahamonde
Carrie Allison
F. Matthews
Simon Baron-Cohen
Carol Brayne
Computing Nash Equilibria for Integer Programming Games
Andrea Lodi
João Pedro Pedroso
The recently defined class of integer programming games (IPG) models situations where multiple self-interested decision makers interact, wit… (see more)h their strategy sets represented by a finite set of linear constraints together with integer requirements. Many real-world problems can suitably be fit in this class, and hence anticipating IPG outcomes is of crucial value for policy makers and regulators. Nash equilibria have been widely accepted as the solution concept of a game. Consequently, their computation provides a reasonable prediction of the games outcome. In this paper, we start by showing the computational complexity of deciding the existence of a Nash equilibrium for an IPG. Then, using sufficient conditions for their existence, we develop two general algorithmic approaches that are guaranteed to approximate an equilibrium under mild conditions. We also showcase how our methodology can be changed to determine other equilibria definitions. The performance of our methods is analyzed through computational experiments in a knapsack game, a competitive lot-sizing game, and a kidney exchange game. To the best of our knowledge, this is the first time that equilibria computation methods for general integer programming games have been designed and computationally tested.
Deep learning-enabled anomaly detection for IoT systems
Adel Abusitta 0001
Adel Abusitta
Glaucio H.S. Carvalho
Omar Abdel Wahab
Talal Halabi
Benjamin C. M. Fung
Saja Al-Mamoori
Detecting Languages Unintelligible to Multilingual Models through Local Structure Probes
Providing better language tools for low-resource and endangered languages is imperative for equitable growth. Recent progress with massively… (see more) multilingual pretrained models has proven surprisingly effective at performing zero-shot transfer to a wide variety of languages. However, this transfer is not universal, with many languages not currently understood by multilingual approaches. It is estimated that only 72 languages possess a "small set of labeled datasets" on which we could test a model's performance, the vast majority of languages not having the resources available to simply evaluate performances on. In this work, we attempt to clarify which languages do and do not currently benefit from such transfer. To that end, we develop a general approach that requires only unlabelled text to detect which languages are not well understood by a cross-lingual model. Our approach is derived from the hypothesis that if a model's understanding is insensitive to perturbations to text in a language, it is likely to have a limited understanding of that language. We construct a cross-lingual sentence similarity task to evaluate our approach empirically on 350, primarily low-resource, languages.
Evaluating the Faithfulness of Importance Measures in NLP by Recursively Masking Allegedly Important Tokens and Retraining
To explain NLP models a popular approach is to use importance measures, such as attention, which inform input tokens are important for makin… (see more)g a prediction. However, an open question is how well these explanations accurately reflect a model's logic, a property called faithfulness. To answer this question, we propose Recursive ROAR, a new faithfulness metric. This works by recursively masking allegedly important tokens and then retraining the model. The principle is that this should result in worse model performance compared to masking random tokens. The result is a performance curve given a masking-ratio. Furthermore, we propose a summarizing metric using relative area-between-curves (RACU), which allows for easy comparison across papers, models, and tasks. We evaluate 4 different importance measures on 8 different datasets, using both LSTM-attention models and RoBERTa models. We find that the faithfulness of importance measures is both model-dependent and task-dependent. This conclusion contradicts previous evaluations in both computer vision and faithfulness of attention literature.
Findings of the WMT’22 Shared Task on Large-Scale Machine Translation Evaluation for African Languages
Md Mahfuz Ibn Alam
Antonios Anastasopoulos
Akshita Bhagia
Marta R. Costa-jussa
Jesse Dodge
Fahim Faisal
Christian Federmann
Natalia N. Fedorova
Francisco S. Guzm'an
Sergey Koshelev
Jean Maillard
Vukosi Marivate
Jonathan Mbuya
Alexandre Mourachko
Safiyyah Saleem
Guillaume Wenzek
We present the results of the WMT’22 SharedTask on Large-Scale Machine Translation Evaluation for African Languages. The shared taskinclud… (see more)ed both a data and a systems track, alongwith additional innovations, such as a focus onAfrican languages and extensive human evaluation of submitted systems. We received 14system submissions from 8 teams, as well as6 data track contributions. We report a largeprogress in the quality of translation for Africanlanguages since the last iteration of this sharedtask: there is an increase of about 7.5 BLEUpoints across 72 language pairs, and the average BLEU scores went from 15.09 to 22.60.
Implementing automation in deep brain stimulation: has the time come?
Alfonso Fasano
Improving Passage Retrieval with Zero-Shot Question Generation
Devendra Singh Sachan
Mike Lewis
Mandar Joshi
Armen Aghajanyan
Wen-tau Yih
Luke Zettlemoyer
We propose a simple and effective re-ranking method for improving passage retrieval in open question answering. The re-ranker re-scores retr… (see more)ieved passages with a zero-shot question generation model, which uses a pre-trained language model to compute the probability of the input question conditioned on a retrieved passage. This approach can be applied on top of any retrieval method (e.g. neural or keyword-based), does not require any domain- or task-specific training (and therefore is expected to generalize better to data distribution shifts), and provides rich cross-attention between query and passage (i.e. it must explain every token in the question). When evaluated on a number of open-domain retrieval datasets, our re-ranker improves strong unsupervised retrieval models by 6%-18% absolute and strong supervised models by up to 12% in terms of top-20 passage retrieval accuracy. We also obtain new state-of-the-art results on full open-domain question answering by simply adding the new re-ranker to existing models with no further changes.
K27M in canonical and noncanonical H3 variants occurs in distinct oligodendroglial cell lineages in brain midline gliomas
Selin Jessa
Abdulshakour Mohammadnia
Ashot S. Harutyunyan
Maud Hulswit
Srinidhi Varadharajan
Hussein Lakkis
Nisha Kabir
Zahedeh Bashardanesh
Steven Hébert
Damien Faury
Maria C. Vladoiu
Samantha Worme
Marie Coutelier
Brian Krug
Augusto Faria Andrade
Manav Pathania
Andrea Bajic
Alexander G. Weil
Benjamin Ellezam
Jeffrey Atkinson … (see 14 more)
Roy W. R. Dudley
Jean-Pierre Farmer
Sebastien Perreault
Benjamin A. Garcia
Valérie Larouche
Livia Garzia
Aparna Bhaduri
Keith L. Ligon
Pratiti Bandopadhayay
Michael D. Taylor
Stephen C. Mack
Nada Jabado
Claudia L. Kleinman
Canonical (H3.1/H3.2) and noncanonical (H3.3) histone 3 K27M-mutant gliomas have unique spatiotemporal distributions, partner alterations, a… (see more)nd molecular profiles. The contribution of the cell-of-origin to these differences has been challenging to uncouple from the oncogenic reprogramming induced by the mutation. Here, we perform an integrated analysis of 116 tumors, including single-cell transcriptome and chromatin accessibility, 3D chromatin architecture and epigenomic profiles, and show that K27M-mutant gliomas faithfully maintain chromatin configuration at developmental genes consistent with anatomically distinct oligodendrocyte-precursor-like cells (OPC). H3.3K27M thalamic gliomas map to prosomere 2-derived lineages. In turn, H3.1K27M ACVR1-mutant pontine gliomas uniformly mirror early ventral NKX6-1+/SHH-dependent brainstem OPCs, while H3.3K27M gliomas frequently resemble dorsal PAX3+/BMP-dependent progenitors. Our data suggest a context-specific vulnerability in H3.1K27M-mutant SHH-dependent ventral OPCs, which rely on acquisition of ACVR1 mutations to drive aberrant BMP signaling required for oncogenesis. The unifying action of K27M mutations is to restrict H3K27me3 at PRC2 landing sites, while other epigenetic changes are mainly contingent on the cell-of-origin chromatin state and cycling rate.