Publications

Tensor of Quantitative Equational Theories.
Giorgio Bacci
Radu Mardare
Gordon Plotkin
Test Sample Accuracy Scales with Training Sample Density in Neural Networks
Andrea Vedaldi
Balaji Lakshminarayanan
Intuitively, one would expect accuracy of a trained neural network's prediction on test samples to correlate with how densely the samples ar… (see more)e surrounded by seen training samples in representation space. We find that a bound on empirical training error smoothed across linear activation regions scales inversely with training sample density in representation space. Empirically, we verify this bound is a strong predictor of the inaccuracy of the network's prediction on test samples. For unseen test sets, including those with out-of-distribution samples, ranking test samples by their local region's error bound and discarding samples with the highest bounds raises prediction accuracy by up to 20% in absolute terms for image classification datasets, on average over thresholds.
The role of case importation in explaining differences in early SARS-CoV-2 transmission dynamics in Canada—A mathematical modeling study of surveillance data
Arnaud Godin
Yiqing Xia
David L Buckeridge
Sharmistha Mishra
Dirk Douwes-Schultz
Maxime Lavigne
Mélanie Drolet
Alexandra M Schmidt
Marc Brisson
Mathieu Maheu-Giroux
The Topic Confusion Task: A Novel Scenario for Authorship Attribution
Jackie CK Cheung
Benjamin C. M. Fung
Authorship attribution is the problem of identifying the most plausible author of an anonymous text from a set of candidate authors. Researc… (see more)hers have investigated same-topic and cross-topic scenarios of authorship attribution, which differ according to whether unseen topics are used in the testing phase. However, neither scenario allows us to explain whether errors are caused by failure to capture authorship style, by the topic shift or by other factors. Motivated by this, we propose the topic confusion task, where we switch the author-topic config-uration between training and testing set. This setup allows us to probe errors in the attribution process. We investigate the accuracy and two error measures: one caused by the models’ confusion by the switch because the features capture the topics, and one caused by the features’ inability to capture the writing styles, leading to weaker models. By evaluating different features, we show that stylometric features with part-of-speech tags are less susceptible to topic variations and can increase the accuracy of the attribution process. We further show that combining them with word-level n - grams can outperform the state-of-the-art technique in the cross-topic scenario. Finally, we show that pretrained language models such as BERT and RoBERTa perform poorly on this task, and are outperformed by simple n -gram features.
Topological Analysis of Single-Cell Hierarchy Reveals Inflammatory Glial Landscape of Macular Degeneration
Manik Kuchroo
Marcello DiStasio
Eric Song
Eda Calapkulu
Maryam Ige
Amar H. Sheth
Madhvi Menon
Abhinav Godavarthi
Yu Xing
Scott Gigante
Holly Steach
Janhavi Narain
George Mourgkos
Rahul M. Dhodapkar
Matthew J. Hirn
Bastian Rieck … (see 3 more)
Brian P. Hafler
Topological Analysis of Single-Cell Hierarchy Reveals Inflammatory Glial Landscape of Macular Degeneration
Manik Kuchroo
Marcello DiStasio
Eric Song
Eda Calapkulu
Maryam Ige
Amar H. Sheth
Madhvi Menon
Abhinav Godavarthi
Yu Xing
Scott Gigante
Holly Steach
Janhavi Narain
George Mourgkos
Rahul M. Dhodapkar
Matthew J. Hirn
Bastian Rieck … (see 3 more)
Brian P. Hafler
Toward Tweet-Mining Framework for Extracting Terrorist Attack-Related Information and Reporting
Farkhund Iqbal
Rabia Batool
Benjamin C. M. Fung
Saiqa Aleem
Ahmed Abbasi
Abdul Rehman Javed
The widespread popularity of social networking is leading to the adoption of Twitter as an information dissemination tool. Existing research… (see more) has shown that information dissemination over Twitter has a much broader reach than traditional media and can be used for effective post-incident measures. People use informal language on Twitter, including acronyms, misspelled words, synonyms, transliteration, and ambiguous terms. This makes incident-related information extraction a non-trivial task. However, this information can be valuable for public safety organizations that need to respond in an emergency. This paper proposes an early event-related information extraction and reporting framework that monitors Twitter streams synthesizes event-specific information, e.g., a terrorist attack, and alerts law enforcement, emergency services, and media outlets. Specifically, the proposed framework, Tweet-to-Act (T2A), employs word embedding to transform tweets into a vector space model and then utilizes the Word Mover’s Distance (WMD) to cluster tweets for the identification of incidents. To extract reliable and valuable information from a large dataset of short and informal tweets, the proposed framework employs sequence labeling with bidirectional Long Short-Term Memory based Recurrent Neural Networks (bLSTM-RNN). Extensive experimental results suggest that our proposed framework, T2A, outperforms other state-of-the-art methods that use vector space modeling and distance calculation techniques, e.g., Euclidean and Cosine distance. T2A achieves an accuracy of 96% and an F1-score of 86.2% on real-life datasets.
Towards a Trace-Preserving Tensor Network Representation of Quantum Channels
Siddarth Srinivasan
Sandesh M. Adhikary
Bibek Pokharel
Byron Boots
The problem of characterizing quantum channels arises in a number of contexts such as quantum process tomography and quantum error correctio… (see more)n. However, direct approaches to parameterizing and optimizing the Choi matrix representation of quantum channels face a curse of dimensionality: the number of parameters scales exponentially in the number of qubits. Recently, Torlai et al. [2020] proposed using locally purified density operators (LPDOs), a tensor network representation of Choi matrices, to overcome the unfavourable scaling in parameters. While the LPDO structure allows it to satisfy a ‘complete positivity’ (CP) constraint required of physically valid quantum channels, it makes no guarantees about a similarly required ‘trace preservation’ (TP) constraint. In practice, the TP constraint is violated, and the learned quantum channel may even be trace-increasing, which is non-physical. In this work, we present the problem of optimizing over TP LPDOs, discuss two approaches to characterizing the TP constraints on LPDOs, and outline the next steps for developing an optimization scheme.
qu an tph ] 10 O ct 2 01 1 Quantum Communication in Rindler Spacetime
Kamil Brádler
P. Hayden
A state that an inertial observer in Minkowski space perceiv es to be the vacuum will appear to an accelerating observer to be a thermal ba … (see more)th of radiation. We study the impact of this Davies-Fulling-Unruh noise on comm unication, particularly quantum communication from an inertial sender to an ac celerating observer and private communication between two inertial observers i n the presence of an accelerating eavesdropper. In both cases, we establish com pact, tractable formulas for the associated communication capacities assuming enco dings that allow a single excitation in one of a fixed number of modes per use of the co mmunications channel. Our contributions include a rigorous presentatio n of the general theory of the private quantum capacity as well as a detailed analysis o f the structure of these channels, including their group-theoretic properties and proof that they are conjugate degradable. Connections between the Unruh channel a d optical amplifiers are also discussed.
Uncovering the Folding Landscape of RNA Secondary Structure Using Deep Graph Embeddings
Egbert Castro
Andrew Benz
Biomolecular graph analysis has recently gained much attention in the emerging field of geometric deep learning. Here we focus on organizing… (see more) biomolecular graphs in ways that expose meaningful relations and variations between them. We propose a geometric scattering autoencoder (GSAE) network for learning such graph embeddings. Our embedding network first extracts rich graph features using the recently proposed geometric scattering transform. Then, it leverages a semi-supervised variational autoencoder to extract a low-dimensional embedding that retains the information in these features that enable prediction of molecular properties as well as characterize graphs. We show that GSAE organizes RNA graphs both by structure and energy, accurately reflecting bistable RNA structures. Also, the model is generative and can sample new folding trajectories.
Understanding Quantum Software Engineering Challenges An Empirical Study on Stack Exchange Forums and GitHub Issues
Mohamed Raed El aoun
Heng Li
Moses Openja
With the advance in quantum computing, quantum software becomes critical for exploring the full potential of quantum computing systems. Rece… (see more)ntly, quantum software engineering (QSE) becomes an emerging area attracting more and more attention. However, it is not clear what are the challenges and opportunities of quantum computing facing the software engineering community. This work aims to understand the QSE-related challenges perceived by developers. We perform an empirical study on Stack Exchange forums where developers post-QSE-related questions & answers and Github issue reports where developers raise QSE-related issues in practical quantum computing projects. Based on an existing taxonomy of question types on Stack Overflow, we first perform a qualitative analysis of the types of QSE-related questions asked on Stack Exchange forums. We then use automated topic modeling to uncover the topics in QSE-related Stack Exchange posts and GitHub issue reports. Our study highlights some particularly challenging areas of QSE that are different from that of traditional software engineering, such as explaining the theory behind quantum computing code, interpreting quantum program outputs, and bridging the knowledge gap between quantum computing and classical computing, as well as their associated opportunities.
Unifying Likelihood-free Inference with Black-box Sequence Design and Beyond