Publications

Heterogeneous Supervised Topic Models
Hal Daumé III
David Blei
Researchers in the social sciences are often interested in the relationship between text and an outcome of interest, where the goal is to bo… (voir plus)th uncover latent patterns in the text and predict outcomes for unseen texts. To this end, this paper develops the heterogeneous supervised topic model (HSTM), a probabilistic approach to text analysis and prediction. HSTMs posit a joint model of text and outcomes to find heterogeneous patterns that help with both text analysis and prediction. The main benefit of HSTMs is that they capture heterogeneity in the relationship between text and the outcome across latent topics. To fit HSTMs, we develop a variational inference algorithm based on the auto-encoding variational Bayes framework. We study the performance of HSTMs on eight datasets and find that they consistently outperform related methods, including fine-tuned black-box models. Finally, we apply HSTMs to analyze news articles labeled with pro- or anti-tone. We find evidence of differing language used to signal a pro- and anti-tone.
How Transferable Are Features in Convolutional Neural Network Acoustic Models across Languages?
Jessica A.F. Thompson
Marc Schönwiesner
Daniel Willett
Characterization of the representations learned in intermediate layers of deep networks can provide valuable insight into the nature of a ta… (voir plus)sk and can guide the development of well-tailored learning strategies. Here we study convolutional neural network (CNN)-based acoustic models in the context of automatic speech recognition. Adapting a method proposed by [1], we measure the transferability of each layer between English, Dutch and German to assess their language-specificity. We observed three distinct regions of transferability: (1) the first two layers were entirely transferable between languages, (2) layers 2–8 were also highly transferable but we found some evidence of language specificity, (3) the subsequent fully connected layers were more language specific but could be successfully finetuned to the target language. To further probe the effect of weight freezing, we performed follow-up experiments using freeze-training [2]. Our results are consistent with the observation that CNNs converge ‘bottom up’ during training and demonstrate the benefit of freeze training, especially for transfer learning.
IIRC: Incremental Implicitly-Refined Classification
We introduce the "Incremental Implicitly-Refined Classi-fication (IIRC)" setup, an extension to the class incremental learning setup where t… (voir plus)he incoming batches of classes have two granularity levels. i.e., each sample could have a high-level (coarse) label like "bear" and a low-level (fine) label like "polar bear". Only one label is provided at a time, and the model has to figure out the other label if it has already learnfed it. This setup is more aligned with real-life scenarios, where a learner usually interacts with the same family of entities multiple times, discovers more granularity about them, while still trying not to forget previous knowledge. Moreover, this setup enables evaluating models for some important lifelong learning challenges that cannot be easily addressed under the existing setups. These challenges can be motivated by the example "if a model was trained on the class bear in one task and on polar bear in another task, will it forget the concept of bear, will it rightfully infer that a polar bear is still a bear? and will it wrongfully associate the label of polar bear to other breeds of bear?". We develop a standardized benchmark that enables evaluating models on the IIRC setup. We evaluate several state-of-the-art lifelong learning algorithms and highlight their strengths and limitations. For example, distillation-based methods perform relatively well but are prone to incorrectly predicting too many labels per image. We hope that the proposed setup, along with the benchmark, would provide a meaningful problem setting to the practitioners
Image Retrieval from Contextual Descriptions
Vibhav Vineet
Edoardo Ponti
The ability to integrate context, including perceptual and temporal cues, plays a pivotal role in grounding the meaning of a linguistic utte… (voir plus)rance. In order to measure to what extent current vision-and-language models master this ability, we devise a new multimodal challenge, Image Retrieval from Contextual Descriptions (ImageCoDe). In particular, models are tasked with retrieving the correct image from a set of 10 minimally contrastive candidates based on a contextual description.As such, each description contains only the details that help distinguish between images.Because of this, descriptions tend to be complex in terms of syntax and discourse and require drawing pragmatic inferences. Images are sourced from both static pictures and video frames.We benchmark several state-of-the-art models, including both cross-encoders such as ViLBERT and bi-encoders such as CLIP, on ImageCoDe.Our results reveal that these models dramatically lag behind human performance: the best variant achieves an accuracy of 20.9 on video frames and 59.4 on static pictures, compared with 90.8 in humans.Furthermore, we experiment with new model variants that are better equipped to incorporate visual and temporal context into their representations, which achieve modest gains. Our hope is that ImageCoDE will foster progress in grounded language understanding by encouraging models to focus on fine-grained visual differences.
Improved DC-Free Run-Length Limited 4B6B Codes for Concatenated Schemes
Elie Ngomseu Mambou
Thibaud Tonnellier
Warren J. Gross
In this letter, we introduce a class of improved DC-free 4B6B codes in terms of error correction capabilities for a serially concatenated ar… (voir plus)chitecture. There are billions of different codebooks that can be derived from the 16 codewords contained in the traditional 4B6B code as per the IEEE 802.15.7 standard for visible light communication (VLC). These codebooks can be classified based on distances properties which determine their error correction performances. The traditional 4B6B code is suitable for hard-decision decoding, however, when a soft decoder is used like in a serially concatenated architecture, that code becomes obsolete. Simulations show that the proposed 4B6B code concatenated with forward error correction (FEC) codes, has better performance compared to state-of-the-art schemes such as the original 4B6B code, the enhanced Miller code, the Manchester code, the 5B10B code and the (0,4) 2/3 RLL code.
Improved Detection of Chronic Obstructive Pulmonary Disease at Chest CT Using the Mean Curvature of Isophotes
Peter Savadjiev
Benoit Gallix
Morteza Rezanejad
Sahir Bhatnagar
Alexandre Semionov
Reza Forghani
Caroline Reinhold
David H. Eidelman
Ronald J. Dandurand
Purpose To determine if the mean curvature of isophotes (MCI), a standard computer vision technique, can be used to improve detection of chr… (voir plus)onic obstructive pulmonary disease (COPD) at chest CT. Materials and Methods In this retrospective study, chest CT scans were obtained in 243 patients with COPD and 31 controls (among all 274: 151 women [mean age, 70 years; range, 44-90 years] and 123 men [mean age, 71 years; range, 29-90 years]) from two community practices between 2006 and 2019. A convolutional neural network (CNN) architecture was trained on either CT images or CT images transformed through the MCI algorithm. Separately, a linear classification based on a single feature derived from the MCI computation (called hMCI1) was also evaluated. All three models were evaluated with cross-validation, using precision-macro and recall-macro metrics, that is, the mean of per-class precision and recall values, respectively (the latter being equivalent to balanced accuracy). Results Linear classification based on hMCI1 resulted in a higher recall-macro relative to the CNN trained and applied on CT images (0.85 [95% CI: 0.84, 0.86] vs 0.77 [95% CI: 0.75, 0.79]) but with a similar reduction in precision-macro (0.66 [95% CI: 0.65, 0.67] vs 0.77 [95% CI: 0.75, 0.79]). The CNN model trained and applied on MCI-transformed images had a higher recall-macro (0.85 [95% CI: 0.83, 0.87] vs 0.77 [95% CI: 0.75, 0.79]) and precision-macro (0.85 [95% CI: 0.83, 0.87] vs 0.77 [95% CI: 0.75, 0.79]) relative to the CNN trained and applied on CT images. Conclusion The MCI algorithm may be valuable toward the automated detection and diagnosis of COPD on chest CT scans as part of a CNN-based pipeline or with stand-alone features.Keywords: Chronic Obstructive Pulmonary Disease, Quantification, Lung, CT Supplemental material is available for this article. See also the invited commentary by Vannier in this issue.© RSNA, 2021.
Inferring global-scale temporal latent topics from news reports to predict public health interventions for COVID-19
Zhi Wen
Guido Powell
Imane Chafi
David L. Buckeridge
The COVID-19 pandemic has highlighted the importance of non-pharmacological interventions (NPI) for controlling epidemics of emerging infect… (voir plus)ious diseases. Despite their importance, NPI have been monitored mainly through the manual efforts of volunteers. This approach hinders measurement of the NPI effectiveness and development of evidence to guide their use to control the global pandemic. We present EpiTopics, a machine learning approach to support automation of the NPI prediction and monitoring at both the document-level and country-level by mining the vast amount of unlabelled news reports on COVID-19. EpiTopics uses a 3-stage, transfer-learning algorithm to classify documents according to NPI categories, relying on topic modelling to support result interpretation. We identified 25 interpretable topics under 4 distinct and coherent COVID-related themes. Importantly, the use of these topics resulted in significant improvements over alternative automated methods in predicting the NPIs in labelled documents and in predicting country-level NPIs for 42 countries.
Interpolation Consistency Training for Semi-Supervised Learning
Juho Kannala
David Lopez-Paz
Arno Solin
Investigating the Performance of Transformer-Based NLI Models on Presuppositional Inferences
Jackie CK Cheung
Presuppositions are assumptions that are taken for granted by an utterance, and identifying them is key to a pragmatic interpretation of lan… (voir plus)guage. In this paper, we investigate the capabilities of transformer models to perform NLI on cases involving presupposition. First, we present simple heuristics to create alternative “contrastive” test cases based on the ImpPres dataset and investigate the model performance on those test cases. Second, to better understand how the model is making its predictions, we analyze samples from sub-datasets of ImpPres and examine model performance on them. Overall, our findings suggest that NLI-trained transformer models seem to be exploiting specific structural and lexical cues as opposed to performing some kind of pragmatic reasoning.
Joint Selection using Deep Reinforcement Learning for Skeleton-based Activity Recognition
Bahareh Nikpour
<div>Skeleton based human activity recognition has attracted lots of attention due to its wide range of applications. Skeleton data in… (voir plus)cludes two or three dimensional coordinates of body joints. All of the body joints are not effective in recognizing different activities, so finding key joints within a video and across different activities has a significant role in improving the performance. In this paper we propose a novel framework that performs joint selection in skeleton video frames for the purpose of human activity recognition. To this end, we formulate the joint selection problem as a Markov Decision Process (MDP) where we employ deep reinforcement learning to find the most informative joints per frame. The proposed joint selection method is a general framework that can be employed to improve human activity classification methods. Experimental results on two benchmark activity recognition data sets using three different classifiers demonstrate effectiveness of the proposed joint selection method.</div>
KNIFE: Kernelized-Neural Differential Entropy Estimation
Georg Pichler
Pierre Colombo
Malik Boudiaf
Gunther Koliander
Mutual Information (MI) has been widely used as a loss regularizer for training neural networks. This has been particularly effective when l… (voir plus)earn dis-entangled or compressed representations of high dimensional data. However, differential entropy (DE), another fundamental measure of information, has not found widespread use in neural network training. Although DE offers a potentially wider range of applications than MI, off-the-shelf DE estimators are either non differentiable, computationally intractable or fail to adapt to changes in the underlying distribution. These drawbacks prevent them from being used as regularizers in neural networks training. To address shortcomings in previously proposed estimators for DE, here we introduce K NIFE , a fully parameterized, differentiable kernel-based estimator of DE. The flexibility of our approach also allows us to construct K NIFE -based estimators for conditional (on either discrete or continuous variables) DE, as well as MI. We empirically validate our method on high-dimensional synthetic data and further apply it to guide the training of neural networks for real-world tasks. Our experiments on a large variety of tasks, including visual domain adaptation, textual fair classification, and textual fine-tuning demonstrate the effectiveness of K NIFE - based estimation. Code can be found at https: //github.com/g-pichler/knife .
LAGr: Label Aligned Graphs for Better Systematic Generalization in Semantic Parsing
Semantic parsing is the task of producing a structured meaning representation for natural language utterances or questions. Recent research … (voir plus)has pointed out that the commonly-used sequence-to-sequence (seq2seq) semantic parsers struggle to generalize systematically, i.e. to handle examples that require recombining known knowledge in novel settings. In this work, we show that better systematic generalization can be achieved by producing the meaning representation (MR) directly as a graph and not as a sequence. To this end we propose LAGr, the Labeling Aligned Graphs algorithm that produces semantic parses by predicting node and edge labels for a complete multi-layer input-aligned graph. The strongly-supervised LAGr algorithm requires aligned graphs as inputs, whereas weakly-supervised LAGr infers alignments for originally unaligned target graphs using an approximate MAP inference procedure. On the COGS and CFQ compositional generalization benchmarks the strongly- and weakly- supervised LAGr algorithms achieve significant improvements upon the baseline seq2seq parsers.