Publications

MOT: A Multi-Omics Transformer for Multiclass Classification Tumour Types Predictions
Mazid Osseni
Franccois Laviolette
J. Corbeil
Refactoring practices in the context of data-intensive systems
Biruk Asmare Muse
Giuliano Antoniol
Learning to Substitute Ingredients in Recipes
Bahare Fatemi
Quentin Duval
Rohit Girdhar
Adriana Romero
Recipe personalization through ingredient substitution has the potential to help people meet their dietary needs and preferences, avoid pote… (voir plus)ntial allergens, and ease culinary exploration in everyone's kitchen. To address ingredient substitution, we build a benchmark, composed of a dataset of substitution pairs with standardized splits, evaluation metrics, and baselines. We further introduce Graph-based Ingredient Substitution Module (GISMo), a novel model that leverages the context of a recipe as well as generic ingredient relational information encoded within a graph to rank plausible substitutions. We show through comprehensive experimental validation that GISMo surpasses the best performing baseline by a large margin in terms of mean reciprocal rank. Finally, we highlight the benefits of GISMo by integrating it in an improved image-to-recipe generation pipeline, enabling recipe personalization through user intervention. Quantitative and qualitative results show the efficacy of our proposed system, paving the road towards truly personalized cooking and tasting experiences.
Score-based Diffusion Models in Function Space
Nikola B. Kovachki
R. Baptista
Kamyar Azizzadenesheli
Jean Kossaifi
Jiaming Song
Karsten Kreis
Jan Kautz
Christopher Pal
Arash Vahdat
Animashree Anandkumar
The Stable Entropy Hypothesis and Entropy-Aware Decoding: An Analysis and Algorithm for Robust Natural Language Generation
Timothy J. O'Donnell
Jason Aaron Edward Weston
Jackie C.K.Cheung
State-of-the-art language generation models can degenerate when applied to open-ended generation problems such as text completion, story gen… (voir plus)eration, or dialog modeling. This degeneration usually shows up in the form of incoherence, lack of vocabulary diversity, and self-repetition or copying from the context. In this paper, we postulate that ``human-like'' generations usually lie in a narrow and nearly flat entropy band, and violation of these entropy bounds correlates with degenerate behavior. Our experiments show that this stable narrow entropy zone exists across models, tasks, and domains and confirm the hypothesis that violations of this zone correlate with degeneration. We then use this insight to propose an entropy-aware decoding algorithm that respects these entropy bounds resulting in less degenerate, more contextual, and"human-like"language generation in open-ended text generation settings.
DEUP: Direct Epistemic Uncertainty Prediction
Epistemic Uncertainty is a measure of the lack of knowledge of a learner which diminishes with more evidence. While existing work focuses on… (voir plus) using the variance of the Bayesian posterior due to parameter uncertainty as a measure of epistemic uncertainty, we argue that this does not capture the part of lack of knowledge induced by model misspecification. We discuss how the excess risk, which is the gap between the generalization error of a predictor and the Bayes predictor, is a sound measure of epistemic uncertainty which captures the effect of model misspecification. We thus propose a principled framework for directly estimating the excess risk by learning a secondary predictor for the generalization error and subtracting an estimate of aleatoric uncertainty, i.e., intrinsic unpredictability. We discuss the merits of this novel measure of epistemic uncertainty, and highlight how it differs from variance-based measures of epistemic uncertainty and addresses its major pitfall. Our framework, Direct Epistemic Uncertainty Prediction (DEUP) is particularly interesting in interactive learning environments, where the learner is allowed to acquire novel examples in each round. Through a wide set of experiments, we illustrate how existing methods in sequential model optimization can be improved with epistemic uncertainty estimates from DEUP, and how DEUP can be used to drive exploration in reinforcement learning. We also evaluate the quality of uncertainty estimates from DEUP for probabilistic image classification and predicting synergies of drug combinations.
Interpersonal attunement in social interactions: from collective psychophysiology to inter-personalized psychiatry and beyond
Dimitris Bolis
Leonhard Schilbach
In this article, we analyse social interactions, drawing on diverse points of views, ranging from dialectics, second-person neuroscience and… (voir plus) enactivism to dynamical systems, active inference and machine learning. To this end, we define interpersonal attunement as a set of multi-scale processes of building up and materializing social expectations—put simply, anticipating and interacting with others and ourselves. While cultivating and negotiating common ground, via communication and culture-building activities, are indispensable for the survival of the individual, the relevant multi-scale mechanisms have been largely considered in isolation. Here, collective psychophysiology, we argue, can lend itself to the fine-tuned analysis of social interactions, without neglecting the individual. On the other hand, an interpersonal mismatch of expectations can lead to a breakdown of communication and social isolation known to negatively affect mental health. In this regard, we review psychopathology in terms of interpersonal misattunement, conceptualizing psychiatric disorders as disorders of social interaction, to describe how individual mental health is inextricably linked to social interaction. By doing so, we foresee avenues for an inter-personalized psychiatry, which moves from a static spectrum of disorders to a dynamic relational space, focusing on how the multi-faceted processes of social interaction can help to promote mental health. This article is part of the theme issue ‘Concepts in interaction: social engagement and inner experiences’.
Limitations of Information-Theoretic Generalization Bounds for Gradient Descent Methods in Stochastic Convex Optimization
MAHDI HAGHIFAM
Borja Rodr'iguez-G'alvez
Ragnar Thobaben
Mikael Skoglund
Daniel M. Roy
Preclinical-to-clinical Anti-cancer Drug Response Prediction and Biomarker Identification Using TINDL
Lixuan Wei
Liewei Wang
Junmei Cairns
Prediction of the response of cancer patients to different treatments and identification of biomarkers of drug response are two major goals … (voir plus)of individualized medicine. Here, we developed a deep learning framework called TINDL, completely trained on preclinical cancer cell lines (CCLs), to predict the response of cancer patients to different treatments. TINDL utilizes a tissue-informed normalization to account for the tissue type and cancer type of the tumors and to reduce the statistical discrepancies between CCLs and patient tumors. Moreover, by making the deep learning black box interpretable, this model identifies a small set of genes whose expression levels are predictive of drug response in the trained model, enabling identification of biomarkers of drug response. Using data from two large databases of CCLs and cancer tumors, we showed that this model can distinguish between sensitive and resistant tumors for 10 (out of 14) drugs, outperforming various other machine learning models. In addition, our small interfering RNA (siRNA) knockdown experiments on 10 genes identified by this model for one of the drugs (tamoxifen) confirmed that tamoxifen sensitivity is substantially influenced by all of these genes in MCF7 cells, and seven of these genes in T47D cells. Furthermore, genes implicated for multiple drugs pointed to shared mechanism of action among drugs and suggested several important signaling pathways. In summary, this study provides a powerful deep learning framework for prediction of drug response and identification of biomarkers of drug response in cancer. The code can be accessed at https://github.com/ddhostallero/tindl.
Characterization of inpaint residuals in interferometric measurements of the epoch of reionization
Michael Pagano
Jing Liu
Adrian Liu
Nicholas S. Kern
Aaron Ewall-Wice
Philip Bull
Robert Pascua
Zara Abdurashidova
Tyrone Adams
James E. Aguirre
Paul Alexander
Zaki S. Ali
Rushelle Baartman
Yanga Balfour
Adam P. Beardsley
Gianni Bernardi
Tashalee S. Billings
Judd D. Bowman
Richard F. Bradley … (voir 58 de plus)
Jacob Burba
Steven Carey
Chris L. Carilli
Carina Cheng
David R. DeBoer
Eloy de Lera Acedo
Matt Dexter
Joshua S. Dillon
Nico Eksteen
John Ely
Nicolas Fagnoni
Randall Fritz
Steven R. Furlanetto
Kingsley Gale-Sides
Brian Glendenning
Deepthi Gorthi
Bradley Greig
Jasper Grobbelaar
Ziyaad Halday
Bryna J. Hazelton
Jacqueline N. Hewitt
Jack Hickish
Daniel C. Jacobs
Austin Julius
MacCalvin Kariseb
Joshua Kerrigan
Piyanat Kittiwisit
Saul A. Kohn
Matthew Kolopanis
Adam Lanman
Paul La Plante
Anita Loots
David Harold Edward MacMahon
Lourence Malan
Cresshim Malgas
Keith Malgas
Bradley Marero
Zachary E. Martinot
Andrei Mesinger
Mathakane Molewa
Miguel F. Morales
Tshegofalang Mosiane
Abraham R. Neben
Bojan Nikolic
Hans Nuwegeld
Aaron R. Parsons
Nipanjana Patra
Samantha Pieterse
Nima Razavi-Ghods
James Robnett
Kathryn Rosie
Peter Sims
Craig Smith
Hilton Swarts
Nithyanandan Thyagarajan
Pieter van Wyngaarden
Peter K. G. Williams
Haoxuan Zheng
Radio Frequency Interference (RFI) is one of the systematic challenges preventing 21cm interferometric instruments from detecting the Epoch … (voir plus)of Reionization. To mitigate the effects of RFI on data analysis pipelines, numerous inpaint techniques have been developed to restore RFI corrupted data. We examine the qualitative and quantitative errors introduced into the visibilities and power spectrum due to inpainting. We perform our analysis on simulated data as well as real data from the Hydrogen Epoch of Reionization Array (HERA) Phase 1 upper limits. We also introduce a convolutional neural network that capable of inpainting RFI corrupted data in interferometric instruments. We train our network on simulated data and show that our network is capable at inpainting real data without requiring to be retrained. We find that techniques that incorporate high wavenumbers in delay space in their modeling are best suited for inpainting over narrowband RFI. We also show that with our fiducial parameters Discrete Prolate Spheroidal Sequences (DPSS) and CLEAN provide the best performance for intermittent ``narrowband'' RFI while Gaussian Progress Regression (GPR) and Least Squares Spectral Analysis (LSSA) provide the best performance for larger RFI gaps. However we caution that these qualitative conclusions are sensitive to the chosen hyperparameters of each inpainting technique. We find these results to be consistent in both simulated and real visibilities. We show that all inpainting techniques reliably reproduce foreground dominated modes in the power spectrum. Since the inpainting techniques should not be capable of reproducing noise realizations, we find that the largest errors occur in the noise dominated delay modes. We show that in the future, as the noise level of the data comes down, CLEAN and DPSS are most capable of reproducing the fine frequency structure in the visibilities of HERA data.
Language Decision Transformers with Exponential Tilt for Interactive Text Environments
Nicolas Gontier
Pau Rodríguez
Issam Hadj Laradji
Christopher Pal
Restoring the missing person to personalized medicine and precision psychiatry
Ana Gómez-Carrillo
Vincent Paquin
Laurence J. Kirmayer
Precision psychiatry has emerged as part of the shift to personalized medicine and builds on frameworks such as the U.S. National Institute … (voir plus)of Mental Health Research Domain Criteria (RDoC), multilevel biological “omics” data and, most recently, computational psychiatry. The shift is prompted by the realization that a one-size-fits all approach is inadequate to guide clinical care because people differ in ways that are not captured by broad diagnostic categories. One of the first steps in developing this personalized approach to treatment was the use of genetic markers to guide pharmacotherapeutics based on predictions of pharmacological response or non-response, and the potential risk of adverse drug reactions. Advances in technology have made a greater degree of specificity or precision potentially more attainable. To date, however, the search for precision has largely focused on biological parameters. Psychiatric disorders involve multi-level dynamics that require measures of phenomenological, psychological, behavioral, social structural, and cultural dimensions. This points to the need to develop more fine-grained analyses of experience, self-construal, illness narratives, interpersonal interactional dynamics, and social contexts and determinants of health. In this paper, we review the limitations of precision psychiatry arguing that it cannot reach its goal if it does not include core elements of the processes that give rise to psychopathological states, which include the agency and experience of the person. Drawing from contemporary systems biology, social epidemiology, developmental psychology, and cognitive science, we propose a cultural-ecosocial approach to integrating precision psychiatry with person-centered care.