Portrait de Fernando Diaz n'est pas disponible

Fernando Diaz

Membre affilié
Professeur agrégé, Carnegie Mellon University, École d'informatique, Language Technologies Institutes
Professeur associé, McGill University, École d'informatique
Chercheur scientifique, Google Pittsburgh
Sujets de recherche
Recherche d'information
Systèmes de recommandation

Biographie

Fernando Diaz est professeur agrégé à l'École d'informatique de l’Université Carnegie Mellon. Il est aussi chercheur scientifique à Google (Pittsburgh) ainsi que membre agrégé de l'École d'informatique de l'Université McGill.

Son principal intérêt de recherche est l’extraction d'information, c'est-à-dire l'étude formelle de la recherche de petits fragments d'information dans de grandes collections de données. L'exemple le plus familier d’extraction d'information est la recherche sur le Web, où les utilisateur·rice·s recherchent à travers une collection de pages Web une ou quelques pages pertinentes. Cependant, la recherche d'information va bien au-delà, et comprend par exemple la recherche interlingue, la personnalisation, la recherche sur le bureau et la recherche interactive. Au fil de ses travaux, Fernando Diaz a exploré les approches distribuées de recherche d'information sur le Web, la recherche interactive et à facettes, les modèles temporels à partir de nouvelles et de requêtes, la recherche d'information multilingue, les méthodes de recherche basées sur des graphiques et l'exploitation d'information à partir de multiples corpus.

Dans sa thèse, il a étudié la relation entre le regroupement de documents et la notation des documents en vue de leur extraction à l'aide de méthodes d'apprentissage automatique et de statistiques. Il a donc mis au point un algorithme d'autoévaluation et d'auto-ajustement du système qui améliore considérablement la performance des algorithmes de récupération dans une variété de corpus.

Étudiants actuels

Doctorat - McGill
Superviseur⋅e principal⋅e :

Publications

Preference-Based Offline Evaluation
C. Clarke
Negar Arabzadeh
A core step in production model research and development involves the offline evaluation of a system before production deployment. Tradition… (voir plus)al offline evaluation of search, recommender, and other systems involves gathering item relevance labels from human editors. These labels can then be used to assess system performance using offline evaluation metrics. Unfortunately, this approach does not work when evaluating highly effective ranking systems, such as those emerging from the advances in machine learning. Recent work demonstrates that moving away from pointwise item and metric evaluation can be a more effective approach to the offline evaluation of systems. This tutorial, intended for both researchers and practitioners, reviews early work in preference-based evaluation and covers recent developments in detail.
Recall as a Measure of Ranking Robustness
Bhaskar Mitra
A Survey of Diversification Metrics and Approaches in Retrieval Systems: From the Perspective of Search and Recommendation
Yansen Zhang
Fuyuan Lyu
Xue Liu
Diversifying search results is an important research topic in retrieval systems in order to satisfy both the various interests of customers … (voir plus)and the equal market exposure of providers. There has been a growing attention on diversity-aware research during recent years, accompanied by a proliferation of literature on methods to promote diversity in search and recommendation. However, the diversity-aware studies in retrieval systems lack a systematic organization and are rather fragmented. In this survey, we are the first to propose a unified taxonomy for classifying the metrics and approaches of diversification in both search and recommendation, which are two of the most extensively researched fields of retrieval systems. We begin the survey with a brief discussion of why diversity is important in retrieval systems
Redefining Relationships in Music
Christian Detweiler
Beth Coleman
Lieke Dom
Chris Donahue
Jesse Engel
Cheng-Zhi Anna Huang
Larry James
Ethan Manilow
Amanda McCroskery
Kyle Pedersen
Pamela Peter-Agbia
Thomas Robert
Marco Zamarato
Ben Zevenbergen
AI tools increasingly shape how we discover, make and experience music. While these tools can have the potential to empower creativity, they… (voir plus) may fundamentally redefine relationships between stakeholders, to the benefit of some and the detriment of others. In this position paper, we argue that these tools will fundamentally reshape our music culture, with profound effects (for better and for worse) on creators, consumers and the commercial enterprises that often connect them. By paying careful attention to emerging Music AI technologies and developments in other creative domains and understanding the implications, people working in this space could decrease the possible negative impacts on the practice, consumption and meaning of music. Given that many of these technologies are already available, there is some urgency in conducting analyses of these technologies now. It is important that people developing and working with these tools address these issues now to help guide their evolution to be equitable and empower creativity. We identify some potential risks and opportunities associated with existing and forthcoming AI tools for music, though more work is needed to identify concrete actions which leverage the opportunities while mitigating risks.
Striving for data-model efficiency: Identifying data externalities on group performance
Esther Rolf
Ben Packer
Alex Beutel
Measuring Commonality in Recommendation of Cultural Content: Recommender Systems to Enhance Cultural Citizenship
Joint Multisided Exposure Fairness for Recommendation
Bhaskar Mitra
Xue Liu
Prior research on exposure fairness in the context of recommender systems has focused mostly on disparities in the exposure of individual or… (voir plus) groups of items to individual users of the system. The problem of how individual or groups of items may be systemically under or over exposed to groups of users, or even all users, has received relatively less attention. However, such systemic disparities in information exposure can result in observable social harms, such as withholding economic opportunities from historically marginalized groups (allocative harm) or amplifying gendered and racialized stereotypes (representational harm). Previously, Diaz et al. developed the expected exposure metric---that incorporates existing user browsing models that have previously been developed for information retrieval---to study fairness of content exposure to individual users. We extend their proposed framework to formalize a family of exposure fairness metrics that model the problem jointly from the perspective of both the consumers and producers. Specifically, we consider group attributes for both types of stakeholders to identify and mitigate fairness concerns that go beyond individual users and items towards more systemic biases in recommendation. Furthermore, we study and discuss the relationships between the different exposure fairness dimensions proposed in this paper, as well as demonstrate how stochastic ranking policies can be optimized towards said fairness goals.
On Natural Language User Profiles for Transparent and Scrutable Recommendation
Filip Radlinski
Krisztian Balog
Lucas Dixon
Ben Wedin
Natural interaction with recommendation and personalized search systems has received tremendous attention in recent years. We focus on the c… (voir plus)hallenge of supporting people's understanding and control of these systems and explore a fundamentally new way of thinking about representation of knowledge in recommendation and personalization systems. Specifically, we argue that it may be both desirable and possible for algorithms that use natural language representations of users' preferences to be developed. We make the case that this could provide significantly greater transparency, as well as affordances for practical actionable interrogation of, and control over, recommendations. Moreover, we argue that such an approach, if successfully applied, may enable a major step towards systems that rely less on noisy implicit observations while increasing portability of knowledge of one's interests.
Retrieval-Enhanced Machine Learning
Hamed Zamani
Mostafa Dehghani
Donald Metzler
Michael Bendersky
Although information access systems have long supportedpeople in accomplishing a wide range of tasks, we propose broadening the scope of use… (voir plus)rs of information access systems to include task-driven machines, such as machine learning models. In this way, the core principles of indexing, representation, retrieval, and ranking can be applied and extended to substantially improve model generalization, scalability, robustness, and interpretability. We describe a generic retrieval-enhanced machine learning (REML) framework, which includes a number of existing models as special cases. REML challenges information retrieval conventions, presenting opportunities for novel advances in core areas, including optimization. The REML research agenda lays a foundation for a new style of information access research and paves a path towards advancing machine learning and artificial intelligence.
Offline Retrieval Evaluation Without Evaluation Metrics
Offline evaluation of information retrieval and recommendation has traditionally focused on distilling the quality of a ranking into a scala… (voir plus)r metric such as average precision or normalized discounted cumulative gain. We can use this metric to compare the performance of multiple systems for the same request. Although evaluation metrics provide a convenient summary of system performance, they also collapse subtle differences across users into a single number and can carry assumptions about user behavior and utility not supported across retrieval scenarios. We propose recall-paired preference (RPP), a metric-free evaluation method based on directly computing a preference between ranked lists. RPP simulates multiple user subpopulations per query and compares systems across these pseudo-populations. Our results across multiple search and recommendation tasks demonstrate that RPP substantially improves discriminative power while correlating well with existing metrics and being equally robust to incomplete data.
Exposing Query Identification for Search Transparency
Ruohan Li
Jianxiang Li
Bhaskar Mitra
Asia J. Biega
Search systems control the exposure of ranked content to searchers. In many cases, creators value not only the exposure of their content but… (voir plus), moreover, an understanding of the specific searches where the content is surfaced. The problem of identifying which queries expose a given piece of content in the ranking results is an important and relatively under-explored search transparency challenge. Exposing queries are useful for quantifying various issues of search bias, privacy, data protection, security, and search engine optimization. Exact identification of exposing queries in a given system is computationally expensive, especially in dynamic contexts such as web search. We explore the feasibility of approximate exposing query identification (EQI) as a retrieval task by reversing the role of queries and documents in two classes of search systems: dense dual-encoder models and traditional BM25 models. We then propose how this approach can be improved through metric learning over the retrieval embedding space. We further derive an evaluation metric to measure the quality of a ranking of exposing queries, as well as conducting an empirical analysis focusing on various practical aspects of approximate EQI. Overall, our work contributes a novel conception of transparency in search systems and computational means of achieving it.
Fairness in Information Access Systems
Michael D. Ekstrand
Anubrata Das
Robin Burke
Recommendation, information retrieval, and other information access systems pose unique challenges for investigating and applying the fairne… (voir plus)ss and non-discrimination concepts that have been developed for studying other machine learning systems. While fair information access shares many commonalities with fair classification, the multistakeholder nature of information access applications, the rank-based problem setting, the centrality of personalization in many cases, and the role of user response complicate the problem of identifying precisely what types and operationalizations of fairness may be relevant, let alone measuring or promoting them. In this monograph, we present a taxonomy of the various dimensions of fair information access and survey the literature to date on this new and rapidly-growing topic. We preface this with brief introductions to information access and algorithmic fairness, to facilitate use of this work by scholars with experience in one (or neither) of these fields who wish to learn about their intersection. We conclude with several open problems in fair information access, along with some suggestions for how to approach research in this space.