Portrait of Justine Zeghal

Justine Zeghal

Postdoctorate - Université de Montréal
Co-supervisor
Research Topics
Bayesian Inference
Cosmology
Generative Models
Probabilistic Models

Publications

MIRA: A Score for Conditional Distribution Accuracy and Model Comparison
We present Mira, a method for estimating the expected probability that samples from a candidate conditional distribution match the true, unk… (see more)nown conditional distribution, for which only data-label pairs are available. We derive theoretical bounds obtained when the candidate distribution matches the true one and when the conditional distributions are independent. This framework thus enables model comparison by quantifying the alignment between the conditional distribution of a candidate model and the data-label pairs of the true model. Consequently, Mira enables Bayesian model comparison through direct posterior validation, bypassing the challenging evidence computation. We demonstrate its effectiveness across several toy problems and Bayesian inference tasks.
Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration
LSST Dark Energy Science Collaboration
Eric Aubourg
Camille Avestruz
Matthew R. Becker
Biswajit Biswas
Rahul Biswas
Boris Bolliet
Adam S. Bolton
Clecio R. Bom
Raphaël Bonnet-Guerrini
Alexandre Boucaud
Jean-Eric Campagne
Chihway Chang
Aleksandra Ćiprijanović
Johann Cohen-Tanugi
Michael W. Coughlin
John Franklin Crenshaw
Juan C. Cuevas-Tello
Juan de Vicente
Seth W. Digel … (see 46 more)
Steven Dillmann
Mariano Javier de León Dominguez Romero
Alex Drlica-Wagner
Sydney Erickson
Alexander T. Gagliano
Christos Georgiou
Aritra Ghosh
Matthew Grayling
Kirill A. Grishin
Alan Heavens
Lindsay R. House
Mustapha Ishak
Wassim Kabalan
Arun Kannawadi
François Lanusse
C. Danielle Leonard
Pierre-François Léget
Michelle Lochner
Yao-Yuan Mao
Peter Melchior
Grant Merz
Martin Millon
Anais Möller
Gautham Narayan
Yuuki Omori
Hiranya Peiris
Andrés A. Plazas Malagón
Nesar Ramachandra
Benjamin Remy
Cécile Roucelle
Jaime Ruiz-Zapatero
Stefan Schuldt
Ignacio Sevilla-Noarbe
Ved G. Shah
Tjitske Starkenburg
Stephen Thorp
Laura Toribio San Cipriano
Tilman Tröster
Roberto Trotta
Padma Venkatraman
Amanda Wasserman
Tim White
Tianqing Zhang
Yuanyuan Zhang
The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data… (see more) (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.
Bridging Simulators with Conditional Optimal Transport
Optimal neural summarization for full-field weak lensing cosmological implicit inference
Denise Lanzieri
T. Lucas Makinen
Alexandre Boucaud
Jean-Luc Starck
François Lanusse
Context. Traditionally, weak lensing cosmological surveys have been analyzed using summary statistics that were either motivated by their an… (see more)alytically tractable likelihoods (e.g., power spectrum) or by their ability to access some higher-order information (e.g., peak counts), but at the cost of requiring a simulation-based inference approach. In both cases, even if the statistics can be very informative, they are not designed nor guaranteed to be statistically sufficient (i.e., to capture all the cosmological information content of the data). With the rise of deep learning, however, it has becomes possible to create summary statistics that are specifically optimized to extract the full cosmological information content of the data. Yet, a fairly wide range of loss functions have been used in practice in the weak lensing literature to train such neural networks, leading to the natural question of whether a given loss should be preferred and whether sufficient statistics can be achieved in theory and in practice under these different choices. Aims. We compare different neural summarization strategies that have been proposed in the literature to identify the loss function that leads to theoretically optimal summary statistics for performing full-field cosmological inference. In doing so, we aim to provide guidelines and insights to the community to help guide future neural network-based cosmological inference analyses. Methods. We designed an experimental setup that allows us to isolate the specific impact of the loss function used to train neural summary statistics on weak lensing data at fixed neural architecture and simulation-based inference pipeline. To achieve this, we developed the sbi_lens JAX package, which implements an automatically differentiable lognormal weak lensing simulator and the tools needed to perform explicit full-field inference with a Hamiltonian Monte Carlo (HMC) sampler over this model. Using sbi_lens , we simulated a w CDM LSST Year 10 weak lensing analysis scenario in which the full-field posterior obtained by HMC sampling gives us a ground truth that can be compared to different neural summarization strategies. Results. We provide theoretical insight into the different loss functions being used in the literature, including mean squared error (MSE) regression, and show that some do not necessarily lead to sufficient statistics, while those motivated by information theory, in particular variational mutual information maximization (VMIM), can in principle lead to sufficient statistics. Our numerical experiments confirm these insights, and we show on our simulated w CDM scenario that the figure of merit (FoM) of an analysis using neural summary statistics optimized under VMIM achieves 100% of the reference Ω c − σ 8 full-field FoM, while an analysis using summary statistics trained under simple MSE achieves only 81% of the same reference FoM.