The upcoming meeting, taking place on November 10 at Mila, will explore how we can collectively develop, govern, and deploy high-performing, reliable, and secure agentic systems by connecting academic researchers, industry experts, and practitioners.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Publications
Dietary Pattern Analysis Adapted to High-Frequency Longitudinal Grocery Purchasing Data: Dynamic Multivariate Poisson Lognormal Network Model
Influenza-like illness (ILI) remains a persistent global health challenge, necessitating accurate forecasting tools for timely public health… (see more) response. This study systematically benchmarks fine-tuned large language models (LLMs), e.g., Llama2 and GPT2, for influenza surveillance forecasting in data-limited time-series settings. We develop a lightweight fine-tuning framework that adapts pre-trained LLMs using compact embedding and prediction layers and evaluate it on seven weekly aggregated real-world surveillance datasets. Despite sample sizes of only ∼523 time points per region and the absence of cloud-based data transfer, fine-tuned LLMs consistently outperform SARIMA, LSTM, PatchTST, CoVTransformer, FEDformer, Time-LLM, and GPT4TS in both accuracy and stability, especially for long-term forecasts across diverse geographic settings. Even in zero-shot settings, pre-trained LLMs capture broad epidemic trends with performance comparable to SARIMA. These findings establish fine-tuned LLMs as efficient and robust forecasting tools suitable for privacy-sensitive, data-scarce public health applications.
LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks … (see more)ship their own code base, hardcode a specific closed-model judge, and support a single evaluation protocol. This fragmentation makes it difficult to study how design choices--the benchmark, the judge model, the prompt, the inference backend--affect the conclusions we draw about model quality. We introduce JudgeArena, an open-source framework that unifies major LLM-judge benchmarks (AlpacaEval, Arena-Hard, MT-Bench, and m-Arena-Hard) under a single interface with swappable judges and comprehensive metadata logging for increased transparency in reporting and reproducibility. It enables systematic studies of judge choices, as any model accessible via vLLM, llama.cpp, or OpenRouter can serve as both candidate and judge. Furthermore, JudgeArena ships with tuned judge configurations for open models that match or outperform closed-model judges, validated on human preference datasets in both English and multilingual settings, reducing the reliance on opaque closed models. Finally, by combining existing human annotations with LLM-judge evaluations of a target model, JudgeArena can simulate LMArena Elo scores with high accuracy offering a practical, open, and low-cost alternative to large-scale human annotation campaigns.
BACKGROUND: Canadian MD-PhD and Clinician Investigator Program (CIP) pathways both train physician-scientists, yet longitudinal data on grad… (see more)uate research productivity are limited. METHODS: We conducted a longitudinal study of McGill University MD-PhD and CIP graduates who completed these programs between 2000 and 2017, using sex- and year-matched MD-only graduates as controls. Outcomes were collected for publications indexed through to December 31, 2022, and included H-index, total publications, total citations, per-paper journal impact factor, per-paper citations, and authorship position (first, second, senior). Productivity was assessed across four time periods: (1) pre-medical/medical training, (2) residency, (3) post-residency/fellowship, and (4) independent practice. RESULTS: Among 549 graduates (MD-PhD = 31; CIP = 97; MD-only = 421), research productivity, research impact, and active research involvement (defined as ≥3 first- or senior-author papers in the prior 5 years) were similar between MD-PhD and CIP graduates; both exceeded MD-only graduates across all metrics. Within-group sex differences were not significant. MD-PhD graduates were more productive during pre-medical/medical training, whereas CIP graduates were comparatively more productive during residency. Sustained productivity in independent practice correlated with research engagement during medical school, residency, and fellowship. In both programs, graduates with active research involvement had higher authorship counts in the 10 years immediately following graduation from medical school. DISCUSSION: Research productivity and impact were comparable between MD-PhD and CIP graduates, and both groups exceeded MD-only peers across measured research metrics. Across both programs, graduates with active research involvement had higher early-career authorship counts, particularly during residency and post-residency/fellowship time periods. These findings provide descriptive benchmarking data for future studies of physician-scientist training pathways in Canada.
Alzheimer's disease (AD) has a higher prevalence in women than men and is more frequently inherited from mothers than fathers. Yet, while ne… (see more)uroimaging and biomarker studies link maternal family history to stronger AD-related alterations, epidemiological studies suggest that paternal history confers comparable or even greater risk. Here, we leverage the deeply profiled PREVENT-AD cohort to derive three intermediate phenotypes of AD susceptibility. Drawing on nearly 1,000 individual study visits, we quantify how these intermediate phenotypes vary as a function of maternal versus paternal AD lineage. We show that lineage-specific differentiation, including both maternal and paternal biases, is reflected in the brain structure and phenome of adult children of AD patients. Cognitive and cardiovascular risk markers, together with associated genetic variants, show the strongest differentiation along the parental-lineage spectrum of disease susceptibility relative to other correlates of AD burden. Our cross-generational analysis ultimately delineates multidimensional parent-of-origin effects in AD genealogy.
Dimensionality reduction-based data visualization is pivotal in comprehending complex biological data. The most common methods, such as PHAT… (see more)E, t-SNE, and UMAP, are unsupervised and therefore reflect the dominant structure in the data, which may be independent of expert-provided labels. Here we introduce a supervised data visualization method called RF-PHATE, which integrates expert knowledge for further exploration of the data. RF-PHATE leverages random forests to capture intricate featurelabel relationships. Extracting information from the forest, RF-PHATE generates low-dimensional visualizations that highlight relevant data relationships while disregarding extraneous features. This approach scales to large datasets and applies to classification and regression. We illustrate RF-PHATE’s prowess through three case studies. In a multiple sclerosis study using longitudinal clinical and imaging data, RF-PHATE unveils a sub-group of patients with non-benign relapsingremitting Multiple Sclerosis, demonstrating its aptitude for time-series data. In the context of Raman spectral data, RF-PHATE effectively showcases the impact of antioxidants on diesel exhaust-exposed lung cells, highlighting its proficiency in noisy environments. Furthermore, RF-PHATE aligns established geometric structures with COVID-19 patient outcomes, enriching interpretability in a hierarchical manner. RF-PHATE bridges expert insights and visualizations, promising knowledge generation. Its adaptability, scalability, and noise tolerance underscore its potential for widespread adoption.
Machine learning models trained on biochemical data are routinely evaluated using splits that fail to account for relational structure, caus… (see more)ing information leakage and over-optimistic performance estimates. Existing splitting methods lack theoretical grounding and scale at best quadratically. We introduce the Relational Generative Process (RGP), a mathematical formalization explaining why relational structure arises in biochemical datasets, and Refnd, a splitting algorithm that leverages a proximity graph computed in loglinear time using Hierarchical Navigable Small World (HNSW). We validate on an antimicrobial peptide dataset, showing that Refnd splits yield lower but more realistic evaluation performance than traditional splits. Refnd is applicable to any dataset arising from an RGP such as protein sequences and structures, small molecules, and nucleotide sequences, and is openly available as a Rust accelerated Python package: pip install refnd.
When shopping for perishable products, consumers typically prefer the freshest items, especially those with a short shelf life. With that in… (see more) mind, retailers establish strict contractual agreements with suppliers to ensure the fulfilment of their orders for perishable products. One key condition in these agreements is the Minimum Life On Receipt (MLOR) rule, which defines the maximum product age that the retailer will accept at full price. In this study, we propose a model that facilitates the negotiation of retailer-supplier terms to increase flexibility. Specifically, we define the share of orders retailers should accept beyond the MLOR at a discounted price. We formulate the problem as a bilevel program considering the individual objectives of the retailer (leader) and the supplier (follower), while also accounting for consumer demand driven by both price and product freshness. To address the bilevel problem, we employ a reformulation-and-decomposition algorithm adapted from the literature. We then compare the supply chain benefits of solving the bilevel program with those of optimising the retailer’s and supplier’s objectives jointly in a centralised approach, as well as to standard contract terms in which products are returned if the supplier does not meet the MLOR requirement. Our results demonstrate that flexible agreements offer significant benefits, with average profit increases of up to 4% for retailers and up to 13% for suppliers. Finally, we provide suggestions for designing new clauses that account for consumer demand variability and retailer’s order frequency.
2026-06-29
International Journal of Production Economics (published)
LiDAR semantic segmentation often degrades under real-world deployment due to evolving sensing conditions, while collecting new annotations … (see more)for retraining is impractical. Test-time adaptation (TTA) updates model parameters online using pseudo-label supervision, but directly applying standard TTA strategies to LiDAR data is challenging. Because pseudo-label reliability is spatially heteroscedastic under range-dependent sparsity and occlusion, uniform updates on globally shared parameters can inject unstable gradients and destabilize adaptation. We propose a geometry-constrained test-time prompt tuning framework for LiDAR semantic segmentation. Our method estimates per-location sensing reliability from depth-consistent beam terminations and neighborhood support, and uses it to reweight spatial supervision. Adaptation is confined to lightweight prompt adapters inserted into a frozen backbone, with spatial gating to prevent unreliable regions from perturbing globally shared representations. A temporally smoothed prototype alignment strategy further stabilizes online updates by accumulating reliable semantic evidence over time. Experiments on standard LiDAR benchmarks demonstrate improved adaptation stability and segmentation performance under deployment variations without additional annotations.