This program supports AI startups at any time of the year. Benefit from cutting-edge resources and tailored support to accelerate your technology's development.
Offered by Mila and the Public Policy Forum, this program is designed to equip policy and decision makers with the tools to navigate the opportunities and risks of AI. The next cohort will be held in French on September 1-2, 2026, at Mila.
Connect with a Mila academic advisor and current student-researchers to learn more about Mila's community and how to join us on August 19, 31 and September 11, 2026.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Although Multi-Objective Reinforcement Learning (MORL) research relies heavily on MO-MuJoCo as its go-to continuous control benchmark, the v… (see more)alidity of the conclusions drawn from it remain underexamined. In this paper, we first discuss three structural limitations of MO-MuJoCo; 1) its objectives are decomposed from pre-existing scalar rewards rather than independently motivated goals; 2) environments repeat the same underlying trade-off structure across varied locomotion morphologies, providing \textit{surface variety} without genuine \textit{problem diversity}; and 3) empirically approximated Pareto fronts appear broadly convex across research, potentially failing to stress-test the limitations of scalarization-based techniques. Setting these concerns aside, we further demonstrate that algorithmic rankings under MO-MuJoCo are highly sensitive to often undocumented evaluation choices in research papers. Across five evaluation axes, including reference point selection, weight distribution, normalization, return type, and front extraction method, pairwise algorithm rankings reverse in up to 47\% of configurations. Variance decomposition reveals that normalization alone accounts for nearly 69\% of hypervolume variance, suppressing the algorithm impact. Ultimately, we argue that progress in MORL research requires not only increased scrutiny of the benchmarks we rely upon, but also greater clarity in how results obtained within them are reported.
2026-08-14
Finding the Frame @ Reinforcement Learning Conference (published)
Modeling user preferences across domains remains a key challenge in slate recommendation (i.e. recommending an ordered sequence of items) re… (see more)search. We investigate how Large Language Models (LLM) can effectively act as world models of user preferences through pairwise reasoning over slates. We conduct an empirical study involving several LLMs on three tasks spanning different datasets. Our results reveal relationships between task performance and properties of the preference function captured by LLMs, hinting towards areas for improvement and highlighting the potential of LLMs as world models in recommender systems.