Ce programme soutient les startups spécialisées en IA à tout moment de l'année. Bénéficiez de ressources de pointe et d'un accompagnement sur mesure pour accélérer le développement de votre technologie.
Offert par Mila et le Forum des politiques publiques, ce programme est conçu pour outiller les décideur·euse·s et les responsables des politiques publiques à naviguer efficacement à travers les opportunités et les risques liés à l'IA. La prochaine cohorte se tiendra en français les 1er et 2 septembre 2026 à Mila.
Échangez avec les conseiller·ère·s académiques de Mila ainsi que des étudiant·e·s-chercheur·euse·s pour en savoir plus sur la communauté de Mila et découvrir comment nous rejoindre les 19 et 31 août et le 11 septembre 2026.
Nous utilisons des témoins pour analyser le trafic et l’utilisation de notre site web, afin de personnaliser votre expérience. Vous pouvez désactiver ces technologies à tout moment, mais cela peut restreindre certaines fonctionnalités du site. Consultez notre Politique de protection de la vie privée pour en savoir plus.
Paramètre des cookies
Vous pouvez activer et désactiver les types de cookies que vous souhaitez accepter. Cependant certains choix que vous ferez pourraient affecter les services proposés sur nos sites (ex : suggestions, annonces personnalisées, etc.).
Cookies essentiels
Ces cookies sont nécessaires au fonctionnement du site et ne peuvent être désactivés. (Toujours actif)
Cookies analyse
Acceptez-vous l'utilisation de cookies pour mesurer l'audience de nos sites ?
Lecteur Multimédia
Acceptez-vous l'utilisation de cookies pour afficher et vous permettre de regarder les contenus vidéo hébergés par nos partenaires (YouTube, etc.) ?
Enning Yang
Alumni
Publications
Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
Most applications of generative AI involve a sequential interaction in which a person inputs a prompt and waits for a response, and where re… (voir plus)action time and adaptivity are not important factors. In contrast, live jamming is a collaborative interaction that requires real-time coordination and adaptation without access to the other player’s future moves, while preserving diversity to sustain a creative flow. Reinforcement learning post-training enables effective adaptation through on-policy interaction, yet it often reduces output diversity by exploiting coherence-based rewards. This collapse, known as ``reward hacking'', affects many RL post-training pipelines, but is especially harmful in live jamming, where musical creativity relies on dynamic variation and mutual responsiveness. In this paper, we propose a novel adversarial training method on policy-generated trajectories to mitigate reward hacking in RL post-training for melody-to-chord accompaniment. A co-evolving discriminator separates policy trajectories from the data distribution, while the policy maximizes the discriminator output in addition to coherence rewards to prevent collapse to trivial outputs. We evaluate accompaniment quality and output diversity in simulation with both fixed test melodies and learned melody agents, and we conduct a user study with the model deployed in a real-time interactive system with expert musicians. Quantitative evaluation and user feedback demonstrate improved output diversity, harmonic coherence, adaptation speed and user agency. Our results demonstrate a simple yet effective method to mitigate reward hacking in RL post-training of generative sequence models.
2025-12-31
International Conference on Learning Representations (Accept (Poster))
Neuroscientific studies exploring real-world dynamic perception often overlook the influence of continuous changes in narrative content. In … (voir plus)our research, we utilize machine learning tools for natural language processing to examine the relationship between movie narratives and neural responses. By analyzing over 50,000 brain images of participants watching Forrest Gump from the studyforrest dataset, we find distinct brain states that capture unique semantic aspects of the unfolding story. The default network, associated with semantic information integration, is the most engaged during movie watching. Furthermore, we identify two mechanisms that underlie how the default network liaises with the amygdala and hippocampus. Our findings demonstrate effective approaches to understanding neural processes in everyday situations and their relation to conscious awareness.
Naturalistic neuroscience opened the door to new insights into neural circuits that serve real-world dynamic perception. Such studies have o… (voir plus)ften neglected the rich texture of the movie narrative itself, but semantic content can be used to contextualize the induced neural responses. Here, we translated natural language processing tools from machine learning to characterize brain states estimated from hidden Markov models. Our analytical strategy allowed pitting shallow unimodal against the deep associative brain network layers in explaining how semantic content of the movie links to observed neural activity. Pooling information across >53,000 brain image time points watching Forrest Gump, we could show that distinct dynamic brain states capture unique semantic facets along the unfolding movie narrative. The spatiotemporal dynamics of brain states explicitly captured subject-level responses throughout the brain network hierarchy. Across all analyses, the default network was most intimately linked to semantic information integration, and this neural system switched online for longest durations during movie watching. Further, we identified and described two mechanisms of how the default network liaises dynamically with microanatomically defined subregion partners: the amygdala and the hippocampus. Our study thus unlocks the potential of natural language processing to explore neural processes in everyday life situations that engage key aspects of conscious awareness.
Naturalistic neuroscience opened the door to new insights into neural circuits that serve real-world dynamic perception. Such studies have o… (voir plus)ften neglected the rich texture of the movie narrative itself, but semantic content can be used to contextualize the induced neural responses. Here, we translated natural language processing tools from machine learning to characterize brain states estimated from hidden Markov models. Our analytical strategy allowed pitting shallow unimodal against the deep associative brain network layers in explaining how semantic content of the movie links to observed neural activity. Pooling information across >53,000 brain image time points watching Forrest Gump, we could show that distinct dynamic brain states capture unique semantic facets along the unfolding movie narrative. The spatiotemporal dynamics of brain states explicitly captured subject-level responses throughout the brain network hierarchy. Across all analyses, the default network was most intimately linked to semantic information integration, and this neural system switched online for longest durations during movie watching. Further, we identified and described two mechanisms of how the default network liaises dynamically with microanatomically defined subregion partners: the amygdala and the hippocampus. Our study thus unlocks the potential of natural language processing to explore neural processes in everyday life situations that engage key aspects of conscious awareness.