The upcoming meeting, taking place on November 10 at Mila, will explore how we can collectively develop, govern, and deploy high-performing, reliable, and secure agentic systems by connecting academic researchers, industry experts, and practitioners.
We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.
Setting cookies
You can enable and disable the types of cookies you wish to accept. However certain choices you make could affect the services offered on our sites (e.g. suggestions, personalised ads, etc.).
Essential cookies
These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
Analytics cookies
Do you accept the use of cookies to measure the audience of our sites?
Multimedia Player
Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
Diffusion and flow models are effective world models for visual reinforcement learning, but existing agents treat them as black-box simulato… (see more)rs, leaving the backbone’s representations unused for control. We introduce DRIFT, an online agent in which a single Flow-Transformer serves as both world model and policy backbone, trained from scratch. We find that denoising features alone are suboptimal for control; DRIFT bridges this gap with a next-latent prediction objective that gives the backbone an explicit dynamics signal. Shortcut flow matching reduces imagination to a single denoising step per frame. Across Atari 100k, Craftium, and Crafter, DRIFT is competitive with both latent-dynamics and diffusion world-model baselines.