Portrait de Gregory Dudek n'est pas disponible

Gregory Dudek

Membre académique associé
Professeur titulaire et Directeur de recherche du laboratoire de robotique mobile, McGill University, École d'informatique
Vice-président et Chef de laboratoire de la recherche du Centre d'intelligence artificielle, Samsung AI Center in Montréal

Biographie

Gregory Dudek est professeur titulaire au Centre sur les machines intelligentes (CIM) de l’École d’informatique et directeur de recherche du Laboratoire de robotique mobile de l’Université McGill. Il est également chef de laboratoire et vice-président de la recherche du Centre d’intelligence artificielle de Samsung à Montréal. Gregory est également un membre académique associé à Mila - Institut québécois d'intelligence artificielle.

Il a écrit, seul ou en collaboration, plus de 300 articles de recherche sur des sujets tels que la description et la reconnaissance d’objets visuels, la localisation de radiofréquences (RF), la navigation et la cartographie robotiques, la conception de systèmes distribués, les télécommunications 5G et la perception biologique. Il a notamment publié le livre Computational Principles of Mobile Robotics, en collaboration avec Michael Jenkin, aux éditions Cambridge University Press. Il a présidé ou a contribué à de nombreuses conférences et activités professionnelles nationales et internationales dans les domaines de la robotique, de la détection par machine et de la vision par ordinateur. Ses recherches portent sur la perception pour la robotique mobile, la navigation et l’estimation de la position, la modélisation de l’environnement et des formes, la vision informatique et le filtrage collaboratif.

Étudiants actuels

Doctorat - McGill
Superviseur⋅e principal⋅e :
Maîtrise recherche - McGill
Superviseur⋅e principal⋅e :

Publications

Latent Attention Augmentation for Robust Autonomous Driving Policies
Ran Cheng
Christopher Agia
Florian Shkurti
Model-free reinforcement learning has become a viable approach for vision-based robot control. However, sample complexity and adaptability t… (voir plus)o domain shifts remain persistent challenges when operating in high-dimensional observation spaces (images, LiDAR), such as those that are involved in autonomous driving. In this paper, we propose a flexible framework by which a policy’s observations are augmented with robust attention representations in the latent space to guide the agent’s attention during training. Our method encodes local and global descriptors of the augmented state representations into a compact latent vector, and scene dynamics are approximated by a recurrent network that processes the latent vectors in sequence. We outline two approaches for constructing attention maps; a supervised pipeline leveraging semantic segmentation networks, and an unsupervised pipeline relying only on classical image processing techniques. We conduct our experiments in simulation and test the learned policy against varying seasonal effects and weather conditions. Our design decisions are supported in a series of ablation studies. The results demonstrate that our state augmentation method both improves learning efficiency and encourages robust domain adaptation when compared to common end-to-end frameworks and methods that learn directly from intermediate representations.
Trajectory-Constrained Deep Latent Visual Attention for Improved Local Planning in Presence of Heterogeneous Terrain
Stefan Wapnick
Travis Manderson
We present a reward-predictive, model-based deep learning method featuring trajectory-constrained visual attention for local planning in vis… (voir plus)ual navigation tasks. Our method learns to place visual attention at locations in latent image space which follow trajectories caused by vehicle control actions to enhance predictive accuracy during planning. The attention model is jointly optimized by the task-specific loss and an additional trajectory-constraint loss, allowing adaptability yet encouraging a regularized structure for improved generalization and reliability. Importantly, visual attention is applied in latent feature map space instead of raw image space to promote efficient planning. We validated our model in visual navigation tasks of planning low turbulence, collision-free trajectories in off-road settings and hill climbing with locking differentials in the presence of slippery terrain. Experiments involved randomized procedural generated simulation and real-world environments. We found our method improved generalization and learning efficiency when compared to no-attention and self-attention alternatives.
An Autonomous Probing System for Collecting Measurements at Depth from Small Surface Vehicles
Yuying Huang
Yiming Yao
Johanna Hansen
Jeremy Mallette
Sandeep Manjanna
This paper presents the portable autonomous probing system (APS), a low-cost robotic design for collecting water quality measurements at tar… (voir plus)geted depths from an autonomous surface vehicle (ASV). This system fills an important but often overlooked niche in marine sampling by enabling mobile sensor observations throughout the near-surface water column without the need for advanced underwater equipment. We present a probe delivery mechanism built with commercially available components and describe the corresponding open-source simulator and winch controller. Finally, we demonstrate the system in a field deployment and discuss design trade-offs and areas for future improvement. Project details are available on https://johannah.github.io/publication/sample-at-depth our website
Multimodal dynamics modeling for off-road autonomous vehicles
Travis Manderson
Aurélio Noca
Dynamics modeling in outdoor and unstructured environments is difficult because different elements in the environment interact with the robo… (voir plus)t in ways that can be hard to predict. Leveraging multiple sensors to perceive maximal information about the robot's environment is thus crucial when building a model to perform predictions about the robot's dynamics with the goal of doing motion planning. We design a model capable of long-horizon motion predictions, leveraging vision, lidar and proprioception, which is robust to arbitrarily missing modalities at test time. We demonstrate in simulation that our model is able to leverage vision to predict traction changes. We then test our model using a real-world challenging dataset of a robot navigating through a forest, performing predictions in trajectories unseen during training. We try different modality combinations at test time and show that, while our model performs best when all modalities are present, it is still able to perform better than the baseline even when receiving only raw vision input and no proprioception, as well as when only receiving proprioception. Overall, our study demonstrates the importance of leveraging multiple sensors when doing dynamics modeling in outdoor conditions.
Learning Intuitive Physics with Multimodal Generative Models
Predicting the future interaction of objects when they come into contact with their environment is key for autonomous agents to take intelli… (voir plus)gent and anticipatory actions. This paper presents a perception framework that fuses visual and tactile feedback to make predictions about the expected motion of objects in dynamic scenes. Visual information captures object properties such as 3D shape and location, while tactile information provides critical cues about interaction forces and resulting object motion when it makes contact with the environment. Utilizing a novel See-Through-your-Skin (STS) sensor that provides high resolution multimodal sensing of contact surfaces, our system captures both the visual appearance and the tactile properties of objects. We interpret the dual stream signals from the sensor using a Multimodal Variational Autoencoder (MVAE), allowing us to capture both modalities of contacting objects and to develop a mapping from visual to tactile interaction and vice-versa. Additionally, the perceptual system can be used to infer the outcome of future physical interactions, which we validate through simulated and real-world experiments in which the resting state of an object is predicted from given initial conditions.
PresSense: Passive Respiration Sensing via Ambient WiFi Signals in Noisy Environments
Yi Tian Xu
X. T. Chen
Xue Liu
Passive sensing with ambient WiFi signals is a promising technique that will enable new types of human-robot interactions while preserving u… (voir plus)sers' privacy. Here, we present PresSense, a system for human respiration sensing in noisy environments. Unlike existing WiFi-based respiration sensors, we employ a human presence detector, improving the robustness in scenarios where no human is present in an Area Of Interest (AOI). We also integrate our novel feature, Peak Distance Histogram (PDH), with other classic WiFi features to achieve better accuracy when someone is present in the AOI. We tested our system using commodity WiFi devices in an office room. Our PresSense outperforms the state of the arts in both respiration rate estimation and presence detection.
Seeing Through your Skin: Recognizing Objects with a Novel Visuotactile Sensor
Francois Hogan
M. Jenkin
Yogesh Girdhar
We introduce a new class of vision-based sensor and associated algorithmic processes that combine visual imaging with high-resolution tactil… (voir plus)e sending, all in a uniform hardware and computational architecture. We demonstrate the sensor’s efficacy for both multi-modal object recognition and metrology. Object recognition is typically formulated as an unimodal task, but by combining two sensor modalities we show that we can achieve several significant performance improvements. This sensor, named the See-Through-your-Skin sensor (STS), is designed to provide rich multi-modal sensing of contact surfaces. Inspired by recent developments in optical tactile sensing technology, we address a key missing feature of these sensors: the ability to capture a visual perspective of the region beyond the contact surface. Whereas optical tactile sensors are typically opaque, we present a sensor with a semitransparent skin that has the dual capabilities of acting as a tactile sensor and/or as a visual camera depending on its internal lighting conditions. This paper details the design of the sensor, showcases its dual sensing capabilities, and presents a deep learning architecture that fuses vision and touch. We validate the ability of the sensor to classify household objects, recognize fine textures, and infer their physical properties both through numerical simulations and experiments with a smart countertop prototype.
MBAIL: Multi-Batch Best Action Imitation Learning utilizing Sample Transfer and Policy Distillation
Dingwei Wu
M. Jenkin
Steve Liu
Batch reinforcement learning (RL) aims to learn a good control policy from a previously collected dataset without requiring additional inter… (voir plus)actions with the environment. Unfortunately, in the real world, we may only have a limited amount of training data for tasks we are interested in. Most batch RL methods are intended to learn a policy over one fixed dataset, and are not intended to learn a policy that can perform well over other tasks. How can we leverage the advantages of batch RL while dealing with limited training data is another challenge in real world. In this work, we propose to add sample transfer and policy distillation to a leading Batch RL approach. The proposed methods are evaluated on multiple control tasks to showcase their effectiveness.
Learning Domain Randomization Distributions for Training Robust Locomotion Policies
Melissa Mozian
Juan Camilo Gamboa Higuera
This paper considers the problem of learning behaviors in simulation without knowledge of the precise dynamical properties of the target rob… (voir plus)ot platform(s). In this context, our learning goal is to mutually maximize task efficacy on each environment considered and generalization across the widest possible range of environmental conditions. The physical parameters of the simulator are modified by a component of our technique that learns the Domain Randomization (DR) that is appropriate at each learning epoch to maximally challenge the current behavior policy, without being overly challenging, which can hinder learning progress. This so-called sweet spot distribution is a selection of simulated domains with the following properties: 1) The trained policy should be successful in environments sampled from the domain randomization distribution; and 2) The DR distribution made as wide as possible, to increase variability in the environments. These properties aim to ensure the trajectories encountered in the target system are close to those observed during training, as existing methods in machine learning are better suited for interpolation than extrapolation. We show how adapting the DR distribution while training context-conditioned policies results in improvements on jump-start and asymptotic performance when transferring a learned policy to the target environment1.
Learning to Drive Off Road on Smooth Terrain in Unstructured Environments Using an On-Board Camera and Sparse Aerial Images
Travis Manderson
Stefan Wapnick
We present a method for learning to drive on smooth terrain while simultaneously avoiding collisions in challenging off-road and unstructure… (voir plus)d outdoor environments using only visual inputs. Our approach applies a hybrid model-based and model-free reinforcement learning method that is entirely self-supervised in labeling terrain roughness and collisions using on-board sensors. Notably, we provide both first-person and overhead aerial image inputs to our model. We find that the fusion of these complementary inputs improves planning foresight and makes the model robust to visual obstructions. Our results show the ability to generalize to environments with plentiful vegetation, various types of rock, and sandy trails. During evaluation, our policy attained 90% smooth terrain traversal and reduced the proportion of rough terrain driven over by 6.1 times compared to a model using only first-person imagery.
View-Invariant Loop Closure with Oriented Semantic Landmarks
Jimmy Li
Karim Koreitem
Recent work on semantic simultaneous localization and mapping (SLAM) have shown the utility of natural objects as landmarks for improving lo… (voir plus)calization accuracy and robustness. In this paper we present a monocular semantic SLAM system that uses object identity and inter-object geometry for view-invariant loop detection and drift correction. Our system's ability to recognize an area of the scene even under large changes in viewing direction allows it to surpass the mapping accuracy of ORB-SLAM, which uses only local appearance-based features that are not robust to large viewpoint changes. Experiments on real indoor scenes show that our method achieves mean drift reduction of 70% when compared directly to ORB-SLAM. Additionally, we propose a method for object orientation estimation, where we leverage the tracked pose of a moving camera under the SLAM setting to overcome ambiguities caused by object symmetry. This allows our SLAM system to produce geometrically detailed semantic maps with object orientation, translation, and scale.
Dynamic planning of redundant robots within a set-based task-priority inverse kinematics framework.
Daniele Di Vito
Mathieux Bergeron
Gianluca Antonelli
This work presents the dynamic planning of redundant robots by merging a global and local planner. The global planner is implemented as a sa… (voir plus)mpling-based algorithm which works in the reduced-dimensionality of the robot workspace applying the Cartesian constraints only. The output trajectory is then checked within a framework of set-based task priority inverse kinematics verifying the fulfillment of the other task constraints. The inverse kinematics framework is used also in real-time as local motion control to ensure a reactive behaviour to address, e.g., mismatch between the apriori information and on-line perception acquisition. During the movement, the motion planner runs in background to adapt to changes in the environment or, in general, to continuously optimize the path. The proposed method is experimentally validated with a Kinova Jaco2 7 degrees of freedom manipulator.