Portrait of Yohaï-Eliel Berreby

Yohaï-Eliel Berreby

PhD - McGill University
Supervisor
Research Topics
Computational Neuroscience
Computer Vision
Deep Learning
NeuroAI
Neuroscience
Reinforcement Learning

Publications

CanViT: Toward Active-Vision Foundation Models
Active computer vision promises efficient, biologically plausible perception through sequential, localized glimpses, but lacks scalable gene… (see more)ral-purpose architectures and pretraining pipelines. As a result, Active-Vision Foundation Models (AVFMs) have remained unexplored. We introduce CanViT, the first task- and policy-agnostic AVFM. CanViT uses scene-relative RoPE to bind a retinotopic Vision Transformer backbone and a spatiotopic scene-wide latent workspace, the canvas. Efficient interaction with this high-capacity working memory is supported by Canvas Attention, a novel asymmetric cross-attention mechanism. We decouple thinking (backbone-level) and memory (canvas-level), eliminating canvas-side self-attention and fully-connected layers to achieve low-latency sequential inference and scalability to large scenes. We propose a label-free active vision pretraining scheme, policy-agnostic passive-to-active dense latent distillation: reconstructing scene-wide DINOv3 embeddings from sequences of low-resolution glimpses with randomized locations, zoom levels, and lengths. We pretrain CanViT-B from a random initialization on 13.2 million ImageNet-21k scenes -- an order of magnitude more than previous active models -- and 1 billion random glimpses, in 166 hours on a single H100. On ADE20K segmentation, a frozen CanViT-B achieves 38.5% mIoU in a single low-resolution glimpse, outperforming the best active model's 27.6% with 19.5x fewer inference FLOPs and no fine-tuning, as well as its FLOP- or input-matched DINOv3 teacher. Given additional glimpses, CanViT-B reaches 45.9% ADE20K mIoU. On ImageNet-1k classification, CanViT-B reaches 81.2% top-1 accuracy with frozen teacher probes. CanViT generalizes to longer rollouts, larger scenes, and new policies. Our work closes the wide gap between passive and active vision on semantic segmentation and demonstrates the potential of AVFMs as a new research axis.
How forward remapping predicts peri-saccadic biphasic mislocalization
B. Suresh Krishna
Neurons in many visual and oculomotor "priority-map" brain areas display forward receptive field (RF) remapping: they respond to stimuli app… (see more)earing before a saccade at the spatial location that their RF will occupy after the saccade. Concurrently, psychophysical studies have shown that flashes around saccade onset are systematically mislocalized in various patterns. One prominent pattern is a biphasic pattern, where flashes right before a saccade are mislocalized in the saccade direction (forward) and flashes right after a saccade are mislocalized opposite to the saccade direction (backward). Although forward RF remapping and biphasic mislocalization have been suspected to be linked, how this works has never been explained. Here, we show how persistent flash-evoked activity and decoding of the flash position after the saccade combine to produce this biphasic mislocalization pattern. We implement a rate model, consistent with the essential properties of RF remapping, and show that biphasic mislocalization results from insufficient remapping before the saccade, and residual/inappropriate remapping after the saccade. Less remapping before the saccade produces larger forward mislocalization of pre-saccadic flashes, and less remapping after the saccade produces smaller backward mislocalization of post-saccadic flashes. Forward RF remapping thus captures a biphasic peri-saccadic flash-mislocalization pattern consistent with behavioral data.