Publications

Engineering TCR-controlled fuzzy logic into CAR T cells enhances therapeutic specificity
Taisuke Kondo
François X.P. Bourassa
Sooraj Achar
MyLinh T. Duong
Anirvan Ghosh
Jérémy Biton
Grégoire Altan-Bonnet
Naomi Taylor
Enhancing Privacy in the Early Detection of Sexual Predators Through Federated Learning and Differential Privacy
The increased screen time and isolation caused by the COVID-19 pandemic have led to a significant surge in cases of online grooming, which i… (see more)s the use of strategies by predators to lure children into sexual exploitation. Previous efforts to detect grooming in industry and academia have involved accessing and monitoring private conversations through centrally-trained models or sending private conversations to a global server. In this work, we implement a privacy-preserving pipeline for the early detection of sexual predators. We leverage federated learning and differential privacy in order to create safer online spaces for children while respecting their privacy. We investigate various privacy-preserving implementations and discuss their benefits and shortcomings. Our extensive evaluation using real-world data proves that privacy and utility can coexist with only a slight reduction in utility.
FoMo: Multi-Modal, Multi-Scale and Multi-Task Remote Sensing Foundation Models for Forest Monitoring
Forests are vital to ecosystems, supporting biodiversity and essential services, but are rapidly changing due to land use and climate change… (see more). Understanding and mitigating negative effects requires parsing data on forests at global scale from a broad array of sensory modalities, and using them in diverse forest monitoring applications. Such diversity in data and applications can be effectively addressed through the development of a large, pre-trained foundation model that serves as a versatile base for various downstream tasks. However, remote sensing modalities, which are an excellent fit for several forest management tasks, are particularly challenging considering the variation in environmental conditions, object scales, image acquisition modes, spatio-temporal resolutions, etc. With that in mind, we present the first unified Forest Monitoring Benchmark (FoMo-Bench), carefully constructed to evaluate foundation models with such flexibility. FoMo-Bench consists of 15 diverse datasets encompassing satellite, aerial, and inventory data, covering a variety of geographical regions, and including multispectral, red-green-blue, synthetic aperture radar and LiDAR data with various temporal, spatial and spectral resolutions. FoMo-Bench includes multiple types of forest-monitoring tasks, spanning classification, segmentation, and object detection. To enhance task and geographic diversity in FoMo-Bench, we introduce TalloS, a global dataset combining satellite imagery with ground-based annotations for tree species classification across 1,000+ categories and hierarchical taxonomic levels. Finally, we propose FoMo-Net, a pre-training framework to develop foundation models with the capacity to process any combination of commonly used modalities and spectral bands in remote sensing.
A Layer Selection Approach to Test Time Adaptation
Mostafa Elaraby
Yann Batiste Pequignot
Frédéric Precioso
Test Time Adaptation (TTA) addresses the problem of distribution shift by adapting a pretrained model to a new domain during inference. When… (see more) faced with challenging shifts, most methods collapse and perform worse than the original pretrained model. In this paper, we find that not all layers are equally receptive to the adaptation, and the layers with the most misaligned gradients often cause performance degradation. To address this, we propose GALA, a novel layer selection criterion to identify the most beneficial updates to perform during test time adaptation. This criterion can also filter out unreliable samples with noisy gradients. Its simplicity allows seamless integration with existing TTA loss functions, thereby preventing degradation and focusing adaptation on the most trainable layers. This approach also helps to regularize adaptation to preserve the pretrained features, which are crucial for handling unseen domains. Through extensive experiments, we demonstrate that the proposed layer selection framework improves the performance of existing TTA approaches across multiple datasets, domain shifts, model architectures, and TTA losses.
AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery
Amirhossein Abaskohi
Amrutha Varshini Ramesh
Shailesh Nanisetty
Christopher Pal
Giuseppe Carenini
Issam Hadj Laradji
Efficient and scalable construction of clinical variable networks for complex diseases with RAMEN.
Yiwei Xiong
Jingtao Wang
Tingting Chen
Douglas D. Fraser
Gregory Fonseca
Simon Rousseau
Genetic modulation of brain dynamics in neurodevelopmental disorders: the impact of copy number variations on resting-state EEG
Adrien E. E. Dubois
Elisabeth Audet-Duchesne
Inga Sophia Knoth
Charles-Olivier Martin
Khadije Jizi
Petra Tamer
Nadine Younis
Sébastien Jacquemont
Sarah Lippé
Research has shown that many copy number variations (CNVs) increase the risk of neurodevelopmental disorders (e.g., autism, ADHD, schizophre… (see more)nia). However, little is known about the effects of CNVs on brain development and function. Resting-state electroencephalography (EEG) is a suitable method to study the disturbances of neuronal functioning in CNVs. We aimed to determine whether there are resting-state EEG signatures that are characteristic of children with pathogenic CNVs. EEG resting-state brain activity of 109 CNV carriers (66 deletion carriers, 43 duplication carriers) aged 3 to 17 years was recorded for 4 minutes. To better account for developmental variations, EEG indices (power spectral density and functional connectivity) were corrected with a normative model estimated from 256 Healthy Brain Network controls. Results showed a decreased exponent of the aperiodic activity and a reduced alpha peak frequency in CNV carriers. Additionally, the study showed altered periodic components and connectivity in several frequency bands. Deletion and duplication carriers exhibited a similar overall pattern of deviations in spectral and connectivity measures, although the significance and effect sizes relative to the control group varied across frequency bands. Deletion and duplication carriers can be differentiated by their periodic power in the gamma band and connectivity in the low alpha band, with duplication carriers showing more disrupted alterations than deletion carriers. The distinctive alterations in spectral patterns were found to be most prominent during adolescence. The results suggest that CNV carriers show electrophysiological alterations compared to neurotypical controls, regardless of the gene dosage effect and their affected genomic region. At the same time, while duplications and deletions share common electrophysiological alterations, each exhibits distinct brain alteration signatures that reflect gene dosage-specific effects.
Leveraging Machine Learning Techniques in Intrusion Detection Systems for Internet of Things
Saeid Jamshidi
Amin Nikanjam
Kawser Wazed Nafi
As the Internet of Things (IoT) continues to expand, ensuring the security of connected devices has become increasingly critical. Traditiona… (see more)l Intrusion Detection Systems (IDS) often fall short in managing the dynamic and large-scale nature of IoT networks. This paper explores how Machine Learning (ML) and Deep Learning (DL) techniques can significantly enhance IDS performance in IoT environments. We provide a thorough overview of various IDS deployment strategies and categorize the types of intrusions common in IoT systems. A range of ML methods -- including Support Vector Machines, Naive Bayes, K-Nearest Neighbors, Decision Trees, and Random Forests -- are examined alongside advanced DL models such as LSTM, CNN, Autoencoders, RNNs, and Deep Belief Networks. Each technique is evaluated based on its accuracy, efficiency, and suitability for real-world IoT applications. We also address major challenges such as high false positive rates, data imbalance, encrypted traffic analysis, and the resource constraints of IoT devices. In addition, we highlight the emerging role of Generative AI and Large Language Models (LLMs) in improving threat detection, automating responses, and generating intelligent security policies. Finally, we discuss ethical and privacy concerns, underscoring the need for responsible and transparent implementation. This paper aims to provide a comprehensive framework for developing adaptive, intelligent, and secure IDS solutions tailored for the evolving landscape of IoT.
Lugha-Llama: Adapting Large Language Models for African Languages
Happy Buzaaba
Alexander Wettig
Christiane Fellbaum
Towards sustainable energy use: Reinforcement learning for demand response in commercial buildings
Seyyedreza Madani
Pierre‐Olivier Pineau
Ysaël Desage
Demand response (DR) is a crucial strategy for balancing electricity demand and reducing environmental impact, especially as power consumpti… (see more)on continues to rise. Small and medium-sized commercial buildings hold significant potential for DR implementation due to their widespread presence and substantial contribution to overall energy use. This study proposes a novel framework to optimize energy management in these buildings by considering three types of loads: non-controllable, controllable with discrete action spaces (HVAC), and controllable with continuous action spaces (lighting). The objective is to minimize costs, reduce CO 2 emissions, improve occupant comfort, and shave peak loads as a unified goal. State-of-the-art reinforcement learning (RL) algorithms are employed and compared against traditional heuristic methods using real-world data. Results show that RL-based approaches can significantly lower energy costs and environmental impacts while maintaining occupant comfort, even under varying outdoor temperature conditions. By incorporating risk assessments through Value at Risk (VaR) and Conditional Value at Risk (CVaR) metrics, this study offers a robust solution for sustainable energy management, providing insights for policymakers and industry practitioners aiming for a more resilient energy future.
Advancing Sustainable Maritime Transport: A Machine Learning Approach to Predict and Mitigate Underwater Radiated Noise from Ships
Soukaina Boujdi
Pierre Cauchy
Alignment of auditory artificial networks with massive individual fMRI brain data leads to generalisable improvements in brain encoding and downstream tasks
Maelle Freteault
Loic Tetrel
Lune P Bellec
Nicolas Farrugia
Artificial neural networks trained in the field of artificial intelligence (AI) have emerged as key tools to model brain processes, sparking… (see more) the idea of aligning network representations with brain dynamics to enhance performance on AI tasks. While this concept has gained support in the visual domain, we investigate here the feasibility of creating auditory artificial neural models directly aligned with individual brain activity. This objective raises major computational challenges, as models have to be trained directly with brain data, which is typically collected at a much smaller scale than data used to train AI models. We aimed to answer two key questions: (1) Can brain alignment of auditory models lead to improved brain encoding for novel, previously unseen stimuli? (2) Can brain alignment lead to generalisable representations of auditory signals that are useful for solving a variety of complex auditory tasks? To answer these questions, we relied on two massive datasets: a deep phenotyping dataset from the Courtois neuronal modelling project, where six subjects watched four seasons (36 hours) of the Friends TV series in functional magnetic resonance imaging and the HEAR benchmark, a large battery of downstream auditory tasks. We fine-tuned SoundNet, a small pretrained convolutional neural network with ∼2.5M parameters. Aligning SoundNet with brain data from three seasons of Friends led to substantial improvement in brain encoding in the fourth season, extending beyond auditory and visual cortices. We also observed consistent performance gains on the HEAR benchmark, particularly for tasks with limited training data, where brain-aligned models performed comparably to the best-performing models regardless of size. We finally compared individual and group models, finding that individual models often matched or outperformed group models in both brain encoding and downstream task performance, highlighting the data efficiency of fine-tuning with individual brain data. Our results demonstrate the feasibility of aligning artificial neural network representations with individual brain activity during auditory processing, and suggest that this alignment is particularly beneficial for tasks with limited training data. Future research is needed to establish whether larger models can achieve even better performance and whether the observed gains extend to other tasks, particularly in the context of few shot learning.