Publications

Deep Spectral Models for Robust Dental Shape Generation
François Guibault
Michal Španěl
Accurate modeling of dental crown morphology is fundamental for diagnosis, orthodontic planning, and computer-aided restoration design. Howe… (see more)ver, datasets suitable for training such models are typically limited in size. We present ToothForge, a deep spectral generative framework that models dental crown geometries from compact, intrinsic representations. By operating in the spectral domain, ToothForge learns a latent manifold of 3D tooth shapes through synchronized spectral embeddings, ensuring consistent modeling across samples with varying connectivity. Spectral synchronization mitigates the instability of Laplace-Beltrami eigenbases and enables efficient learning in a low-dimensional space. The framework is thoroughly evaluated through robustness analysis, ablation studies, and benchmarking against PCA-based statistical shape models and point-based generative frameworks. Results show that synchronized spectral modeling achieves reconstruction and generative performance comparable to or exceeding spatial approaches, while maintaining compactness and geometric interpretability. Together, the compact synchronized coefficients and low-dimensional learning space make the framework particularly suitable for limited datasets, as often encountered in dental and medical domains, and applicable in real-world scenarios where guaranteeing consistent connectivity across shapes from various clinics is unrealistic.
Influence of Autonomy and Interfaces on Human and Multi-Robot Teams: a Study on Planetary Exploration
Marcel Kaufmann
Advances in robotic autonomy and interfaces have transformed human-robot teaming across domains, from disaster response to planetary science… (see more). However, critical gaps remain in understanding how autonomy and interface design affect human performance and cognitive demands, especially in large-scale, unstructured environments such as those on the Moon or Mars. We present a human-in-the-loop system comprising two (semi-)autonomous robots supervised by a single human operator. The system was evaluated in real caves at Lava Beds National Monument (California) and in a controlled within-subject study (n=38) at Polytechnique Montréal exploring both real and simulated caves. Participants interacted using either a traditional screen interface or a novel real-time, immersive VR interface, developed for this study and field-tested during NASA's BRAILLE campaign. We find that continuous physiological measurements (HRV) align with subjective NASA TLX scores in the context of human and multi-robot planetary exploration. Compared to benchmarks from prior studies, the screen interface resulted in low workload, while VR was rated in the low-to-moderate range. The low-autonomy VR-waypoint condition resulted in the least effective performance, with the fewest automated science detections, whereas both the full-autonomy VR and screen-based conditions yielded comparably higher exploration and detection performance. Both interfaces supported high situational awareness, with accuracy measures near 90%. Autonomy did not significantly affect situational awareness, but full-autonomy did reduce operator input effort. Trust levels did not significantly vary across conditions, motivating more detailed assessment methods in future studies. The results inform how to align interface design and autonomy to support effective multi-robot supervision in future missions.
RAG-Safe: A Recall-First Safety Framework Comparing Open-Source and Commercial LLM Moderation Pipelines
False negatives—missed detections of harmful content—remain the dominant risk in safety-critical moderation pipelines. We introduce RAG-… (see more)Safe, a recall-first framework that integrates distribution-preserving contrastive augmentation, committee-diverse retrieval, and a recall-oriented decision policy into a unified moderation architecture. The framework is evaluated using a compact, fully auditable testbed designed to enforce strict leakage control: original samples alone determine the train–test split, and all paraphrases inherit their parent assignment. Within this controlled setting, conventional retrieval-augmented pipelines—both commercial (API embeddings + hosted LLM) and open-source (FAISS + local LLaMA-3)—consistently under-detect unsafe content (FLAGGED recall 0.44). Applying RAG-Safe raises FLAGGED recall to approximately 0.56 across both stacks while preserving overall accuracy ( 0.66) and macro-F1 ( 0.65). A non-RAG classifier baseline provided in our public repository shows similar recallfirst behaviour, reinforcing that these gains are not architecture-specific. Rather than comparing individual model components, we interpret the results as pipeline-level evidence that boundary-focused augmentation, retrieval diversity, and calibrated thresholds jointly shift LLM moderation into a safer operating regime. We conclude by discussing limitations—particularly domain transferability and adversarial robustness—and outline directions for scaling RAG-Safe to broader moderation contexts. Keywords: Content moderation, Recall-first classification, Distribution-preserving data augmentation, Committee-based retrieval, Retrieval-augmented large language models, Safety-critical AI
Special Dedication to Music and Songs Proliferation - [A short Note]
Nonvikan Karl-Augustt Alahassa
Bidossessi R.U. Alahassa
Christiane Rousseau
J. Tossa
Bakary Manga
Suljo Linic
Samuel Bassetto
Dimitrios Koukoulopoulos
Maciej Augustyniak
Bruno Rémillard
Nathalie Lacelle
Leonard Wantchekon
Emmanuel Stip
Jérôme Théau
Damien Échevin
Daniel F. Nadeau
Victor M. Panaretos
Marlène Frigon
David Haziza … (see 2 more)
Julie Carrier
Mylène Bédard
This is a Special Dedication to Music and Songs Proliferation - [A short Note].
Stabilizing Black-Box Prompt Optimization with Textual Regularization and Signal Aggregation
MohammadReza Davari
Utkarsh Garg
Weixin Cai
An increasing number of NLP applications interact with large language models (LLMs) through black-box APIs, making prompt engineering critic… (see more)al for controlling model behavior. Recent Automatic Prompt Optimization (APO) methods iteratively refine prompts using model-generated critiques (often called as textual gradients), but they predominantly optimize from failures and underutilize information contained in correct predictions, leading to instability and semantic drift. We propose TRAS (Textual Regularization with Aggregated Signals), a feedback-centric framework that is plug-and-play with existing APO search backbones. It retains the standard textual gradient signal from prior work for error correction, and introduces a complementary textual regularizer derived from successful predictions to preserve beneficial prompt components. Because both signals are stochastic and can be noisy, we further introduce Monte Carlo Signal Aggregation (MCSA), which samples multiple gradients or regularizers and aggregates them into a single actionable directive, emphasizing consistent, actionable advice while filtering out outliers. Motivated by rapid model churn, we also formalize Automatic Prompt Migration (APM), the practical problem of adapting an expert prompt across model versions or API providers without losing critical instructions. Across standard APO and APM scenarios, our approach consistently outperforms strong baselines, yielding higher accuracy, faster convergence, and lower query cost, while substantially reducing the degradation observed under naive prompt migration.
A Stochastic--Geometric Theory of Scaling Laws in Grokking
Róisín Luo
Ihsan Ullah
Karyn Morrissey
Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only b… (see more)egins to generalize after a prolonged delay, often through an abrupt transition. Despite extensive empirical study, its underlying mechanism remains poorly understood. In this work, we first theoretically characterize a shell--core topological configuration of the reachable solution space induced by Adam's optimization dynamics with weight-shrinkage regularization, supported by empirical evidence. This optimization-induced topological configuration gives rise to grokking. In model's parameter space, random initialization solutions concentrate on a thin outer spherical shell, enclosing another spherical shell of memorization solutions, which in turn contains a core corresponding to the generalization solutions. Leveraging stopping-time theory, we then analyze the geometry of this topological configuration and the solution transition time at which optimization trajectories escape the memorization manifold and first reach the boundary of the generalization manifold. Our theoretical analysis derives grokking scaling laws for the learning rate, batch size, and
ToxiSight: Leveraging Moderator Expertise Through Behavioral Measurement in Gaming Toxicity Annotation
Vicki Chen
Domenico Tullo
Content moderation systems commonly treat human annotators as interchangeable label sources, resolving disagreements through majority voting… (see more) or expert arbitration. We present ToxiSight, an annotation platform that reframes this assumption: rather than extracting consensus, the system supports moderator reasoning by treating hesitation, revision, and disagreement as signals revealing where content is genuinely ambiguous and where taxonomic guidelines fail. ToxiSight integrates gaming-specific contextual widgets with behavioral telemetry, capturing the cognitive processes underlying toxicity validation decisions. Through deployment with 10 professional moderators across 60,000 lines of gaming chat, we demonstrate that behavioral patterns expose systematic category failures invisible to traditional inter-annotator metrics. The Controversial category shows 72% revision rates with fast processing times, indicating immediate recognition of definitional breakdown, while Threats (Life-Threatening) exhibits 75% revisions with slow processing, signaling genuine interpretive complexity. Completion rates improved from 60% to 95%, and moderators reported reduced decision stress when permitted to express uncertainty. This case study demonstrates that trustworthy toxicity detection requires annotation systems designed around the irreducible complexity of human judgment, not against it.
Depth Exploration for LLM Decoding
Weisi Yang
Stephen Xia
Autoregressive LLM decoding evaluates every generated token through the full layer stack, even though many tokens become predictable at inte… (see more)rmediate depths. Existing lossless depth-adaptive methods exploit this redundancy by choosing a single non-final exit depth and verifying its prediction with the final-depth model. However, our measurements show that this selection-based strategy leaves substantial headroom: choosing an exit too late wastes computation, while choosing one too early triggers fallback and discards dependent drafts. We propose Depth Exploration Decoding (DEX), a lossless decoding algorithm that replaces single-depth selection with parallel exploration over multiple candidate depths. At each commit position, DEX validates candidates against the final-depth reference, commits exactly the final-depth token, and collapses the exploration lattice to retain only reusable branch states. This expand--commit--collapse procedure preserves equivalence to standard autoregressive decoding while reducing the cost of committing each token. Across early-exit-trained and standard LLMs, DEX outperforms representative depth-selection baselines and achieves competitive end-to-end throughput against speculative and distributed decoding methods. Moreover, DEX improves as the explored depths become finer, showing that parallel depth exploration provides a scalable way to exploit the underused depth axis of LLM decoding.
Safety from Honesty in a Disinterested AI Predictor
Oliver Richardson
Tomáš Gavenčiak
Michael Cohen
Rory Svarc
Gaël Gendron
David Hyland
Aton Kamanda
Adam Oberman
Francis Rhys Ward
Anna Gavenčiak
Jacob Livingston Slosser
Vincent Mai
Iulian Serban
Joumana Ghosn
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed… (see more) behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of"epistemically contextualized"natural-language statements. We argue that such a Predictor can honestly predict agents, actions, and their consequences without itself being an agent that selects outputs to achieve goals. This rests on data representation and on the training procedure. Epistemic contextualization of text distinguishes latent factual claims from communication acts, so expressions of goals are treated as evidence to be explained rather than drives the model adopts. With a posterior-seeking training objective, this is intended to drive the Predictor toward calibrated, cautious predictions. Training proceeds so downstream effects of deploying a prediction never serve as a reward signal; any agency the system needs is supplied by explicit scaffolding constrained by guardrails. We prove that, under assumptions on the training dynamics and on the argued sparsity of dangerous Predictors, the probability that training produces a Predictor whose guarded deployment carries residual harm above a specified threshold is small: a dangerous Predictor would have to underestimate harm in a coordinated way across many queries while such coordinated patterns are rare under the initialization distribution and receive no direct training signal. Safety and accuracy are jointly supported in this framework, since the constraints that secure accuracy are the same ones that make coordinated deception costly. These guarantees against misalignment and agency arising from within the Predictor itself do not preclude the use of the Predictor as part of an agentic system.
Formal Epistemic Structure Leaves Geometric Traces in LLMs: Evidence from S5 Multi-Agent Logic and RoBERTa
David John Lemay
Emil Sayilov
We ask whether the formal structure of S5 multi-agent epistemic logic leaves recoverable geometric traces in a fine-tuned language model. Va… (see more)n Benthem's product topology for S5 predicts that the state space of
HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models
Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabilitie… (see more)s into Large Language Models (LLMs). However, numerous studies report performance degradation on various downstream tasks due to information loss during discretization. To address this, we propose a novel approach combining temporally compressed discrete tokens with dimensionality-reduced continuous residuals. Our framework consists of a hybridized discrete-continuous focal modulation codec and a hybrid Transformer. This architecture performs autoregressive inference in the discrete domain, coupled with non-autoregressive prediction and continuous residual upsampling. Experimental results show that our approach significantly improves the retention of speaker characteristics compared to discrete-only methods, while simultaneously reducing the number of required autoregressive steps.
Can LLMs Cook Jamaican Couscous? A Study of Cultural Novelty in Recipe Generation