Publications

Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning
Charles Edward Gagnon
Steven H. H. Ding
Philippe Charland
Benjamin C. M. Fung
We present a practical pipeline for recovering source code from stripped binary functions by combining reverse engineering, anchor-based sou… (see more)rce code retrieval, and large language model reasoning. Our binary-to-source-code retrieval method attempts to identify the source function from a source code database, rather than generating approximate decompiled pseudocode. It extracts anchors such as strings, constants, external calls, and available function names using Ghidra, retrieves candidate files via an inverted-index search database, narrows candidates to likely function snippets, and re-ranks them with a large language model (LLM) based on disassembly, decompiled code, and source metadata. Confident matches can also serve as anchors in later passes. In an evaluation backed by our high-fidelity source code database on a stripped, optimized tcpdump binary, our proposed binary-to-source matching method achieves 95.2% assembly instruction coverage. Experiments on a GitHub-based retrieval database showed lower performance with 35.5% instruction coverage on average, mainly due to retrieval misses. These results show that source-level binary recovery excels with high-quality databases and remains a useful tool in noisy environments.
D-CLIPSE: Distributed Consensus-based Localization with Passive Listening on Shared State Exchange
Kyle Biron-Gricken
Multi-robot localization that is accurate and consistent is imperative for downstream tasks such as planning and control. Centralized filter… (see more)ing approaches optimally fuse all available sensor measurements of the team. However, a centralized solution is rarely implementable due to hardware, communication, and computational constraints. Distributed approaches deploy a filter on each robot to estimate their own state and neighbours' states using inter-robot communication. This paper proposes a consistent, communication-efficient, and consensus-based distributed filtering framework that shares both preintegrated odometry and relevant shared states among communicating robots. The proposed method is validated in simulated and experimental scenarios, showing near centralized performance in accuracy, and especially in consistency, compared to the current state-of-the-art decentralized approach.
Parallel versions of the mesh adaptive direct search algorithm
Sébastien Le Digabel
Christophe Tribes
This work surveys the different parallel variants of the mesh adaptive direct search (MADS) algorithm for constrained blackbox optimization.… (see more) These problems can inherently imply high computational costs due to the possible large number of variables and multi-modality of the search space. In addition, the potential time-intensive nature and time heterogeneity of the blackboxes defining the problem prompts the need for efficient implementations. Parallelism emerges as an actionable solution to mitigate computation time, as modern computer systems rely on multi-core architecture. The reviewed methods employ diverse levels of parallelism and distinct parallel strategies to effectively tackle each aspect outlined above. The manuscript details the practical implementations, provides computational results, and offers insights into the advantages and limitations of each MADS parallel method.
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions
Training reinforcement-learning agents directly on physical robots makes every fall costly, since a fall can damage the platform and cannot … (see more)be undone like a simulator reset; the goal is therefore to minimize falls during training rather than trade them off against return, as constrained Markov decision process (MDP) formulations do. A standard mitigation hands control to a separate recovery policy whenever the agent leaves a designer-specified safe region (a subset of state space it should stay within), but the resulting mixed-policy rollouts silently bias every on-policy update, and the importance-sampling correction that would remove this bias is ill-defined whenever the recovery policy is deterministic. We address this bias with a drop-in modification of proximal policy optimization (PPO). Its core is an unbiased policy-gradient estimator that uses the score function only at safe timesteps and never evaluates the recovery policy's density, so it stays valid even when the recovery policy is deterministic, exactly where importance sampling breaks, and it empirically dominates importance sampling even when the recovery policy is stochastic. Because the recovery policy still makes credit assignment slow near the safe-region boundary, two further components accelerate learning: a closed-form value for recovery-triggering states when dynamics and recovery are deterministic, and an imitation loss that copies recovery actions only when recovery succeeds. On a three-environment, five-seed benchmark, the resulting algorithm reduces training-time falls by factors of 233x, 48x, and 26x on HalfCheetah, Ant, and Unitree Go1 over standard PPO, while matching or exceeding PPO's final reward, and on Ant, where the recovery policy is unreliable, it is the only method that reaches 80% of the best final reward.
The 2026 Singapore Consensus on Global AI Safety Research Priorities
Mohan Kankanhalli
Lee Wan Sie
Chris Meserole
Luke Ong
Stuart Russell
Dawn Song
Max Tegmark
Brian Tse
Xue Lan
Andrew Yao
Zhang Ya-Qin
Zhou Bowen
Stephen Casper
Oskar Galeev
Ima (Imane) Bello
Kwan Yee Ng
Vanessa Wilfred … (see 100 more)
Erica Liaw
Lee Chein Inn
Lin Wanxuan
Ng En Qi
Jonathan Lee
José Villalobos
Abhishek Aggarwal
Adam Gleave
Alex Leung
Alvin Kwock
Anthony Tung
Arisa Siong
Arthur Tea
BEN BUCKNALL
Benjamin Weinstein-Raun
He Bing Sheng
Liu Bo
Bryan Kian Hsiang Low
Chris Ngo
Clement Neo
Cyrus Hodes
Dan Hendrycks
Daniel Ross
Liu Dapeng
Denise Wong
Djordje Žikelić
Elham Tabassi
Fabien Le Voyer
Fazl Barez
Gabriel Nicholas
Henry Papadatos
Jaan Tallinn
James Petrie
Xu Jia
Shao Jing
Jonathan Barry
Julia Chen
Sun Jun
Karson Elmgren
Kat Lyness
Katherine Lee
Kristy Loke
Lee Kwee Geak
Leslie Teo
Meng Ling Yu
Lisa Soder
Madhulika Srikumar
Malcolm Murray
Mark Brakel
Mark Nitzberg
Mary Phuong
Matthew Jagielski
Max Fenkell
Miro Plueckebaum
Kim Myuhng Joo
Hu Naying
Neil Davison
Nicolas Miailhe
Niki Iliadis
Nur Syahidah Sahrom
Ong Chen Hui
Pradeep Varakantham
Rebecca Finlay
Renata Dwan
Robert Opp
Rumman Chowdhury
Saad Siddiqui
Sabina Nong
Sam Ramadori
Sami Jawhar
Samuel Boger
Sara Hooker
Ying Shao Wei
Sebastian Hallensleben
Shinyuk Kang
Sophie Toura
Sreejith Balakrishnan
Stephanie Kasaon
Stephen Clare
Summer Yue
Sunny Yuqing Sun
Supheakmungkol Sarin
Tian Tian
Tim Schreier
Tori Westerhoff
Urvashi Aneja
Wayne Tee
Lu Wei
Xu Wei
Zhang Wenxuan
Hu Xia
Yang Xiaofang
Pan Xudong
Xiao Yajun
Yifan Jia
Tan Yong Khiam
Yuejin Du
Yuma Kurihara
Tan Zhi Xuan
Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential … (see more)to embracing AI with confidence. The 2026 Singapore Consensus is an outcome of the second International Scientific Exchange on AI Safety, bringing together over 100 contributors spanning 13 countries from frontier developers, government safety institutes, academia, and civil society. Building on the 2025 report, it presents a global understanding of technical AI safety research problems of top priority, now with a dedicated focus on societal resilience and on managing the risks of increasingly autonomous AI agents.
Vision Transformers Learn Gestalt-Like Figure-Ground Cues from Natural Images
Matthias Tangemann
Benjamin Lo
Zygmunt Pizlo
Dirk B. Walther
Sven Dickinson
Figure-ground organization in the human visual system relies on several shape-based cues, including surroundedness, convexity, and symmetry.… (see more) While these cues have been extensively studied using abstract stimuli, little is known about how they operate under natural conditions or how they arise from the statistics of natural scenes. Deep neural networks offer a promising path forward: a model that relies on the same figure-ground cues as humans would provide tractable experimental access to the underlying mechanisms. In this study, we evaluate shape-based figure-ground organization in Vision Transformers (ViTs), for which prior work has demonstrated the emergence of object-based grouping. We test 25 ViTs spanning supervised and self-supervised training objectives, by fitting linear probes to predict figure-ground assignment from intermediate patch representations using both natural images and controlled artificial stimuli that isolate individual cues. Our results show that ViTs robustly encode surroundedness and convexity, and that probes trained on natural images generalize zero-shot to artificial stimuli across several models. For symmetry we observe mixed results: the cue is encoded for uniformly colored but not for textured regions. Taken together, our findings demonstrate that Gestalt-like figure-ground cues can be learned from natural scene statistics and position ViTs as a compelling model system for studying the computational mechanisms of perceptual organization. Code and data is available at https://github.com/mtangemann/mlvbench
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence
Martin Zborowski
Alberto Tosato
Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent mi… (see more)salignment, and goal misgeneralization. Recent evidence suggests that some misalignment behaviors are encoded as linear structure in activation space, making it tractable via activation steering, which could be used as a lightweight runtime defense. We implement three methods: Steer-With-Fixed-Coefficient (SwFC), which applies uniform additive steering, and two novel projection-aware methods, Steer-to-Target-Projection (StTP) and Steer-to-Mirror-Projection (StMP), that use a logistic regression decision boundary to selectively intervene only on tokens whose activations fall below the threshold. We evaluate these methods on two threat models, dishonesty and dismissiveness, using malicious system prompts as a controlled proxy for misalignment. We conduct our experiments on two architectures (Llama-3.3-70B-Instruct and Qwen3.6-27B). All methods substantially recover alignment. StTP and StMP preserve general capabilities (MMLU, MT-Bench, AlpacaEval) better than uniform steering. Finally, we show that our honesty steering generalizes to out-of-distribution scenarios: a single honesty direction extracted from the aligned model significantly raises scores on the MASK benchmark, suppresses deception in multi-agent settings (Among Us), doubles the hidden-behavior discovery rate on AuditBench, and restores honesty in an emergently misaligned model
Apriel-Reasoner: RL Post-Training for General-Purpose and Efficient Reasoning
Rafael Pardinas
Ehsan Kamalloo
Building general-purpose reasoning models using reinforcement learning with verifiable rewards (RLVR) across diverse domains has been widely… (see more) adopted by frontier open-weight models. However, their training recipes and domain mixtures are often not disclosed. Joint optimization across domains poses significant challenges: domains vary widely in rollout length, problem difficulty, and sample efficiency. Further, models with long chain-of-thought traces increase inference cost and latency, making efficiency critical for practical deployment. We present Apriel-Reasoner, trained with a reproducible multi-domain RL post-training recipe on Apriel-Base, a 15B-parameter open-weight LLM, across five domains using public datasets: mathematics, code generation, instruction following, logical puzzles, and function calling. We introduce adaptive domain sampling that preserves target completed-rollout ratios despite heterogeneous rollout dynamics, and a difficulty-aware length penalty that, at no additional training overhead, encourages longer reasoning for difficult problems and shorter traces for easy ones. Trained with a strict 16K-token output budget, Apriel-Reasoner remains effective at a 32K output budget and improves over Apriel-Base on AIME 2025, GPQA, MMLU-Pro, and LiveCodeBench while producing 30-50% shorter reasoning traces. Among the evaluated open-weight models of similar scale, it improves the accuracy-token trade-off using a fully public-data recipe.
Efficient Safety Alignment of Language Models via Latent Personality Traits
David Williams-King
Adam Oberman
Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternat… (see more)ives. Latent Adversarial Training (LAT) is among the most effective defenses, but can degrade utility and requires training on large datasets of harmful prompts. We introduce Latent Personality Alignment (LPA), which replaces explicit harm refusal with adversarial training on just 66 harm-agnostic statements drawn from psychometric personality literature. We hypothesize that personality-anchored representations share latent structure with harm avoidance, so adversarially stabilizing them implicitly constrains the subspace exploited by jailbreak attacks. LPA achieves near-zero attack success rates on HarmBench across direct requests and five jailbreak methods, despite never seeing harmful content during training and no loss of performance on standard benchmarks. Moreover, the training process is lightweight; the entire procedure completes in minutes on a single GPU and uses 75x fewer examples than standard LAT. Extensive ablations demonstrate the robustness, efficiency, and generalization of our method.
Empirical Comparison of Unified Benders Cuts for Multi-Commodity Fixed-Charge Network Design
Eric Larsen
Jean-François Cordeau
Antonio Frangioni
Among the many types of acceleration techniques designed to improve the performance of Benders decomposition, unified cut generation schemes… (see more) have recently attracted a keen interest. Unified cuts aim for a better balance between the generation of optimality and feasibility cuts, while also providing a way to compare the strengths of different feasibility cuts. Our goal is to assess the experimental performance of a broad selection of unified and distinct Benders cuts in the context of the multi-commodity fixed-charge network design problem (MCFNDP). We express under a common mathematical structure and notation the construction of each unified or distinct Benders cut considered. We also explain how the generic formulations of the Benders cuts can be specialized to conform to the specifications of the MCFNDP. In addition, we suggest bespoke methods for comparing the performance of several solution methods when the benchmark is made up of heterogeneous problem instances. We report the results of a systematic empirical analysis comparing the performances of 50 Benders methods involving unified or distinct cuts in applications to a common testing bench made up of standardized MCFNDP instances. The analysis identifies a small number of leading Benders methods, namely those featuring the static Brandenberg-Stursberg cuts and the Hosseini-Turner l1-deepest cuts. In addition, we also report results obtained by using both Gurobi and CPLEX as the supporting solver to the SMS++ computation library.
LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning
LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc r… (see more)emoval methods. Unlearning has emerged as a promising solution, with state-of-the-art(SOTA) methods often following a localize-first, unlearn-second paradigm that targets specific model parameters. However, existing benchmarks evaluate unlearning solely at the output level, leaving open the question of whether unlearning truly erases knowledge from a model's parameters or merely obfuscates it, a concern reinforced by the success of resurfacing attacks. To bridge this gap, we introduce LACUNA: the first unlearning testbed with ground-truth parameter-level localization. LACUNA injects PII of synthetic individuals into predefined parameters of 1B and 7B OLMo-based models via masked continual pretraining, enabling direct evaluation of whether unlearning targets the weights responsible for knowledge storage. We use LACUNA to benchmark current SOTA unlearning methods and find that, despite strong output-level performance, existing methods are highly imprecise and susceptible to resurfacing attacks. We further show that when localization is successful, even a simple gradient-based unlearning method achieves strong erasure and robustness to resurfacing attacks, highlighting the importance of precise unlearning. We release LACUNA to complement behavioral evaluations and drive further advances in robust, localization-based unlearning.
LLM2Vec-Gen: Generative Embeddings from Large Language Models
LLM-based text embedders typically encode the semantic content of their input. However, embedding tasks require mapping diverse inputs to si… (see more)milar outputs. Typically, this input-output is addressed by training embedding models with paired data using contrastive learning. In this work, we propose a novel self-supervised approach, LLM2Vec-Gen, which adopts a different paradigm: rather than encoding the input, we learn to represent the model's potential response. Specifically, we add trainable special tokens to the LLM's vocabulary, append them to input, and optimize them to represent the LLM's response in a fixed-length sequence. Training is guided by the LLM's own completion for the query, along with an unsupervised embedding teacher that provides distillation targets. This formulation helps to bridge the input-output gap and transfers LLM capabilities such as safety alignment and reasoning to embedding tasks. Crucially, the LLM backbone remains frozen and training requires only unlabeled queries. LLM2Vec-Gen achieves state-of-the-art self-supervised performance on the Massive Text Embedding Benchmark (MTEB), improving by 9.3% over the best unsupervised embedding teacher. We also observe up to 43.2% reduction in harmful content retrieval and 29.3% improvement in reasoning capabilities for embedding tasks. Finally, the learned embeddings are interpretable and can be decoded into text to reveal their semantic content.