Publications

Learning Implicit Feasibility Constraints for Real-World Routing and Scheduling: Application to Log Transportation
Abdelhakim Abdellaoui
Ayoub Boufous
Issmail El Hallaoui
François Aubé
Mouloud Amazouz
Real-world vehicle routing and scheduling problems involve complex operational rules and feasibility constraints typically formulated as mix… (see more)ed-integer linear programs (MILP). However, optimization tools are built around a fixed set of hard-coded constraints, while in practice this set evolves as new rules or preferences emerge, seasonally or permanently. Updating it requires modeling and operations research skills that planners rarely have, so generated plans are routinely adjusted by hand based on practical knowledge. Building on recent work that uses machine learning to recover such hidden constraints, we propose a data-driven constraint-learning approach that trains three complementary predictors, a Graph Neural Network (GNN), a decision tree, and a linear regression, on historical execution data from a log-truck routing and scheduling problem (
The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show
Parsa Esmati
Katja Hofmann
Majid Mirmehdi
Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simula… (see more)tors. Yet it remains unclear whether these models internally encode physical structure, or merely reproduce motion patterns seen during training. We study this question by probing video diffusion models along latent trajectories corresponding to real videos with known physical plausibility. To obtain such trajectories, we approximately invert the deterministic sampling process by integrating the learned velocity field backward from a clean video latent to noise, giving access to the model's intermediate states and attention maps. Using these recovered trajectories, we show that physical plausibility is linearly decodable from diffusion transformer states across IntPhys and InfLevel, reaching around 81.27% average accuracy and outperforming dedicated representation-learning baselines such as V-JEPA and VideoMAE. Surprisingly, this signal is absent from the VAE latent input and emerges inside the denoising transformer itself, despite the model not being trained with a self-supervised predictive objective. These findings suggest that physically meaningful representations can arise as a byproduct of generative denoising.
Would you still call this Dax? Novel Visual References in VLMs and Humans
Vision-language models (VLMs), like human learners, are frequently exposed to new visual concepts, but how they map novel visual references … (see more)to language after exposure remains largely underexplored, particularly when those references contradict prior knowledge from pre-training. To study this, we present the Novel Visual References Dataset (NVRD): 19,176 images spanning 90 visual concepts across different levels of visual novelty, each with up to 20 increasingly perturbed versions of the original object to probe generalization. Unlike prior work on visual augmentations of familiar concepts, NVRD comprises entirely novel, open-ended stimuli constructed from scratch, mirroring how humans encounter genuinely new concepts. We evaluate 3 open- and 2 closed-source models alongside 2,400 human judgments for direct human-model comparison, and find that (i) models struggle to acquire novel concepts in-context when they contradict prior knowledge, and (ii) while models and humans show correlated sensitivity to visual perturbations, models significantly overgeneralize, extending learned labels to stimuli that humans reject. We contribute NVRD as a corpus and benchmark for research on visual concept learning in both humans and machines.
EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents
Tong Nie
Yuewen Mei
Yihong Tang
Junlin He
Jie Deng
Jian Sun
Wei Ma
Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximi… (see more)zing adversariality to expose failures while preserving realism. Existing methods usually manage this trade-off with handcrafted heuristics, confining generation to known priors and overlooking underexplored patterns. While recent open-ended agentic evolution can push this limit, unconstrained general agents lack strict simulator grounding and tend to collapse the multi-objective tension into single-scalar maximization. Here we present EvoDrive, the first automated, LLM-based agentic evolution framework for multi-objective scenario generation. EvoDrive employs a simulator-grounded actor-critic architecture where a memory-driven actor iteratively proposes improvements to the generators and critics filter out implausible candidates, and a self-evolving world evaluator routes promising proposals to optimize simulation budgets. EvoDrive further maintains a Pareto archive of evaluated candidates to preserve diverse attack-realism trade-offs and guide future evolution via simulation feedback. Benchmark results on MetaDrive and CARLA show that EvoDrive not only significantly expands the Pareto frontier across various generators, but also produces valuable scenarios for policy training.
FOLD: Fuzzy Online Deduplication for Very Large Evolving Datasets via Approximate Nearest Neighbor Search
Nelson Bore
Pritish Mishra
Constantin Adam
Eyal de Lara
Fuzzy deduplication is key to constructing large language model training corpora. However, classic Locality-Sensitive Hashing pipelines scal… (see more)e poorly as corpora grow and are ill-suited to continuous ingestion. We present FOLD (Fuzzy Online Deduplication), an online fuzzy deduplication system that delivers high recall and throughput for evolving datasets. FOLD maintains an incrementally updated HNSW index over admitted documents, retrieving a small, high-quality candidate neighborhood for each incoming document instead of repeatedly rebuilding global buckets or rescanning the accumulated corpus. To our knowledge, FOLD is the first online fuzzy deduplication system to use HNSW. However, applying Jaccard similarity out of the box causes score crowding, making graph traversal unreliable within a small number of steps. FOLD addresses this with a bitmap representation that provides a more discriminative, Jaccard-aligned signal during HNSW search. Across four LLM-scale datasets (LM1B, C4, RealNews, and Common Crawl), FOLD stays fast and accurate as the corpus grows: at the largest evaluated scales, it maintains 93-97% recall and achieves up to 2.09x higher throughput than competing alternatives, whose best recall reaches only 76%.
PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification
Simulation environments are useful for both robot policy learning and planning verification and validation. Traditionally, the process of cr… (see more)eating a simulation was onerous. Creating a bespoke simulation environment for each individual environment that a robot would operate in was simply infeasible. In this work, we introduce PerceptTwin, a fully automatic pipeline that constructs interactive simulations directly from semantic scene representations produced by a robot's perception stack. PerceptTwin combines open-vocabulary object maps with 3D asset generation, affordance prediction, and commonsense condition checking. These interactive simulations can be used to validate and refine plans before they are executed on the robot hardware. Borrowing from the AI alignment literature, we also introduce an LLM judge that verifies plan correctness and alignment with human preferences. Experiments show that PerceptTwin feedback allows LLM planners to refine plans, enhance safety, and resist harmful black-box prompting attacks. In our suite of tasks, PerceptTwin improves plan success by an average of approximately 39% for GPT5, GPT5Mini, and GPT5Nano planners. Additionally, PerceptTwin also improves human plan verification by up to 18% on average for plans that fail due to unfilled skill preconditions. Our results demonstrate the potential of open-vocabulary scene simulation from robot perception as a foundation for safer, more reliable robot planning.
Rotate2Think: Geometric Priming via Orthogonal Rotation to Improve Language Model Reasoning
Christopher Pal
Reasoning models achieve strong performance on challenging tasks by generating explicit intermediate reasoning traces before producing a fin… (see more)al answer. Yet the internal structure of representation space when reasoning remains poorly understood: how do a model's hidden representations differ during thinking versus the embeddings of the input prompt, and can this structure be exploited to elicit stronger reasoning at inference time? We show that both input embeddings and thinking embeddings (mean-pooled last-layer hidden states over the prompt and reasoning trace, respectively) exhibit extremely high conicity, with all vectors clustering tightly around a single mean direction. Crucially, these mean input and thinking directions are non-collinear, with thinking embeddings occupying a geometrically distinct region of embedding space across many different models and benchmark tasks. This observation motivates casting the input-to-thinking transition as a rotation problem admitting a closed-form solution via orthogonal Procrustes analysis. We propose Rotate2Think, a training-free method that estimates this rotation from a small set of correctly solved examples and injects the resulting synthetic thinking vector between thinking delimiters at inference time, providing a geometric primer at the onset of the reasoning trace. Evaluated across multiple benchmarks and model families, Rotate2Think improves accuracy in 30 of 32 model-benchmark configurations across mathematics, science, and code tasks, and generalizes zero-shot to multimodal reasoning on MATH-Vision.
Self-assembled chamber-like cardiac organoids for modeling cardiac chamber formation and cardiotoxicity assessment
Xinle Zou
Fanwen Wang
Huilin Zheng
Xianzhuang Liu
Tianci Kong
Rui Jiang
Yingying Guo
Yu Liang
Bo Wang
Duanqing Pei
$\textsf{SKILL.nb}$: Selective Formalization and Gated Execution for Durable Agent Workflows
Amine El hattami
Christopher Pal
AI agents increasingly convert past experience into reusable artifacts such as code, workflows, and procedural memories. Reuse improves effi… (see more)ciency, but these artifacts can also carry obsolete assumptions across interface drift, repeated repairs, or changing task distributions, especially in web automation. We introduce SKILL.nb, a framework for governing reusable agent workflows through evidence-calibrated lifecycle policies. Its key mechanism is *selective formalization*: execution evidence decides which workflow steps should become executable code, which should remain natural-language-guided, and when those choices should be revised. SKILL.nb stores workflows as auditable, versioned notebooks that interleave natural-language guidance, multi-language executable cells, validation gates, fallback paths, and multimodal evidence such as outputs, screenshots, and error traces. At runtime, SKILL.nb performs *gate-conditioned execution*: unlike all-or-nothing scripts, each step can execute code when its gates validate, or fall back locally to an NL procedure or step intent when drift invalidates the executable realization. Cell-level records of attempted realizations, gate outcomes, outputs, screenshots, and fallbacks make both workflow updates and executions auditable. On WebArena-Verified, SKILL.nb achieves 53.7% single-round success, improving over the strongest baseline by 3.9 percentage points. Across three re-executions, it retains 91.7% of initially successful tasks, 15.5 points above the next best method. Under bounded repair, it recovers 72.9% of subsequent failures while limiting post-repair regressions to 4.2%, compared with 15.0-17.0% regression rates for persistent baselines. It also leads the compared methods on Mind2Web cross-website and cross-domain splits. In a realistic GitLab migration test, SKILL.nb preserves performance when reusing frozen state learned on GitLab 15.7, with frozen-versus-fresh target-version gaps of only -1.7 points on GitLab 16.11 and +0.6 points on GitLab 18.9; the least-degraded persistent baseline drops by 10.6-11.1 points. These results identify lifecycle governance and gate-conditioned execution as reliability axes beyond one-shot task success. Code, data, and evaluation scripts will be released after review.
Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning
Anthony GX-Chen
Gheorghe Comanici
Zaheer Abbas
Eser Aygün
David Smalling
Shibl Mourad
Andre Barreto
Mark Rowland
Classical reinforcement learning (RL) typically seeks a deterministic policy that maximizes the expected sum of a scalar reward. Yet, modern… (see more) applications such as language model fine-tuning or scientific discovery demand diversity. Existing remedies such as entropy regularization or diversity bonuses often require fragile trade-offs that sacrifice performance for stochasticity or rely on heuristic metrics that can misalign policy rankings. We argue that diversity is more naturally understood as the rational response to uncertainty in the reward. When the reward function is not perfectly known--as is the case with ambiguous preferences or imperfect reward models--committing to a single action can be sub-optimal. Building on this, we propose a fundamental reformulation of the RL objective by replacing the scalar reward with a distribution over reward functions, and applying a non-linear objective over sets of actions. The result is a framework in which calibrated behavioural diversity emerges naturally, remains controllable through the reward function distribution, and is obtained without sacrificing expected reward. Focusing on the contextual bandit setting, we derive a principled gradient estimator for this objective and prove that our formulation naturally generalizes both vanilla policy gradient and more recently developed action-set approaches. Our empirical results demonstrate that this framework offers a robust and theoretically grounded alternative for complex RL tasks where the traditional formulation of the problem fails to induce the desired breadth of agent behaviour.
Accelerated green material and solvent discovery with chemistry- and physics-guided generative AI
Eslam G. Al-Sakkari
Marzouk Benali
Olumoye Ajao
Daria C. Boffito
Antibiotic dispensing practices and determinants among informal healthcare providers in low- and middle-income countries: a mixed-methods scoping review
M V Pai
Poshan Thapa
Giorgia Sulis
Charity Oga-Omenka
Buna Bhandari
Sumanth Gandra
Genevieve Gore
Surbhi Sheokand
Prachi Shukla
Meera Tandan
Diwash Timalsina
Shweta Bohora
Swostika Thapaliya
Anupama Bhusal
Chandrashekhar Joshi
Md Asadullah
Mili Dutta
Samira Abbasgholizadeh Rahimi
Introduction Antimicrobial stewardship efforts in low- and middle-income countries (LMICs) largely focus on qualified practitioners, yet inf… (see more)ormal healthcare providers (IPs) deliver much of the primary care. Although these providers frequently dispense antibiotics, their practices remain poorly documented and are not captured in existing surveillance systems.Methods Using the Joanna Briggs Institute methodology, this scoping review synthesised evidence on antibiotic dispensing and its determinants among IPs in LMICs. Nine databases (MEDLINE, EMBASE, SCOPUS, Global Health, CINAHL, Web of Science, LILACS, African Journals Online via Africa-Wide Information and Index Medicus for the South-East Asia Region) were searched, yielding 12 095 records, of which 37 studies met the inclusion criteria.Results Across 31 studies reporting dispensing practices, antibiotic use by IPs varied widely: 18%–74% in studies using standardised methods, 5%–96% in provider-reported studies and 2%–86% in consumer-reported studies. Eight qualitative studies identified key behavioural and contextual determinants shaping dispensing, including limited knowledge, experience-based learning, patient expectations, peer and pharmaceutical influence, perceived consequences of withholding antibiotics and economic incentives.Conclusion Antibiotic dispensing by IPs is widespread and represents a large but unmeasured component of antibiotic use in LMICs. These findings highlight a critical gap in antimicrobial resistance surveillance and highlight the need for stewardship strategies that effectively engage this provider group.