Publications

GeCo: Fine-Grained Cross-View UAV Geo-Localization via Geometric Consistency
Tianqi Ying
Kaikai Pan
Qidi Zhong
Wenyuan Xu
This technical report presents GeCo, our solution for the challenge at the 4th Workshop on UAVs in Multimedia (UAVM 2026), ACM Multimedia 20… (voir plus)26. We use a partially fine-tuned DINOv2 to extract features and fuse via cross-attention. The relative pose (heading and translation) is then regressed and regularized by geometric consistency loss. GeCo achieves a 0.00583 final score on the PairUAV dataset. Code is available at https://github.com/Link-ying/GeCo.
Confidence Intervals for the Return Process in Markov Decision Processes
In this work, we derive confidence intervals for the return process in discounted reward Markov Decision Processes with continuous state and… (voir plus) action spaces. These confidence bounds depend only on the statistics of the value function, which may be derived using dynamic programming. In the special case of MDPs with uniformly bounded value functions, simpler confidence intervals are provided for the return process. Finally, we study the effect of epistemic uncertainty on the derived confidence intervals. Numerical examples are provided to show how these bounds may be used in practice.
TEA: Text Encoder Alignment for Robust Concept Erasure in Text-to-Image Models
Text-to-image diffusion models can be misused to generate harmful content through adversarial or paraphrased prompts that bypass built-in sa… (voir plus)fety mechanisms. Existing concept erasure methods often suffer from limited robustness against adversarial prompts, degradation of benign generation quality, or reliance on inference-time interventions that introduce persistent computational overhead. To address these limitations, we formulate concept erasure as a domain alignment problem in the text representation space. We propose a lightweight Text Encoder Alignment framework (TEA) that fine-tunes only the text encoder while keeping the generative backbone fully frozen. Given concept--anchor prompt pairs, our method trains a discriminator to distinguish token-level representations of concept-containing prompts from those of safe anchor prompts, while updating the text encoder to make these representations indistinguishable. TEA introduces zero inference-time overhead and requires only a small number of fine-tuning steps, making it highly efficient to deploy at scale. Despite this efficiency, TEA achieves state-of-the-art erasure robustness against black-box and white-box adversarial attacks on Stable Diffusion v1.4, while preserving generation quality on benign prompts. Furthermore, TEA is model-agnostic and achieves the lowest attack success rate on Stable Diffusion v3.5, extending concept erasure to a Rectified Flow Transformer architecture with T5 conditioning where prior methods remain largely unexplored. Code is available at \href{https://github.com/alirezafarashah/TEA.git}{https://github.com/alirezafarashah/TEA.git}
Three Winning AI Strategies
Maxime C. Cohen
Eddy Hage-Youssef
Daniel McCarthy
D. Daniel Sokol
AI competition is widely framed as a winner-take-all race. Mobile consumer data for the AI assistant category from 2023 through 2025 suggest… (voir plus)s otherwise: three strategies coexist, and all are growing. Scale (ChatGPT), adjacency (Gemini), and premium specialization (Claude) capture value in strategically distinct ways. Major launches reward the launcher without measurably harming rivals. These archetypes are not unique to AI; they recur whenever a disruptive technology opens a market before competition settles. For managers, the most important question is which strategy their distribution, economics, and customer base can sustain.
What are we measuring? A critical examination of MO-MuJoCo and evaluation consistency in continuous multi-objective reinforcement learning
Although Multi-Objective Reinforcement Learning (MORL) research relies heavily on MO-MuJoCo as its go-to continuous control benchmark, the v… (voir plus)alidity of the conclusions drawn from it remain underexamined. In this paper, we first discuss three structural limitations of MO-MuJoCo; 1) its objectives are decomposed from pre-existing scalar rewards rather than independently motivated goals; 2) environments repeat the same underlying trade-off structure across varied locomotion morphologies, providing \textit{surface variety} without genuine \textit{problem diversity}; and 3) empirically approximated Pareto fronts appear broadly convex across research, potentially failing to stress-test the limitations of scalarization-based techniques. Setting these concerns aside, we further demonstrate that algorithmic rankings under MO-MuJoCo are highly sensitive to often undocumented evaluation choices in research papers. Across five evaluation axes, including reference point selection, weight distribution, normalization, return type, and front extraction method, pairwise algorithm rankings reverse in up to 47\% of configurations. Variance decomposition reveals that normalization alone accounts for nearly 69\% of hypervolume variance, suppressing the algorithm impact. Ultimately, we argue that progress in MORL research requires not only increased scrutiny of the benchmarks we rely upon, but also greater clarity in how results obtained within them are reported.
Artificial intelligence in implementation science and practice: a living scoping review
Guillaume Fontaine
Olivia Di Lalla
Rachael Laritz
Jeremiah Durran
Chi Zhang
Jeffery Chan
Laura Crump
Alenda Dwiadila Matra Putra
Ruopeng An
Rinad S. Beidas
Christine Fahim
Elvin Geng
Ian D. Graham
Janna Hastings
Sylvie D. Lambert
France Légaré
Susan Michie
Byron J. Powell
Justin Presseau … (voir 6 de plus)
Joseph Elias
Thomas Rudge
Sharon E. Straus
James Thomas
Vivian Welch
Natalie Taylor
Artificial intelligence (AI) encompasses computational systems that perform tasks typically requiring human intelligence, including machine … (voir plus)learning, generative AI, and agentic applications. The use of AI to support implementation activities is growing, but its applications and evaluation remain poorly characterised. We conducted the first cycle of a living scoping review to identify how AI is being used and evaluated across implementation science and practice. We followed JBI and Cochrane guidance and reported findings per PRISMA-ScR and PRISMA-LSR. We searched six databases through April 6, 2026, and included sources describing or evaluating AI to support implementation research or practice activities. We excluded adjacent uses such as knowledge synthesis automation. We extracted study characteristics, AI approaches, implementation tasks, evaluation methods, outcomes, and risks, and synthesised findings descriptively. We identified 7,203 records and included 40 sources, 34 (85%) of which were published since 2021. Thirty-one sources described, developed, or evaluated a specific AI system, comprising 22 primary research articles, three protocols, three conference abstracts, and three other sources. The remaining nine conceptual, methodological, framework or review articles discussed potential applications of AI in implementation science. Across all 40 sources, AI was most often used or proposed to support evaluation (25/40), implementation strategy selection and tailoring (23/40), barrier and facilitator assessment (21/40), and implementation monitoring (20/40). Among the 31 sources involving a specific AI system, the use of conventional machine learning was most common (10/31), followed by generative AI (7/31), multicomponent AI systems or studies comparing AI approaches (5/31), and non-generative deep learning (3/31). Implementation-related outcomes were reported or prospectively specified in 25/31 sources, technical performance in 19/31, human-centred outcomes in 12/31, time or efficiency in 10/31, clinical or health system outcomes in 9/31, equity in 3/31, and cost or resource outcomes in 2/31. Risks were discussed in about half of sources, although assessment of harms and environmental or societal consequences was rare. AI is in a nascent stage of supporting implementation science and practice. Reported use is narrow and methodologically underdeveloped, and much routine use is likely unpublished. The field needs rigorous, comparative, prospective, and equity-attentive research to establish whether and how AI improves implementation methods, processes, and outcomes. Open Science Framework, May 2025: https://doi.org/10.17605/OSF.IO/2Q5DV
Determinants of functional burden pleiotropy and gene dosage responses across human traits
Sayeh Kazem
Kuldeep Kumar
Jane Yang
Florian Benitiere
Josephine Mollon
Thomas Renne
Laura M. Schultz
Emma E.M. Knowles
Worrawat Engchuan
Omar Shanta
Bhooma Thiruvahindrapuram
Jeffrey R. MacDonald
Celia M. T. Greenwood
Stephen W. Scherer
Laura Almasy
Jonathan Sebat
David C. Glahn
Sébastien Jacquemont
Pleiotropic and monotonic effects of gene dosage are central to understanding comorbidities in developmental pediatric and psychiatric disor… (voir plus)ders, yet the underlying biological processes are not well characterized. Here we develop a functional burden analysis to investigate the association of all protein-coding copy-number variants, genome-wide, with 43 complex traits in approximately 500,000 UK Biobank participants. We test variant associations disrupting 172 tissue or cell-type gene sets, finding associations for all traits, which we replicate in the All of Us cohort. Functional burden pleiotropy, defined as the number of traits significantly associated with a gene set, correlates with genetic constraint and is higher for brain than non-brain functions, even after normalizing for genetic constraint. Levels of pleiotropy, measured by burden correlation, are similar in deletions and loss-of-function single-nucleotide variants, and higher than in common variants and duplications. Most gene dosage responses are non-monotonic, with deletions and duplications showing same-direction effects, and monotonic responses decrease with genetic constraint. We observe associations between functional gene sets and traits for either deletions or duplications, but rarely both, with negatively correlated effect sizes. Together, these results link genetic constraint and brain-specific mechanisms to the whole-body multimorbidity of neurodevelopmental and psychiatric conditions. Gene dosage can help explain comorbidities in developmental pediatric and psychiatric disorders. Here, the authors map how rare copy-number variants disrupting tissue and cell-type gene sets shape 43 human traits.
Analyzing Flexible Search Distributions in Black-Box Optimization with Normalizing Flow-based Estimation of Distribution Algorithms
Black-box optimization often requires search distributions that can adapt to complex geometric structures under limited evaluation budgets. … (voir plus)We propose NF-EDA, a Normalizing Flow-based Estimation of Distribution Algorithm that replaces fixed Gaussian models with a learned, flexible search distribution. Beyond optimization performance, our goal is to better understand how increased distributional expressiveness affects search behavior. In contrast to classical Gaussian-based methods, NF-EDA can adapt to curved, asymmetric, and non-elliptical regions of the search space, enabling broader yet structured exploration during early stages of optimization. By tracking the evolution of the learned distribution over iterations, we analyze how NF-EDA reshapes its sampling behavior compared to predefined parametric approaches such as CMA-ES and Gaussian EDAs. Experimental results on selected COCO BBOB functions, including Rastrigin, Schwefel, Lunacek bi-Rastrigin, and Rosenbrock, show that NF-EDA achieves faster early progress and reduced variability across runs, particularly in higher-dimensional settings. An ablation against a Gaussian EDA with matching update rules further demonstrates that these effects arise from the learned flow transformation rather than from the surrounding EDA procedure alone. These findings highlight the importance of flexible search distributions for understanding and improving model-based black-box optimization.
Conductivity Preservation of Aryl Diazonium-Functionalized Bilayer Graphene Probed by Operando Hall
Bahar Molavi
Claudia M. Bazán
Tony Vuu
Jade Cimmino
Sébastien Côté
Delphine Bouilly
Thomas Szkopek
Abstract Covalent functionalization of graphene is a means to achieve robust immobilization of functional groups, but it results in a signif… (voir plus)icant loss of electrical conductivity in monolayer graphene (MonoG) due to the introduction of strong scattering by point defects. This work presents gate-activated covalent functionalization of bilayer graphene (BiG) integrated with an operando Hall characterization to measure the charge carrier density and mobility in real-time. Using an integrated Ag/AgCl gate electrode to modulate the BiG Fermi level, we achieve precise control over the covalent grafting of aryl diazonium groups on BiG. We demonstrate that BiG preserves the majority of its conductivity after functionalization by losing only 20% of its conductivity, whereas MonoG experiences 80% reduction in conductivity under similar conditions. BiG enables a significantly wider tuning range for surface coverage while preserving the conductivity. Operando Hall measurements of BiG functionalization reveal that the observed conductivity decrease is primarily driven by a reduction in charge mobility due to short-range scattering, while the charge carrier density changes only modestly. Furthermore, we characterize the impact of covalent attachments on the graphene density of states (DOS) and interfacial charge storage through Hall-derived quantum capacitance and electrochemical impedance spectroscopy (EIS). Finally, we demonstrate the application of carboxyphenyl functionalized BiG for pH sensing and present a site-binding model that describes the electrostatic coupling between site-binding coverage and the charge density within the conducting channel. This architecture provides a robust and tunable platform for graphene FET sensors.
Customers’ Multihoming Behavior in Ride-Hailing: Empirical Evidence from Uber and Lyft
Sandeep Chitla
Maxime C. Cohen
Srikanth Jagabathula
Dmitry Mitrofanov
Problem definition: Are customers loyal to a ride-hailing platform or they see this service as a commodity and multihome (i.e., check severa… (voir plus)l platforms before booking a ride)? Using a large panel dataset on ride-hailing transactions, we investigate to what extent customers multihome. Our dataset offers a unique opportunity to study this question as we observe the repeated choices of riders for both Uber and Lyft. Our dataset comprises more than 1.4 million rides completed by 162 thousand riders in NYC in 2018. Methodology/results: We develop a comprehensive structural model that incorporates both operational (price and waiting time) and behavioral factors (e.g., platform stickiness) to explain riders’ choices. Our model also accounts for the dynamic interactions between customers and platforms by assuming that riders update their beliefs on price and waiting time in a Bayesian fashion. Finally, the riders’ propensity to multihome is modeled by incorporating the consideration set formation of customers into our framework. We find that riders’ choices are not fully explained by operational factors, hence indicating that customers view the platforms as differentiated service providers. While 83.4% of riders took rides with a single platform, our model shows that even the remaining 16.6%, who used both Uber and Lyft at least once, considered both platforms only 43.4% of the time. Managerial implications: It is crucial for ride-hailing platforms to capture this single (or multi)-homing behavior while designing price discounts. Specifically, personalized discounts may be ineffective if the platform is not part of the customer’s consideration set. Our results show that targeting customers earlier in their lifecycle can enhance the platform’s market share by 77.56% more than their current discounting strategy. We also find that targeting customers with low search friction results in a 24.78% increase in market share relative to targeting customers with high search friction.
IP Protection in the Era of Visual Generative AI: A Survey
Shunchang Liu
Han Yu
Cao Yang
Chaochao Chen
Yuping Yan
Yaochu Jin
Lingjuan Lyu
The rapid evolution of visual generative AI has introduced a wide range of intellectual property risks, spanning the unauthorized learning, … (voir plus)reproduction, extraction, misuse, and redistribution of protected data and model assets. To address these risks, a growing body of technical defenses has been proposed. However, existing surveys typically organize this literature by lifecycle stage or technical mechanism, which can obscure the protective intent of different methods. This survey presents a two-dimensional taxonomy for IP protection in visual generative models. The primary axis is a Control Logic View, which classifies methods into Information Exposure Control, Generative Behavior Constraint, and Attribution&Accountability according to the risk variable they regulate. The secondary axis distinguishes Data IP from Model IP as cross-cutting asset dimensions. Under this framework, we systematically review protection methods, align evaluation protocols with protection objectives, and discuss open challenges including proactive model-level safeguards, standardized evaluation, robustness against adaptive attacks, and explainable evidence. This survey aims to offer a principled, systematic, and easy-to-follow overview for both new and experienced researchers in visual generative AI IP protection.
Divergent specializations for motion-driven representations in higher lateral and dorsal visual areas
Sophia Robert
Maryam Vaziri-Pashkam
The human visual system integrates both static and dynamic information to support form and shape perception, yet the computational principle… (voir plus)s underlying the integration of motion for object recognition remain unclear. Artificial neural networks (ANNs) offer a computational framework for developing and testing hypotheses about these principles: if ANNs trained on motion-related tasks develop representations that align with brain activity and support object categorization, this would suggest that the training objectives and architectural constraints of these networks may capture key aspects of motion processing in biological visual systems in general, and motion processing for object recognition, in particular. Here, we investigated this question using “object kinematograms”, stimuli in which object form is conveyed solely through motion cues. We measured neural responses of two higher regions of the lateral and the dorsal visual pathways, respectively, with strong sensitivity to dynamic cues from objects: lateral occipitotemporal cortex (LOT bio ), and left supramarginal gyrus (SMG lh ), as well as primary visual cortex (V1). We compared brain responses to representations extracted from two neural networks: SlowFast, a dual-pathway architecture trained on action recognition that processes slow- and fast-varying visual information with cross-pathway integration, and DorsalNet, a model of the primate dorsal visual pathway trained on embodied self-motion estimation. Representational similarity analysis revealed distinct representational profiles across brain areas, demonstrating functional specialization in motion-based form processing. LOT bio was best characterized by the slow pathway of the SlowFast model, whereas SMG lh showed strong similarity to both models. Critically, we found that representations aligned with brain activity also better supported behavioral function: the full SlowFast model, incorporating both slow and fast pathways, outperformed other models in few-shot categorization of object kinematograms and showed the highest similarity to human perceptual judgments. These findings demonstrate that with appropriate inductive biases, specifically, dual-pathway architectures for multi-scale motion processing and training objectives focused on dynamic visual tasks, ANNs can develop functionally useful representations of motion-defined forms that exhibit better alignment with the visual regions involved in processing dynamic visual signals.