Publications

FedMLAC: Mutual learning driven heterogeneous federated audio classification
Rajib Rana
Youyang Qu
Xiaohui Tao
Ji Zhang
Carlos Busso
Shivakumara Palaiahnakote
Federated Learning (FL) offers a privacy-preserving framework for training audio classification (AC) models across decentralized clients wit… (voir plus)hout sharing raw data. However, Federated Audio Classification faces three major challenges: data heterogeneity , model heterogeneity , and data corruption , which degrade performance in real-world settings. While existing methods often address these issues separately, a unified solution remains underexplored. We propose FedMLAC, a mutual learning-based FL framework that tackles all three challenges simultaneously. Each client maintains a personalized local AC model and a lightweight, globally shared Plug-in model. These models interact via bidirectional knowledge distillation, enabling global knowledge sharing while adapting to local data distributions, thus addressing both data and model heterogeneity. To counter data corruption, we introduce a Layer-wise Pruning Aggregation (LPA) strategy that filters anomalous Plug-in updates based on parameter deviations during aggregation. Extensive experiments on four diverse AC benchmarks, including both speech and non-speech tasks, show that FedMLAC consistently outperforms state-of-the-art baselines in classification accuracy and robustness to noisy data.
MambAdapter: Lightweight Mamba-Based Adapters for Parameter-Efficient Transfer Learning in Speech and Audio
Salman Hussain Ali
Umberto Cappellazzo
Fine-tuning Transformer-based foundation models has become the dominant strategy for domain adaptation in audio and speech processing. To re… (voir plus)duce the computational and memory costs of this process, parameter-efficient transfer learning (PETL) methods have been widely explored. Meanwhile, Mamba, a recent state-space model, has emerged as a promising alternative to Transformers for sequence modeling. In this work, we present MambAdapter, a parameter-efficient transfer learning approach that integrates Mamba into low-rank bottleneck adapters. Our design combines parameter sharing across adapters with the injection of a lightweight Mamba module, enabling more effective modeling of audio features. We demonstrate that MambAdapter matches or outperforms strong PETL baselines on four audio classification tasks and five speech recognition languages, even when operating under reduced parameter budgets.
Radiomic prediction of substantial LVSI in endometrial cancer using reduced field of view DWI- a feasibility study
Akiyo Takada
Daniel A. Di Giovanni
Takuro Horikoshi
Takahiro Tsuboyama
Hajime Yokota
Sakurako Harada‐Kagitani
Evan McNabb
Jérémy Dana
Rita Zakarian
Haruto Sugawara
Yuka Matsumoto
Yuji Habu
Hirokazu Usui
Kaori Koga
Katsuhiro Nasu
Takashi Uno
Caroline Reinhold
Engineered Nonheme Iron Enzymes Enable Asymmetric Hydrogenation of Alkenes
Yunfei He
Shuang-Yu Dai
Mei‐Yan Xu
Baixu Ma
Lizhi Tao
Developing biocatalytic systems capable of reducing simple alkenes is highly desirable for synthetic chemistry and biosynthesis, yet existin… (voir plus)g enzymes remain largely restricted to their ability to convert polarized, electron-deficient substrates. Here, we present a nonheme iron metalloenzyme platform that enables hydrogenation of styrenes, conjugated nitriles and amides, and nonconjugated olefins through a putative iron–hydride mechanism. Starting from the Fe(II)/ α -ketoglutarate-dependent dioxygenase GOX, iterative rounds of directed evolution produced an engineered “alkene hydrogenase” (AHase-6) containing 16 mutations and promoting NaBH 4 -driven reduction across diverse C═C bond motifs. Kinetic analysis indicates that this enzymatic hydrogenation process proceeds via formation of an enzyme–substrate ternary complex through a sequential mechanism. Mechanistic studies further reveal that alkene insertion occurs with regioselectivity governed primarily by substrate electronics and sterics. These findings establish nonheme iron enzymes as an unrecognized scaffold for metal–hydride-based hydrogenation and highlight their potential as sustainable, tunable alternatives to traditional catalytic systems.
S$^2$COPE: Self-Supervised Concept Discovery via Preference Learning
Shilong Xiang
Zirui Zhang
Chengzhi Mao
Current representation learning paradigms force a fundamental compromise: self-supervised methods scale to massive datasets but yield opaque… (voir plus) features, whereas interpretable models remain bottlenecked by the need for dense human annotation. We introduce Self-Supervised Concept discOvery via Preference lEarning (\model), a label-free framework that resolves this dilemma. Instead of treating Vision-Large-Language Models (VLLMs) as static feature extractors, \model leverages them as active participants in a self-supervised preference optimization loop. By autonomously hypothesizing, validating, and reinforcing candidate visual attributes directly from raw imagery, our framework discovers novel, structured concepts without a single label. Extensive experiments across natural, medical, and physics domains demonstrate that \model successfully extracts domain-specific concepts where standard VLLMs often fail to generate. By amortizing concept discovery directly into the VLLM backbone through our self-supervised preference objective -- rather than relying on static generation and disjoint filtering -- we achieve up to a 24-point absolute improvement in downstream top-1 classification accuracy on unseen data. Our work suggest that interpretability can emerge through a model's autonomous interaction with incidental visual structures, without any human supervision.
Artificial intelligence-assisted ganglion cell detection in Hirschsprung's disease: A comparative evaluation of two deep learning approaches
E Wang
Karl Grenier
Peter Savadjiev
Background. Definitive diagnosis of Hirschsprung's disease (HD) requires pathological identification of enteric ganglion cells. This process… (voir plus) is time-consuming and subject to inter-observer variability. Artificial intelligence (AI) tools have the potential to standardize and accelerate this workflow, but no study has determined which AI approach best serves intraoperative HD pathology diagnostics. Method. This study compared the U-Net and You Only Look Once version 26 (YOLO26) frameworks for ganglion cell detection using a single-centre retrospective dataset of 54 whole-slide images (WSIs) from rectal biopsies. WSIs were tiled into 397,731 image patches (128x128 pixels), further partitioned into training (70%), validation (15%), and testing (15%) sets. Models were evaluated on tile- and patient-level diagnostic metrics and processing latency. Results. The U-Net achieved a tile-level sensitivity of 82.9%, showing no statistically significant difference compared to YOLO26 (79.1%; p = 0.097). However, YOLO26 demonstrated a statistically significant advantage in tile-level specificity (96.1% vs. 93.9%; p < 0.001) and reduced mean inference latency (7.64 ms vs. 11.57 ms/tile). At the patient level, both models achieved 100% diagnostic sensitivity. Despite low patient-level specificity (0.0% U-Net; 11.8% YOLO26), the tissue-level diagnostic burden of false positives was 6.00% for U-Net and 3.50% for YOLO26. Conclusion. The U-Net is preferred when nominal gains in sensitivity are prioritized, while the YOLO26 is an alternative that optimizes efficiency and false positive suppression. Both models serve as robust screening filters to augment the pathologist's workflow and should be selected based on workflow requirements. Prospective validation on larger, multi-centre datasets is required before clinical implementation.
Characterizing Cultural Localization in AI-Generated Stories
Shaily Bhatt
Supriti Vijay
Jeremiah Milbauer
The global use of artificial intelligence has increased interest in assessing the ability to generate culturally localized content, includin… (voir plus)g stories. Cultural localization in stories often occurs through either templated localization -- the use of cultural markers (e.g., names, locations) in a generic narrative -- or holistic localization -- the variation of plots, values, and themes, in addition to cultural markers. We propose a method to measure the degree to which content was generated through templated localization. Specifically, we identify the lexical tokens that distinguish stories across nationalities and measure the similarity of the narratives that remain after removing them. In stories generated by five models on 125 topics for 193 nationalities, our method is able to detect that only a small subset (9-17%) of the vocabulary accounts for the variation across nationalities and that the narratives that remain after removing them contain repeated multi-word sequences, suggesting the presence of a shared culturally-agnostic narrative template. Finally, we characterize the cultural markers for their stereotypicality and offensiveness, finding that markers from 19 countries, mostly located in the Global South, are on average offensive.
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner
Sree Harsha Nelaturu
Damian Stachura
Anastassia Kornilova
Jon Crall
Tommaso Cerruti
Yanan Long
Yifan Mai
Sanchit Ahuja
Asaf Yehudai
Marek Šuppa
John P. Lalor
Oluwagbemike Olowe
Jatin Ganhotra
Brian H. Hu
Eliya Habba
Andrew M. Bean
Chang Liu
Sander Land
Steven Dillmann … (voir 28 de plus)
Aniketh Garikaparthi
Elron Bandel
Saki Imai
James Edgell
Wm. Matthew Kennedy
Jenny Chim
Patrick Meusling
Asteria Kaeberlein
Venkata Ramachandra Karthik Chundi
Manasi Patwardhan
Martin Ku
Austin Meek
Leon Knauer
Brian Wingenroth
Usman Gohar
Felix Friedrich
Jennifer Mickel
Arman Cohan
Stella Biderman
Irene Solaiman
Zeerak Talat
Anka Reuel
Mubashara Akhtar
Gjergji Kasneci
Avijit Ghosh
Leshem Choshen
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that … (voir plus)challenge analysis and comparison. First, results are saved in incompatible formats, scattered across leaderboards, papers, blog posts, evaluation harness logs, and custom repositories. Second, results are created by different evaluation frameworks, which produce divergent scores for nominally identical evaluations and record metadata inconsistently, hindering comparison, cross-community evaluation science, cost reduction, and reuse. We introduce Every Eval Ever, the first shared schema and community-crowdsourced repository for AI evaluation results. The schema standardizes how evaluations are represented in a unified, single JSON document. It is source-agnostic by design, ingesting results from evaluation harnesses and papers alike, and optionally stores per-instance outputs for fine-grained analysis. We contribute: (i) a community-governed metadata schema with a companion instance-level schema, the first standardization effort of its kind; (ii) automatic converters from popular formats, evaluation harnesses, and leaderboards to the unified schema; and (iii) a crowdsourced community database hosted on Hugging Face, currently spanning to date 22,235 models, 2,273 unique benchmarks, and 31 evaluation formats.
Simplifying the Modeling of Arbitrary Conditionals in Natural Language
Causal Transformers model sequences through an autoregressive factorization of the joint distribution, which enables efficient left-to-right… (voir plus) decoding and conditional likelihood computation. However, they cannot tractably sample from or evaluate arbitrary conditionals -- e.g., a block of text conditioned on past and future tokens. Recent work aims to solve this problem through novel architectures, but they often lead to sub-optimal modeling of such conditionals and degraded generations. We propose Arbitrary Conditionals GPT (AC-GPT) which introduces a simple modification to standard causal Transformers to enable evaluating and sampling from arbitrary conditionals -- including past, future, and mixed contexts -- within a single forward pass. Unlike prior approaches, our method preserves the standard left-to-right ordering and next-token prediction objective essential for both strong performance and efficient training on natural language. Crucially, this compatibility allows existing LLMs to be fine-tuned for arbitrary conditioning. Our empirical results indicate that our method outperforms baselines on modeling arbitrary conditionals, without degrading standard left-to-right performance.
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
Xiaoyuan Liu
Jianhong Tu
Yuqi Chen
Siyuan Xie
Sihan Ren
Tianneng Shi
Gal Gantar
Evan Sandoval
D Lee
Daniel Miao
Peter J. Gilbert
Nick Hynes
Mauro Staver
Warren He
David Marn
Andrew Low
Xi Zhang
Elron Bandel
Michal Shmueli-Scheuer
Somasekhar Reddy … (voir 9 de plus)
Alexandre Lacoste
R Radha Krishnan
Elham Tabassi
Yu Su
Victor Barres
Chenguang Wang
Wenbo Guo
Dawn Song
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harn… (voir plus)esses that require heavy integration, create test-production mismatch, and limit fair comparison across diverse agent designs. The root problem is the lack of an open, agent-agnostic assessment interface. We advocate Agentified Agent Assessment (AAA), where evaluation is performed by judge agents and all participants interact through standardized protocols: A2A for task management and MCP for tool access. Conventional benchmarking defines two separate interfaces, one for the benchmark and one for the agent, while AAA only needs one; this yields a generic, unified framework that separates assessment logic from agent implementation and enables reproducible, interoperable, and multi-agent evaluation. We further introduce AgentBeats as a concrete realization of AAA: we identify five practical operation modes that make standardized assessment compatible with real-world constraints on openness, privacy, and reproducibility. To evaluate our design at scale, we conduct two studies: a five-month open competition that drew 298 judge agents across 12 categories together with 467 subject agents from independent participants, showing that AAA applies across a heterogeneous range of benchmarks; and a case study on coding agents that confirms agentified evaluation preserves fidelity with the public record while surfacing previously missing head-to-head results, yielding research insights about agent design. Combining a community-scale field study and a controlled coding case study, we verify that AAA delivers coverage, practicality, and fidelity across heterogeneous scenarios at scale. Together, AAA and AgentBeats offer a clear path toward open, standardized, and reproducible agent assessment.
Data-Driven Stochastic Vehicle Routing Problems with Deadlines Under Decision-Dependent Travel Time
Shanshan Wang
Leandro C. Coelho
Problem definition: Vehicle routing problems (VRPs) with deadlines have received significant attention around the world. Motivated by a real… (voir plus)-world food delivery problem, we assume that the travel time depends on the routing decisions, and we study a data-driven stochastic VRP with deadlines and endogenous uncertainty. Methodology/results: We use the nonparametric approaches, including k-nearest neighbor (kNN) and kernel density estimation (KDE), to estimate the decision-dependent probability distribution of travel time. To solve the resulting problem efficiently, we employ a logic-based Benders decomposition (LBBD) algorithm with several algorithmic enhancements. In particular, we propose a novel family of optimality cuts that includes the expected delay for all the subroutes. Moreover, we solve a total travel cost minimization problem to warm start the algorithm. We also use a local search procedure to improve the current routing decision and propose a machine learning–based lower bound heuristic to efficiently solve problems of realistic size. A practical case study for a food delivery routing problem using real-world data is conducted to show the efficiency of the proposed techniques and the advantage of the data-driven stochastic VRP in reducing the expected delay. Managerial implications: In our case study, we show that incorporating routing decisions into a nonparametric model outperforms a state-of-the-art data-driven parametric model by 23% on average in terms of the expected delay and the order-assignment decisions obtained from a robust model with travel-time predictors by 26% on average. Moreover, compared with the drivers’ actual routes and arc-based VRP models that ignore the endogenous uncertainty, our suggested routes can significantly improve the on-time performance of delivery services. We also quantify the value of the proposed routes with different service deadlines. Funding: S. Wang was partially supported by the Natural Sciences and Engineering Research Council of Canada [Grant RGPIN-2016-05208], IVADO, and a joint project between the Fonds de Recherche du Québec - Société et Culture (FRQSC) and the National Natural Science Foundation of China (NSFC) [Grant 295837]. She is also most recently supported by the National Natural Science Foundation of China [Grants 72501014, 72371022, and 72272014]. Supplemental Material: The online appendix is available at https://doi.org/10.1287/msom.2024.0899 .
Feature Geometry of Language Models Transfer Across Modalities to Time Series
Language models transfer to time-series forecasting, but it is unclear whether this reflects reusable internal structure or rapid relearning… (voir plus) under a familiar architecture. We study this transfer directly by comparing pretrained and randomly initialized versions of the same model on a forecasting objective whose inputs have little semantic overlap with text but still require autoregressive sequential structure. Across Qwen3-0.6B finetuning experiments, language initialization gives coherent per-example gradients from the first update, while random initialization first passes through a low-alignment warmup phase. Effective-rank and hidden-state analyses show that finetuning selectively reshapes an existing representation geometry rather than constructing the simpler temporal geometry found by models trained from scratch. Cross-domain sparse features and causal ablations then expose candidate transferred primitives, including a Layer~1 head--MLP circuit whose ablation selectively increases loss on periodic forecasting and repetitive language passages. These results support an account of cross-modal transfer in which autoregressive pretraining creates temporal feature geometry that can be selected and specialized outside language.