Publications

Perplexity Cannot Always Tell Right from Wrong
Federico Barbero
Christos Perivolaropoulos
Simon Kayode Osindero
Perplexity -- a function measuring a model's overall level of"surprise"when encountering a particular output -- has gained significant tract… (voir plus)ion in recent years, both as a loss function and as a simple-to-compute metric of model quality. Prior studies have pointed out several limitations of perplexity, often from an empirical manner. Here we leverage recent results on Transformer continuity to show in a rigorous manner how perplexity may be an unsuitable metric for model selection. Specifically, we prove that, if there is any sequence that a compact decoder-only Transformer model predicts accurately and confidently -- a necessary pre-requisite for strong generalisation -- it must imply existence of another sequence with very low perplexity, but not predicted correctly by that same model. Further, by analytically studying iso-perplexity plots, we find that perplexity will not always select for the more accurate model -- rather, any increase in model confidence must be accompanied by a commensurate rise in accuracy for the new model to be selected.
Secure Tool Manifest and Digital Signing Solution for Verifiable MCP and LLM Pipelines
Saeid Jamshidi
Kawser Wazed Nafi
Amin Nikanjam
Mohammad Hamdaqa
Securing Time in Energy IoT: A Clock-Dynamics-Aware Spatio-Temporal Graph Attention Network for Clock Drift Attacks and Y2K38 Failures
Saeid Jamshidi
Omar Abdel Wahab
Rolando Herrero
The integrity of time in distributed Internet of Things (IoT) devices is crucial for reliable operation in energy cyber-physical systems, su… (voir plus)ch as smart grids and microgrids. However, IoT systems are vulnerable to clock drift, time-synchronization manipulation, and timestamp discontinuities, such as the Year 2038 (Y2K38) Unix overflow, all of which disrupt temporal ordering. Conventional anomaly-detection models, which assume reliable timestamps, fail to capture temporal inconsistencies. This paper introduces STGAT (Spatio-Temporal Graph Attention Network), a framework that models both temporal distortion and inter-device consistency in energy IoT systems. STGAT combines drift-aware temporal embeddings and temporal self-attention to capture corrupted time evolution at individual devices, and uses graph attention to model spatial propagation of timing errors. A curvature-regularized latent representation geometrically separates normal clock evolution from anomalies caused by drift, synchronization offsets, and overflow events. Experimental results on energy IoT telemetry with controlled timing perturbations show that STGAT achieves 95.7% accuracy, outperforming recurrent, transformer, and graph-based baselines with significant improvements (d>1.8, p0.001). Additionally, STGAT reduces detection delay by 26%, achieving a 2.3-time-step delay while maintaining stable performance under over
Tri-LLM Cooperative Federated Zero-Shot Intrusion Detection with Semantic Disagreement and Trust-Aware Aggregation
Saeid Jamshidi
Omar Abdel Wahab
Kawser Wazed Nafi
Federated learning (FL) has become an effective paradigm for privacy-preserving, distributed Intrusion Detection Systems (IDS) in cyber-phys… (voir plus)ical and Internet of Things (IoT) networks, where centralized data aggregation is often infeasible due to privacy and bandwidth constraints. Despite its advantages, most existing FL-based IDS assume closed-set learning and lack mechanisms such as uncertainty estimation, semantic generalization, and explicit modeling of epistemic ambiguity in zero-day attack scenarios. Additionally, robustness to heterogeneous and unreliable clients remains a challenge in practical applications. This paper introduces a semantics-driven federated IDS framework that incorporates language-derived semantic supervision into federated optimization, enabling open-set and zero-shot intrusion detection for previously unseen attack behaviors. The approach constructs semantic attack prototypes using a Tri-LLM ensemble of GPT-4o, DeepSeek-V3, and LLaMA-3-8B, aligning distributed telemetry features with high-level attack concepts. Inter-LLM semantic disagreement is modeled as epistemic uncertainty for zero-day risk estimation, while a trust-aware aggregation mechanism dynamically weights client updates based on reliability. Experimental results show stable semantic alignment across heterogeneous clients and consistent convergence. The framework achieves over 80% zero-shot detection accuracy on unseen attack patterns, improving zero-day discrimination by more than 10% compared to similarity-based baselines, while maintaining low aggregation instability in the presence of unreliable or compromised clients.
Boosting CVaR Policy Optimization with Quantile Gradients
Optimizing Conditional Value-at-risk (CVaR) using policy gradient (a.k.a CVaR-PG) faces significant challenges of sample inefficiency. This … (voir plus)inefficiency stems from the fact that it focuses on tail-end performance and overlooks many sampled trajectories. We address this problem by augmenting CVaR with an expected quantile term. Quantile optimization admits a dynamic programming formulation that leverages all sampled data, thus improves sample efficiency. This does not alter the CVaR objective since CVaR corresponds to the expectation of quantile over the tail. Empirical results in domains with verifiable risk-averse behavior show that our algorithm within the Markovian policy class substantially improves upon CVaR-PG and consistently outperforms other existing methods.
MGSM-Pro: A Simple Strategy for Robust Multilingual Mathematical Reasoning Evaluation
Tianyi Xu
Alfred Malengo Kondoro
Tadesse Destaw Belay
Catherine Nana Nyaah Essuman
Ifeoma Okoh
Ganiyat Afolabi
Ayodele Awokoya
Large language models have made substantial progress in mathematical reasoning. However, benchmark development for multilingual evaluation h… (voir plus)as lagged behind English in both difficulty and recency. Recently, GSM-Symbolic showed a strong evidence of high variance when models are evaluated on different instantiations of the same question; however, the evaluation was conducted only in English. In this paper, we introduce MGSM-Pro, an extension of MGSM dataset with GSM-Symbolic approach. Our dataset provides five instantiations per MGSM question by varying names, digits and irrelevant context. Evaluations across nine languages reveal that many low-resource languages suffer large performance drops when tested on digit instantiations different from those in the original test set. We further find that models robustness in HRL setting do not necessarily translate to LRL. Moreover, proprietary models, such as Gemini 2.5 Flash and GPT-4.1 are less robust to digit, whereas Gemini 3.0 Pro is more robust. Among open models, GPT-OSS 120B and DeepSeek v3 show stronger robustness. Based on these findings, we recommend evaluating each problem using at least five digit-varying instantiations to obtain a more robust and realistic assessment of math reasoning.
Patient safety culture in the operating room of African hospitals: a systematic review
Jacques Fadhili Bake
Naïcen Ghanmi
Elena Guadagno
K. M. Claude
Tsongo Kibendelwa Zacharie
Patient safety in operating rooms has globally improved through interventions such as the World Health Organization (WHO) Surgical Safety Ch… (voir plus)ecklist and multidisciplinary team training. However, while evidence from high-income countries is well documented, there remains limited consolidated knowledge on the understanding, application, and effectiveness of safety culture interventions in African surgical settings, which this review seeks to address. This systematic review examined factors and protocols affecting surgical safety in African operating rooms. We hypothesized that persistent systemic barriers undermine safety culture despite adoption of global measures. Following PRISMA 2020, we searched eight databases (Medline, Embase, Cochrane, Africa-Wide, CINAHL, Global Health, Global Index Medicus, Web of Science) from inception to 5 December 2024, using variations of text words present in the title, abstract, or keyword fields, alongside relevant subject headings, to identify articles addressing surgical safety and culture throughout Africa. Included studies involved operating room professionals in African countries and used quantitative, qualitative, or mixed-methods designs. We excluded non-operating room settings, patient-only studies, inaccessible full texts, reviews, editorials, letters, conference abstracts, and duplicates. Two reviewers independently screened and appraised studies using the Mixed Methods Appraisal Tool. Findings were synthesized narratively with subgroup analysis by study type and theme. Out of 9,875 identified records, 22 studies from 12 African countries (2014–2024) met inclusion criteria, with Ethiopia contributing the highest number (n = 4). Various assessment tools, including the Hospital Survey on Patient Safety Culture, the Safety Attitudes Questionnaire, and the National Surgical, Obstetric, and Anaesthesia Plans interview manual, revealed recurring challenges: inadequate non-punitive responses to errors, communication barriers, hierarchical structures, and resource constraints. Four interventions showed promise: implementation and training on the WHO Surgical Safety Checklist, Safe Surgery 2020 initiatives, Non-Technical Skills for Surgeons training, and multidisciplinary training. The heterogeneity of study designs, sample sizes, and outcome measures limited direct comparisons and precluded meta-analysis. Nonetheless, the review highlights persistent barriers and emerging opportunities to strengthen patient safety culture in African operating rooms. While the WHO Surgical Safety Checklist remains valuable, sustainable progress requires multi-level strategies that address systemic constraints and incorporate context-sensitive adaptations. PROSPERO, CRD42024627076.
Parallel and Customizable Equality Saturation
Abd-El-Aziz Zayed
Mai Jacob Peng
Equality saturation enables compilers to explore many semantically equivalent program variants, deferring optimization decisions to a final … (voir plus)extraction phase. However, existing frameworks exhibit sequential execution and hard-coded saturation loops. This limits scalability and requires significant engineering effort to customize saturation behavior. This paper addresses these limitations using three novel techniques. First, it shows how saturation can be parallelized thanks to the use of thread-safe data structures and the notion of deferred e-graph updates. Second, it provides an extensible mechanism to express custom and composable saturation strategies. Third, it generalizes e-graph metadata to support custom e-graph annotations. The implementation, written in Scala, is evaluated on four use-cases: classical program optimization, idiom recognition, scalability strategies and incremental equality saturation. The results show that it outperforms several existing equality saturation engines, including the highly optimized egglog library. When used to reimplement an existing idiom recognition technique, the new design finds higher-quality idioms, 16× faster. Additionally, the design is able to natively express state-of-the-art custom equality saturation behavior such as incremental equality saturation and multi-phase rewriting strategies without any modification to the core library.
Signal from Structure: Exploiting Submodular Upper Bounds in Generative Flow Networks
Generative Flow Networks (GFlowNets; GFNs) are a class of generative models that learn to sample compositional objects proportionally to the… (voir plus)ir a priori unknown value, their reward. We focus on the case where the reward has a specified, actionable structure, namely that it is submodular. We show submodularity can be harnessed to retrieve upper bounds on the reward of compositional objects that have not yet been observed. We provide in-depth analyses of the probability of such bounds occurring, as well as how many unobserved compositional objects can be covered by a bound. Following the Optimism in the Face of Uncertainty principle, we then introduce SUBo-GFN, which uses the submodular upper bounds to train a GFN. We show that SUBo-GFN generates orders of magnitude more training data than classical GFNs for the same number of queries to the reward function. We demonstrate the effectiveness of SUBo-GFN in terms of distribution matching and high-quality candidate generation on synthetic and real-world submodular tasks.
Benchmarking the geographic generalization of deep learning models for precipitation downscaling
Luca Schmidt
Nicole Ludwig
Matthew Chantry
Christian Lessig
Earth System Models (ESM) are our main tool for projecting the impacts of climate change. However, running these models at sufficient resolu… (voir plus)tion for local-scale risk-assessments is not computationally feasible. Deep learning-based super-resolution models offer a promising solution to downscale ESM outputs to higher resolutions by learning from data. Yet, due to regional variations in climatic processes, these models typically require retraining for each geographical area–demanding high-resolution observational data, which is unevenly available across the globe. This highlights the need to assess how well these models generalize across geographic regions. To address this, we introduce RainShift, a dataset and benchmark for evaluating downscaling under geographic distribution shifts. We evaluate state-of-the-art downscaling approaches including GANs and diffusion models in generalizing across data gaps between the Global North and Global South. Our findings reveal substantial performance drops in out-of-distribution regions, depending on model and geographic area. While expanding the training domain generally improves generalization, it is insufficient to overcome shifts between geographically distinct regions. We show that addressing these shifts through, for example, domain adaptation can improve spatial generalization. Our work advances the global applicability of downscaling methods and represents a step toward reducing inequities in access to high-resolution climate information.
Sudanese-Flores: Extending FLORES+ to Sudanese Arabic Dialect
Hadia Mohmmedosman Ahmed Samil
In this work, we introduce Sudanese-Flores, an extension of the popular Flores+ machine translation (MT) benchmark to the Sudanese Arabic di… (voir plus)alect. We translate both the DEV and DEVTEST splits of the Modern Standard Arabic dataset into the corresponding Sudanese dialect, resulting in a total of 2,009 sentences. While the dialect was recently introduced in Google Translate, there are no available benchmark in this dialect despite spoken by over 40 million people. Our evaluation on two leading LLMs such as GPT-4.1 and Gemini 2.5 Flash showed that while the performance English to Arabic is impressive (more than 23 BLEU), they struggle on Sudanese dialect (less than 11 BLEU) in zero-shot settings. In few-shot scenario, we achieved only a slight boost in performance.
Anatomically-aware conformal prediction for medical image segmentation with random walks
Christian Desrosiers