Portrait of Srishti Yadav is unavailable

Srishti Yadav

Alumni

Publications

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner
Sree Harsha Nelaturu
Damian Stachura
Anastassia Kornilova
Jon Crall
Tommaso Cerruti
Yanan Long
Yifan Mai
Sanchit Ahuja
Asaf Yehudai
Marek Šuppa
John P. Lalor
Oluwagbemike Olowe
Jatin Ganhotra
Brian H. Hu
Eliya Habba
Andrew M. Bean
Chang Liu
Sander Land
Steven Dillmann … (see 28 more)
Aniketh Garikaparthi
Elron Bandel
Saki Imai
James Edgell
Wm. Matthew Kennedy
Jenny Chim
Patrick Meusling
Asteria Kaeberlein
Venkata Ramachandra Karthik Chundi
Manasi Patwardhan
Martin Ku
Austin Meek
Leon Knauer
Brian Wingenroth
Usman Gohar
Felix Friedrich
Jennifer Mickel
Arman Cohan
Stella Biderman
Irene Solaiman
Zeerak Talat
Anka Reuel
Mubashara Akhtar
Gjergji Kasneci
Avijit Ghosh
Leshem Choshen
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that … (see more)challenge analysis and comparison. First, results are saved in incompatible formats, scattered across leaderboards, papers, blog posts, evaluation harness logs, and custom repositories. Second, results are created by different evaluation frameworks, which produce divergent scores for nominally identical evaluations and record metadata inconsistently, hindering comparison, cross-community evaluation science, cost reduction, and reuse. We introduce Every Eval Ever, the first shared schema and community-crowdsourced repository for AI evaluation results. The schema standardizes how evaluations are represented in a unified, single JSON document. It is source-agnostic by design, ingesting results from evaluation harnesses and papers alike, and optionally stores per-instance outputs for fine-grained analysis. We contribute: (i) a community-governed metadata schema with a companion instance-level schema, the first standardization effort of its kind; (ii) automatic converters from popular formats, evaluation harnesses, and leaderboards to the unified schema; and (iii) a crowdsourced community database hosted on Hugging Face, currently spanning to date 22,235 models, 2,273 unique benchmarks, and 31 evaluation formats.