Patrik Reizinger

Visiteur de recherche indépendant - Univeristy of Tübingen

Superviseur⋅e principal⋅e

Simon Lacoste-Julien

Sujets de recherche

Apprentissage de représentations

Apprentissage profond

Causalité

Généralisation hors distribution (OOD)

Site web

Google Scholar

GitHub

Publications

Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations

Wieland Brendel

Identifiability in representation learning is commonly evaluated using standard metrics (e.g., MCC, DCI, R^2) on synthetic benchmarks with k… (voir plus)nown ground-truth factors. These metrics are assumed to reflect recovery up to the equivalence class guaranteed by identifiability theory. We show that this assumption holds only under specific structural conditions: each metric implicitly encodes assumptions about both the data-generating process (DGP) and the encoder. When these assumptions are violated, metrics become misspecified and can produce systematic false positives and false negatives. Such failures occur both within classical identifiability regimes and in post-hoc settings where identifiability is most needed. We introduce a taxonomy separating DGP assumptions from encoder geometry, use it to characterise the validity domains of existing metrics, and release an evaluation suite for reproducible stress testing and comparison.

2026-02-26

arXiv (prépublication)

doi.org

arxiv.org

Causality is Key for Interpretability Claims to Generalise

Shruti Joshi

Aaron Mueller

David Klindt

Wieland Brendel

Patrik Reizinger

Dhanya Sridhar

Interpretability research on large language models (LLMs) has yielded important insights into model behaviour, yet recurring pitfalls persis… (voir plus)t: findings that do not generalise, and causal interpretations that outrun the evidence. Our position is that causal inference specifies what constitutes a valid mapping from model activations to invariant high-level structures, the data or assumptions needed to achieve it, and the inferences it can support. Specifically, Pearl's causal hierarchy clarifies what an interpretability study can justify. Observations establish associations between model behaviour and internal components. Interventions (e.g., ablations or activation patching) support claims how these edits affect a behavioural metric (e.g., average change in token probabilities) over a set of prompts. However, counterfactual claims -- i.e., asking what the model output would have been for the same prompt under an unobserved intervention -- remain largely unverifiable without controlled supervision. We show how causal representation learning (CRL) operationalises this hierarchy, specifying which variables are recoverable from activations and under what assumptions. Together, these motivate a diagnostic framework that helps practitioners select methods and evaluations matching claims to evidence such that findings generalise.

2026-02-17

arXiv (prépublication)

doi.org

arxiv.org

Mila Techaide 2026

Propulsion d'entrepreneurs scientifiques

Avantage IA : productivité dans la fonction publique

Patrik Reizinger

Publications

Mila Techaide 2026

Propulsion d'entrepreneurs scientifiques

Avantage IA : productivité dans la fonction publique

Mots-clés populaires:

Patrik Reizinger

Publications