Оценка интерпретируемости SAE: пайплайн измерения важнее архитектуры модели
Команда исследователей проверила надёжность autointerpretability-оценок для разреженных автоэнкодеров (SAE) — метода интерпретации нейросетей. Разброс из-за методологии измерения превысил разброс между архитектурами по всем четырём метрикам. Detection оказалась самой стабильной метрикой, fuzzing — ненадёжной во всех условиях. Значит, часть межстатейных сравнений отражает различия пайплайна оценки, а не самих моделей.
AI-processed from arXiv cs.LG; edited by Hamidun News
Researchers published a paper on arXiv in July 2026 showing that when comparing the interpretability of sparse autoencoders (SAE), the result depends more on the chosen evaluation pipeline than on the architecture of the models themselves.
What was tested and how
The authors checked whether autointerpretability scores reflect stable properties of a model's features or side effects of the measurement procedure. In the standard pipeline, one language model explains each SAE feature, and a second language model evaluates the quality of that explanation. It is exactly these scores that are relied upon when comparing SAE interpretability across different papers.
The experiment covered four metrics, two models, and four axes of methodological variation — and across all conditions the initial assumption of score stability was not confirmed.
- Four evaluation metrics: simulation, detection, fuzzing, and purity
- Two models: Pythia-160M and Apertus-8B
- Four axes of methodological variation in the evaluation pipeline
- Detection is the most stable metric, fuzzing is unreliable under all conditions
- Three tools from the authors: variance decomposition, Stability Check, and Minimum Reporting Checklist
Why methodology matters more than architecture
Methodological variability, taken together, exceeds architectural variability across all metrics and both tested models — this is the paper's main conclusion. In simpler terms, the spread in scores caused by the design of the measurement procedure itself turned out to be larger than the spread between the different SAE architectures that the procedure is supposed to distinguish.
"Methodological variability in aggregate exceeds architectural
variability across all metrics for all tested models," the paper's abstract on arXiv states.
This undermines the very logic of cross-paper comparisons: if two teams obtained different autointerpretability scores, the difference may be explained by their pipeline settings rather than by the quality of the autoencoders themselves.
Where the evaluations break down
Each metric showed its own instability profile, with detection holding up most reliably; fuzzing, by contrast, proved unreliable under all tested conditions. Different metrics cannot be treated as interchangeable — the choice of a specific metric by itself changes the final conclusion.
A separate problem is feature ranking. According to the study, top-k lists of the most interpretable features do not remain stable when the corpus or sampling conditions change. At the same time, average scores remain stable and mask instability at the level of individual features — a failure that cannot be noticed by tracking only the similarity of explanations.
What this means
Unreliable evaluation slows progress in interpretability exactly when reliable tools for understanding AI systems are especially needed. So that part of the comparisons in the SAE literature stops reflecting differences in pipelines instead of differences in models, the authors propose a variance decomposition method, a Stability Check procedure, and a Minimum Reporting Checklist — a minimal set of parameters that papers should disclose.
Frequently asked questions
Which evaluation metric turned out to be the most reliable?
According to the results of the study, detection turned out to be the most stable of the four metrics. The fuzzing metric, in contrast, was found to be unreliable under all tested conditions, while simulation and purity showed their own instability profiles.
What do the authors propose for more reliable evaluation?
The authors propose three tools: a variance decomposition method, a Stability Check procedure to verify the stability of scores, and a Minimum Reporting Checklist — a list of pipeline parameters that need to be disclosed in papers so that comparisons remain comparable.
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.