TopK Sparse Autoencoders Across Three Model Architectures: Dictionary Collapse, Dense-Feature Degeneracy, and the Limits of Activation-Pattern Feature Matching

J. Melton · American Code Labs · Preprint, August 2026

Download PDF

Abstract

TopK sparse autoencoders trained at matched sparsity and dictionary size on Llama-3.2-3B, Mistral-7B-v0.3 and Qwen2.5-3B. Reports abrupt dictionary collapse followed by partial recovery, dense-feature degeneracy severe enough to invalidate frequency-based feature selection, and a null cross-architecture matching result whose detection floor is measured directly by planted-signal power analysis at correlation 0.95 to 1.00.

Code and data

github.com/american-code/ResearchPapersdata/sae-runs/

Licensed CC BY 4.0.