Sparse autoencoders are increasingly used as interpretability interfaces for vision representations, where individual sparse latents can act as handles for recurring visual features. We study a deliberately narrow setting: TopK sparse autoencoders trained on DINOv2 ViT-S/14 layer-9 patch features from one COCO train2017 subset and evaluated on COCO val2017 objects. We ask whether training on mixed input resolutions improves the stability of sparse feature handles across resolutions. Multi-resolution training gives competitive reconstruction transfer, but does not preserve object-local sparse feature identity as well as high-resolution single-resolution baselines. A three-seed core check supports the main high-resolution baseline gap, while focused controls for per-resolution exposure, dictionary size, and active TopK budget improve reconstruction more readily than object-level sparse-support stability. The result is a scoped cautionary finding: for TopK SAEs on this backbone and layer, reconstruction transfer and object-level sparse identity should be evaluated separately when using sparse latents as feature-tracking tools.
Arsenii Shvetsov (Wed,) studied this question.