Grammar categories sit in groups of interpretability features, not single ones

single source· 1 articles · confidence: medium · first seen 2026-09-23 20:00 UTC

What this means for you

Nothing to act on. This is a result about what one interpretability tool exposes, not a release — no weights, no API, and one preprint behind it. If you rely on sparse-autoencoder features for grammar or syntax work, the paper's finding is that single features will not line up with single categories.

Part-of-speech categories — noun, verb, preposition — are highly recoverable from the internal activations a sparse autoencoder exposes, but they do not map one to one onto individual features. A sparse autoencoder is an interpretability tool that splits a model's activations into many separate units researchers can inspect. In this study each grammatical category was carried by a compact group of sparse features, with wide variation between tags, overlap between related categories, and stability on held-out data. The authors say the recoverability is not explained by memorised words, and that open and closed classes behave differently. It is an arXiv preprint dated 23 September 2026, not peer reviewed, and no evaluation harness is reported.

Key facts

  • ·The study tested whether part-of-speech distinctions map one-to-one onto individual sparse-autoencoder latents and found they do not. source
  • ·Part-of-speech categories were highly recoverable from SAE activations, with each category supported by a compact group of sparse latents. source
  • ·The authors report the recoverability is not reducible to lexical memorisation. source
  • ·Open and closed part-of-speech classes differed substantially in how they were represented. source
  • ·The feature groups remained stable on held-out data and showed overlap between related categories. source
  • ·The paper is arXiv preprint 2609.29362, published 23 September 2026. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire