Compositional Interpretability

CompInterp uncovers the structure of neural representations, showing how simple features compose into complex behaviours. By unifying tensor and neural network paradigms, model weights and data are treated as a single modality. This compositional lens on design, analysis and control paves the way for inherently interpretable AI without compromising performance.

Compositional architectures capture rich non-linear (polynomial) relationships between representation spaces. Instead of masking them through linear approximations, CompInterp methods expose their inherent hierarchical structure accross levels of abstraction. This allows for weight-based subcircuit analysis, grounding interpretability in formal (de)compositions rather than post-hoc activation-based heuristics.

We’re now scaling compositional interpretability to transformers and CNNs by leveraging their low-rank structure through tensor decomposition and information theory. Learn more in our latest talk!

news

Sep 07, 2026 We organised a tutorial on the evolution of interpretability at the AIMLAI workshop (ECML PKDD’26)! :teacher:
Aug 21, 2026 Our work on tensor similarity is at the Tractable Probabilistic Modeling workshop (UAI’26)! :triangular_ruler:
Aug 12, 2026 We gave a talk on moving from mechanistic to compositional interpretability at the Interpretable Deep Learning seminar! :satellite:
Jul 11, 2026 Our compositional interpretability poster is at the Compositional Learning Workshop (ICML’26)! :tada:
Jun 29, 2026 We presented compositional interpretability at the Theory of CS & Computational Creativity workshop (ICCC’26)! :art:

selected publications

  1. From Mechanistic to Compositional Interpretability
    Ward Gauderis*, Thomas Dooms*, Steven T. Homer, Kola Ayonrinde, and 1 more author
    In 2nd Workshop on Compositional Learning: Safety, Interpretability, and Agents At the Forty-Third International Conference on Machine Learning, Jul 2026
  2. Bilinear autoencoders find interpretable manifolds
    May 2026
  3. Compositionality Unlocks Deep Interpretable Models
    In Connecting Low-Rank Representations in AI: At the 39th Annual AAAI Conference on Artificial Intelligence, Nov 2024