Self-Attention And Beyond the Infinite: Towards Linear Transformers with Infinite Self-Attention
Softmax attention costs O(N²), which caps how much context a Transformer can see. We reframe each attention layer as a diffusion step on a content-adaptive graph of tokens, accumulating multi-hop interactions through a discounted Neumann series. This connects attention directly to classical graph centrality — Katz, PageRank, eigenvector centrality — so token importance becomes interpretable rather than a black box. From that foundation we derive Linear-InfSA, an O(N) variant that approximates the principal eigenvector of the attention operator without ever building the N × N matrix — a drop-in replacement for standard Vision Transformer attention.