Attention is all you need (really)
The Transformer architecture, introduced by Vaswani et al. in 2017, replaced recurrence with self attention. The core operation is scaled dot product attention: given queries Q, keys K, and values V (all linear projections of the input), the output is softmax(QK^T / √d)V.