Attention Is All You Need
Vaswani et al. (2017) — introduces the Transformer architecture and self-attention as a scalable sequence modeling approach.
Why I liked it: It changed how modern NLP systems are designed and laid the foundation for today’s Gen AI models.