Module 2: Scaled Dot-Product & Self-Attention

Welcome to Module 2: Scaled Dot-Product & Self-Attention. In this module, we dissect the mathematical heart of the Transformer architecture—how tokens dynamically route information and attend to context across a sequence.


Concepts in this Module


← Module 1: Embeddings
LLM Axon Home
Concept 01: Scaled Dot-Product →