Axon 03: Large Language Models & Transformers

Welcome to the Large Language Models & Transformers Axon. This track develops generative AI and modern transformer architectures from first principles—moving from subword tokenization and vector embedding geometries to multi-head self-attention, deep transformer decoder blocks, rotary position embeddings (RoPE), and autoregressive sampling strategies.


Modules in this Axon

1. Tokenization & Vector Embeddings


2. Scaled Dot-Product & Self-Attention


3. The Transformer Architecture


4. Generation, RoPE & Sampling


← Previous Axon: Machine Learning & Vision
Curriculum Home
Next Axon: Physics, Dynamics & Actuation →