Axon 03: Large Language Models & Transformers
Welcome to the Large Language Models & Transformers Axon. This track develops generative AI and modern transformer architectures from first principles—moving from subword tokenization and vector embedding geometries to multi-head self-attention, deep transformer decoder blocks, rotary position embeddings (RoPE), and autoregressive sampling strategies.
Modules in this Axon
1. Tokenization & Vector Embeddings
- Concept 01: Byte-Pair Encoding (BPE) & Vocabulary Lookups
- Concept 02: High-Dimensional Semantic Vectors & Cosine Distance
2. Scaled Dot-Product & Self-Attention
- Concept 01: Scaled Dot-Product & Self-Attention (Q, K, V)
- Concept 02: Multi-Head Attention & Feature Subspaces
3. The Transformer Architecture
- Concept 01: Residual Skip Connections & RMSNorm
- Concept 02: The Transformer Decoder Block (SwiGLU & Feed-Forward)
4. Generation, RoPE & Sampling
- Concept 01: Rotary Position Embeddings (RoPE) & Context Length
- Concept 02: Autoregressive Next-Token Sampling (Temperature, Top-k, Top-p)