Transformer Decoder Block Architecture & Tensor Flow
Concept 02 Demo
1. Input Stream
2. Multi-Head Attention
3. Pre-RMSNorm
4. SwiGLU MLP Block
5. Output Stream
Active Component
Multi-Head Self-Attention
Tensor Dimension Shape
(Batch, SeqLen=9, Dim=768)
Mathematical Role
Dynamic Context Mixing between Tokens