Transformer architecture, taken apart piece by piece — attention, positional encoding, feedforward layers, normalization, and everything else under the hood. Each post is one concept, written the way I wish someone had explained it to me: intuition first, then the mechanics, then the math.
Posts in this series appear below as I write them.