LLMs, the whole thing · Part 9
Multi-head attention, one shape at a time (Part 9)
Part 9 of LLMs, the whole thing: two heads, one token, and the difference between computing attention and organizing its tensors. A concrete PyTorch walkthrough of view, transpose, concatenation and the output projection.