Attention Head
An independent parallel compute sub-unit inside self-attention layers that tracks relationships between words.
Attention Heads allow transformer models to process different representation subspaces in parallel. By having multiple heads, a model can associate a single word with several different contextual dimensions simultaneously.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.