Skip to content
← All tracks

Attention Variants

PyTorch

After Attention 101 — the real-world variants that actually run in modern LLMs.

0 / 14 solved
  1. 1. Not solved yet. Cross Attention
  2. 2. Not solved yet. Causal Attention Mask
  3. 3. Not solved yet. Grouped-Query Attention
  4. 4. Not solved yet. Sliding Window Attention
  5. 5. Not solved yet. Efficient Attention with Masking
  6. 6. Not solved yet. KV Cache for Autoregressive Decoding
  7. 7. Not solved yet. Flash Attention Score Computation
  8. 8. Not solved yet. Causal Self-Attention Block
  9. 9. Not solved yet. Cross-Attention Block
  10. 10. Not solved yet. LLaMA-Style Transformer Block
  11. 11. Not solved yet. Encoder-Decoder Transformer Forward Pass
  12. 12. Not solved yet. Train Encoder-Decoder Seq2Seq Step
  13. 13. Not solved yet. Encoder-Decoder Greedy Decode
  14. 14. Not solved yet. Encoder-Decoder Beam Search

Check yourself

4 questions · one attempt each

These do not count toward finishing the track. They are here to catch the things that are easy to read past.

0 / 4

MQA uses a single KV head. Why is GQA usually preferred over it?

MHA:   kv_heads = num_heads
GQA:   1 < kv_heads < num_heads
MQA:   kv_heads = 1
Question 1 of 4