We can't find the internet
Attempting to reconnect
Something went wrong!
Attempting to reconnect
← All tracks
0 / 4
Attention Variants
PyTorchAfter Attention 101 — the real-world variants that actually run in modern LLMs.
0
/ 14 solved
- 1. Not solved yet. Cross Attention
- 2. Not solved yet. Causal Attention Mask
- 3. Not solved yet. Grouped-Query Attention
- 4. Not solved yet. Sliding Window Attention
- 5. Not solved yet. Efficient Attention with Masking
- 6. Not solved yet. KV Cache for Autoregressive Decoding
- 7. Not solved yet. Flash Attention Score Computation
- 8. Not solved yet. Causal Self-Attention Block
- 9. Not solved yet. Cross-Attention Block
- 10. Not solved yet. LLaMA-Style Transformer Block
- 11. Not solved yet. Encoder-Decoder Transformer Forward Pass
- 12. Not solved yet. Train Encoder-Decoder Seq2Seq Step
- 13. Not solved yet. Encoder-Decoder Greedy Decode
- 14. Not solved yet. Encoder-Decoder Beam Search
Check yourself
4 questions · one attempt eachThese do not count toward finishing the track. They are here to catch the things that are easy to read past.
MQA uses a single KV head. Why is GQA usually preferred over it?
MHA: kv_heads = num_heads
GQA: 1 < kv_heads < num_heads
MQA: kv_heads = 1
Question 1 of 4