Skip to content
← All tracks

Attention 101

PyTorch

From dot products to multi-head transformers. Each step composes onto the next.

0 / 5 solved
  1. 1. Not solved yet. Implement Scaled Dot-Product Attention
  2. 2. Not solved yet. Self-Attention Layer
  3. 3. Not solved yet. Multi-Query Attention
  4. 4. Not solved yet. Transformer Encoder Block
  5. 5. Not solved yet. Multi-Head Attention Block

Check yourself

4 questions · one attempt each

These do not count toward finishing the track. They are here to catch the things that are easy to read past.

0 / 4

Attention scores are computed without the 1/sqrt(d_k) division, with d_k = 128 and unit-variance q and k. What actually goes wrong?

scores = q @ k.transpose(-2, -1)        # no scaling
weights = scores.softmax(dim=-1)
Question 1 of 4