Skip to content
← All tracks

Performance

PyTorch

Everything before this, applied and measured. Fusing an elementwise chain, deleting a Python loop, preallocating, and not calling `.item()` in one. Benchmarked, so the improvement is a number rather than a claim.

0 / 7 solved
  1. 1. Not solved yet. Two hundred kernels, or one
  2. 2. Not solved yet. Four allocations, or one
  3. 3. Not solved yet. Rebuilt every step
  4. 4. Not solved yet. One kernel for every head
  5. 5. Not solved yet. Sort once
  6. 6. Not solved yet. Time it properly
  7. 7. Not solved yet. Everything at once

Check yourself

4 questions · one attempt each

These do not count toward finishing the track. They are here to catch the things that are easy to read past.

0 / 4

What does torch.no_grad() change about a forward pass?

lin(torch.ones(1, 2)).requires_grad

with torch.no_grad():
    lin(torch.ones(1, 2)).requires_grad
Question 1 of 4