Skip to content
← All tracks

Optimizers from Scratch

PyTorch

Build the workhorse optimizers byte by byte. Start with vanilla SGD and end at Adam.

0 / 4 solved
  1. 1. Not solved yet. Implement Gradient Descent Step
  2. 2. Not solved yet. Implement Momentum Update
  3. 3. Not solved yet. Implement Adam Optimizer Step
  4. 4. Not solved yet. Train with Adam End-to-End

Check yourself

4 questions · one attempt each

These do not count toward finishing the track. They are here to catch the things that are easy to read past.

0 / 4

With a constant gradient of 1.0, lr 0.1 and momentum 0.9, why does the parameter move further on each step?

p = torch.tensor([1.0], requires_grad=True)
opt = torch.optim.SGD([p], lr=0.1, momentum=0.9)
# three steps, gradient is 1.0 each time
Question 1 of 4