Skip to content
← All tracks

Optax

JAX

Gradient transforms, optimizer chains, schedules, weight decay, EMA, masking. The production optimizer library for JAX.

This track is written for JAX, which isn't the mode you're browsing in.

0 / 25 solved
  1. 1. Not solved yet. Optax SGD Step
  2. 2. Not solved yet. Optax SGD with Momentum
  3. 3. Not solved yet. Optax Adam Step
  4. 4. Not solved yet. AdamW with Decoupled Weight Decay
  5. 5. Not solved yet. Optax RMSprop Step
  6. 6. Not solved yet. Constant Schedule
  7. 7. Not solved yet. Linear Schedule
  8. 8. Not solved yet. Warmup + Cosine Decay
  9. 9. Not solved yet. Piecewise Constant Schedule
  10. 10. Not solved yet. Exponential Decay Schedule
  11. 11. Not solved yet. Chain: Clip + SGD
  12. 12. Not solved yet. Adam + Weight Decay (Chain)
  13. 13. Not solved yet. Global-Norm Gradient Clipping
  14. 14. Not solved yet. Lookahead Optimizer Wrapper
  15. 15. Not solved yet. multi_transform per-Param Group
  16. 16. Not solved yet. Optax EMA on Params
  17. 17. Not solved yet. Gradient Accumulation via MultiSteps
  18. 18. Not solved yet. masked: Apply WD Only To Certain Params
  19. 19. Not solved yet. inject_hyperparams for Runtime LR
  20. 20. Not solved yet. zero_nans for NaN-Safe Training
  21. 21. Not solved yet. Full Training Step (Loss + Grad + Update)
  22. 22. Not solved yet. Train Step with Frozen Params (Mask)
  23. 23. Not solved yet. Train Step with Global-Norm Clipping
  24. 24. Not solved yet. Train Step with Warmup Schedule
  25. 25. Not solved yet. 4-Step Training Loop with Scan + Loss Curve

Check yourself

4 questions · one attempt each

These do not count toward finishing the track. They are here to catch the things that are easy to read past.

0 / 4

chain(clip_by_global_norm(1.0), sgd(0.1)) sees a gradient of 100. What update comes out, and would swapping the order matter?

optax.chain(
    optax.clip_by_global_norm(1.0),
    optax.sgd(0.1),
)
Question 1 of 4