We can't find the internet
Attempting to reconnect
Something went wrong!
Attempting to reconnect
← All tracks
0 / 4
Flax
JAXModules, layers from scratch, attention, transformer architectures, training loops, lifted transforms. Production model code in JAX (Linen API).
This track is written for JAX, which isn't the mode you're browsing in.
0
/ 100 solved
- 1. Not solved yet. Module with @nn.compact
- 2. Not solved yet. Module with setup() (alternative to compact)
- 3. Not solved yet. Custom Parameter Initializer
- 4. Not solved yet. init() and apply() Round-Trip
- 5. Not solved yet. nn.Sequential Composition
- 6. Not solved yet. Module with Multiple Named Sub-Modules
- 7. Not solved yet. Module That Branches on a Config Flag
- 8. Not solved yet. Three Levels of Module Nesting
- 9. Not solved yet. Multiple PRNG Streams (params, dropout)
- 10. Not solved yet. Train vs Eval Branches via train Flag
- 11. Not solved yet. Implement Dense from Scratch
- 12. Not solved yet. Implement Conv1D from Scratch
- 13. Not solved yet. Implement Conv2D with Stride and Padding
- 14. Not solved yet. Implement Transposed Convolution
- 15. Not solved yet. Implement Depthwise-Separable Convolution
- 16. Not solved yet. Implement LayerNorm with γ/β
- 17. Not solved yet. Implement BatchNorm with Mutable batch_stats
- 18. Not solved yet. Implement GroupNorm
- 19. Not solved yet. Implement RMSNorm (Modern LLM Norm)
- 20. Not solved yet. Implement Dropout with RNG Threading
- 21. Not solved yet. Scaled Dot-Product Attention
- 22. Not solved yet. Multi-Head Self-Attention with Flax
- 23. Not solved yet. Causal Multi-Head Self-Attention
- 24. Not solved yet. Cross-Attention with Flax MHA
- 25. Not solved yet. Multi-Head Attention with KV Cache
- 26. Not solved yet. Grouped-Query Attention (GQA)
- 27. Not solved yet. Multi-Query Attention (MQA)
- 28. Not solved yet. Sliding-Window Attention (Mistral-style)
- 29. Not solved yet. ALiBi: Attention with Linear Biases
- 30. Not solved yet. Block-Diagonal Attention Mask
- 31. Not solved yet. Token Embedding with Flax
- 32. Not solved yet. Sinusoidal Position Encoding
- 33. Not solved yet. Learned Position Embedding
- 34. Not solved yet. Rotary Position Embedding (RoPE)
- 35. Not solved yet. ALiBi Bias Matrix
- 36. Not solved yet. T5 Relative Position Bucketing
- 37. Not solved yet. Tied Input/Output Embedding
- 38. Not solved yet. ViT Patch Embedding
- 39. Not solved yet. Transformer Encoder Block (Pre-LN)
- 40. Not solved yet. Transformer Decoder Block (Pre-LN)
- 41. Not solved yet. Pre-LN vs Post-LN Residual Pattern
- 42. Not solved yet. Mini GPT — Decoder-Only Language Model
- 43. Not solved yet. Mini BERT — Encoder-Only Hidden States
- 44. Not solved yet. Mini T5 — Encoder-Decoder with RMSNorm and Tied Embeddings
- 45. Not solved yet. Vision Transformer (Mean-Pool Variant)
- 46. Not solved yet. Vision Transformer with [CLS] Token
- 47. Not solved yet. DeiT — Data-Efficient Image Transformer
- 48. Not solved yet. SwiGLU Feed-Forward Network
- 49. Not solved yet. ResNet Basic Block
- 50. Not solved yet. ResNet Bottleneck Block
- 51. Not solved yet. Tiny ResNet Classifier
- 52. Not solved yet. Tiny U-Net
- 53. Not solved yet. GRU Cell Step
- 54. Not solved yet. LSTM Cell Step
- 55. Not solved yet. Bidirectional RNN (Flax)
- 56. Not solved yet. Mixture-of-Experts FFN
- 57. Not solved yet. Squeeze-and-Excitation Block (Flax)
- 58. Not solved yet. Vision-Language Fusion
- 59. Not solved yet. TrainState — One Step
- 60. Not solved yet. train_step with value_and_grad
- 61. Not solved yet. eval_step — Forward + Metrics
- 62. Not solved yet. Label-Smoothed Cross-Entropy
- 63. Not solved yet. Mixed-Precision Training Step
- 64. Not solved yet. Train with Mutable batch_stats
- 65. Not solved yet. Multi-Task Two-Head Loss
- 66. Not solved yet. Sharded Eval Loss
- 67. Not solved yet. Warmup-Cosine LR at Step
- 68. Not solved yet. Gradient Accumulation Step
- 69. Not solved yet. EMA of Parameters
- 70. Not solved yet. Orbax Save (Tree-Leaf Count)
- 71. Not solved yet. Orbax Load (Restore via Template)
- 72. Not solved yet. HF Weight Load (Kernel Transpose)
- 73. Not solved yet. Pre-train then Fine-tune (Frozen Trunk)
- 74. Not solved yet. Per-Param Weight Decay Mask
- 75. Not solved yet. Per-Param Learning Rate Multipliers
- 76. Not solved yet. Param Freezing via Grad Zeroing
- 77. Not solved yet. Test-Time Augmentation Aggregation
- 78. Not solved yet. Distributed Checkpoint (Sharding Math)
- 79. Not solved yet. nn.scan over an RNN cell
- 80. Not solved yet. nn.scan over layers
- 81. Not solved yet. nn.vmap with shared params
- 82. Not solved yet. nn.checkpoint (gradient checkpointing)
- 83. Not solved yet. nn.jit (Flax-aware JIT lift)
- 84. Not solved yet. nn.remat with checkpoint policies
- 85. Not solved yet. Composed lifts: nn.scan + nn.vmap
- 86. Not solved yet. Batched init via jax.vmap
- 87. Not solved yet. jax.lax.scan inside a Flax Module
- 88. Not solved yet. Custom lift: roll your own ensemble
- 89. Not solved yet. PartitionSpec Layout
- 90. Not solved yet. with_sharding_constraint Annotation
- 91. Not solved yet. nn.with_partitioning Annotation
- 92. Not solved yet. flax.struct.dataclass — Pytree-Friendly State
- 93. Not solved yet. Param Surgery — Kernel Replace
- 94. Not solved yet. Param Surgery — Zero Last Layer
- 95. Not solved yet. Param Surgery — Freeze First Dense
- 96. Not solved yet. Partial Init — Warm-Start From Smaller Checkpoint
- 97. Not solved yet. Multiple Mutable Collections
- 98. Not solved yet. Param Sharing — One Module, Two Call Sites
- 99. Not solved yet. shard_map Simulation — Manual SPMD
- 100. Not solved yet. Mini LM Capstone — Putting It All Together
Check yourself
4 questions · one attempt eachThese do not count toward finishing the track. They are here to catch the things that are easy to read past.
A Linen model with BatchNorm is applied in training mode. Why does it raise without mutable=?
BN().apply(v, x, train=True)
BN().apply(v, x, train=True, mutable=["batch_stats"])
Question 1 of 4