We can't find the internet
Attempting to reconnect
Something went wrong!
Attempting to reconnect
PyTorch learning tracks
PyTorchGuided sequences that interleave written explainers with the problems that apply them. Tracks remember how far you got, and you can work one with friends as a squad. For a plain ordered run of problems, see the roadmap.
-
Optimizers from Scratch
0 / 4
Build the workhorse optimizers byte by byte. Start with vanilla SGD and end at Adam.
-
Tensor Foundations
0 / 7
The basics of working with tensors — shapes, indexing, broadcasting, and the operations you'll reach for daily.
-
Activations
0 / 9
From the classic non-linearities to modern gated activations. End with implementing a custom gradient.
-
Loss Functions
0 / 8
The training objectives that shape every model — regression, classification, distribution-matching, and contrastive.
-
Regression from Scratch
0 / 4
Linear, polynomial, regularized, and logistic — the classics every interview revisits.
-
Classifiers from Scratch
0 / 5
Build the simplest classifiers end-to-end — binary, multi-class, and sequence.
-
Normalization
0 / 6
BatchNorm, LayerNorm, GroupNorm, RMSNorm, dropout — when to use which, and how each shifts gradients.
-
CNNs from Scratch
0 / 15
Build convolutional networks one block at a time — from raw conv to skip connections and squeeze-excitation.
-
Recurrent Networks
0 / 6
RNN, LSTM, GRU, bidirectional — the architectures that ruled NLP before transformers.
-
Tokenization & Embeddings
0 / 15
From raw text to token tensors — BPE, subword, and the embedding matrices that turn ids into vectors.
-
Attention 101
0 / 5
From dot products to multi-head transformers. Each step composes onto the next.
-
Attention Variants
0 / 14
After Attention 101 — the real-world variants that actually run in modern LLMs.
-
Position Encodings
0 / 5
Sinusoidal, learned, relative, RoPE, ALiBi — how transformers know where tokens live.
-
Decoding Strategies
0 / 8
Greedy through speculative — every way to turn logits into tokens.
-
Parameter-Efficient FT
0 / 4
Adapt large models without touching most of their weights — LoRA, adapters, prefix-tuning.
-
Production ML — Training Stack
0 / 12
From minimal training loops to gradient accumulation, EMAs, distributed primitives, and checkpoints. Train models at scale.
-
Metrics & Evaluation
0 / 7
The numbers that tell you whether your model is actually working — accuracy, F1, AUC, perplexity, BLEU.
-
Generative Models
0 / 11
From autoencoders through VAEs to modern diffusion. The math behind generating images, audio, and text.