We can't find the internet
Attempting to reconnect
Something went wrong!
Attempting to reconnect
← All tracks
0 / 4
Tokenization & Embeddings
PyTorchFrom raw text to token tensors — BPE, subword, and the embedding matrices that turn ids into vectors.
0
/ 15 solved
- 1. Not solved yet. Implement Embedding Lookup
- 2. Not solved yet. BPE Merge Step
- 3. Not solved yet. BPE Encode Text
- 4. Not solved yet. Tokenize and Pad Batch
- 5. Not solved yet. Learned Absolute Position Embedding
- 6. Not solved yet. Tied Input/Output Embeddings
- 7. Not solved yet. Subword Tokenizer: Greedy Longest-Prefix-Match
- 8. Not solved yet. MLM Masking Strategy
- 9. Not solved yet. MLM Forward Pass
- 10. Not solved yet. Train MLM Pretraining Step
- 11. Not solved yet. MLM Forward with Tied Output Head
- 12. Not solved yet. MLM Eval — Masked Accuracy
- 13. Not solved yet. Causal LM Forward Pass
- 14. Not solved yet. Train Causal LM Pretraining Step
- 15. Not solved yet. Train Tiny GPT End-to-End
Check yourself
4 questions · one attempt eachThese do not count toward finishing the track. They are here to catch the things that are easy to read past.
A batch touches 3 of a 50,000-row embedding table. Which rows get a gradient?
emb = nn.Embedding(50_000, 256)
loss = emb(torch.tensor([7, 42, 7])).sum()
loss.backward()
Question 1 of 4