March 2, 2025
92,014
1,702
112
1.97%
Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.
Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.
Try a name, a topic, or a quoted line
1:28:01Now PlayingLinks to the book:
Link to the GitHub repository
This is a supplementary video going over text data preparations steps (tokenization, byte pair encoding, data loaders, etc.) for LLM training.
You can find additional bonus materials on GitHub:
Byte Pair Encoding (BPE) Tokenizer From Scratch,
Comparing Various Byte Pair Encoding (BPE) Implementations,
Understanding the Difference Between Embedding Layers and Linear Layers,
Data sampling with a sliding window with number data,
A video on the effect of random seeds
Sentinel Indexing in Progress
Metadata and chapters are available. Claim extraction for this episode is pending.
All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.