Sebastian Raschka
Sebastian Raschka
@sebastianraschka·88.6K subscribers·306 videos

Build an LLM from Scratch 3: Coding attention mechanisms

Posted

March 11, 2025

Views

58,921

Likes

1,092

Comments

68

Engagement

1.97%

Search the Record

Indexed

Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.

Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.

Try a name, a topic, or a quoted line

Chapters

YouTube Description

as posted by the channel

Links to the book:

Link to the GitHub repository

This is a supplementary video explaining how attention mechanisms (self-attention, causal attention, multi-head attention) work by coding them from scratch.

You can find additional bonus materials on GitHub:

Comparing Efficient MultiHead Attention Implementations,

Understanding PyTorch Buffers,

Guests & Subjects Covered

Build an LLM from Scratch 3Manning Link

Sentinel Indexing in Progress

Metadata and chapters are available. Claim extraction for this episode is pending.

All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.