Sebastian Raschka
Sebastian Raschka
@sebastianraschka·88.6K subscribers·306 videos

What I Learned From Implementing LLM Architectures From Scratch (And How to Get Started)

Posted

May 12, 2026

Views

23,297

Likes

1,070

Comments

66

Engagement

4.88%

Search the Record

Indexed

Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.

Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.

Try a name, a topic, or a quoted line

Chapters

YouTube Description

as posted by the channel

LLM Architecture Gallery

In this talk, I discuss what we can learn from implementing LLM architectures from scratch in Python and PyTorch.

The main idea is that to really understand how modern LLMs work, it helps to inspect the actual implementation details: attention variants, normalization layers, configuration files, KV cache optimizations, and the small architectural choices that often make a model work correctly.

I also walk through how I approach new open-weight models, how I compare them against reference implementations, and what broader architecture trends emerge from looking at many recent LLMs.

Chapters:

Links & Promotions

Guests & Subjects Covered

LLM Architecture Gallery InPyTorch ThePython What PythonThe LLMLLMs KVReasoning Model From Scratch

Sentinel Indexing in Progress

Metadata and chapters are available. Claim extraction for this episode is pending.

All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.

What I Learned From Implementing LLM Architectures From Scratch (And How to Get Started) · Sebastian Raschka · Sentinel