Sebastian Raschka
Sebastian Raschka
@sebastianraschka·88.6K subscribers·306 videos

LLM Building Blocks & Transformer Alternatives

Posted

October 27, 2025

Views

19,077

Likes

661

Comments

37

Engagement

3.66%

Search the Record

Indexed

Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.

Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.

Try a name, a topic, or a quoted line

Chapters

YouTube Description

as posted by the channel

Resources:

Understanding and Coding the KV Cache in LLMs from Scratch

The Big Architecture Comparison

- Beyond Standard LLMs: Linear Attention Hybrids, Text Diffusion, Code World Models, and Small Recursive Transformers

Reasoning From Scratch book

Description:

Learn the core components of modern transformer-based large language models (LLMs) and the practical techniques that make inference faster and cheaper.

We walk through Grouped-Query Attention (GQA), Multi-Head Latent Attention (MLA), and Sliding Window Attention (SWA), show where Mixture of Experts (MoE) fits into today’s architectures, and finish with a look at promising alternatives and hybrid models beyond standard transformers.

Chapters:

#LLM #Transformers #DeepLearning #MachineLearning #Inference #MoE #GQA #MLA #SWA #KVCache

Guests & Subjects Covered

KV CacheThe Big Architecture ComparisonSmall Recursive TransformersReasoning From ScratchDescription LearnSliding Window Attention SWAExperts MoEChapters Intro Main

Sentinel Indexing in Progress

Metadata and chapters are available. Claim extraction for this episode is pending.

All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.