Nvidia
Nvidia
@nvidia·2.2M subscribers·2.7K videos

Extreme Co-Design for Efficient Tokenomics and AI at Scale

Posted

February 12, 2026

Views

6,481

Likes

267

Engagement

4.12%

Search the Record

Indexed

Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.

Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.

Try a name, a topic, or a quoted line

YouTube Description

as posted by the channel

As AI enters the era of real-time reasoning, the key metric for deploying AI at scale is now cost per token — how much it costs to generate intelligence.

Reasoning models like mixture-of-experts (MoE) generate massive volumes of tokens to deliver higher-quality results, placing pressure on the entire system — from compute and memory to networking, storage, and software.

Featuring insights from NVIDIA, Signal65, Microsoft Azure, and CoreWeave, this discussion explains why extreme co-design — optimizing the full stack as a unified system — is essential to lowering cost per token and maximizing AI ROI, making end-to-end system design the most powerful lever for scaling efficient AI.

Guests & Subjects Covered

As AINVIDIA Signal Microsoft AzureAI ROIAI Learn

Sentinel Indexing in Progress

Metadata and chapters are available. Claim extraction for this episode is pending.

All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.

Extreme Co-Design for Efficient Tokenomics and AI at Scale · Nvidia · Sentinel