Inferact: Building the Infrastructure That Runs Modern AI
Inferact is a new AI infrastructure company founded by the creators and core maintainers of vLLM. Its mission is to build a universal, open-source inference layer that makes large AI models faster, cheaper, and more reliable to run across any hardware, model architecture, or deployment environment. Together, they broke down how modern AI models are actually run in production, why “inference” has quietly become one of the hardest problems in AI infrastructure, and how the open-source project vLLM emerged to solve it. The conversation also looked at why the vLLM team started Inferact and their vision for a universal inference layer that can run any model, on any chip, efficiently.
January 22, 2026
43 min
78
Full
78 of 100
Episode 78 · January 22, 2026
Inferact: Building the Infrastructure That Runs Modern AI
0:00 / 43:37
Excerpt playback isn’t available for this episode — use the full-episode link.
Search the Record
Moment search isn’t open for this episode yet
When it opens, you’ll be able to type any phrase and jump straight to the second it was said.
A16z Episodes Around January 22, 2026
Episode 78 of 100 — walk the feed in the order it was published.
Inferact: Building the Infrastructure That Runs Modern AI
Summary
Inferact is a new AI infrastructure company founded by the creators and core maintainers of vLLM. Its mission is to build a universal, open-source inference layer that makes large AI models faster, cheaper, and more reliable to run across any hardware, model architecture, or deployment environment. Together, they broke down how modern AI models are actually run in production, why “inference” has quietly become one of the hardest problems in AI infrastructure, and how the open-source project vLLM emerged to solve it. The conversation also looked at why the vLLM team started Inferact and their vision for a universal inference layer that can run any model, on any chip, efficiently.
Topics
Episode details and artwork are published by the show’s own feed. Sentinel plays short cited excerpts and links every episode back to its publisher.









