#86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning
April 3, 2020
1 hr 48 min
86
Full
85 of 498
Episode 86 · April 3, 2020
#86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning
0:00 / 1:48:28
Excerpt playback isn’t available for this episode — use the full-episode link.
Search the Record
Moment search isn’t open for this episode yet
When it opens, you’ll be able to type any phrase and jump straight to the second it was said.
Charles Hoskinson Episodes Around April 3, 2020
Episode 85 of 498 — walk the feed in the order it was published.
#81 – Anca Dragan: Human-Robot Interaction and Reward Engineering
#86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning
Summary
<p>David Silver leads the reinforcement learning research group at DeepMind and was lead researcher on AlphaGo, AlphaZero and co-lead on AlphaStar, and MuZero and lot of important work in reinforcement learning.</p>
<p>Support this podcast by signing up with these sponsors:<br />
– MasterClass: <a href="https://masterclass.com/lex">https://masterclass.com/lex</a><br />
– Cash App – use code “LexPodcast” and download:<br />
– Cash App (App Store): <a href="https://apple.co/2sPrUHe">https://apple.co/2sPrUHe</a><br />
– Cash App (Google Play): <a href="https://bit.ly/2MlvP5w">https://bit.ly/2MlvP5w</a></p>
<p>EPISODE LINKS:<br />
Reinforcement learning (book): https://amzn.to/2Jwp5zG</p>
<p><span style="font-weight: 400;">This conversation is part of the Artificial Intelligence podcast.</span> If you would like to get more information about this podcast go to <a href="https://lexfridman.com/ai">https://lexfridman.com/ai</a> or connect with @lexfridman on <a href="https://twitter.com/lexfridman">Twitter</a>, <a href="https://www.linkedin.com/in/lexfridman/">LinkedIn</a>, <a href="https://www.facebook.com/lexfridman">Facebook</a>, <a href="https://medium.com/@lexfridman">Medium</a>, or <a href="https://www.youtube.com/lexfridman">YouTube</a> where you can watch the video versions of these conversations. If you enjoy the podcast, please rate it 5 stars on <a href="https://podcasts.apple.com/us/podcast/artificial-intelligence/id1434243584">Apple Podcasts</a>, follow on <a href="https://open.spotify.com/show/2MAi0BvDc6GTFvKFPXnkCL">Spotify</a>, or support it on <a href="https://www.patreon.com/lexfridman">Patreon</a>.</p>
<p>Here’s the outline of the episode. On some podcast players you should be able to click the timestamp to jump to that time.</p>
<p>OUTLINE:<br />
00:00 – Introduction<br />
04:09 – First program<br />
11:11 – AlphaGo<br />
21:42 – Rule of the game of Go<br />
25:37 – Reinforcement learning: personal journey<br />
30:15 – What is reinforcement learning?<br />
43:51 – AlphaGo (continued)<br />
53:40 – Supervised learning and self play in AlphaGo<br />
1:06:12 – Lee Sedol retirement from Go play<br />
1:08:57 – Garry Kasparov<br />
1:14:10 – Alpha Zero and self play<br />
1:31:29 – Creativity in AlphaZero<br />
1:35:21 – AlphaZero applications<br />
1:37:59 – Reward functions<br />
1:40:51 – Meaning of life</p>