Artwork

内容由LessWrong提供。所有播客内容(包括剧集、图形和播客描述)均由 LessWrong 或其播客平台合作伙伴直接上传和提供。如果您认为有人在未经您许可的情况下使用您的受版权保护的作品,您可以按照此处概述的流程进行操作https://zh.player.fm/legal
Player FM -播客应用
使用Player FM应用程序离线!

“What Is The Alignment Problem?” by johnswentworth

46:26
 
分享
 

Manage episode 461626068 series 3364760
内容由LessWrong提供。所有播客内容(包括剧集、图形和播客描述)均由 LessWrong 或其播客平台合作伙伴直接上传和提供。如果您认为有人在未经您许可的情况下使用您的受版权保护的作品,您可以按照此处概述的流程进行操作https://zh.player.fm/legal
So we want to align future AGIs. Ultimately we’d like to align them to human values, but in the shorter term we might start with other targets, like e.g. corrigibility.
That problem description all makes sense on a hand-wavy intuitive level, but once we get concrete and dig into technical details… wait, what exactly is the goal again? When we say we want to “align AGI”, what does that mean? And what about these “human values” - it's easy to list things which are importantly not human values (like stated preferences, revealed preferences, etc), but what are we talking about? And don’t even get me started on corrigibility!
Turns out, it's surprisingly tricky to explain what exactly “the alignment problem” refers to. And there's good reasons for that! In this post, I’ll give my current best explanation of what the alignment problem is (including a few variants and the [...]
---
Outline:
(01:27) The Difficulty of Specifying Problems
(01:50) Toy Problem 1: Old MacDonald's New Hen
(04:08) Toy Problem 2: Sorting Bleggs and Rubes
(06:55) Generalization to Alignment
(08:54) But What If The Patterns Don't Hold?
(13:06) Alignment of What?
(14:01) Alignment of a Goal or Purpose
(19:47) Alignment of Basic Agents
(23:51) Alignment of General Intelligence
(27:40) How Does All That Relate To Todays AI?
(31:03) Alignment to What?
(32:01) What are a Humans Values?
(36:14) Other targets
(36:43) Paul!Corrigibility
(39:11) Eliezer!Corrigibility
(40:52) Subproblem!Corrigibility
(42:55) Exercise: Do What I Mean (DWIM)
(43:26) Putting It All Together, and Takeaways
The original text contained 10 footnotes which were omitted from this narration.
---
First published:
January 16th, 2025
Source:
https://www.lesswrong.com/posts/dHNKtQ3vTBxTfTPxu/what-is-the-alignment-problem
---
Narrated by TYPE III AUDIO.
---
Images from the article:
undefined
undefined
  continue reading

425集单集

Artwork
icon分享
 
Manage episode 461626068 series 3364760
内容由LessWrong提供。所有播客内容(包括剧集、图形和播客描述)均由 LessWrong 或其播客平台合作伙伴直接上传和提供。如果您认为有人在未经您许可的情况下使用您的受版权保护的作品,您可以按照此处概述的流程进行操作https://zh.player.fm/legal
So we want to align future AGIs. Ultimately we’d like to align them to human values, but in the shorter term we might start with other targets, like e.g. corrigibility.
That problem description all makes sense on a hand-wavy intuitive level, but once we get concrete and dig into technical details… wait, what exactly is the goal again? When we say we want to “align AGI”, what does that mean? And what about these “human values” - it's easy to list things which are importantly not human values (like stated preferences, revealed preferences, etc), but what are we talking about? And don’t even get me started on corrigibility!
Turns out, it's surprisingly tricky to explain what exactly “the alignment problem” refers to. And there's good reasons for that! In this post, I’ll give my current best explanation of what the alignment problem is (including a few variants and the [...]
---
Outline:
(01:27) The Difficulty of Specifying Problems
(01:50) Toy Problem 1: Old MacDonald's New Hen
(04:08) Toy Problem 2: Sorting Bleggs and Rubes
(06:55) Generalization to Alignment
(08:54) But What If The Patterns Don't Hold?
(13:06) Alignment of What?
(14:01) Alignment of a Goal or Purpose
(19:47) Alignment of Basic Agents
(23:51) Alignment of General Intelligence
(27:40) How Does All That Relate To Todays AI?
(31:03) Alignment to What?
(32:01) What are a Humans Values?
(36:14) Other targets
(36:43) Paul!Corrigibility
(39:11) Eliezer!Corrigibility
(40:52) Subproblem!Corrigibility
(42:55) Exercise: Do What I Mean (DWIM)
(43:26) Putting It All Together, and Takeaways
The original text contained 10 footnotes which were omitted from this narration.
---
First published:
January 16th, 2025
Source:
https://www.lesswrong.com/posts/dHNKtQ3vTBxTfTPxu/what-is-the-alignment-problem
---
Narrated by TYPE III AUDIO.
---
Images from the article:
undefined
undefined
  continue reading

425集单集

所有剧集

×
 
Loading …

欢迎使用Player FM

Player FM正在网上搜索高质量的播客,以便您现在享受。它是最好的播客应用程序,适用于安卓、iPhone和网络。注册以跨设备同步订阅。

 

快速参考指南

边探索边听这个节目
播放