Data, data, everywhere - enough for AGI?

"The Cognitive Revolution" | AI Builders, Researchers, and Live Player Analysis

内容由Turpentine, Erik Torenberg, and Nathan Labenz提供。所有播客内容（包括剧集、图形和播客描述）均由 Turpentine, Erik Torenberg, and Nathan Labenz 或其播客平台合作伙伴直接上传和提供。如果您认为有人在未经您许可的情况下使用您的受版权保护的作品，您可以按照此处概述的流程进行操作https://zh.player.fm/legal。

18d ago 1:01:40

MP3•单集首页

In this podcast, Nathan and Nick dive deep into the data requirements for achieving Artificial General Intelligence. They explore the current paradigms, the role of data in approximating intelligence, and the scaling trends for GPT models. The discussion covers various datasets, from email and Twitter to YouTube and genomic data, as they analyze the feasibility of reaching the target of 100 trillion high-quality tokens. While the bull case suggests an abundance of data, the bear case highlights the limits on high-quality data, prompting a fascinating exploration of what makes data good for AI and the potential for AI to generate its own data.

Chapters

(00:00) Introduction

(05:04) Scaling Hypothesis of Intelligence

(07:32) Is There Enough High Quality Data?

(10:19) Algorithms Impacting Data Requirements

(17:42) Sponsor : Omneky

(18:04) Estimating High Quality Token Requirements

(24:07) Astronomy and YouTube Data Scale

(29:42) Genomics Data

(37:58) Sponsors : Brave / Plumb / Squad

(41:16) Code Datasets and Synthetic Data

(45:48) The Bear Case: Quality and Usability of Data

(50:54) Investment Trends and Compute Efficiency

(54:19) Training Run

(57:21) Synthetic Data Generation and Self-Play

129集单集