Multimodal Video Understanding Infinite Machine Learning: Artificial Intelligence | Startups | Technology podcast

Artwork

Tech Prateek Joshi

内容由Prateek Joshi提供。所有播客内容（包括剧集、图形和播客描述）均由 Prateek Joshi 或其播客平台合作伙伴直接上传和提供。如果您认为有人在未经您许可的情况下使用您的受版权保护的作品，您可以按照此处概述的流程进行操作https://zh.player.fm/legal。

Infinite Machine Learning: Artificial Intelligence | Startups | Technology « »
Multimodal Video Understanding

6M ago 42:12

分享

MP3•单集首页

内容由Prateek Joshi提供。所有播客内容（包括剧集、图形和播客描述）均由 Prateek Joshi 或其播客平台合作伙伴直接上传和提供。如果您认为有人在未经您许可的情况下使用您的受版权保护的作品，您可以按照此处概述的流程进行操作https://zh.player.fm/legal。

Jae Lee is the cofounder and CEO of Twelve Labs, where they are building video understanding infrastructure to help developers build programs that can see, hear, and understand the world. He was previously the Lead Data Scientist at the Ministry of National Defense in South Korea. He has a bachelors in computer science from UC Berkeley.
In this episode, we cover a range of topics including:
- What is multimodal video understanding
- State of play in multimodal video
- The founding of Twelve Labs
- The launch of Pegasus-1
- Four core principles: Efficient Long-form Video Processing, Multimodal Understanding, Video-native Embeddings, Deep Alignment between Video and Language Embeddings
- Differences between multimodal vs traditional video analysis
- In what ways can malicious actors misuse this technology?
- The future of multimodal video understanding
Jae's favorite books:
- Deep Learning (Authors: Ian Goodfellow, Yoshua Bengio, Aaron Courville)
- The Giving Tree (Author: Shel Silverstein)
--------
Where to find Prateek Joshi:
Newsletter: https://prateekjoshi.substack.com
Website: https://prateekj.com
LinkedIn: https://www.linkedin.com/in/prateek-joshi-91047b19
Twitter: https://twitter.com/prateekvjoshi

… continue reading

143集单集

#Tech #Prateek Joshi

Artwork

Multimodal Video Understanding

Infinite Machine Learning: Artificial Intelligence | Startups | Technology

13 subscribers

published 6M ago

分享

MP3•单集首页

内容由Prateek Joshi提供。所有播客内容（包括剧集、图形和播客描述）均由 Prateek Joshi 或其播客平台合作伙伴直接上传和提供。如果您认为有人在未经您许可的情况下使用您的受版权保护的作品，您可以按照此处概述的流程进行操作https://zh.player.fm/legal。

Jae Lee is the cofounder and CEO of Twelve Labs, where they are building video understanding infrastructure to help developers build programs that can see, hear, and understand the world. He was previously the Lead Data Scientist at the Ministry of National Defense in South Korea. He has a bachelors in computer science from UC Berkeley.
In this episode, we cover a range of topics including:
- What is multimodal video understanding
- State of play in multimodal video
- The founding of Twelve Labs
- The launch of Pegasus-1
- Four core principles: Efficient Long-form Video Processing, Multimodal Understanding, Video-native Embeddings, Deep Alignment between Video and Language Embeddings
- Differences between multimodal vs traditional video analysis
- In what ways can malicious actors misuse this technology?
- The future of multimodal video understanding
Jae's favorite books:
- Deep Learning (Authors: Ian Goodfellow, Yoshua Bengio, Aaron Courville)
- The Giving Tree (Author: Shel Silverstein)
--------
Where to find Prateek Joshi:
Newsletter: https://prateekjoshi.substack.com
Website: https://prateekj.com
LinkedIn: https://www.linkedin.com/in/prateek-joshi-91047b19
Twitter: https://twitter.com/prateekvjoshi

… continue reading

143集单集

#Tech #Prateek Joshi

所有剧集

×

欢迎使用Player FM

Player FM正在网上搜索高质量的播客，以便您现在享受。它是最好的播客应用程序，适用于安卓、iPhone和网络。注册以跨设备同步订阅。

收听超过500个主题