Amazon Teams Up with Twelve Labs to Revolutionize Video Search with AI
On December 3, 2024, Amazon Web Services (AWS), a subsidiary of Amazon.com Inc. (NASDAQ: AMZN), made waves at the AWS re:Invent conference by announcing a strategic partnership with Twelve Labs, a pioneering startup specializing in multimodal artificial intelligence (AI). This collaboration aims to enhance the searchability of video content, making it as accessible as text.
1. Transforming Video Content with Multimodal AI
Twelve Labs is set to leverage AWS’s robust infrastructure to develop and scale its proprietary foundation models. These models will enable applications that can interpret and analyze various elements within video content, such as actions, objects, and background sounds. This advancement will allow developers to create tools that can search, classify, summarize, and segment video clips into chapters.
Unlocking New Possibilities for Developers
The availability of these foundation models on the AWS Marketplace opens the door for developers across various industries—including media, entertainment, sports, and gaming—to harness advanced video search capabilities. For instance, sports organizations can utilize this technology to efficiently catalog extensive libraries of game footage, facilitating the retrieval of specific frames for live broadcasts. Additionally, coaches can analyze athletes' techniques, leading to actionable performance improvements. Media companies can personalize viewer experiences by generating tailored highlight reels based on individual preferences.
2. A Vision for the Future of Video Content
Jae Lee, co-founder and CEO of Twelve Labs, articulated the mission behind the startup: “Nearly 80% of the world’s data is in video, yet most of it is unsearchable. We are now able to address this challenge, surfacing highly contextual videos to bring experiences to life.” He emphasized the role of AWS in providing the computational power necessary to push the boundaries of video understanding, allowing Twelve Labs to accelerate model training and deliver solutions to developers worldwide.
Comprehensive Video Analysis and Summarization
Twelve Labs’ Marengo and Pegasus foundation models offer groundbreaking video analysis capabilities. These models deliver not only text summaries but also audio translations in over 100 languages. They analyze the relationships between words, images, and sounds, allowing users to perform natural language searches to access specific moments in video content. This technology is already being employed by major sports leagues to automatically produce highlight reels, enhancing the viewing experience and boosting fan engagement.
3. Reducing Costs and Accelerating Model Training
To optimize its operations, Twelve Labs employs Amazon SageMaker HyperPod for training its foundation models. This innovative solution allows the startup to process various data formats—including videos, images, speech, and text—simultaneously. By distributing training workloads across multiple AWS compute instances, Twelve Labs can conduct uninterrupted training sessions that last weeks or even months. This approach significantly reduces costs and accelerates the time-to-market for their advanced AI models.
Strategic Collaboration for Global Expansion
In conjunction with a three-year Strategic Collaboration Agreement (SCA), Twelve Labs will work closely with AWS to deploy its video understanding models across new sectors. The partnership also aims to enhance Twelve Labs’ model training capabilities. AWS Activate, a program designed to support startup growth, has empowered Twelve Labs to expand its generative AI technology globally, enabling them to analyze vast quantities of video data with remarkable precision.
4. Conclusion: A New Era for Video Content Accessibility
With this collaboration, AWS and Twelve Labs are poised to transform the way industries interact with video content. By making video data searchable and actionable, they are unlocking new opportunities for creativity, analysis, and engagement across diverse sectors. As the partnership unfolds, it will be interesting to observe how these advancements shape the future of video technology.