
Multimodal training datasets
sievedata.com (opens in a new tab)Sieve is an AI research lab building high-quality multimodal datasets spanning video, audio, images, text and 3D, combining exabyte-scale data infrastructure, novel multimodal understanding techniques and dozens of proprietary data sources. Video alone accounts for 80% of internet traffic, and across modalities data has become the enabling medium behind creativity, communication, gaming, AR/VR and robotics — which makes training data, rather than model architecture, the binding constraint on those applications. Sieve partners with the world's top AI labs and does it with a team of roughly 25 people.
