
Multimodal video understanding models
twelvelabs.io (opens in a new tab)TwelveLabs builds the intelligence layer for video, which accounts for roughly 90% of the world's data and remains largely unsearchable by its contents. Its multimodal AI models understand video the way people do — across sight, sound and motion together rather than by transcribing audio and treating frames separately — and power production-scale workloads across media, entertainment, sports, security and government. The company is headquartered in San Francisco with offices in Seoul, New York and London, and employees distributed globally.