
Real-time voice and audio foundation models
cartesia.ai (opens in a new tab)Cartesia builds the model architectures underlying real-time multimodal AI. Its founding team met as PhDs at the Stanford AI Lab, where they invented State Space Models — a new primitive for training efficient large-scale foundation models that handles long sequences differently from the transformer architectures dominating the field. The company pairs deep model research with systems engineering and a design-minded product team, so the architectures reach shipped experiences rather than remaining papers. Its work is novel throughout and the team prizes execution speed alongside a high technical bar.