
Full-duplex speech-to-speech AI models
misolabs.ai (opens in a new tab)Miso Labs trains full-duplex speech-to-speech AI models, which can listen and speak at the same time, interrupting or talking over a user the way people do in real conversation. Building them takes enormous volumes of audio, and the company's speech dataset now exceeds 100 million hours, roughly 11,000 years of recordings, alongside human feedback data used to refine the models. To produce that data, Miso Labs runs a 7,000-square-foot recording studio in North Hollywood and works with a growing roster of voice actors who record from their own home studios.