My team and I build voice isolation tech at Krisp. Originally, our sole focus was human-to-human communication, where we invested heavily in representative test datasets and automated testing to ensure reliability in production.
With the recent rise of voice agents, we've expanded our scope to another critical use case: delivering intelligible audio to the STT models powering these bots.
Today, we're releasing a carefully curated benchmark dataset to test how voice isolation models actually perform in the wild. The data covers a wide range of real-world scenarios: from open offices and call centers to phone calls and car microphones. Crucially for voice isolation, the dataset includes audio with competing background voices, where modern STT engines struggle the most.
On this benchmark, Krisp Voice Isolation models demonstrate a strong recovery in STT performance, yielding an average 73% relative reduction in WER across multiple STT engines. Importantly, the dataset reflects real-world difficulty, including challenging samples where our own models fall short.
Would love to hear your feedback (in particular, scenarios we missed, tests we didn't do, etc.) and see how the community uses the data.
Hi HN, Hayk here.
My team and I build voice isolation tech at Krisp. Originally, our sole focus was human-to-human communication, where we invested heavily in representative test datasets and automated testing to ensure reliability in production.
With the recent rise of voice agents, we've expanded our scope to another critical use case: delivering intelligible audio to the STT models powering these bots.
Today, we're releasing a carefully curated benchmark dataset to test how voice isolation models actually perform in the wild. The data covers a wide range of real-world scenarios: from open offices and call centers to phone calls and car microphones. Crucially for voice isolation, the dataset includes audio with competing background voices, where modern STT engines struggle the most.
On this benchmark, Krisp Voice Isolation models demonstrate a strong recovery in STT performance, yielding an average 73% relative reduction in WER across multiple STT engines. Importantly, the dataset reflects real-world difficulty, including challenging samples where our own models fall short.
Would love to hear your feedback (in particular, scenarios we missed, tests we didn't do, etc.) and see how the community uses the data.
For the dataset and test results, please see here
Dataset: [https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benc...]
Results: [https://huggingface.co/spaces/Krisp-AI/VoiceIsolation-Benchm...]