Latest News

Xiaomi’s New Open-Source AI Can Isolate One Speaker From Overlapping Voices

Xiaomi has released and open-sourced Xiaomi-CocktailASR-1, a speech recognition model designed to identify and transcribe one person’s voice even when several people are speaking at the same time.

Xiaomi developed the model to address what is commonly known as the “cocktail party problem,” where overlapping voices can confuse conventional automatic speech recognition systems.

It Can Recognize Voices

CocktailASR-1 works by first taking a short audio sample of the person a user wants to follow.

The model then uses that sample as a voice reference while processing a recording containing multiple speakers. It attempts to identify only the selected person’s speech and transcribe it while ignoring the others.

This could be useful for meetings, interviews, group conversations, and other recordings where several people speak over one another.

Built Around an LLM Architecture

Xiaomi says CocktailASR-1 uses an end-to-end large language model architecture.

According to the company, the model achieved state-of-the-art results across several multi-speaker speech recognition benchmarks and outperformed existing systems designed for similar tasks.

Xiaomi also says the model maintains competitive performance in recordings containing only one speaker, meaning the multi-speaker capabilities do not significantly reduce standard transcription quality.

Avoids Transcribing the Wrong Person

The model is also designed to avoid transcribing the wrong person.

If the selected speaker is not present in a recording, CocktailASR-1 can return an empty result instead of attempting to assign another person’s speech to the target speaker.

The model also includes a reasoning mode that can provide additional information associated with how it arrived at a transcription.

Available on GitHub and Hugging Face

CocktailASR-1 is the latest AI model Xiaomi has made available to developers.

It follows other open-source releases, including MiMo-V2-Flash and Xiaomi Robotics-0.

Developers can now access Xiaomi-CocktailASR-1 through GitHub and Hugging Face for testing and further development.

The post Xiaomi’s New Open-Source AI Can Isolate One Speaker From Overlapping Voices appeared first on ProPakistani.

Show More

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button

Adblock Detected

Please consider supporting us by disabling your ad blocker