AI Headphones Can Now Translate Multiple Voices in Real Time

Imagine sitting around a dinner table with friends who slip in and out of Spanish, French, or German—and still being able to follow every word in your native language.

2 mins read
A representational image [Mark Paton/Unsplash]

In a breakthrough that could reshape multilingual communication, researchers have developed a new AI-powered headphone system that can translate multiple people speaking at once in real time—complete with voice cloning and spatial audio cues. The technology, dubbed Spatial Speech Translation, was recently featured in MIT Technology Review for its innovative approach to one of automatic translation’s biggest challenges: group conversations involving multiple languages and speakers.

Imagine sitting around a dinner table with friends who slip in and out of Spanish, French, or German—and still being able to follow every word in your native language. That’s the experience the Spatial Speech Translation system aims to deliver.

Developed by a team led by Professor Shyam Gollakota at the University of Washington, the system uses artificial intelligence to isolate, interpret, and translate the voices of several speakers simultaneously. Unlike existing translation technologies—such as Meta’s AI features in its Ray-Ban smart glasses, which typically only handle one voice at a time—this system works in dynamic, multi-speaker environments and offers voice cloning that mimics each person’s tone and speech direction.

“There are so many smart people across the world, and the language barrier prevents them from having the confidence to communicate,” Gollakota told MIT Technology Review. “We think this kind of system could be transformative.” He shared a personal anecdote, noting how his mother, who speaks Telugu, struggles to communicate when visiting the U.S., despite having great ideas.

The prototype runs on off-the-shelf noise-canceling headphones equipped with microphones and is powered by a laptop using Apple’s M2 silicon chip—the same chip found in Apple’s Vision Pro headset. The research was presented at the ACM CHI Conference on Human Factors in Computing Systems held in Yokohama, Japan.

Spatial Speech Translation relies on two main AI models. The first maps the acoustic space around the listener to identify who is speaking and from which direction. The second model translates the detected speech into English text from Spanish, French, or German. Crucially, it also clones the speaker’s vocal tone and emotion, which helps produce a translation that sounds authentic rather than robotic. The result: the translated voice appears to come from the speaker’s actual position and sounds remarkably like them.

“This is a useful application. It can help people,” said Alina Karakanta, a computational linguist from Leiden University, who was not involved in the research.

Experts agree that what the team has accomplished is technically impressive, especially given the notoriously difficult task of separating human voices in noisy, real-world environments. “Real-time speech-to-speech translation is incredibly hard,” noted Samuele Cornell of Carnegie Mellon University. He cautioned, however, that scaling the system to real-world conditions will require more robust training data—including recordings with background noise.

The biggest challenge the team now faces is latency. Gollakota’s team is working to reduce the delay between a speaker’s utterance and the translated output to under a second, aiming for smoother, real-time conversation flow. However, linguistic structure complicates the effort. For example, German, which often places critical sentence elements at the end, takes longer to translate accurately compared to French or Spanish.

Claudio Fantinuoli, a translation researcher at Johannes Gutenberg University, warns that pushing for faster translations might compromise quality. “The longer you wait [before translating], the more context you have, and the better the translation will be. It’s a balancing act,” he said.

Still, if successful, the technology could significantly lower communication barriers in multilingual settings—from business meetings to family gatherings—bringing the long-standing dream of real-time, natural-sounding translation closer to reality.

Sri Lanka Guardian

The Sri Lanka Guardian is an online web portal founded in August 2007 by a group of concerned Sri Lankan citizens including journalists, activists, academics and retired civil servants. We are independent and non-profit. Email: editor@slguardian.org

Leave a Reply

Your email address will not be published.

Latest from Blog