Skip to main content
Public defence

Yuzhu Wang: Giving AI a microphone is easy, but teaching it to listen is harder

Tampere University
LocationKorkeakoulunkatu 7, Tampere
Hervanta Campus, Kampusareena, auditorium A223 and remote connection (link TBA)
Date6.10.2026 12.00–16.00 (UTC+3)
LanguageEnglish
Entrance feeFree of charge
A man stands in front of a purple wall with the text Tampere University Alumni.
Photo: Feixiang Xiao
In his doctoral dissertation, MSc Yuzhu Wang investigated neural speech separation in real-world conversational scenarios. Taking inspiration from the remarkable adaptability of biological hearing, his research asks how machines can separate and follow individual voices when the acoustic environment is complex and constantly changing. Wang developed speech separation methods for situations in which the number of speakers is unknown, conversations continue far beyond the short recordings used to train the system, and speakers move while talking. The proposed methods enabled neural systems to estimate how many people were speaking, keep track of the same speakers over extended recordings, and adapt as speakers changed position. Together, the results move machine hearing closer to the conditions in which real conversations actually happen. The research can support meeting transcription, hearing devices, teleconferencing and voice assistants.

The doctoral dissertation titled Neural Speech Separation in Real-world Conversational Scenarios by MSc Yuzhu Wang will be publicly examined in the Faculty of Information Technology and Communication Sciences at Tampere University on 6 October 2026. The dissertation is in the fields of signal processing and machine learning.

The Opponents will be Professor Reinhold Häb-Umbach from Paderborn University, Germany, and Associate Professor Tom Bäckström from Aalto University. The Custos will be Professor Tuomas Virtanen from Tampere University.