
Photo: Feixiang Xiao
In his doctoral dissertation, MSc Yuzhu Wang investigated neural speech separation in real-world conversational scenarios. Taking inspiration from the remarkable adaptability of biological hearing, his research asks how machines can separate and follow individual voices when the acoustic environment is complex and constantly changing. Wang developed speech separation methods for situations in which the number of speakers is unknown, conversations continue far beyond the short recordings used to train the system, and speakers move while talking. The proposed methods enabled neural systems to estimate how many people were speaking, keep track of the same speakers over extended recordings, and adapt as speakers changed position. Together, the results move machine hearing closer to the conditions in which real conversations actually happen. The research can support meeting transcription, hearing devices, teleconferencing and voice assistants.
The doctoral dissertation titled Neural Speech Separation in Real-world Conversational Scenarios by MSc Yuzhu Wang will be publicly examined in the Faculty of Information Technology and Communication Sciences at Tampere University on 6 October 2026. The dissertation is in the fields of signal processing and machine learning.
The Opponents will be Professor Reinhold Häb-Umbach from Paderborn University, Germany, and Associate Professor Tom Bäckström from Aalto University. The Custos will be Professor Tuomas Virtanen from Tampere University.
