My name is Zilong (Leon) Huang. I am a Ph.D. student in the Department of Electrical and Electronic Engineering at The Hong Kong Polytechnic University, advised by Prof. Kong Aik Lee and Prof. Man-Wai Mak. I am currently a visiting researcher at Kyoto University’s Speech and Audio Processing Laboratory, working with Prof. Tatsuya Kawahara.
My research interests include:
- Multimodal Large Language Models
- Affective Computing
- Emotion Recognition in Conversation
- Preference Alignment
🔥 News
2026.08: 🎉 One paper has been accepted by the 4th MRAC Workshop at ACM Multimedia 2026! [pdf] [code]
2026.07: 🎤 Tutorial speaker for Speech Large Language Models: Architectures, Efficient Adaptation, and Applications at ICME 2026.
2026.07: 🎉 Our team won the sixth place in MER2026 Challenge EmoPrefer Track (ACM Multimedia 2026).
2026.07: 🎉 Three papers have been accepted by INTERSPEECH 2026! [pdf] [pdf]
2026.03: 🎉 Started a visiting research attachment at Kyoto University’s Speech and Audio Processing Laboratory.
2026.01: 🎉 One paper has been accepted by ICASSP 2026! [paper]
2025.05: 🎉 Two papers have been accepted by INTERSPEECH 2025! [pdf] [pdf]
2025.01: 🎉 One paper has been accepted by ICASSP 2025! [paper]
2024.12: 🎉 Our team won 2nd place in the NIST SRE 2024 Audio-Visual Track.
2024.05: 🎉 One paper has been accepted by INTERSPEECH 2024! [pdf]
📝 Publications
- ACM MM 2026 Workshop Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion, Zilong Huang, J. Peng, J. Li, K. Li, W. Ren, K.-A. Lee, M.-W. Mak, and T. Kawahara. [code]
Error-augmented preference optimization with calibrated multi-model fusion for reliable emotion preference learning. - INTERSPEECH 2026 EII-SCL: Harnessing Emotional Inertia for Multimodal Emotion Recognition in Conversation, Zilong Huang, K.-A. Lee, C.-X. Gan, Z. Jin, R. Zuo, and M.-W. Mak.
Models speaker-specific emotional inertia through supervised contrastive learning to improve emotion-shift understanding. - INTERSPEECH 2026 EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation, Zilong Huang, K.-A. Lee, J. Li, Z. Li, and M.-W. Mak.
Supervises modality uncertainty and dynamically weights conflicting cues for robust multimodal emotion recognition. - ICASSP 2026 Distilling Attention Knowledge for Speaker Verification, Z. Jin, S. Liu, Z. Li, C.-X. Gan, Zilong Huang, M.-W. Mak, and K.-A. Lee.
- arXiv 2026 UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition, C.-X. Gan, P. Bell, M.-W. Mak, Z. Li, Z. Jin, Zilong Huang, and K.-A. Lee.
- INTERSPEECH 2025 The Sub-3Sec Problem: From Text-Independent to Text-Dependent Corpus, R. Zuo, K.-A. Lee, Zilong Huang, and M.-W. Mak.
- INTERSPEECH 2025 IDIR: Identifying and Distilling Informative Relations for Speaker Verification, C.-X. Gan, Z. Li, Z. Jin, Zilong Huang, M.-W. Mak, and K.-A. Lee.
- ICASSP 2025 Denoising Student Features with Diffusion Models for Knowledge Distillation in Speaker Verification, Z. Jin, Y. Tu, Z. Li, Zilong Huang, C.-X. Gan, and M.-W. Mak.
- INTERSPEECH 2024 MM-NodeFormer: Node Transformer Multimodal Fusion for Emotion Recognition in Conversation, Zilong Huang, M.-W. Mak, and K.-A. Lee.
Fuses text, audio, and visual emotion features according to their modality-specific emotional richness.
💻 Research & Challenge Experience
- Visiting Researcher, Kyoto University (Mar 2026 – Sep 2026). Personality-aware LLM-based emotion recognition with speaker traits, conversational context, and multimodal affective cues.
- ACM Multimedia 2026 MER Challenge. Our team won sixth place in Track 3 (MER-Prefer).
- NIST SRE 2024 Audio-Visual Track. Developed audio-visual representations and fusion; team won 2nd place.
- INTERSPEECH 2025 AVSEC-4. Contributed to BAV-MossFormer2 for binaural audio-visual speech enhancement. [team report]
📖 Education
May 2024 – Present, The Hong Kong Polytechnic University
Ph.D. in Electrical and Electronic Engineering
September 2022 – March 2024, The Hong Kong Polytechnic University
M.Sc. in Electrical and Information Engineering
Distinction (GPA: 3.86/4.3)
July 2023 – August 2023, Peking University
Summer School in Data Analysis and Visualization
September 2018 – June 2022, Fujian Normal University / University of Huddersfield
B.Eng. in Communication Engineering (Sino-British)
First-Class Honours
🧑🔬 Services
Reviewer:
- AAAI 2026, 2027
- ISCSLP 2026
- ICME 2026
- ICASSP 2025, 2026 (Outstanding Reviewer, ICASSP 2025)
- IJCNN 2025
Other service:
- Tutorial Speaker, “Speech Large Language Models for Under-Resourced Languages,” Interspeech 2026, Sydney, Australia.
- Tutorial Speaker, “Speech Large Language Models: Architectures, Efficient Adaptation, and Applications,” ICME 2026, Bangkok, Thailand.
- Invited Talk, “Emotion Recognition in Conversation,” NCMMSC Student Forum, Xinjiang, China, August 2024.
- Teaching Assistant, “ENG2003 Information Technology,” Hong Kong SAR, China, September–November 2024.
🎓 Academic Activities
- IEEE International Conference on Multimedia & Expo (ICME 2026), Thailand, July 2026 (Tutorial Speaker).
- ICASSP 2025 Satellite Event, Suzhou, China, May 2025.
- IEEE Spoken Language Technology Workshop (SLT 2024), Macao, China, December 2024.
- PolyU Research Student Conference, Hong Kong SAR, China, August 2024.
- INTERSPEECH 2024, Kos Island, Greece, September 2024.
- ICASSP 2024, Seoul, Korea, April 2024.
🎖 Honors and Awards
- PolyU RSAP-Outgoing, 2026.
- PolyU Outstanding Graduate Scholarship Award, 2024.
- National Second Prize, MathorCup, 2021.
- University of Huddersfield Scholarships, 2018–2022.
🌟 Extracurricular Activities
- Hall Tutor, Homantin Student Hall (Red Hall), 1 April 2025 – 1 March 2026.
- Gold Award and Audience’s Choice Award, National Landmarks CUS Postcard Design Competition 2025, January 2026.
- Volunteer, The 15th National Games Hong Kong Zone Volunteer Program (「第十五屆全國運動會香港賽區」義工), July 2024 – February 2026.
- Classic Chinese Cuisine Course, Shenzhen New Oriental Culinary School (深圳新东方烹饪学校), January–March 2024.