I am a tenure-track Assistant Professor in the School of Artificial Intelligence (SAI) at The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China. Before joining CUHK-Shenzhen, I was a Research Scientist at A*STAR, Singapore, an Associate Faculty at the Singapore Institute of Technology (SIT), teaching the Large Language Models module, and a Research Fellow in the Department of Electrical and Computer Engineering (ECE) at the National University of Singapore (NUS). I have received a Ph.D. degree from the National University of Singapore, supervised by Prof. Haizhou Li (IEEE Fellow) and Prof. Shuzhi Sam Ge (IEEE Fellow). During my PhD studies, I was a visiting research scholar at National Institute of Informatics (Japan), supervised by Prof. Junichi Yamagishi. I received a B.Sc degree from Nanjing University, Nanjing, China in 2017.

My research interest includes multilingual speech and audio intelligence, multimodal language models, music and singing processing, and trustworthy AI. I have published high-impact top-tier AI journals and conferences, including IEEE Transaction on Audio, Speech and Language Processing (TALSP), IEEE Transactions on Multimedia (TMM), IEEE Signal Processing Letters (SPL), Speech Communications, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), ACL, EMNLP, AAAI, INTERSPEECH, IEEE ASRU, IEEE APSIPA ASC, IEEE Spoken Language Technology Workshop (SLT), Transactions of the Association for Computational Linguistics (TACL) and Pattern Recognition Letters.

πŸ”₯πŸ”₯ I am recruiting fully funded PhD students, Research Assistants/Research Engineers, Master’s students, Research interns (on-site or remote) and Visiting students to work on research in speech, singing, multimodal AI, and trustworthy AI. Students and researchers will have access to extensive GPU resources and opportunities to collaborate with leading researchers such as Prof. Haizhou Li (IEEE Fellow). Outstanding candidates may also be recommended for PhD programs and research internships at top universities and research labs in Singapore, the US, Japan, and the UK, as well as opportunities at top-tier technology companies and industrial research labs such as Google DeepMind, Meta, and Microsoft. If you are interested, email me at gaoxiaoxue@cuhk.edu.cn with your CV and indicate the position you are applying for. Please review CUHK-Shenzhen SAI’s admission requirements on the official website before contacting me.πŸ”₯πŸ”₯

πŸ”₯πŸ”₯ Master of Philosophy and PhD Application Deadlines:πŸ”₯πŸ”₯

πŸ“…πŸŒ±Spring 2027 Intake: 31 October 2026,apply at https://sai.cuhk.edu.cn/en/node/35

πŸ“…πŸ‚Fall 2027 Intake: 31 May 2027, apply at https://sai.cuhk.edu.cn/en/node/35

⭐ Applications are reviewed on a rolling basis. Early applications are strongly encouraged!

πŸ”₯ News

  • 2026: Β πŸŽ‰πŸŽ‰ Dr. Gao serves as one of the organizers of AACL 2026 TrustAudio workshop! More details for paper submissions in the TrustAudio website!
  • 2026: Β πŸŽ‰πŸŽ‰ One Signal Processing Letters paper has been accepted for publication!
  • 2026: Β πŸŽ‰πŸŽ‰ Two APSIPA papers have been accepted for publication!
  • 2026: Β πŸŽ‰πŸŽ‰ Two Interspeech papers have been accepted for publication!
  • 2026: Β πŸŽ‰πŸŽ‰ Our IJCAI paper has been accepted for publication!
  • 2026: Β πŸŽ‰πŸŽ‰ Dr. Gao serves as one of the organizers of VoiceMOS Challenge 2026! Join us through the VoiceMOS 2026 website.
  • 2026: Β πŸŽ‰πŸŽ‰ Our ACL paper has been accepted for publication!
  • 2026: Β πŸŽ‰πŸŽ‰ Our Pattern Recognition Letters has been accepted for publication!
  • 2026: Β πŸŽ‰πŸŽ‰ Our TACL paper has been accepted for publication!
  • 2025: Β πŸŽ‰πŸŽ‰ Dr. Gao serves as one of the organizers of AAAI 2026 worshop on audio AI. Our AAAI 2026 workshop has been accepted and is now open for submissions! 🌟 Check it in audio-aaai.
  • 2025: Β πŸŽ‰πŸŽ‰ Dr. Gao serves as one of the advisory board of ICASSP 2026 grand challenge. Our ICASSP 2026 grand challenge has been accepted and is now open for submissions! 🌟 Check it in ICASSP 2026 Cadenza Challenge website.
  • 2025: Β πŸŽ‰πŸŽ‰ Our IEEE ASRU paper has been accepted for publication!
  • 2025: Β πŸŽ‰πŸŽ‰ Our ACL paper has been accepted for publication!
  • 2025: Β πŸŽ‰πŸŽ‰ Our TALSP regular paper has been accepted for publication!
  • 2024: Β πŸŽ‰πŸŽ‰ Our ICASSP paper has been accepted for publication!
  • 2024: Β πŸŽ‰πŸŽ‰ Our AAAI has been accepted for publication!
  • 2024: Β πŸŽ‰πŸŽ‰ Our TMM has been accepted for publication!
  • 2024: Β πŸŽ‰πŸŽ‰ Our EMNLP has been accepted for publication!
  • 2024: Β πŸŽ‰πŸŽ‰ Two Signal Processing Letters have been accepted for publication!
  • 2023: Β πŸŽ‰πŸŽ‰ Dr. Gao was invited as the leading Guest Editor of the special issue β€œModeling of Multimodal Speech Recognition and Language Processing” in Electronics (IF:2.9, ISSN 2079-9292).
  • 2023: Β πŸŽ‰πŸŽ‰ Our TALSP regular paper has been accepted for publication!
  • 2023: Β πŸŽ‰πŸŽ‰ Two papers have been accepted by ICASSP 2023!
  • 2020: Β πŸŽ‰πŸŽ‰ Won first places for two tasks in Automatic Lyrics-to-Audio Alignment Task in Music Information Retreval Evaluation eXchange International Benchmarking Competition 2020. Check it in NUS ECE news.
  • 2019: Β πŸŽ‰πŸŽ‰ Received Best Poster Award Runner Up Prize at 4th Workshop for Young Female Researchers in INTERSPEECH, Graz, Austria. Check it in NUS ECE news.

πŸ“œ Research Area

Speech and Audio Intelligence :
   Automatic speech recognitionοΌ›Speech Emotion Recognition; Speech synthesis; Audio captioning; Agentic AI in speech and audio.
Singing and Music Processing :
   Speech-to-singing conversion; Singing voice conversion; Singing generation; Lyrics-to-audio alignment; Chord transcription; Music source separation; Musical genre recognition.
Trustworthy AI :
   Audio security and safety; Robustness and adversarial security; Fairness and bias in speech AI; LLM Jailbreak and safety alignment.
Multi-modal Processing:
   Audio-visual active speaker detection; Large audio language models; Multimodal generation.
Self-supervised Learning :
   Self-supervised speech processing; Self-supervised language processing.
Large Language Models :
   Large audio language models; Speech LLMs; Speech synthesis with large language models; Multilingual LLMs.

πŸ’» Research Experiences

  • 2026.08 - Present, Assistant Professor, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China.
  • 2024.02 - 2026.08, Research Scientist, A*STAR, Singapore.
  • 2023.11 - 2024.01, Visiting Researcher, Academia Sinica.
  • 2022.11 - 2023.11, Research Fellow, National University of Singapore (NUS), Singapore.
  • 2022.07 - 2022.08, Research Scholar, National Institute of Informatics, Japan.

πŸ“– Educations

  • 2017.08 - 2022.10, Ph.D. in Electrical and Computer Engineering, National University of Singapore (NUS), Singapore.
  • 2013.09 - 2017.07, B.Sc. in Electronic Information Science and Technology, Nanjing University, Nanjing, China.

πŸ“ Publications

– Journal Papers –

– Conference Papers –

πŸŽ– Honors and Awards

  • 2020 Ranked first in Automatic Lyrics-to-Audio Alignment Task in Music Information Retreval Evaluation eXchange International Benchmarking Competition 2020. The winning Lyrics-to-Audio Alignment system NUS Auto Lyrix Align is now available online as an interactive web interface: The winning Lyrics-to-Audio Alignment system NUS Auto Lyrix Align is now available online as an interactive web interface: https://autolyrixalign.hltnus.org/
  • 2020 Ranked first in Automatic Lyrics Transcription Task in Music Information Retreval Evaluation eXchange International Benchmarking Competition 2020.
  • 2019 Best Poster Award Runner Up Prize, β€œSpeech-to-Singing Conversion and Synthesis” at 4th Workshop for Young Female Researchers in INTERSPEECH, Graz, Austria.
  • 2019 ISCA Grants,β€œAverage Modeling for Spectral Mapping in Speech-to-Singing Conversion” at 2019 Speech Processing Courses in Crete Conversational Speech Synthesis: from design to evaluation, University of Crete, Heraklion Crete, Greece.
  • 2016 Meritorious Winner (Top 8% winner), American Mathematical Contest in Modeling.
  • 2015 National Second Prize, National Undergraduate Electronic Design Contest.

πŸ’¬ Talks

  • 2025.06, Generative AI in Speech: from Muscle Signals to Text, SINFRA 2025, France.
  • 2022.08, Automatic Lyrics Transcription of Polyphonic Music, National Institute of Informatics, Japan.

πŸ’» Internships

  • 2023.11 - 2024.01, Visiting Researcher, Academia Sinica.
  • 2022.07 - 2022.08, National Institute of Informatics, Japan.

πŸ“š Research Web Platform

πŸ‘” Projects

  • Synthetic Data Generation for Scaling Multilingual Models, A*STAR, Singapore
  • Expressive and Empathetic Human-AI Interaction by Enhancing Multilingual, Multi-modal Large Language Model, A*STAR, Singapore.
  • Heterogeneous Preference Learning for Value-Aligned Large Language Models in Multicultural Societies: A Resource-Efficient Approach, A*STAR, Singapore.
  • Intelligent Modelling for Decision-making in Critical Urban Systems, A*STAR, Singapore
  • Building National-Level Capability in Large Language Models β€œNational LLM”, A*STAR, Singapore
  • SpeechEval Phase II: SHE4EDU (Speech Highlighter and Evaluation for Education), A*STAR, Singapore
  • SingaKids Pic2Speak: Multilingual AI Tutor – Uplifting Singapore’s Bilingual Edge, A*STAR, Singapore
  • Human-Robot Collaborative AI for Advanced Manufacturing And Engineering, NUS, Singapore.
  • Perfect Singing Vocals, NUS, Singapore.