Work & Research / 02

Building tools to understand voice.

My work connects speech technology, music, voice physiology, and machine learning. The projects start with a question: what can sound tell us about the people making it?

Experience

Research and engineering across academia and industry.

  • Co-Founder & Research Scientist·Nettverk
    Sweden / Germany·Jul. 2025 – Mar. 2026
    • Co-founded a voice-first AI platform for operating rooms, enabling passive clinical documentation and real-time decision support.
    • Selected for the Johnson & Johnson AI Health QuickFire program; established research collaboration with Karolinska Institutet (KI).
    • Designed and led development of speech-driven ML systems for low-latency inference in noisy, multi-speaker surgical environments.
    • Built hardware-free, scalable pipelines integrating ASR, speaker understanding, and context-aware information retrieval.
  • Associate Researcher·KTH Speech, Music and Hearing Lab
    Stockholm, Sweden·Mar. 2025 – Present
    • Contributed to Språkbanken Tal (CLARIN SPEECH), Sweden's national speech technology infrastructure, building scalable tools and datasets for the European Language Grid.
    • Pioneered WaveEGG, the first ML model to predict physiological voice signals (EGG) directly from acoustic input — eliminating contact-based hardware and enabling non-invasive, real-time clinical voice assessment at scale.
    • Integrated ML-based TTS evaluation into clinical-grade software, optimizing runtime performance.
  • Doctoral Researcher·KTH Speech, Music and Hearing Lab
    Stockholm, Sweden·Jan. 2021 – Mar. 2025
    • Invented and led the development of VoiceMap, a hybrid DSP + ML benchmarking framework for speech synthesis, supporting large-scale multi-model comparison and voice system characterization and classification.
    • Built automated pipelines for 500+ hours of synthetic and pathological speech data, enabling real-time voice quality mapping and load-aware batch inference.
    • Collaborated with clinical and software engineers to co-design signal evaluation tools with low-latency processing and multi-platform deployment.
  • TTS Engineer·TikTok / ByteDance AI Lab
    Beijing, China·Jan. 2020 – Jan. 2021
    • Engineered and productionized backend systems for TikTok/Douyin's automated voice-over tool, supporting daily generation of millions of user videos.
    • Designed an A/B testing platform for perceptual TTS quality, integrating real-time performance tracking and workload-aware model selection.
    • Reduced inference latency by 30% through model quantization and batch-serving optimization across GPU clusters under global user load.
  • Lab Assistant·Peking University — Linguistics Lab
    Beijing, China·Sep. 2017 – Jan. 2020
    • Architected a web-based speech data infrastructure with structured metadata, enabling query-efficient access to 1000+ hours of annotated corpus for national-scale phonetic research.
    • Engineered automated signal processing pipelines in Python for large-scale acoustic data collection, annotation, and post-processing.

All publications

Peer-reviewed work and preprints.

2026 · Journal of Voice

A Voice Mapping Based Comparative Study of Vocal Characteristics in Bel Canto, Chinese National Singing and Popular Singing

2026 · arXiv

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment

Read paper ↗
2025 · Journal of the Acoustical Society of America (JASA)

A WaveNet-Based Model for Predicting the Electroglottographic Signal from the Acoustic Voice Signal

Read paper ↗
2024 · Journal of Speech, Language, and Hearing Research

Effects of Speech Characteristics on Electroglottographic and Instrumental Acoustic Voice Analysis Metrics in Women with Structural Dysphonia Before and After Treatment

Read paper ↗
2024 · Journal of Voice

Effects on Voice Quality of Thyroidectomy: A Qualitative and Quantitative Study Using Voice Maps

Read paper ↗
2022 · Applied Sciences

Mapping Phonation Types by Clustering of Multiple Metrics

Read paper ↗
2020 · ISAIMS '20: Proceedings of the 1st International Symposium on Artificial Intelligence in Medical Sciences

Spectrum Analysis of Bone-conducted Speech: A Study Based on Intelligibility

Read paper ↗
2019 · CISP-BMEI 2019 12th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics

Acoustic Analysis of Resonance Characteristics of Head Voice and Chest Voice

Read paper ↗

Service

  • Peer Reviewer · JASA, Journal of Voice, Laryngoscope, Scientific Reports, IEEE/ACM TASLP (20+ reviews) ·(Ongoing)
  • Organizing Committee Member · World Voice Day (Världsröstdagen), Sweden ·(May 2024)
  • Member · Voice Foundation, Philadelphia, USA ·(May 2023)

Recognition

  • Research Fellow · Karl Engvar Foundation Research Fellowship ·(2024)
  • Grant Recipient · KTH Travel Research Grant ·(2022)
  • International Fellowship Recipient · Doctoral Study Support Program for Overseas Researchers ·(2021)
  • Academic Excellence Award · Peking University ·(2019)
  • National Scholarship · Ministry of Education, China ·(2017)
  • Industrial Scholarship · Yiyang Industry Fund ·(2016)
  • Innovation Award 1st Prize · College Student Innovation and Entrepreneurship Project ·(2015)