Outstanding Reviewer Award; one co-authored paper accepted.
Generative modeling · Multimodal intelligence · AI systems
Huadai Liu刘华岱
Generative intelligence across modalities.
I am Huadai Liu 刘华岱, a researcher and engineer working on generative and multimodal intelligence. I study how models learn expressive representations, reason across modalities and generate efficiently at scale.
My current work is grounded in speech, sound and music while addressing broader questions in representation learning and generative modeling. My first-author research has appeared at NeurIPS, ICML, ICLR, ACL, EMNLP and ACM MM.

News
Recent updates
Area Chair for ACL Rolling Review.
Two papers accepted, including first-author STAR-VAE.
One co-authored paper accepted to Findings.
Two papers accepted, including first-author PrismAudio.
First-author ThinkSound accepted.
Oral presentation for first-author MEDIC.
Oral · SAC Highlight for first-author FlashAudio.
First-author OmniAudio accepted.
Selected research
2024 — 2026
AudioCALM
Continuous autoregressive modeling for universal audio generation.
- Method
- Extends next-token prediction to continuous audio latents: a flow-matching head replaces the discrete softmax, block-causal AR-Flow supports streaming and arbitrary length, and A-MoME adds speech-specific capacity without extra inference cost for sound or music.
- Evidence
- One jointly trained model matches modality-specific state of the art across speech, sound and music.

STAR-VAE
Reshaping continuous audio latents around the structure of sound.
- Method
- Formalizes the rate–distortion–regularity trilemma, then uses a growth-based constraint field to route semantic structure and stochastic texture into capacity-matched channel subspaces. A CNN–Mamba backbone captures local detail and long-range context.
- Evidence
- State-of-the-art reconstruction fidelity and text-to-audio generation, with stronger semantic preservation.

PrismAudio
Reinforcement learning for video-to-audio across four perceptual axes.
- Method
- Decomposes planning into semantic, temporal, aesthetic and spatial chains of thought, each paired with a targeted reward. Fast-GRPO uses hybrid ODE–SDE sampling to make multi-objective RL post-training practical.
- Evidence
- State of the art on all four dimensions across VGGSound and the out-of-domain AudioCanvas benchmark.

ThinkSound
Reasoning before synthesis for interactive audio generation and editing.
- Method
- A multimodal LLM produces AudioCoT plans for foundational foley, object-centric refinement and language-guided editing; a unified audio foundation model executes each stage without fragmenting the workflow.
- Evidence
- State of the art on video-to-audio and reasoning metrics, with strong out-of-distribution performance and 1.4K+ GitHub stars.

FlashAudio
Straightening generative trajectories to make one-step synthesis work.
- Method
- Rectified flow is strengthened with bifocal timestep sampling, immiscible data–noise pairing and anchored classifier-free guidance—three interventions aimed at stable, high-fidelity one-step text-to-audio.
- Evidence
- Outperforms diffusion baselines using hundreds of steps at 400× real time; ACL 2025 Oral and SAC Highlight.

AudioLCM
High-fidelity text-to-audio in two to four sampling steps.
- Method
- Guided latent consistency distillation maps intermediate states directly toward the trajectory origin, while a multi-step ODE solver and an efficient transformer design preserve quality under aggressive step reduction.
- Evidence
- Competitive with hundreds-step systems at 333× real time on a single 4090Ti; 1.2K+ GitHub stars.
Research systems
View open-source workCosyVoice 2 at production scale
Designed and deployed Rectified Flow inference for CosyVoice 2, accelerating production within an open-source ecosystem with 22.5K+ GitHub stars and 2.6K forks.
Hours of speech at training scale
Architected a distributed training stack adopted across the Tongyi audio model family and Alibaba Cloud workloads.
Open-source reach for ThinkSound
Led the model, AudioCoT dataset and complete training and inference release; 81 forks, with weights and demos distributed through GitHub, Hugging Face and ModelScope.
Publications
Selected first-author work.
Academic service
Area Chair
Outstanding Reviewer
ICML · NeurIPS · ICLR · AAAI · ACL · EMNLP · ECCV · ACM Multimedia
IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)
Contact
Huadai Liu · 刘华岱