Generative modeling · Multimodal intelligence · AI systems

Huadai Liu刘华岱

Generative intelligence across modalities.

I am Huadai Liu 刘华岱, a researcher and engineer working on generative and multimodal intelligence. I study how models learn expressive representations, reason across modalities and generate efficiently at scale.

My current work is grounded in speech, sound and music while addressing broader questions in representation learning and generative modeling. My first-author research has appeared at NeurIPS, ICML, ICLR, ACL, EMNLP and ACM MM.

Huadai Liu by the water

News

Recent updates

ECCV 2026

Outstanding Reviewer Award; one co-authored paper accepted.

ACL Rolling Review

Area Chair for ACL Rolling Review.

ICML 2026

Two papers accepted, including first-author STAR-VAE.

ACL 2026 · Findings

One co-authored paper accepted to Findings.

ICLR 2026

Two papers accepted, including first-author PrismAudio.

ACM MM 2025

Oral presentation for first-author MEDIC.

ACL 2025

Oral · SAC Highlight for first-author FlashAudio.

ICML 2025

First-author OmniAudio accepted.

Selected research

2024 — 2026
AudioCALM architecture with continuous autoregressive flow matching, block-causal attention and A-MoME
Architecture · continuous AR-Flow
01Universal Audio · 2026

AudioCALM

Continuous autoregressive modeling for universal audio generation.

Method
Extends next-token prediction to continuous audio latents: a flow-matching head replaces the discrete softmax, block-causal AR-Flow supports streaming and arbitrary length, and A-MoME adds speech-specific capacity without extra inference cost for sound or music.
Evidence
One jointly trained model matches modality-specific state of the art across speech, sound and music.
02ICML · 2026

STAR-VAE

Reshaping continuous audio latents around the structure of sound.

Method
Formalizes the rate–distortion–regularity trilemma, then uses a growth-based constraint field to route semantic structure and stochastic texture into capacity-matched channel subspaces. A CNN–Mamba backbone captures local detail and long-range context.
Evidence
State-of-the-art reconstruction fidelity and text-to-audio generation, with stronger semantic preservation.
03ICLR · 2026

PrismAudio

Reinforcement learning for video-to-audio across four perceptual axes.

Method
Decomposes planning into semantic, temporal, aesthetic and spatial chains of thought, each paired with a targeted reward. Fast-GRPO uses hybrid ODE–SDE sampling to make multi-objective RL post-training practical.
Evidence
State of the art on all four dimensions across VGGSound and the out-of-domain AudioCanvas benchmark.
ThinkSound teaser showing step-by-step chain-of-thought audio generation and editing
AudioCoT · generation and editing
04NeurIPS · 2025

ThinkSound

Reasoning before synthesis for interactive audio generation and editing.

Method
A multimodal LLM produces AudioCoT plans for foundational foley, object-centric refinement and language-guided editing; a unified audio foundation model executes each stage without fragmenting the workflow.
Evidence
State of the art on video-to-audio and reasoning metrics, with strong out-of-distribution performance and 1.4K+ GitHub stars.
FlashAudio official comparison of generation speed and audio quality
Quality–speed frontier · one step
05ACL Oral · 2025

FlashAudio

Straightening generative trajectories to make one-step synthesis work.

Method
Rectified flow is strengthened with bifocal timestep sampling, immiscible data–noise pairing and anchored classifier-free guidance—three interventions aimed at stable, high-fidelity one-step text-to-audio.
Evidence
Outperforms diffusion baselines using hundreds of steps at 400× real time; ACL 2025 Oral and SAC Highlight.
AudioLCM architecture showing guided latent consistency distillation
Guided latent consistency distillation
06ACM MM · 2024

AudioLCM

High-fidelity text-to-audio in two to four sampling steps.

Method
Guided latent consistency distillation maps intermediate states directly toward the trajectory origin, while a multi-step ODE solver and an efficient transformer design preserve quality under aggressive step reduction.
Evidence
Competitive with hundreds-step systems at 333× real time on a single 4090Ti; 1.2K+ GitHub stars.
5–10×

CosyVoice 2 at production scale

Designed and deployed Rectified Flow inference for CosyVoice 2, accelerating production within an open-source ecosystem with 22.5K+ GitHub stars and 2.6K forks.

100K+

Hours of speech at training scale

Architected a distributed training stack adopted across the Tongyi audio model family and Alibaba Cloud workloads.

1.4K+

Open-source reach for ThinkSound

Led the model, AudioCoT dataset and complete training and inference release; 81 forks, with weights and demos distributed through GitHub, Hugging Face and ModelScope.

Publications

Selected first-author work.

Full publication list

Academic service

ACL Rolling Review · 2026

Area Chair

ECCV · 2026

Outstanding Reviewer

Conference reviewer

ICML · NeurIPS · ICLR · AAAI · ACL · EMNLP · ECCV · ACM Multimedia

Journal reviewer

IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)

Contact

Huadai Liu · 刘华岱

Open to new research directions.

I welcome thoughtful conversations around research, collaboration and technically ambitious problems.

Start a conversation

Email Huadai
Research collaborationAcademic exchangeApplied research