SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation

Youliang Zhang(Hainan University), Li Xiu, Zhaoyang Li, Deyu Zhou(Southeast University), Jiahe Zhang, Zixin Yin, Gang Yu(Shandong University)
arXiv (Cornell University)
July 14, 2025
Cited by 0


Related Papers

Self-training from labeled features for sentiment analysis
|Information Processing & Management|2010|190
ATM: Adversarial-neural Topic Model
|Information Processing & Management|2019|95
Biomedical Relation Extraction: From Binary to Complex
|Computational and Mathematical Methods in Medicine|2014|93