UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
Huaishao Luo(Southwest Jiaotong University), Ming Zhou, Botian Shi, Taroon Bharti, Tianrui Li(Southwest Jiaotong University), Lei Ji, Haoyang Huang, Jason Li(Queen's University), Nan Duan(Microsoft Research Asia (China))
Cited by 169
Related Papers
CodeBERT: A Pre-Trained Model for Programming and Natural Languages
|Unknown|2020|2.5k
CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning
|Neurocomputing|2022|676
UniXcoder: Unified Cross-Modal Pre-training for Code Representation
|Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)|2022|554
Predicting citywide crowd flows using deep spatio-temporal residual networks
|Artificial Intelligence|2018|517
Forecasting Fine-Grained Air Quality Based on Big Data
|Unknown|2015|472