REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
Sofian Chaybouti(Goethe University Frankfurt), Hilde Kuehne(University of Bonn), Moritz Wolter, Walid Bousselham(IBM (United States))
Cited by 0
Related Papers
Grounding Everything: Emerging Localization Properties in Vision-Language Transformers
|Unknown|2024|38
Efficient Self-Ensemble for Semantic Segmentation
|arXiv (Cornell University)|2021|15
Efficient Self-Ensemble Framework for Semantic Segmentation
|arXiv (Cornell University)|2021|12
Learning Situation Hyper-Graphs for Video Question Answering
|Unknown|2023|8