Learning Situation Hyper-Graphs for Video Question Answering
Aisha Urooj Khan(University of Central Florida), Mubarak Shah(University of Central Florida), Walid Bousselham(IBM (United States)), Niels da Vitoria Lobo(University of Central Florida), Chuang Gan(University of Massachusetts Amherst), Hilde Kuehne(University of Bonn), Bo Wu(IBM (United States)), Kim Chheu(Western Michigan University)
Cited by 4
Related Papers
Visual Tracking: An Experimental Survey
|IEEE Transactions on Pattern Analysis and Machine Intelligence|2014|1.6k
Grounding Everything: Emerging Localization Properties in Vision-Language Transformers
|Unknown|2024|38
Efficient Self-Ensemble for Semantic Segmentation
|arXiv (Cornell University)|2021|15
Efficient Self-Ensemble Framework for Semantic Segmentation
|arXiv (Cornell University)|2021|12
Look, Listen, and Act: Towards Audio-Visual Embodied Navigation
|arXiv (Cornell University)|2019|10