CAST: Cross-Attention in Space and Time for Video Action RecognitionDong‐Ho Lee, Jinwoo Choi, Jongseo Lee|arXiv (Cornell University)|2023Cited by 5