Generalized pyramid co-attention with learnable aggregation net for video question answering. (December 2021)
- Record Type:
- Journal Article
- Title:
- Generalized pyramid co-attention with learnable aggregation net for video question answering. (December 2021)
- Main Title:
- Generalized pyramid co-attention with learnable aggregation net for video question answering
- Authors:
- Gao, Lianli
Chen, Tangming
Li, Xiangpeng
Zeng, Pengpeng
Zhao, Lei
Li, Yuan-Fang - Abstract:
- Highlights: To handle the complexity of videos in V-VQA, we propose a generalized pyramid co-attention mechanism with diversity learning to explicitly encourage accuracy and diverse attention maps. For this generalized module, two possible ways are tried, Multi-path Pyramid Co-attention with diversity learning (MPC) and Cascaded Pyramid Transformer Co-attention with diversity learning (CPTC). This strategy benefits the capturing of distinct, complementary and informative features. To aggregate the sequential features without destroying the feature distributions and temporal information, we propose a new learnable aggregation component. It imitates Bags-of-Words (BoW) quantization mechanism to automatically aggregate adaptively-weighted frame-level feature (or word-level feature). We extensively evaluate the effectiveness of the overall model on two publicly available datasets (i.e., TGIF-QA and TVQA) for V-VQA task. The experimental results demonstrate that our model outperforms the existing state-of-the-art by a large margin and our extended CPTC performs better than MPC. Code and model have been released at: https://github.com/lixiangpengcs/LAD-Net-for-VideoQA . Abstract: Video based visual question answering (V-VQA) remains challenging at the intersection of vision and language. In this paper, we propose a novel architecture, namely Generalized Pyramid Co-attention with Learnable Aggregation Net (GPC) to address two central problems: 1) how to deploy co-attention to V-VQAHighlights: To handle the complexity of videos in V-VQA, we propose a generalized pyramid co-attention mechanism with diversity learning to explicitly encourage accuracy and diverse attention maps. For this generalized module, two possible ways are tried, Multi-path Pyramid Co-attention with diversity learning (MPC) and Cascaded Pyramid Transformer Co-attention with diversity learning (CPTC). This strategy benefits the capturing of distinct, complementary and informative features. To aggregate the sequential features without destroying the feature distributions and temporal information, we propose a new learnable aggregation component. It imitates Bags-of-Words (BoW) quantization mechanism to automatically aggregate adaptively-weighted frame-level feature (or word-level feature). We extensively evaluate the effectiveness of the overall model on two publicly available datasets (i.e., TGIF-QA and TVQA) for V-VQA task. The experimental results demonstrate that our model outperforms the existing state-of-the-art by a large margin and our extended CPTC performs better than MPC. Code and model have been released at: https://github.com/lixiangpengcs/LAD-Net-for-VideoQA . Abstract: Video based visual question answering (V-VQA) remains challenging at the intersection of vision and language. In this paper, we propose a novel architecture, namely Generalized Pyramid Co-attention with Learnable Aggregation Net (GPC) to address two central problems: 1) how to deploy co-attention to V-VQA task considering the complex and diverse content of videos; and 2) how to aggregate the frame-level features (or word-level features) without destroying the feature distributions and temporal information. To solve the first problem, we propose a Generalized Pyramid Co-attention structure with a novel diversity learning module to explicitly encourage attention accuracy and diversity. And we first instantiate it into a Multi-path Pyramid Co-attention (MPC) to capture diverse feature. Then we find each attention branch of original co-attention mechanism does not interact with the others, which results in coarse attention maps. So we extend the MPC structure to a Cascaded Pyramid Transformer Co-attention (CPTC) module in which we replace co-attention with transformer co-attention. To solve the second problem, we propose a new learnable aggregation method with a set of evidence gates. It automatically aggregates adaptively-weighted frame-level features (or word-level features) to extract rich video (or question) context semantic information. With evidence gates, it then further chooses the most related signals representing the evidence information to predict the answer. Extensive validations on the two V-VQA datasets, TGIF-QA and TVQA show that both our proposed MPC and CPTC achieve the state-of-the-art performance and CPTC performs better under various settings and metrics. Code and model have been released at:https://github.com/lixiangpengcs/LAD-Net-for-VideoQA . … (more)
- Is Part Of:
- Pattern recognition. Volume 120(2021)
- Journal:
- Pattern recognition
- Issue:
- Volume 120(2021)
- Issue Display:
- Volume 120, Issue 2021 (2021)
- Year:
- 2021
- Volume:
- 120
- Issue:
- 2021
- Issue Sort Value:
- 2021-0120-2021-0000
- Page Start:
- Page End:
- Publication Date:
- 2021-12
- Subjects:
- Video question answering -- Diversity learning -- Learnable aggregation -- Cascaded pyramid transformer co-attention
Pattern perception -- Periodicals
Perception des structures -- Périodiques
Patroonherkenning
006.4 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00313203 ↗
http://www.sciencedirect.com/ ↗ - DOI:
- 10.1016/j.patcog.2021.108145 ↗
- Languages:
- English
- ISSNs:
- 0031-3203
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 18489.xml