Enhancing the alignment between target words and corresponding frames for video captioning. (March 2021)
- Record Type:
- Journal Article
- Title:
- Enhancing the alignment between target words and corresponding frames for video captioning. (March 2021)
- Main Title:
- Enhancing the alignment between target words and corresponding frames for video captioning
- Authors:
- Tu, Yunbin
Zhou, Chang
Guo, Junjun
Gao, Shengxiang
Yu, Zhengtao - Abstract:
- Highlights: Visual tags are introduced to bridge the gap between vision and language. A textual-temporal attention model is devised and incorporated into the decoder to build exact alignment between target words and corresponding frames. Extensive experiments on two well-known datasets, i.e., MSVD and MSR-VTT, demonstrate that our proposed approach achieves remarkable improvements over the state-of-the-art methods. Abstract: Video captioning aims at translating from a sequence of video frames into a sequence of words with the encoder-decoder framework. Hence, it is critical to align these two different sequences. Most existing methods exploit soft-attention (temporal attention) mechanism to align target words with corresponding frames, where the relevance of them merely depends on the previously generated words (i.e., language context). As we know, however, there is an inherent gap between vision and language, and most of the words in a caption belong to non-visual words (e.g. "a", "is", and "in"). Hence, merely with the guidance of the language context, existing temporal attention-based methods cannot exactly align target words with corresponding frames. In order to address this problem, we first introduce pre-detected visual tags from the video to bridge the gap between vision and language. The reason is that visual tags not only belong to textual modality, but also can convey visual information. Then, we present a Textual-Temporal Attention Model (TTA) to exactly alignHighlights: Visual tags are introduced to bridge the gap between vision and language. A textual-temporal attention model is devised and incorporated into the decoder to build exact alignment between target words and corresponding frames. Extensive experiments on two well-known datasets, i.e., MSVD and MSR-VTT, demonstrate that our proposed approach achieves remarkable improvements over the state-of-the-art methods. Abstract: Video captioning aims at translating from a sequence of video frames into a sequence of words with the encoder-decoder framework. Hence, it is critical to align these two different sequences. Most existing methods exploit soft-attention (temporal attention) mechanism to align target words with corresponding frames, where the relevance of them merely depends on the previously generated words (i.e., language context). As we know, however, there is an inherent gap between vision and language, and most of the words in a caption belong to non-visual words (e.g. "a", "is", and "in"). Hence, merely with the guidance of the language context, existing temporal attention-based methods cannot exactly align target words with corresponding frames. In order to address this problem, we first introduce pre-detected visual tags from the video to bridge the gap between vision and language. The reason is that visual tags not only belong to textual modality, but also can convey visual information. Then, we present a Textual-Temporal Attention Model (TTA) to exactly align the target words with corresponding frames. The experimental results show that our proposed method outperforms the state-of-the-art methods on two well known datasets, i.e., MSVD and MSR-VTT. 1 … (more)
- Is Part Of:
- Pattern recognition. Volume 111(2021)
- Journal:
- Pattern recognition
- Issue:
- Volume 111(2021)
- Issue Display:
- Volume 111, Issue 2021 (2021)
- Year:
- 2021
- Volume:
- 111
- Issue:
- 2021
- Issue Sort Value:
- 2021-0111-2021-0000
- Page Start:
- Page End:
- Publication Date:
- 2021-03
- Subjects:
- Video captioning -- Alignment -- Visual tags -- Textual-temporal attention
Pattern perception -- Periodicals
Perception des structures -- Périodiques
Patroonherkenning
006.4 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00313203 ↗
http://www.sciencedirect.com/ ↗ - DOI:
- 10.1016/j.patcog.2020.107702 ↗
- Languages:
- English
- ISSNs:
- 0031-3203
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 14935.xml