DevsNet: Deep Video Saliency Network using Short-term and Long-term Cues. (July 2020)
- Record Type:
- Journal Article
- Title:
- DevsNet: Deep Video Saliency Network using Short-term and Long-term Cues. (July 2020)
- Main Title:
- DevsNet: Deep Video Saliency Network using Short-term and Long-term Cues
- Authors:
- Fang, Yuming
Zhang, Chi
Min, Xiongkuo
Huang, Hanqin
Yi, Yugen
Zhai, Guangtao
Lin, Chia-Wen - Abstract:
- Highlights: We design a novel video saliency detection model by design the new 3-D ConvNet and B-ConvLSTM to extract short-term and long-term spatiotemporal cues, respectively. Through combining short-term and long-term spatiotemporal features, the proposed model can obtain promising performance for video saliency prediction. We design a new two-layer B-ConvLSTM structure for long-term spatiotemporal feature extraction for video saliency detection. The proposed B-ConvLSTM can extract the temporal information not just from the previous video frames but also from the next frames, which demonstrates that the proposed network takes both the forward and backward temporal features into account. Abstract: Recently, there have been various saliency detection methods proposed for still images based on deep learning techniques. However, the research on saliency detection for video sequences is still limited. In this study, we introduce a novel deep learning framework of saliency detection for video sequences, namely Deep Video Saliency Network (DevsNet). DevsNet mainly consists of two components: 3D Convolutional Network (3D-ConvNet) and Bidirectional Convolutional Long-Short Term Memory Network (B-ConvLSTM). 3D-ConvNet is constructed to learn short-term spatiotemporal information and the long-term spatiotemporal features are learned by B-ConvLSTM. The designed B-ConvLSTM can extract the temporal information not just from the previous video frames but also from the next frames, whichHighlights: We design a novel video saliency detection model by design the new 3-D ConvNet and B-ConvLSTM to extract short-term and long-term spatiotemporal cues, respectively. Through combining short-term and long-term spatiotemporal features, the proposed model can obtain promising performance for video saliency prediction. We design a new two-layer B-ConvLSTM structure for long-term spatiotemporal feature extraction for video saliency detection. The proposed B-ConvLSTM can extract the temporal information not just from the previous video frames but also from the next frames, which demonstrates that the proposed network takes both the forward and backward temporal features into account. Abstract: Recently, there have been various saliency detection methods proposed for still images based on deep learning techniques. However, the research on saliency detection for video sequences is still limited. In this study, we introduce a novel deep learning framework of saliency detection for video sequences, namely Deep Video Saliency Network (DevsNet). DevsNet mainly consists of two components: 3D Convolutional Network (3D-ConvNet) and Bidirectional Convolutional Long-Short Term Memory Network (B-ConvLSTM). 3D-ConvNet is constructed to learn short-term spatiotemporal information and the long-term spatiotemporal features are learned by B-ConvLSTM. The designed B-ConvLSTM can extract the temporal information not just from the previous video frames but also from the next frames, which demonstrates that the proposed model considers both the forward and backward temporal information. By combining the short-term and long-term spatiotemporal cues, the proposed DevsNet can extract saliency information for video sequences effectively and efficiently. Extensive experiments have been conducted to show that the proposed model can obtain better performance in spatiotemporal saliency prediction than the state-of-the-art models. … (more)
- Is Part Of:
- Pattern recognition. Volume 103(2020:Jul.)
- Journal:
- Pattern recognition
- Issue:
- Volume 103(2020:Jul.)
- Issue Display:
- Volume 103 (2020)
- Year:
- 2020
- Volume:
- 103
- Issue Sort Value:
- 2020-0103-0000-0000
- Page Start:
- Page End:
- Publication Date:
- 2020-07
- Subjects:
- Video saliency detection -- Spatiotemporal saliency -- 3D convolution network (3D-ConvNet) -- Bidirectional convolutional long-short term memory network (B-ConvLSTM)
Pattern perception -- Periodicals
Perception des structures -- Périodiques
Patroonherkenning
006.4 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00313203 ↗
http://www.sciencedirect.com/ ↗ - DOI:
- 10.1016/j.patcog.2020.107294 ↗
- Languages:
- English
- ISSNs:
- 0031-3203
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 13547.xml