Exploiting appearance transfer and multi-scale context for efficient person image generation. (April 2022)
- Record Type:
- Journal Article
- Title:
- Exploiting appearance transfer and multi-scale context for efficient person image generation. (April 2022)
- Main Title:
- Exploiting appearance transfer and multi-scale context for efficient person image generation
- Authors:
- Shen, Chengkang
Wang, Peiyan
Tang, Wei - Abstract:
- Highlights: We introduce a novel two-stream context-aware appearance transfer network for efficient person image generation. It progressively transfers the appearance from the source stream to the target stream guided by their dense spatial correspondence and multi-scale context. The proposed appearance transfer module is the first of its kind to use the target stream to query and transfer the source stream. It effectively handles the difficulty of large motion. The proposed multi-scale context module is the first attempt to apply atrous convolutions for contextual modeling in person image generation. Multi-scale context helps the network recover occluded pixels. Compared with state-of-the-art methods, our network achieves comparable or superior performance using much fewer parameters while being significantly faster. We also show our network has a great advantage when large pose transform occurs. Abstract: Pose guided person image generation means to generate a photo-realistic person image conditioned on an input person image and a desired pose. This task requires spatial manipulation of the source image according to the target pose. However, convolutional neural networks (CNNs) are inherently limited to geometric transformations due to the fixed geometric structures in their building modules, i.e., convolution, pooling and unpooling, which cannot handle large motion and occlusions caused by large pose transform. This paper introduces a novel two-stream context-awareHighlights: We introduce a novel two-stream context-aware appearance transfer network for efficient person image generation. It progressively transfers the appearance from the source stream to the target stream guided by their dense spatial correspondence and multi-scale context. The proposed appearance transfer module is the first of its kind to use the target stream to query and transfer the source stream. It effectively handles the difficulty of large motion. The proposed multi-scale context module is the first attempt to apply atrous convolutions for contextual modeling in person image generation. Multi-scale context helps the network recover occluded pixels. Compared with state-of-the-art methods, our network achieves comparable or superior performance using much fewer parameters while being significantly faster. We also show our network has a great advantage when large pose transform occurs. Abstract: Pose guided person image generation means to generate a photo-realistic person image conditioned on an input person image and a desired pose. This task requires spatial manipulation of the source image according to the target pose. However, convolutional neural networks (CNNs) are inherently limited to geometric transformations due to the fixed geometric structures in their building modules, i.e., convolution, pooling and unpooling, which cannot handle large motion and occlusions caused by large pose transform. This paper introduces a novel two-stream context-aware appearance transfer network to address these challenges. It is a three-stage architecture consisting of a source stream and a target stream. Each stage features an appearance transfer module, a multi-scale context module and two-stream feature fusion modules. The appearance transfer module handles large motion by finding the dense correspondence between the two-stream feature maps and then transferring the appearance information from the source stream to the target stream. The multi-scale context module handles occlusion via contextual modeling, which is achieved by atrous convolutions of different sampling rates. Both quantitative and qualitative results indicate the proposed network can effectively handle challenging cases of large pose transform while retaining the appearance details. Compared with state-of-the-art approaches, it achieves comparable or superior performance using much fewer parameters while being significantly faster. … (more)
- Is Part Of:
- Pattern recognition. Volume 124(2022)
- Journal:
- Pattern recognition
- Issue:
- Volume 124(2022)
- Issue Display:
- Volume 124, Issue 2022 (2022)
- Year:
- 2022
- Volume:
- 124
- Issue:
- 2022
- Issue Sort Value:
- 2022-0124-2022-0000
- Page Start:
- Page End:
- Publication Date:
- 2022-04
- Subjects:
- Person image generation -- Appearance transfer -- Multi-scale context -- Efficient image generation
Pattern perception -- Periodicals
Perception des structures -- Périodiques
Patroonherkenning
006.4 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00313203 ↗
http://www.sciencedirect.com/ ↗ - DOI:
- 10.1016/j.patcog.2021.108451 ↗
- Languages:
- English
- ISSNs:
- 0031-3203
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 22256.xml