Learning to transfer focus of graph neural network for scene graph parsing. (April 2021)
- Record Type:
- Journal Article
- Title:
- Learning to transfer focus of graph neural network for scene graph parsing. (April 2021)
- Main Title:
- Learning to transfer focus of graph neural network for scene graph parsing
- Authors:
- Jiang, Junjie
He, Zaixing
Zhang, Shuyou
Zhao, Xinyue
Tan, Jianrong - Abstract:
- Highlights: A new neural network architecture, the graphical focal network, is proposed to improve the recognition rate of semantic relationship in scene graph parsing task. The proposed graphical focal loss transfers the focus of network learning to the semantic relationship types with high value but limited instances. The proposed relative depth encoding module and regional layout encoding module introduce effective 3D spatial layout information. On the two evaluation metrics of scene graph parsing tasks, our method has achieved new advanced performance on the Visual Genome dataset. Abstract: Scene graph parsing has become a new challenge in the field of image understanding and pattern recognition in recent years. It captures objects and their relationships, and provides a structured representation of the visual scene. Among the three types of high-level relationships of scene graphs, semantic relationships, which contain the global understanding of the scene, are the core and the most valuable, while geometric and possessive relationships contain local and limited information. However, semantic relationships have the characteristics of multiple types and fewer instances, leading to a low recognition rate of most semantic relationships by existing detectors. To address this issue, this paper proposes a new architecture, the graphical focal network, which uses a decision-level global detector to capture the dependencies between object and relationship local detectors. WeHighlights: A new neural network architecture, the graphical focal network, is proposed to improve the recognition rate of semantic relationship in scene graph parsing task. The proposed graphical focal loss transfers the focus of network learning to the semantic relationship types with high value but limited instances. The proposed relative depth encoding module and regional layout encoding module introduce effective 3D spatial layout information. On the two evaluation metrics of scene graph parsing tasks, our method has achieved new advanced performance on the Visual Genome dataset. Abstract: Scene graph parsing has become a new challenge in the field of image understanding and pattern recognition in recent years. It captures objects and their relationships, and provides a structured representation of the visual scene. Among the three types of high-level relationships of scene graphs, semantic relationships, which contain the global understanding of the scene, are the core and the most valuable, while geometric and possessive relationships contain local and limited information. However, semantic relationships have the characteristics of multiple types and fewer instances, leading to a low recognition rate of most semantic relationships by existing detectors. To address this issue, this paper proposes a new architecture, the graphical focal network, which uses a decision-level global detector to capture the dependencies between object and relationship local detectors. We construct a graphical focal loss, which overcomes the lack of semantic relationship instances by adjusting the proportion of relationship loss based on the degree of relationship rarity and learning difficulty, and improves the stability of key object recognition by adjusting the proportion of object loss based on the degree of node connectivity and the value of neighborhood relationships. The proposed relative depth encoding module and regional layout encoding module, respectively, introduce relative depth information and more effective geometric layout information between objects, thereby further improving the performance. Experiments using the Visual Genome benchmark show that our method outperforms the most advanced competitors in two types of performance metrics. For semantic types, the recognition rate of our method is 2.0 times that of the baseline. … (more)
- Is Part Of:
- Pattern recognition. Volume 112(2021)
- Journal:
- Pattern recognition
- Issue:
- Volume 112(2021)
- Issue Display:
- Volume 112, Issue 2021 (2021)
- Year:
- 2021
- Volume:
- 112
- Issue:
- 2021
- Issue Sort Value:
- 2021-0112-2021-0000
- Page Start:
- Page End:
- Publication Date:
- 2021-04
- Subjects:
- Semantic relationship -- Graphical focus -- Scene graph -- Class imbalance -- Image understanding
Pattern perception -- Periodicals
Perception des structures -- Périodiques
Patroonherkenning
006.4 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00313203 ↗
http://www.sciencedirect.com/ ↗ - DOI:
- 10.1016/j.patcog.2020.107707 ↗
- Languages:
- English
- ISSNs:
- 0031-3203
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 15745.xml