GuessWhich? Visual dialog with attentive memory network. (June 2021)
- Record Type:
- Journal Article
- Title:
- GuessWhich? Visual dialog with attentive memory network. (June 2021)
- Main Title:
- GuessWhich? Visual dialog with attentive memory network
- Authors:
- Zhao, Lei
Lyu, Xinyu
Song, Jingkuan
Gao, Lianli - Abstract:
- Highlights: We use memory network in the cooperative 'GuessWhich' game between Q-BOT and A-BOT. It reduces the repetition of the generated dialogs and makes image retrieval efficient. We propose a novel Attentive Memory Network that adds a fusion model to the memory network. The fusion model can effectively use the manually labeled caption and the image. Thus the generated dialogs and the predicted image representation can be visually grounded. Experiments conducted on VisDial 1.0 datasets demonstrate that our generated dialogs are natural and precise, and the results exceed the state-of-the-art 'GuessWhich' based visual dialog algorithms. Extensive image retrieval experiments prove that our method also can generate more accurate results compared to the benchmarks. Abstract: Visual dialog is a task that two agents: Question-BOT (Q-BOT) and Answer-BOT (A-BOT), which communicate in natural language on the situation of information asymmetry. Q-BOT generates questions based on an image caption and a historical dialog. A-BOT answers the questions grounded on the image. Moreover, we play a cooperative 'image guessing' game between Q-BOT and A-BOT, so that Q-BOT can select an unseen image from a set of images. However, as the valid information of the image caption and the historical dialog fades along the interaction, existing methods usually generate irrelevant and homogenous questions, which are worthless to the visual dialog system. To tackle this issue, we propose an A ttentiveHighlights: We use memory network in the cooperative 'GuessWhich' game between Q-BOT and A-BOT. It reduces the repetition of the generated dialogs and makes image retrieval efficient. We propose a novel Attentive Memory Network that adds a fusion model to the memory network. The fusion model can effectively use the manually labeled caption and the image. Thus the generated dialogs and the predicted image representation can be visually grounded. Experiments conducted on VisDial 1.0 datasets demonstrate that our generated dialogs are natural and precise, and the results exceed the state-of-the-art 'GuessWhich' based visual dialog algorithms. Extensive image retrieval experiments prove that our method also can generate more accurate results compared to the benchmarks. Abstract: Visual dialog is a task that two agents: Question-BOT (Q-BOT) and Answer-BOT (A-BOT), which communicate in natural language on the situation of information asymmetry. Q-BOT generates questions based on an image caption and a historical dialog. A-BOT answers the questions grounded on the image. Moreover, we play a cooperative 'image guessing' game between Q-BOT and A-BOT, so that Q-BOT can select an unseen image from a set of images. However, as the valid information of the image caption and the historical dialog fades along the interaction, existing methods usually generate irrelevant and homogenous questions, which are worthless to the visual dialog system. To tackle this issue, we propose an A ttentive M emory N etwork (AMN) to fully exploit the image caption and historical dialog information. Specifically, the attentive memory network mainly consists of a memory network and a fusion module. The memory network holds long term historical dialog information and gives each round of the dialog a different weight. Aside from the historical dialog information, the fusion module in Q-BOT and A-BOT further uses the image caption and the image feature, respectively. The caption information assists Q-BOT with the attentive generation of the questions, and the image feature helps A-BOT produce precise answers. With the AMN, the generated questions are diverse and concentrated, and the corresponding answers are accurate. The experimental results on VisDial v1.0 show the effectiveness of our proposed model, which outperforms the state-of-the-art methods. … (more)
- Is Part Of:
- Pattern recognition. Volume 114(2021)
- Journal:
- Pattern recognition
- Issue:
- Volume 114(2021)
- Issue Display:
- Volume 114, Issue 2021 (2021)
- Year:
- 2021
- Volume:
- 114
- Issue:
- 2021
- Issue Sort Value:
- 2021-0114-2021-0000
- Page Start:
- Page End:
- Publication Date:
- 2021-06
- Subjects:
- Visual dialog -- Attentive memory network -- Reinforcement learning
Pattern perception -- Periodicals
Perception des structures -- Périodiques
Patroonherkenning
006.4 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00313203 ↗
http://www.sciencedirect.com/ ↗ - DOI:
- 10.1016/j.patcog.2021.107823 ↗
- Languages:
- English
- ISSNs:
- 0031-3203
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 15940.xml