Open-ended remote sensing visual question answering with transformers. Issue 18 (17th September 2022)
- Record Type:
- Journal Article
- Title:
- Open-ended remote sensing visual question answering with transformers. Issue 18 (17th September 2022)
- Main Title:
- Open-ended remote sensing visual question answering with transformers
- Authors:
- Al Rahhal, Mohamad M.
Bazi, Yakoub
Alsaleh, Sara O.
Al-Razgan, Muna
Mekhalfi, Mohamed Lamine
Al Zuair, Mansour
Alajlan, Naif - Abstract:
- ABSTRACT: Visual question answering (VQA) has been attracting attention in remote sensing very recently. However, the proposed solutions remain rather limited in the sense that the existing VQA datasets address closed-ended question-answer queries, which may not necessarily reflect real open-ended scenarios. In this paper, we propose a new dataset named VQA-TextRS that was built manually with human annotations and considers various forms of open-ended question-answer pairs. Moreover, we propose an encoder-decoder architecture via transformers on account of their self-attention property that allows relational learning of different positions of the same sequence without the need of typical recurrence operations. Thus, we employed vision and natural language processing (NLP) transformers respectively to draw visual and textual cues from the image and respective question. Afterwards, we applied a transformer decoder, which enables the cross-attention mechanism to fuse the earlier two modalities. The fusion vectors correlate with the process of answer generation to produce the final form of the output. We demonstrate that plausible results can be obtained in open-ended VQA. For instance, the proposed architecture scores an accuracy of 84.01% on questions related to the presence of objects in the query images.
- Is Part Of:
- International journal of remote sensing. Volume 43:Issue 18(2022)
- Journal:
- International journal of remote sensing
- Issue:
- Volume 43:Issue 18(2022)
- Issue Display:
- Volume 43, Issue 18 (2022)
- Year:
- 2022
- Volume:
- 43
- Issue:
- 18
- Issue Sort Value:
- 2022-0043-0018-0000
- Page Start:
- 6809
- Page End:
- 6823
- Publication Date:
- 2022-09-17
- Subjects:
- Visual question answering -- remote sensing -- open-set dataset -- vision transformers -- encoder-decoder architecture
Remote sensing -- Periodicals
Télédétection -- Périodiques
621.3678 - Journal URLs:
- http://www.tandfonline.com/toc/tres20/current ↗
http://www.tandfonline.com/ ↗ - DOI:
- 10.1080/01431161.2022.2145583 ↗
- Languages:
- English
- ISSNs:
- 0143-1161
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 4542.528000
British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 24596.xml