Visual question answering from another perspective: CLEVR mental rotation tests. (April 2023)
- Record Type:
- Journal Article
- Title:
- Visual question answering from another perspective: CLEVR mental rotation tests. (April 2023)
- Main Title:
- Visual question answering from another perspective: CLEVR mental rotation tests
- Authors:
- Beckham, Christopher
Weiss, Martin
Golemo, Florian
Honari, Sina
Nowrouzezahrai, Derek
Pal, Christopher - Abstract:
- Highlights: We propose a version of CLEVR that is inspired by mental rotation tests. Latent feature volumes can be used instead of feature maps for VQA tasks grounded in 3D. Contrastive learning can be used to learn an encoder that maps images to latent volumes. Abstract: Different types of mental rotation tests have been used extensively in psychology to understand human visual reasoning and perception. Understanding what an object or visual scene would look like from another viewpoint is a challenging problem that is made even harder if it must be performed from a single image. We explore a controlled setting whereby questions are posed about the properties of a scene if that scene was observed from another viewpoint. To do this we have created a new version of the CLEVR dataset that we call CLEVR Mental Rotation Tests (CLEVR-MRT). Using CLEVR-MRT we examine standard methods, show how they fall short, then explore novel neural architectures that involve inferring volumetric representations of a scene. These volumes can be manipulated via camera-conditioned transformations to answer the question. We examine the efficacy of different model variants through rigorous ablations and demonstrate the efficacy of volumetric representations.
- Is Part Of:
- Pattern recognition. Volume 136(2023)
- Journal:
- Pattern recognition
- Issue:
- Volume 136(2023)
- Issue Display:
- Volume 136, Issue 2023 (2023)
- Year:
- 2023
- Volume:
- 136
- Issue:
- 2023
- Issue Sort Value:
- 2023-0136-2023-0000
- Page Start:
- Page End:
- Publication Date:
- 2023-04
- Subjects:
- Deep learning -- Computer vision -- Visual question answering -- Contrastive learning -- Clevr
Pattern perception -- Periodicals
Perception des structures -- Périodiques
Patroonherkenning
006.4 - Journal URLs:
- http://www.sciencedirect.com/science/journal/00313203 ↗
http://www.sciencedirect.com/ ↗ - DOI:
- 10.1016/j.patcog.2022.109209 ↗
- Languages:
- English
- ISSNs:
- 0031-3203
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 25681.xml