Benchmarking unsupervised near-duplicate image detection. (30th November 2019)
- Record Type:
- Journal Article
- Title:
- Benchmarking unsupervised near-duplicate image detection. (30th November 2019)
- Main Title:
- Benchmarking unsupervised near-duplicate image detection
- Authors:
- Morra, Lia
Lamberti, Fabrizio - Abstract:
- Highlights: Unsupervised near-duplicate image detection requires high specificity up to 10 − 6 – 10 − 9 . Empirical comparison of CNN-based descriptors for near-duplicate image detection. Validated, principled methodology to estimate sensitivity and estimate false alarms. Fine-tuning CNNs for retrieval is beneficial but may suffer in specificity. New set of annotations released for near-duplicate detection benchmarking. Graphical abstract: Abstract: Unsupervised near-duplicate detection has many practical applications ranging from social media analysis and web-scale retrieval, to digital image forensics. It entails running a threshold-limited query on a set of descriptors extracted from the images, with the goal of identifying all possible near-duplicates, while limiting the false positives due to visually similar images. Since the rate of false alarms grows with the dataset size, a very high specificity is thus required, up to 1–10 − 9 for realistic use cases; this important requirement, however, is often overlooked in literature. In recent years, descriptors based on deep convolutional neural networks have matched or surpassed traditional feature extraction methods in content-based image retrieval tasks. To the best of our knowledge, ours is the first attempt to establish the performance range of deep learning-based descriptors for unsupervised near-duplicate detection on a range of datasets, encompassing a broad spectrum of near-duplicate definitions. We leverage bothHighlights: Unsupervised near-duplicate image detection requires high specificity up to 10 − 6 – 10 − 9 . Empirical comparison of CNN-based descriptors for near-duplicate image detection. Validated, principled methodology to estimate sensitivity and estimate false alarms. Fine-tuning CNNs for retrieval is beneficial but may suffer in specificity. New set of annotations released for near-duplicate detection benchmarking. Graphical abstract: Abstract: Unsupervised near-duplicate detection has many practical applications ranging from social media analysis and web-scale retrieval, to digital image forensics. It entails running a threshold-limited query on a set of descriptors extracted from the images, with the goal of identifying all possible near-duplicates, while limiting the false positives due to visually similar images. Since the rate of false alarms grows with the dataset size, a very high specificity is thus required, up to 1–10 − 9 for realistic use cases; this important requirement, however, is often overlooked in literature. In recent years, descriptors based on deep convolutional neural networks have matched or surpassed traditional feature extraction methods in content-based image retrieval tasks. To the best of our knowledge, ours is the first attempt to establish the performance range of deep learning-based descriptors for unsupervised near-duplicate detection on a range of datasets, encompassing a broad spectrum of near-duplicate definitions. We leverage both established and new benchmarks, such as the Mir-Flick Near-Duplicate (MFND) dataset, in which a known ground truth is provided for all possible pairs over a general, large scale image collection. To compare the specificity of different descriptors, we reduce the problem of unsupervised detection to that of binary classification of near-duplicate vs. not-near-duplicate images. The latter can be conveniently characterized using Receiver Operating Curve (ROC). Our findings in general favor the choice of fine-tuning deep convolutional networks, as opposed to using off-the-shelf features, but differences at high specificity settings depend on the dataset and are often small. The best performance was observed on the MFND benchmark, achieving 96% sensitivity at a false positive rate of 1.43 × 10 − 6 . … (more)
- Is Part Of:
- Expert systems with applications. Volume 135(2019)
- Journal:
- Expert systems with applications
- Issue:
- Volume 135(2019)
- Issue Display:
- Volume 135, Issue 2019 (2019)
- Year:
- 2019
- Volume:
- 135
- Issue:
- 2019
- Issue Sort Value:
- 2019-0135-2019-0000
- Page Start:
- 313
- Page End:
- 326
- Publication Date:
- 2019-11-30
- Subjects:
- Near-duplicate detection -- Convolutional neural networks -- Instance-level retrieval -- Unsupervised detection -- Performance analysis -- Image forensics
Expert systems (Computer science) -- Periodicals
Systèmes experts (Informatique) -- Périodiques
Electronic journals
006.33 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09574174 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.eswa.2019.05.002 ↗
- Languages:
- English
- ISSNs:
- 0957-4174
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3842.004220
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 11148.xml