ConvPatchTrans: A script identification network with global and local semantics deeply integrated. (August 2022)
- Record Type:
- Journal Article
- Title:
- ConvPatchTrans: A script identification network with global and local semantics deeply integrated. (August 2022)
- Main Title:
- ConvPatchTrans: A script identification network with global and local semantics deeply integrated
- Authors:
- Yang, Ke
Yi, Jizheng
Chen, Aibin
Liu, Jiaqi
Chen, Wenjie
Jin, Ze - Abstract:
- Abstract: Optical Character Recognition (OCR) system serves the need of reading text from images. Script identification that identifies the language of the text in the image is an important part of OCR technology and an indispensable role in the stability and accuracy of the OCR system. The most challenging for script identification is the interference caused by similarities between texts in different languages. In this paper, a two-branch network named ConvPatchTrans is designed to process global and local semantic features separately, focusing on the text and each word in a picture. The ConvPatchTrans extracts feature from different stages of the Visual Geometry Group network (VGGNet) as global and local semantics. For the global branch, the linear classifier is recommended. For the local branch, text image data is converted to image sequence data. Then, multi-layers convolution-enhanced Transformer (MCET) is proposed to bring about the deep fusion of sequence. Finally, the global and local branches are fused by an adaptive weighted fusion method to get the best result. In order to verify the effectiveness of our proposed method, four public script identification datasets are used for comparative experiments. Our method has obtained the highest values among currently published methods on the CVSI2015 and MLE2E datasets, which are 98.90% and 97.50%, respectively. At the same time, satisfactory results are also obtained on the other two datasets. Highlights: A scene textAbstract: Optical Character Recognition (OCR) system serves the need of reading text from images. Script identification that identifies the language of the text in the image is an important part of OCR technology and an indispensable role in the stability and accuracy of the OCR system. The most challenging for script identification is the interference caused by similarities between texts in different languages. In this paper, a two-branch network named ConvPatchTrans is designed to process global and local semantic features separately, focusing on the text and each word in a picture. The ConvPatchTrans extracts feature from different stages of the Visual Geometry Group network (VGGNet) as global and local semantics. For the global branch, the linear classifier is recommended. For the local branch, text image data is converted to image sequence data. Then, multi-layers convolution-enhanced Transformer (MCET) is proposed to bring about the deep fusion of sequence. Finally, the global and local branches are fused by an adaptive weighted fusion method to get the best result. In order to verify the effectiveness of our proposed method, four public script identification datasets are used for comparative experiments. Our method has obtained the highest values among currently published methods on the CVSI2015 and MLE2E datasets, which are 98.90% and 97.50%, respectively. At the same time, satisfactory results are also obtained on the other two datasets. Highlights: A scene text script identification network with the deep fusion of global and local semantic features. Obtaining global and local semantic information from different stages of VGGNet loaded with pre-trained parameters. Multi-layer convolutional enhancement Transformer for local semantics. Decision fusion with adaptive weights. … (more)
- Is Part Of:
- Engineering applications of artificial intelligence. Volume 113(2022)
- Journal:
- Engineering applications of artificial intelligence
- Issue:
- Volume 113(2022)
- Issue Display:
- Volume 113, Issue 2022 (2022)
- Year:
- 2022
- Volume:
- 113
- Issue:
- 2022
- Issue Sort Value:
- 2022-0113-2022-0000
- Page Start:
- Page End:
- Publication Date:
- 2022-08
- Subjects:
- Script identification -- Dual network -- Transformer -- Feature fusion -- Deeply integrated -- OCR (Optical Character Recognition)
Engineering -- Data processing -- Periodicals
Artificial intelligence -- Periodicals
Expert systems (Computer science) -- Periodicals
Ingénierie -- Informatique -- Périodiques
Intelligence artificielle -- Périodiques
Systèmes experts (Informatique) -- Périodiques
Artificial intelligence
Engineering -- Data processing
Expert systems (Computer science)
Periodicals
620.00285 - Journal URLs:
- http://www.sciencedirect.com/science/journal/09521976 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.engappai.2022.104916 ↗
- Languages:
- English
- ISSNs:
- 0952-1976
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3755.704500
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 21869.xml