Deep learning-based NLP data pipeline for EHR-scanned document information extraction. Issue 2 (11th June 2022)
- Record Type:
- Journal Article
- Title:
- Deep learning-based NLP data pipeline for EHR-scanned document information extraction. Issue 2 (11th June 2022)
- Main Title:
- Deep learning-based NLP data pipeline for EHR-scanned document information extraction
- Authors:
- Hsu, Enshuo
Malagaris, Ioannis
Kuo, Yong-Fang
Sultana, Rizwana
Roberts, Kirk - Abstract:
- Abstract: Objective: Scanned documents in electronic health records (EHR) have been a challenge for decades, and are expected to stay in the foreseeable future. Current approaches for processing include image preprocessing, optical character recognition (OCR), and natural language processing (NLP). However, there is limited work evaluating the interaction of image preprocessing methods, NLP models, and document layout. Materials and Methods: We evaluated 2 key indicators for sleep apnea, Apnea hypopnea index (AHI) and oxygen saturation (SaO2 ), from 955 scanned sleep study reports. Image preprocessing methods include gray-scaling, dilating, eroding, and contrast. OCR was implemented with Tesseract. Seven traditional machine learning models and 3 deep learning models were evaluated. We also evaluated combinations of image preprocessing methods, and 2 deep learning architectures (with and without structured input providing document layout information), with the goal of optimizing end-to-end performance. Results: Our proposed method using ClinicalBERT reached an AUROC of 0.9743 and document accuracy of 94.76% for AHI, and an AUROC of 0.9523 and document accuracy of 91.61% for SaO2 . Discussion: There are multiple, inter-related steps to extract meaningful information from scanned reports. While it would be infeasible to experiment with all possible option combinations, we experimented with several of the most critical steps for information extraction, including image processingAbstract: Objective: Scanned documents in electronic health records (EHR) have been a challenge for decades, and are expected to stay in the foreseeable future. Current approaches for processing include image preprocessing, optical character recognition (OCR), and natural language processing (NLP). However, there is limited work evaluating the interaction of image preprocessing methods, NLP models, and document layout. Materials and Methods: We evaluated 2 key indicators for sleep apnea, Apnea hypopnea index (AHI) and oxygen saturation (SaO2 ), from 955 scanned sleep study reports. Image preprocessing methods include gray-scaling, dilating, eroding, and contrast. OCR was implemented with Tesseract. Seven traditional machine learning models and 3 deep learning models were evaluated. We also evaluated combinations of image preprocessing methods, and 2 deep learning architectures (with and without structured input providing document layout information), with the goal of optimizing end-to-end performance. Results: Our proposed method using ClinicalBERT reached an AUROC of 0.9743 and document accuracy of 94.76% for AHI, and an AUROC of 0.9523 and document accuracy of 91.61% for SaO2 . Discussion: There are multiple, inter-related steps to extract meaningful information from scanned reports. While it would be infeasible to experiment with all possible option combinations, we experimented with several of the most critical steps for information extraction, including image processing and NLP. Given that scanned documents will likely be part of healthcare for years to come, it is critical to develop NLP systems to extract key information from this data. Conclusion: We demonstrated the proper use of image preprocessing and document layout could be beneficial to scanned document processing. Lay Summary: Electronic health records frequently contain scanned documents, usually the result of faxed reports from other providers. Scanned documents are a challenge to processing, being images and not text, yet they frequently contain important information. Automatically extracting information from scanned documents, therefore, not only requires normal natural language processing (NLP) methods, but also additional steps such as optical character recognition (OCR) that converts images to text. Given that scanned documents will likely be part of healthcare for years to come, it is critical to develop NLP systems to extract key information from this data. This paper evaluates a battery of methods for extracting information from sleep study reports, though the methods should generalize to many other clinical NLP tasks related to scanned documents. Specifically, we focus on 2 key indicators for sleep apnea, Apnea hypopnea index (AHI) and oxygen saturation (SaO2 ). We experiment with several image preprocessing methods and 7 machine learning-based NLP models. Our best-performing method achieves a document-level accuracy of 94.8% for identifying AHI values and 91.6% for identifying SaO2 values. Overall, we demonstrate the proper use of image preprocessing and document layout could be beneficial to scanned document processing. … (more)
- Is Part Of:
- JAMIA open. Volume 5:Issue 2(2022)
- Journal:
- JAMIA open
- Issue:
- Volume 5:Issue 2(2022)
- Issue Display:
- Volume 5, Issue 2 (2022)
- Year:
- 2022
- Volume:
- 5
- Issue:
- 2
- Issue Sort Value:
- 2022-0005-0002-0000
- Page Start:
- Page End:
- Publication Date:
- 2022-06-11
- Subjects:
- scanned document -- optical character recognition -- natural language processing -- electronic health records -- polysomnography
Medical informatics -- Periodicals
610.285 - Journal URLs:
- http://www.oxfordjournals.org/ ↗
https://academic.oup.com/jamiaopen ↗ - DOI:
- 10.1093/jamiaopen/ooac045 ↗
- Languages:
- English
- ISSNs:
- 2574-2531
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 21808.xml