Correlating mammographic and pathologic findings in clinical decision support using natural language processing and data mining methods. Issue 1 (29th August 2016)
- Record Type:
- Journal Article
- Title:
- Correlating mammographic and pathologic findings in clinical decision support using natural language processing and data mining methods. Issue 1 (29th August 2016)
- Main Title:
- Correlating mammographic and pathologic findings in clinical decision support using natural language processing and data mining methods
- Authors:
- Patel, Tejal A.
Puppala, Mamta
Ogunti, Richard O.
Ensor, Joe E.
He, Tiancheng
Shewale, Jitesh B.
Ankerst, Donna P.
Kaklamani, Virginia G.
Rodriguez, Angel A.
Wong, Stephen T. C.
Chang, Jenny C. - Abstract:
- Abstract : BACKGROUND: A key challenge to mining electronic health records for mammography research is the preponderance of unstructured narrative text, which strikingly limits usable output. The imaging characteristics of breast cancer subtypes have been described previously, but without standardization of parameters for data mining. METHODS: The authors searched the enterprise‐wide data warehouse at the Houston Methodist Hospital, the Methodist Environment for Translational Enhancement and Outcomes Research (METEOR), for patients with Breast Imaging Reporting and Data System (BI‐RADS) category 5 mammogram readings performed between January 2006 and May 2015 and an available pathology report. The authors developed natural language processing (NLP) software algorithms to automatically extract mammographic and pathologic findings from free text mammogram and pathology reports. The correlation between mammographic imaging features and breast cancer subtype was analyzed using one‐way analysis of variance and the Fisher exact test. RESULTS: The NLP algorithm was able to obtain key characteristics for 543 patients who met the inclusion criteria. Patients with estrogen receptor‐positive tumors were more likely to have spiculated margins ( P = .0008), and those with tumors that overexpressed human epidermal growth factor receptor 2 (HER2) were more likely to have heterogeneous and pleomorphic calcifications ( P = .0078 and P = .0002, respectively). CONCLUSIONS: MammographicAbstract : BACKGROUND: A key challenge to mining electronic health records for mammography research is the preponderance of unstructured narrative text, which strikingly limits usable output. The imaging characteristics of breast cancer subtypes have been described previously, but without standardization of parameters for data mining. METHODS: The authors searched the enterprise‐wide data warehouse at the Houston Methodist Hospital, the Methodist Environment for Translational Enhancement and Outcomes Research (METEOR), for patients with Breast Imaging Reporting and Data System (BI‐RADS) category 5 mammogram readings performed between January 2006 and May 2015 and an available pathology report. The authors developed natural language processing (NLP) software algorithms to automatically extract mammographic and pathologic findings from free text mammogram and pathology reports. The correlation between mammographic imaging features and breast cancer subtype was analyzed using one‐way analysis of variance and the Fisher exact test. RESULTS: The NLP algorithm was able to obtain key characteristics for 543 patients who met the inclusion criteria. Patients with estrogen receptor‐positive tumors were more likely to have spiculated margins ( P = .0008), and those with tumors that overexpressed human epidermal growth factor receptor 2 (HER2) were more likely to have heterogeneous and pleomorphic calcifications ( P = .0078 and P = .0002, respectively). CONCLUSIONS: Mammographic imaging characteristics, obtained from an automated text search and the extraction of mammogram reports using NLP techniques, correlated with pathologic breast cancer subtype. The results of the current study validate previously reported trends assessed by manual data collection. Furthermore, NLP provides an automated means with which to scale up data extraction and analysis for clinical decision support. Cancer 2017;114–121. © 2016 American Cancer Society. Abstract : Mammography research is inherently limited due to the time and expense required to manually extract data from unstructured free text mammogram reports. The authors have developed a natural language processing system that accurately extracts mammographic findings from free text reports. In the current study, they use this novel natural language processing system and statistical data mining to demonstrate the relationship between mammographic imaging characteristics and breast cancer subtype. … (more)
- Is Part Of:
- Cancer. Volume 123:Issue 1(2017)
- Journal:
- Cancer
- Issue:
- Volume 123:Issue 1(2017)
- Issue Display:
- Volume 123, Issue 1 (2017)
- Year:
- 2017
- Volume:
- 123
- Issue:
- 1
- Issue Sort Value:
- 2017-0123-0001-0000
- Page Start:
- 114
- Page End:
- 121
- Publication Date:
- 2016-08-29
- Subjects:
- data mining -- imaging characteristics -- mammographic to pathologic correlation -- natural language processing -- subtypes of breast cancer
Cancer -- Periodicals
Cancer -- Cytopathology -- Periodicals
616.99405 - Journal URLs:
- http://onlinelibrary.wiley.com/journal/10.1002/(ISSN)1097-0142 ↗
http://onlinelibrary.wiley.com/ ↗ - DOI:
- 10.1002/cncr.30245 ↗
- Languages:
- English
- ISSNs:
- 0008-543X
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3046.450000
British Library DSC - BLDSS-3PM
British Library STI - ELD Digital store - Ingest File:
- 1637.xml