Development of an information retrieval tool for biomedical patents. (June 2018)
- Record Type:
- Journal Article
- Title:
- Development of an information retrieval tool for biomedical patents. (June 2018)
- Main Title:
- Development of an information retrieval tool for biomedical patents
- Authors:
- Alves, Tiago
Rodrigues, Rúben
Costa, Hugo
Rocha, Miguel - Abstract:
- Highlights: The paper proposes a patent information retrieval pipeline, able to search and retrieve patent data from relevant databases. The proposed pipeline can recover both metadata and full texts associated to patents. The patent pipeline is integrated within @Note2, an open-source user- friendly software framework for biomedical text mining, in the form of a new plug-in. Integration with @Note allows a number of text mining tools to be applied to abstracts and full texts of patents Two case studies were used to demonstrate the pipeline's main features and show its potential uses. Abstract: Background and objective : The volume of biomedical literature has been increasing in the last years. Patent documents have also followed this trend, being important sources of biomedical knowledge, technical details and curated data, which are put together along the granting process. The field of Biomedical text mining (BioTM) has been creating solutions for the problems posed by the unstructured nature of natural language, which makes the search of information a challenging task. Several BioTM techniques can be applied to patents. From those, Information Retrieval (IR) includes processes where relevant data are obtained from collections of documents. In this work, the main goal was to build a patent pipeline addressing IR tasks over patent repositories to make these documents amenable to BioTM tasks. Methods : The pipeline was developed within @Note2, an open-source computationalHighlights: The paper proposes a patent information retrieval pipeline, able to search and retrieve patent data from relevant databases. The proposed pipeline can recover both metadata and full texts associated to patents. The patent pipeline is integrated within @Note2, an open-source user- friendly software framework for biomedical text mining, in the form of a new plug-in. Integration with @Note allows a number of text mining tools to be applied to abstracts and full texts of patents Two case studies were used to demonstrate the pipeline's main features and show its potential uses. Abstract: Background and objective : The volume of biomedical literature has been increasing in the last years. Patent documents have also followed this trend, being important sources of biomedical knowledge, technical details and curated data, which are put together along the granting process. The field of Biomedical text mining (BioTM) has been creating solutions for the problems posed by the unstructured nature of natural language, which makes the search of information a challenging task. Several BioTM techniques can be applied to patents. From those, Information Retrieval (IR) includes processes where relevant data are obtained from collections of documents. In this work, the main goal was to build a patent pipeline addressing IR tasks over patent repositories to make these documents amenable to BioTM tasks. Methods : The pipeline was developed within @Note2, an open-source computational framework for BioTM, adding a number of modules to the core libraries, including patent metadata and full text retrieval, PDF to text conversion and optical character recognition. Also, user interfaces were developed for the main operations materialized in a new @Note2 plug-in. Results : The integration of these tools in @Note2 opens opportunities to run BioTM tools over patent texts, including tasks from Information Extraction, such as Named Entity Recognition or Relation Extraction. We demonstrated the pipeline's main functions with a case study, using an available benchmark dataset from BioCreative challenges. Also, we show the use of the plug-in with a user query related to the production of vanillin. Conclusions : This work makes available all the relevant content from patents to the scientific community, decreasing drastically the time required for this task, and provides graphical interfaces to ease the use of these tools. … (more)
- Is Part Of:
- Computer methods and programs in biomedicine. Volume 159(2018)
- Journal:
- Computer methods and programs in biomedicine
- Issue:
- Volume 159(2018)
- Issue Display:
- Volume 159, Issue 2018 (2018)
- Year:
- 2018
- Volume:
- 159
- Issue:
- 2018
- Issue Sort Value:
- 2018-0159-2018-0000
- Page Start:
- 125
- Page End:
- 134
- Publication Date:
- 2018-06
- Subjects:
- Biomedical text mining -- Information retrieval -- Information extraction -- Patents -- PDF to text conversion
Medicine -- Computer programs -- Periodicals
Biology -- Computer programs -- Periodicals
Computers -- Periodicals
Medicine -- Periodicals
Médecine -- Logiciels -- Périodiques
Biologie -- Logiciels -- Périodiques
Biology -- Computer programs
Medicine -- Computer programs
Periodicals
Electronic journals
610.28 - Journal URLs:
- http://www.sciencedirect.com/science/journal/01692607 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.cmpb.2018.03.012 ↗
- Languages:
- English
- ISSNs:
- 0169-2607
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 3394.095000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 6300.xml