Sampling strategies for information extraction over the deep web. Issue 2 (March 2017)
- Record Type:
- Journal Article
- Title:
- Sampling strategies for information extraction over the deep web. Issue 2 (March 2017)
- Main Title:
- Sampling strategies for information extraction over the deep web
- Authors:
- Barrio, Pablo
Gravano, Luis - Abstract:
- Highlights: First large-scale and fine-grained evaluation of query-based sampling techniques. Learned keyword queries perform substantially better than queries derived from tuples. Focusing on—and processing exhaustively—effective queries leads to high efficiency. Focusing on—and processing in rounds—less-effective queries favors quality. Filtering underperforming queries favors sampling efficiency but hurts quality. Abstract: Information extraction systems discover structured information in natural language text. Having information in structured form enables much richer querying and data mining than possible over the natural language text. However, information extraction is a computationally expensive task, and hence improving the efficiency of the extraction process over large text collections is of critical interest. In this paper, we focus on an especially valuable family of text collections, namely, the so-called deep-web text collections, whose contents are not crawlable and are only available via querying. Important steps for efficient information extraction over deep-web text collections (e.g., selecting the collections on which to focus the extraction effort, based on their contents; or learning which documents within these collections—and in which order—to process, based on their words and phrases) require having a representative document sample from each collection. These document samples have to be collected by querying the deep-web text collections, an expensiveHighlights: First large-scale and fine-grained evaluation of query-based sampling techniques. Learned keyword queries perform substantially better than queries derived from tuples. Focusing on—and processing exhaustively—effective queries leads to high efficiency. Focusing on—and processing in rounds—less-effective queries favors quality. Filtering underperforming queries favors sampling efficiency but hurts quality. Abstract: Information extraction systems discover structured information in natural language text. Having information in structured form enables much richer querying and data mining than possible over the natural language text. However, information extraction is a computationally expensive task, and hence improving the efficiency of the extraction process over large text collections is of critical interest. In this paper, we focus on an especially valuable family of text collections, namely, the so-called deep-web text collections, whose contents are not crawlable and are only available via querying. Important steps for efficient information extraction over deep-web text collections (e.g., selecting the collections on which to focus the extraction effort, based on their contents; or learning which documents within these collections—and in which order—to process, based on their words and phrases) require having a representative document sample from each collection. These document samples have to be collected by querying the deep-web text collections, an expensive process that renders impractical the existing sampling approaches developed for other data scenarios. In this paper, we systematically study the space of query-based document sampling techniques for information extraction over the deep web. Specifically, we consider (i) alternative query execution schedules, which vary on how they account for the query effectiveness, and (ii) alternative document retrieval and processing schedules, which vary on how they distribute the extraction effort over documents. We report the results of the first large-scale experimental evaluation of sampling techniques for information extraction over the deep web. Our results show the merits and limitations of the alternative query execution and document retrieval and processing strategies, and provide a roadmap for addressing this critically important building block for efficient, scalable information extraction. … (more)
- Is Part Of:
- Information processing & management. Volume 53:Issue 2(2017:Mar.)
- Journal:
- Information processing & management
- Issue:
- Volume 53:Issue 2(2017:Mar.)
- Issue Display:
- Volume 53, Issue 2 (2017)
- Year:
- 2017
- Volume:
- 53
- Issue:
- 2
- Issue Sort Value:
- 2017-0053-0002-0000
- Page Start:
- 309
- Page End:
- 331
- Publication Date:
- 2017-03
- Subjects:
- Information extraction -- Sampling -- Deep web -- Text mining -- Scalability
Information storage and retrieval systems -- Periodicals
Information science -- Periodicals
Systèmes d'information -- Périodiques
Sciences de l'information -- Périodiques
Information science
Information storage and retrieval systems
Periodicals
658.4038 - Journal URLs:
- http://www.sciencedirect.com/science/journal/03064573 ↗
http://www.elsevier.com/journals ↗ - DOI:
- 10.1016/j.ipm.2016.11.006 ↗
- Languages:
- English
- ISSNs:
- 0306-4573
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - 4493.893000
British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 639.xml