Corpus processing service: A Knowledge Graph platform to perform deep data exploration on corpora. Issue 2 (27th January 2021)
- Record Type:
- Journal Article
- Title:
- Corpus processing service: A Knowledge Graph platform to perform deep data exploration on corpora. Issue 2 (27th January 2021)
- Main Title:
- Corpus processing service: A Knowledge Graph platform to perform deep data exploration on corpora
- Authors:
- Staar, Peter W. J.
Dolfi, Michele
Auer, Christoph - Abstract:
- Abstract: Knowledge Graphs have been fast emerging as the de facto standard to model and explore knowledge in weakly structured data. Large corpora of documents constitute a source of weakly structured data of particular interest for both the academic and business world. Key examples include scientific publications, technical reports, manuals, patents, regulations, etc. Such corpora embed many facts that are elementary to critical decision making or enabling new discoveries. In this paper, we present a scalable cloud platform to create and serve Knowledge Graphs, which we named corpus processing service (CPS). Its purpose is to process large document corpora, extract the content and embedded facts, and ultimately represent these in a consistent knowledge graph that can be intuitively queried. To accomplish this, we use state‐of‐the‐art natural language understanding models to extract entities and relationships from documents converted with our previously presented corpus conversion service platform. This pipeline is complemented with a newly developed graph engine which ensures extremely performant graph queries and provides powerful graph analytics capabilities. Both components are tightly integrated and can be easily consumed through REST APIs. Additionally, we provide user interfaces to control the data ingestion flow and formulate queries using a visual programming approach. The CPS platform is designed as a modular microservice system operating on Kubernetes clusters.Abstract: Knowledge Graphs have been fast emerging as the de facto standard to model and explore knowledge in weakly structured data. Large corpora of documents constitute a source of weakly structured data of particular interest for both the academic and business world. Key examples include scientific publications, technical reports, manuals, patents, regulations, etc. Such corpora embed many facts that are elementary to critical decision making or enabling new discoveries. In this paper, we present a scalable cloud platform to create and serve Knowledge Graphs, which we named corpus processing service (CPS). Its purpose is to process large document corpora, extract the content and embedded facts, and ultimately represent these in a consistent knowledge graph that can be intuitively queried. To accomplish this, we use state‐of‐the‐art natural language understanding models to extract entities and relationships from documents converted with our previously presented corpus conversion service platform. This pipeline is complemented with a newly developed graph engine which ensures extremely performant graph queries and provides powerful graph analytics capabilities. Both components are tightly integrated and can be easily consumed through REST APIs. Additionally, we provide user interfaces to control the data ingestion flow and formulate queries using a visual programming approach. The CPS platform is designed as a modular microservice system operating on Kubernetes clusters. Finally, we validate the quality of queries on our end‐to‐end knowledge pipeline in a real‐world application in the oil and gas industry. … (more)
- Is Part Of:
- Applied AI Letters. Volume 1:Issue 2(2020)
- Journal:
- Applied AI Letters
- Issue:
- Volume 1:Issue 2(2020)
- Issue Display:
- Volume 1, Issue 2 (2020)
- Year:
- 2020
- Volume:
- 1
- Issue:
- 2
- Issue Sort Value:
- 2020-0001-0002-0000
- Page Start:
- n/a
- Page End:
- n/a
- Publication Date:
- 2021-01-27
- Subjects:
- document processing -- knowledge graph -- semantic search
006.3 - Journal URLs:
- http://onlinelibrary.wiley.com/ ↗
- DOI:
- 10.1002/ail2.20 ↗
- Languages:
- English
- ISSNs:
- 2689-5595
- Deposit Type:
- Legaldeposit
- View Content:
- Available online (eLD content is only available in our Reading Rooms) ↗
- Physical Locations:
- British Library DSC - BLDSS-3PM
British Library HMNTS - ELD Digital store - Ingest File:
- 15564.xml