Real-time RDF extraction from unstructured data streams

Research output: Contributions to collected editions/worksArticle in conference proceedingsResearchpeer-review

Authors

  • Daniel Gerber
  • Sebastian Hellmann
  • Lorenz Bühmann
  • Tommaso Soru
  • Ricardo Usbeck
  • Axel Cyrille Ngonga Ngomo

The vision behind the Web of Data is to extend the current document-oriented Web with machine-readable facts and structured data, thus creating a representation of general knowledge. However, most of the Web of Data is limited to being a large compendium of encyclopedic knowledge describing entities. A huge challenge, the timely and massive extraction of RDF facts from unstructured data, has remained open so far. The availability of such knowledge on the Web of Data would provide significant benefits to manifold applications including news retrieval, sentiment analysis and business intelligence. In this paper, we address the problem of the actuality of the Web of Data by presenting an approach that allows extracting RDF triples from unstructured data streams. We employ statistical methods in combination with deduplication, disambiguation and unsupervised as well as supervised machine learning techniques to create a knowledge base that reflects the content of the input streams. We evaluate a sample of the RDF we generate against a large corpus of news streams and show that we achieve a precision of more than 85%.

Original languageEnglish
Title of host publicationThe Semantic Web, ISWC 2013 : 12th International Semantic Web Conference, Proceedings
EditorsHarith Alani, Lalana Kagal, Achille Fokoue, Paul Groth, Chris Biemann, Josiane Xavier Parreira, Lora Aroyo, Natasha Noy, Chris Welty, Krzyztof Janowicz
Number of pages16
PublisherSpringer Verlag
Publication date2013
Pages135-150
ISBN (print)9783642413346
DOIs
Publication statusPublished - 2013
Externally publishedYes
Event12th International Semantic Web Conference, ISWC 2013 - Sydney Convention Centre , Sydney, NSW, Australia
Duration: 21.10.201325.10.2013
http://iswc2013.semanticweb.org

Recently viewed

Publications

  1. Automatic Error Detection in Gaussian Processes Regression Modeling for Production Scheduling
  2. A Wavelet Based Algorithm without a Priori Knowledge of Noise Level for Gross Errors Detection
  3. Soft Optimal Computing Methods to Identify Surface Roughness in Manufacturing Using a Monotonic Regressor
  4. Modeling items for text comprehension assessment using confirmatory factor analysis
  5. Validation of an open source, remote web-based eye-tracking method (WebGazer) for research in early childhood
  6. A Column Generation Approach for Bus Driver Rostering Problems
  7. Vielfalt des Alterns - Differenz oder Integration?
  8. Towards productive functions?
  9. Homogenization methods for multi-phase elastic composites with non-elliptical reinforcements
  10. Activity–rest schedules in physically demanding work and the variation of responses with age
  11. Digital Business Transformation and the Changing Role of the IT Function
  12. High-precision frequency measurements: indispensable tools at the core of the molecular-level analysis of complex systems.
  13. Determinants in the online distribution of digital content
  14. Covert and overt automatic imitation are correlated
  15. Mathematical Modelling of molecular adsorption in zeolite coated frequency domain sensors
  16. An antisaturating adaptive preaction and a slide surface to achieve soft landing control for electromagnetic actuators
  17. Archives
  18. Ionic liquids vs. ethanol as extraction media of algicidal compounds from mango processing waste
  19. To Row Together or Paddle One's Own Canoe? Simulating Strategies to Spur Digital Platform Growth
  20. Knowledge Production in Consulting Teams: A Self-Organization Approach
  21. Landscape moderation of biodiversity patterns and processes - eight hypotheses
  22. The Case of Willetta Huggins