Real-time RDF extraction from unstructured data streams

Research output: Contributions to collected editions/worksArticle in conference proceedingsResearchpeer-review

Authors

  • Daniel Gerber
  • Sebastian Hellmann
  • Lorenz Bühmann
  • Tommaso Soru
  • Ricardo Usbeck
  • Axel Cyrille Ngonga Ngomo

The vision behind the Web of Data is to extend the current document-oriented Web with machine-readable facts and structured data, thus creating a representation of general knowledge. However, most of the Web of Data is limited to being a large compendium of encyclopedic knowledge describing entities. A huge challenge, the timely and massive extraction of RDF facts from unstructured data, has remained open so far. The availability of such knowledge on the Web of Data would provide significant benefits to manifold applications including news retrieval, sentiment analysis and business intelligence. In this paper, we address the problem of the actuality of the Web of Data by presenting an approach that allows extracting RDF triples from unstructured data streams. We employ statistical methods in combination with deduplication, disambiguation and unsupervised as well as supervised machine learning techniques to create a knowledge base that reflects the content of the input streams. We evaluate a sample of the RDF we generate against a large corpus of news streams and show that we achieve a precision of more than 85%.

Original languageEnglish
Title of host publicationThe Semantic Web, ISWC 2013 : 12th International Semantic Web Conference, Proceedings
EditorsHarith Alani, Lalana Kagal, Achille Fokoue, Paul Groth, Chris Biemann, Josiane Xavier Parreira, Lora Aroyo, Natasha Noy, Chris Welty, Krzyztof Janowicz
Number of pages16
PublisherSpringer Verlag
Publication date2013
Pages135-150
ISBN (print)9783642413346
DOIs
Publication statusPublished - 2013
Externally publishedYes
Event12th International Semantic Web Conference, ISWC 2013 - Sydney Convention Centre , Sydney, NSW, Australia
Duration: 21.10.201325.10.2013
http://iswc2013.semanticweb.org

Recently viewed

Publications

  1. Simple saturated relay non-linear PD control for uncertain motion systems with friction and actuator constraint
  2. Fast, Fully Automated Analysis of Voriconazole from Serum by LC-LC-ESI-MS-MS with Parallel Column-Switching Technique
  3. A geometric approach for controlling an electromagnetic actuator with the help of a linear Model Predictive Control
  4. Toward Application and Implementation of in Silico Tools and Workflows within Benign by Design Approaches
  5. Using learning protocols for knowledge acquisition and problem solving with individual and group incentives
  6. Accounting and Modeling as Design Metaphors for CEMIS
  7. Universal Threshold Calculation for Fingerprinting Decoders using Mixture Models
  8. Using complexity metrics with R-R intervals and BPM heart rate measures
  9. Recurrence quantificationanalysis as a general-purpose tool for bridging the gap between qualitative and quantitative analysis
  10. Understanding reading as a form of language-use
  11. An extended analytical approach to evaluating monotonic functions of fuzzy numbers
  12. FaST: A linear time stack trace alignment heuristic for crash report deduplication
  13. Constrained Independence for Detecting Interesting Patterns
  14. A localized boundary element method for the floating body problem
  15. Multidimensional recurrence quantification analysis (MdRQA) for the analysis of multidimensional time-series
  16. A Quadrant Approach of Camera Calibration Method for Depth Estimation Using a Stereo Vision System
  17. Mapping interest rate projections using neural networks under cointegration
  18. Towards a Global Script?
  19. What does it mean to be sensitive for the complexity of (problem oriented) teaching?
  20. The Influence of Note-taking on Mathematical Solution Processes while Working on Reality-Based Tasks
  21. Microstructural development of as-cast AM50 during Constrained Friction Processing: grain refinement and influence of process parameters
  22. Distributed robust Gaussian Process regression
  23. Gain Scheduling Controller for Improving Level Control Performance
  24. Machine Learning and Knowledge Discovery in Databases
  25. Paraphrasing Method for Controlling a Robotic Arm Using a Large Language Model
  26. Problem solving in mathematics education