Real-time RDF extraction from unstructured data streams

Publikation: Beiträge in SammelwerkenAufsätze in KonferenzbändenForschungbegutachtet

Authors

  • Daniel Gerber
  • Sebastian Hellmann
  • Lorenz Bühmann
  • Tommaso Soru
  • Ricardo Usbeck
  • Axel Cyrille Ngonga Ngomo

The vision behind the Web of Data is to extend the current document-oriented Web with machine-readable facts and structured data, thus creating a representation of general knowledge. However, most of the Web of Data is limited to being a large compendium of encyclopedic knowledge describing entities. A huge challenge, the timely and massive extraction of RDF facts from unstructured data, has remained open so far. The availability of such knowledge on the Web of Data would provide significant benefits to manifold applications including news retrieval, sentiment analysis and business intelligence. In this paper, we address the problem of the actuality of the Web of Data by presenting an approach that allows extracting RDF triples from unstructured data streams. We employ statistical methods in combination with deduplication, disambiguation and unsupervised as well as supervised machine learning techniques to create a knowledge base that reflects the content of the input streams. We evaluate a sample of the RDF we generate against a large corpus of news streams and show that we achieve a precision of more than 85%.

OriginalspracheEnglisch
TitelThe Semantic Web, ISWC 2013 : 12th International Semantic Web Conference, Proceedings
HerausgeberHarith Alani, Lalana Kagal, Achille Fokoue, Paul Groth, Chris Biemann, Josiane Xavier Parreira, Lora Aroyo, Natasha Noy, Chris Welty, Krzyztof Janowicz
Anzahl der Seiten16
VerlagSpringer Verlag
Erscheinungsdatum2013
Seiten135-150
ISBN (Print)9783642413346
DOIs
PublikationsstatusErschienen - 2013
Extern publiziertJa
Veranstaltung12th International Semantic Web Conference, ISWC 2013 - Sydney Convention Centre , Sydney, NSW, Australien
Dauer: 21.10.201325.10.2013
http://iswc2013.semanticweb.org

DOI

Zuletzt angesehen

Publikationen

  1. What does it mean to be sensitive for the complexity of (problem oriented) teaching?
  2. Combining a PI Controller with an Adaptive Feedforward Control in PMSM
  3. Improving students’ science text comprehension through metacognitive self-regulation when applying learning strategies
  4. “Ideation is Fine, but Execution is Key”
  5. Age effects on controlling tools with sensorimotor transformations
  6. A new way of assessing the interaction of a metallic phase precursor with a modified oxide support substrate as a source of information for predicting metal dispersion
  7. Computing regression statistics from grouped data
  8. Performance analysis for loss systems with many subscribers and concurrent services
  9. Stimulating Computing
  10. Explaining and controlling for the psychometric properties of computer-generated figural matrix items
  11. Scaffolding argumentation in mathematics with CSCL scripts
  12. Foundations and applications of computer based material flow networks for einvironmental management
  13. A localized boundary element method for the floating body problem
  14. Robust feedback linearization control of a throttle plate by using an approximated pd regulator
  15. TARGET SETTING FOR OPERATIONAL PERFORMANCE IMPROVEMENTS - STUDY CASE -
  16. Integration of laser scanning and projection speckle pattern for advanced pipeline monitoring
  17. Partitioned beta diversity patterns of plants across sharp and distinct boundaries of quartz habitat islands
  18. Computer als Medium
  19. OKBQA framework towards an open collaboration for development of natural language question-answering systems over knowledge bases
  20. Learning from Erroneous Examples: When and How do Students Benefit from them?
  21. Analysis of PI controllers with anti-windup techniques on level systems
  22. An Adaptive and Optimized Switching Observer for Sensorless Control of an Electromagnetic Valve Actuator in Camless Internal Combustion Engines
  23. Gaussian processes for dispatching rule selection in production scheduling
  24. Learning Analytics with Matlab Grader in Undergraduate Engineering Courses
  25. TRY plant trait database – enhanced coverage and open access
  26. An evaluation of BPR methodologies adopting NIMSAD: A systematic framework for understanding and evaluating methodologies
  27. An expert-based reference list of variables for characterizing and monitoring social-ecological systems
  28. Practical guide to SAP Netweaver PI-development
  29. Two models for gradient inelasticity based on non-convex energy
  30. Modelling and implementation of an Order2Cash Process in distributed systems
  31. Preventive Diagnostics for cardiovascular diseases based on probabilistic methods and description logic
  32. An Orthogonal Wavelet Denoising Algorithm for Surface Images of Atomic Force Microscopy
  33. A Multilevel Inverter Bridge Control Structure with Energy Storage Using Model Predictive Control for Flat Systems
  34. Mirrored piezo servo hydraulic actuators for use in camless combustion engines and its Control with mirrored inputs and MPC
  35. Data-driven and physics-based modelling of process behaviour and deposit geometry for friction surfacing