Real-time RDF extraction from unstructured data streams

Publikation: Beiträge in SammelwerkenAufsätze in KonferenzbändenForschungbegutachtet

Authors

  • Daniel Gerber
  • Sebastian Hellmann
  • Lorenz Bühmann
  • Tommaso Soru
  • Ricardo Usbeck
  • Axel Cyrille Ngonga Ngomo

The vision behind the Web of Data is to extend the current document-oriented Web with machine-readable facts and structured data, thus creating a representation of general knowledge. However, most of the Web of Data is limited to being a large compendium of encyclopedic knowledge describing entities. A huge challenge, the timely and massive extraction of RDF facts from unstructured data, has remained open so far. The availability of such knowledge on the Web of Data would provide significant benefits to manifold applications including news retrieval, sentiment analysis and business intelligence. In this paper, we address the problem of the actuality of the Web of Data by presenting an approach that allows extracting RDF triples from unstructured data streams. We employ statistical methods in combination with deduplication, disambiguation and unsupervised as well as supervised machine learning techniques to create a knowledge base that reflects the content of the input streams. We evaluate a sample of the RDF we generate against a large corpus of news streams and show that we achieve a precision of more than 85%.

OriginalspracheEnglisch
TitelThe Semantic Web, ISWC 2013 : 12th International Semantic Web Conference, Proceedings
HerausgeberHarith Alani, Lalana Kagal, Achille Fokoue, Paul Groth, Chris Biemann, Josiane Xavier Parreira, Lora Aroyo, Natasha Noy, Chris Welty, Krzyztof Janowicz
Anzahl der Seiten16
VerlagSpringer Verlag
Erscheinungsdatum2013
Seiten135-150
ISBN (Print)9783642413346
DOIs
PublikationsstatusErschienen - 2013
Extern publiziertJa
Veranstaltung12th International Semantic Web Conference, ISWC 2013 - Sydney Convention Centre , Sydney, NSW, Australien
Dauer: 21.10.201325.10.2013
http://iswc2013.semanticweb.org

DOI

Zuletzt angesehen

Publikationen

  1. “Ideation is Fine, but Execution is Key”
  2. Supporting the Development and Realization of Data-Driven Business Models with Enterprise Architecture Modeling and Management
  3. Considerations on efficient touch interfaces - How display size influences the performance in an applied pointing task
  4. A new way of assessing the interaction of a metallic phase precursor with a modified oxide support substrate as a source of information for predicting metal dispersion
  5. Computing regression statistics from grouped data
  6. Foundations and applications of computer based material flow networks for einvironmental management
  7. Mapping interest rate projections using neural networks under cointegration
  8. Partitioned beta diversity patterns of plants across sharp and distinct boundaries of quartz habitat islands
  9. Analysis of PI controllers with anti-windup techniques on level systems
  10. Using Fuzzy PD Controllers for Soft Motions in a Car-like Robot
  11. An expert-based reference list of variables for characterizing and monitoring social-ecological systems
  12. The fuzzy relationship of intelligence and problem solving in computer simulations
  13. Neural network-based estimation and compensation of friction for enhanced deep drawing process control
  14. Resolving the Complexity-Flexibility Dilemma in Multi-Issue Negotiations: Nested Bracketing as a Strategy to Enhance Negotiation Outcomes
  15. Changes of Perception
  16. Self-regulation in error management training: emotion control and metacognition as mediators of performance effects
  17. Resource extraction technologies - is a more responsible path of development possible?
  18. GENESIS - A generic RDF data access interface
  19. In-Vehicle Sensor System for Monitoring Efficiency of Vehicle E/E Architectures
  20. Semantic Evaluation Services for Web-Based Exercises
  21. Emergency detection based on probabilistic modeling in AAL-environments
  22. Functional Richness and Relative Resilience of Bird Communities in Regions with Different Land Use Intensities
  23. Dimension estimates for certain sets of infinite complex continued fractions
  24. The effects of different on-line adaptive response time limits on speed and amount of learning in computer assisted instruction and intelligent tutoring
  25. Effectiveness of a Web-Based Cognitive Behavioural Intervention for Subthreshold Depression
  26. Understanding Low-Code Evolution, Adoption and Ecosystem for Software Development
  27. Towards Advanced Learning in Dispatching Rule-Based Scheuling
  28. Offline question answering over linked data using limited resources
  29. The professional context as a predictor for response distortion in the Adaption-Innovation-Inventory – An investigation using mixture-distribution item-response theory models
  30. Biodegradation screening of chemicals in an artificial matrix simulating the water-sediment interface
  31. Loss systems in a random environment: steady state analysis