Real-time RDF extraction from unstructured data streams

Research output: Contributions to collected editions/worksArticle in conference proceedingsResearchpeer-review

Authors

  • Daniel Gerber
  • Sebastian Hellmann
  • Lorenz Bühmann
  • Tommaso Soru
  • Ricardo Usbeck
  • Axel Cyrille Ngonga Ngomo

The vision behind the Web of Data is to extend the current document-oriented Web with machine-readable facts and structured data, thus creating a representation of general knowledge. However, most of the Web of Data is limited to being a large compendium of encyclopedic knowledge describing entities. A huge challenge, the timely and massive extraction of RDF facts from unstructured data, has remained open so far. The availability of such knowledge on the Web of Data would provide significant benefits to manifold applications including news retrieval, sentiment analysis and business intelligence. In this paper, we address the problem of the actuality of the Web of Data by presenting an approach that allows extracting RDF triples from unstructured data streams. We employ statistical methods in combination with deduplication, disambiguation and unsupervised as well as supervised machine learning techniques to create a knowledge base that reflects the content of the input streams. We evaluate a sample of the RDF we generate against a large corpus of news streams and show that we achieve a precision of more than 85%.

Original languageEnglish
Title of host publicationThe Semantic Web, ISWC 2013 : 12th International Semantic Web Conference, Proceedings
EditorsHarith Alani, Lalana Kagal, Achille Fokoue, Paul Groth, Chris Biemann, Josiane Xavier Parreira, Lora Aroyo, Natasha Noy, Chris Welty, Krzyztof Janowicz
Number of pages16
PublisherSpringer Verlag
Publication date2013
Pages135-150
ISBN (print)9783642413346
DOIs
Publication statusPublished - 2013
Externally publishedYes
Event12th International Semantic Web Conference, ISWC 2013 - Sydney Convention Centre , Sydney, NSW, Australia
Duration: 21.10.201325.10.2013
http://iswc2013.semanticweb.org

Recently viewed

Publications

  1. A Service-oriented Search framework for full text, geospatial and semantic search
  2. 7th open challenge on question answering over linked data (QALD-7)
  3. An expert-based reference list of variables for characterizing and monitoring social-ecological systems
  4. AGDISTIS - Graph-based disambiguation of named entities using linked data
  5. OKBQA framework towards an open collaboration for development of natural language question-answering systems over knowledge bases
  6. Holistic and scalable ranking of RDF data
  7. HAWK - hybrid question answering using linked data
  8. ASSESS — automatic self-assessment using linked data
  9. GENESIS - A generic RDF data access interface
  10. Treating dialogue quality evaluation as an anomaly detection problem
  11. Semantic Answer Type and Relation Prediction Task (SMART 2021)
  12. Towards an open question answering architecture
  13. Offline question answering over linked data using limited resources
  14. GERBIL - General entity annotator benchmarking framework
  15. Mathematical relation between extended connectivity and eigenvector coefficients.
  16. 8th challenge on question answering over linked data (QALD-8)
  17. Entity linking in 40 languages using MAG
  18. On the distinctiveness of tags in collaborative tagging systems
  19. Developing a sustainable platform for entity annotation benchmarks
  20. Proceedings of the 7th Natural Language Interfaces for the Web of Data (NLIWoD)
  21. German Utilities and distributed PV
  22. Enhancing Community Interactions with Data-Driven Chatbots - The DBpedia Chatbot
  23. German Utilities and Distributed PV
  24. Analyzing Talk and Text II: Thematic Analysis
  25. Canopy leaf traits, basal area, and age predict functional patterns of regenerating communities in secondary subtropical forests
  26. Investigating quality raters' performance using interface evaluation methods
  27. NIF4OGGD - NLP interchange format for open German governmental data
  28. CETUS – a baseline approach to type extraction
  29. Question answering over linked data
  30. Support from the Internet for Individuals with Mental Disorders