Web-scale extension of RDF knowledge bases from templated websites

Research output: Contributions to collected editions/worksArticle in conference proceedingsResearchpeer-review

Authors

  • Lorenz Bühmann
  • Ricardo Usbeck
  • Axel Cyrille Ngonga Ngomo
  • Muhammad Saleem
  • Andreas Both
  • Valter Crescenzi
  • Paolo Merialdo
  • Disheng Qiu

Only a small fraction of the information on the Web is represented as Linked Data. This lack of coverage is partly due to the paradigms followed so far to extract Linked Data.While converting structured data to RDF is well supported by tools, most approaches to extract RDF from semi-structured data rely on extraction methods based on ad-hoc solutions. In this paper, we present a holistic and open-source framework for the extraction of RDF from templated websites. We discuss the architecture of the framework and the initial implementation of each of its components. In particular, we present a novel wrapper induction technique that does not require any human supervision to detect wrappers for web sites. Our framework also includes a consistency layer with which the data extracted by the wrappers can be checked for logical consistency. We evaluate the initial version of REX on three different datasets. Our results clearly show the potential of using templated Web pages to extend the Linked Data Cloud. Moreover, our results indicate the weaknesses of our current implementations and how they can be extended.

Original languageEnglish
Title of host publicationThe SemanticWeb - ISWC 2014 - 13th International SemanticWeb Conference, Proceedings
EditorsTania Tudorache, Craig Knoblock, Paul Groth, Carole Goble, Chris Welty, Abraham Bernstein, Peter Mika, Denny Vrandečić, Natasha Noy, Krzysztof Janowicz
Number of pages16
PublisherSpringer Nature Switzerland AG
Publication date2014
Pages66-81
ISBN (print)978-3-319-11963-2
ISBN (electronic)978-3-319-11964-9
DOIs
Publication statusPublished - 2014
Externally publishedYes
Event13th International Semantic Web Conference, ISWC 2014 - Riva del Garda, Italy
Duration: 19.10.201423.10.2014
Conference number: 13
https://search.worldcat.org/de/title/semantic-web-iswc-2014-13th-international-semantic-web-conference-riva-del-garda-italy-october-19-23-2014-proceedings-part-i/oclc/941304230

Bibliographical note

Publisher Copyright:
© Springer International Publishing Switzerland 2014.

Recently viewed

Publications

  1. Clause identification using entropy guided transformation learning
  2. Intellectual property issues in the use and distribution of remote sensing data
  3. Mathematical Modeling for Robot 3D Laser Scanning in Complete Darkness Environments to Advance Pipeline Inspection
  4. Constraints are the solution, not the problem
  5. Investigation and modeling of the material behavior due to evolving dislocation microstructures in fcc and bcc metals
  6. A Service-oriented Search framework for full text, geospatial and semantic search
  7. Parameters Estimation of a Lotka-Volterra Model in an Application for Market Graphics Processing Units
  8. Empowering materials processing and performance from data and AI
  9. Changes in the Complexity of Limb Movements during the First Year of Life across Different Tasks
  10. Estimation and interpretation of a Heckman selection model with endogenous covariates
  11. Comparison of Bio-Inspired Algorithms in a Case Study for Optimizing Capacitor Bank Allocation in Electrical Power Distribution
  12. The signal location task as a method quantifying the distribution of attention
  13. Who can receive the pass? A computational model for quantifying availability in soccer
  14. Changing the Administration from within:
  15. FaST: A linear time stack trace alignment heuristic for crash report deduplication
  16. Towards a Bayesian Student Model for Detecting Decimal Misconceptions
  17. Mining positional data streams
  18. Universal Threshold Calculation for Fingerprinting Decoders using Mixture Models
  19. Analyzing math teacher students' sensitivity for aspects of the complexity of problem oriented mathematics instruction
  20. Real-time RDF extraction from unstructured data streams
  21. Combining a PI Controller with an Adaptive Feedforward Control in PMSM
  22. “Ideation is Fine, but Execution is Key”
  23. Age effects on controlling tools with sensorimotor transformations
  24. Applications of the Simultaneous Modular Approach in the Field of Material Flow Analysis
  25. Assessing Effects Through Semi-Field and Field Toxicity Testing
  26. Understanding reading as a form of language-use
  27. A new way of assessing the interaction of a metallic phase precursor with a modified oxide support substrate as a source of information for predicting metal dispersion
  28. Computing regression statistics from grouped data
  29. HAWK - hybrid question answering using linked data
  30. A Line with Variable Direction, which Traces No Contour, and Delimits No Form
  31. Identification of conductive fiber parameters with transcutaneous electrical nerve stimulation signal using RLS algorithm
  32. Explaining and controlling for the psychometric properties of computer-generated figural matrix items
  33. Scaffolding argumentation in mathematics with CSCL scripts