The qualitative extension only evaluates the quality of the semantic elements but it
does not deal with the correction of the original ambiguous records. In Section6.2.1,
we presented a simple and a complex use case. The simple use case is tackled rel- atively easily because only one FRBR entity of each type is present in the record.
However, the complex use case requires more effort to be solved. Figure 6.9depicts
the different transformations applied to the original MARC record up to the enriched one obtained by using our approach. The top part of the figure shows the initial MARC record and its FRBRized representation from the FRBRizer tool. We notice that both of them suffer from the same problems, i.e., the translators (Eckerhard Schultz, and Hans- Joachim Maass) and the second creator (Per Wahlöö) are not included in the output of the transformation because it is not clear how they are related to the entities found in the record.
The bottom right part of the figure illustrates the FRBR-ML based representation in which the missing semantic information about persons is enhanced and corrected. For instance, “Hans-Joachim Maass”, “Eckerhard Schultz” and “Per Wahlöö” have
been identified as persons by the knowledge-based matching method (Section6.4.1).
The second problem is tackled using the correction method (Section 6.4.2). For
example, “Hans-Joachim Maass” is linked to the German expression “Verschlossen und verriegelt” that he has translated. The local collection lookup did not return any match for the expression title but querying z3950.l i br is.k b.se with the query
“” f ind @at t rset bi b− 1 @at t r 1 = 4 Verschlossen und ver riegel t” enabled to
discover the correct relationship. On the other hand, the relationship between “Per Wahlöö” and “Det slutna rummet” was found during the intra-collection search since there is a record which has only this work. From this FRBR-ML format, it is possible to convert to RDF, OWL, ORE and back to MARC.
The bottom-left part of the figure shows the results of transformation from the FRBR- ML to MARC. We notice that it is the corrected and enhanced version of the original record. Entities are grouped by the $8 linking field. In our example, “$8 1” groups the work “Det slutna rummet”, the expression “Verschlossen und verriegelt”, the creators
“Per Wahlöö” and “Maj Sjöwall”, and the translator “Hans-Joachim Maass”. For the
second work, we applied “$8 2” as linking field. Indicators11 W and E were adopted
to denote whether the entity is related to work or expression. As an example, “Per Wahlöö” is related to both works since both indicators are W . In addition, the correct
r el at or cod es “$4 t r l” for translators and “$4 aut” for creators are used to denote
their roles.
This new representation is inspired by both UNIMARC and MARC 21. The separation between names and titles comes from UNIMARC format while the use of the $8 link- ing field is common in MARC 21. Thus, our format has been adapted to fulfill our requirements, it is still compatible with with the ISO MARC standard.
6.7
Summary
Experience and user feedback in the cultural heritage community has shown that the adoption of new semantic technologies is slow, mainly because traditional library cat- alogs are still employed with records stored in the legacy format. Thus, we have pre- sented in this chapter FRBR-ML – an approach for extracting entity information from bibliographic data to populate a knowledge base. The format included in FRBR-ML can be used as an intermediary format to easily transform from/to MARC, RDF/XML, OWL and ORE. By writing an appropriate converter, one may also convert to other popular formats such as Dublin Core or ONIX.
The enrichment step in the metadata knowledge base population of FRBR-ML con- sists of different strategies to tackle issues related to the lack of semantics in MARC records and to the identification of basic relationships between entities. We have studied novel techniques for disambiguating obscure entries in original records, thus allowing to correct the initial input data. In addition, we have designed new metrics to check the quantity and quality of the transformation. These metrics evaluate the completeness, the percentage of duplicates and the amount of extra information due to the enrichment process.
The results of this chapters’ experiments are promising. The merging process effec- tively removes duplicate entities, thus substantially reducing the size of the knowl- edge base. However, the format included in FRBR-ML contains redundant properties,
but these redundancies can be easily removed during transformation to other for- mats. Additionally, it ensures a very high rate of completeness while allowing to correct and enhance ambiguous records with semantic information. This chapter also demonstrated that this semantic enrichment minimizes the rate of potential incorrect information.
In the future, there are several areas for improvement. The user feedback from librar- ians is important and should help detect the potential weaknesses and advantages of the approach in real world settings. This feedback mechanism could be integrated into the system when running in continuous mode. Although the approach was pre- sented in the context of MARC-based information and the FRBR conceptual model, the solution is generic that can be deployed for other types of information migration as well.
Another interesting issue to address is discovering complex relationships between entities. To fulfill this goal, a possible solution is to use pattern matching between
involved entities. This issue is addressed in Chapter8in the context natural language
documents.
Although, we have seen several research projects on the use of FRBR as an underlying model recently, these projects seldom explore the full potential of FRBR. In particular, no in-depth studies have been performed on how to actually deploy FRBR-inspired applications to end-user applications such as library catalogs available online. The issues concerned are related to the FRBR user tasks which basically deals with the presentation that should be entity-centric rather than simply displaying a groupings of works discovered in other records.
Other sources of structured data are websites that sell various products, such as Ama- zon12, Flickr13, Twitter14. It is common for these large Web sites to expose their data
for reuse via Web APIs. Programmable Web15 provides a comprehensive list of such
APIs. These APIs provide various query interfaces and return results using a number of different formats such as XML, JSON or ATOM. Therefore, an interesting question is: Can we apply conceptual domain model to product descriptions exposed via Web APIs? In the next chapter, an experimental study is presented on the example of using FRBR to describe product information.
12http://www.amazon.com(Last checked January 2012) 13http://www.flickr.com(Last checked January 2012) 14http://www.twitter.com(Last checked January 2012)
7
Exploiting Metadata on the Web
7.1
Introduction
In the metadata extraction approach presented in the previous chapter, we have con- sidered existing legacy metadata as the source of input data. However, the Web still remains the primary source of information for many users. The amount of data avail- able online is far larger than the one stored in library catalogs. However, this Web data is not well structured and not machine-interpretable, although the emergence
of the Semantic Web aims at tackling this issue[25]. For instance, representing facts
with triples enables computers to understand and use reasoning to answer complex queries. Libraries are also increasingly interested in linking their data to the LOD
cloud[131] as it enables semantic reuse of data providing a basis for new and inno-
vative services.
In this chapter, we closely examine application of a semantic model model to de- scriptions of Web product metadata. As the Web contains a lot of resources that may represent products of creative or artistic endeavor, we present an approach to trans- form the information about these products into the FRBR model. As the FRBR model focuses on modeling creative works in multiple levels of abstraction (work, expres-
sion, manifestation, and item), an example of such hierarchy is depicted in Figure7.1.
In this example, the three-part epic by J.R.R. Tolkien “The Lord of the Rings” is an ab- stract work encompassing “The Fellowship of the Ring”, “The Two Towers”, and “The Return of the King”. We advocate that such a representation would enable websites (e.g. e-commerce) to better organize and exploit those products.
Tolkien J.R.R
The Fellowship
of the Ring The Two Towers The Return of the King
The Lord of the Rings
The Two Towers Paperback edition,
September 2003
The Two Towers
(Eng.)
To tårn
(norw. translation)
The Two Towers Paperback edition, June 2005 Ringenes Herre - To tårn (2003) Person Work Work Expression Work Expression
Manifestation Manifestation Manifestation
Work has created has created is realized in is embodied in is embodied in has pa rt - Person
- Work - Expression - Manifestation
Legends
Figure 7.1: A Fragment of Lord of the Rings FRBR Work by J.R.R. Tolkien.