Towards Automatic Cataloging of Image and Textual Collections with Wikipedia

Tokinori Suzuki, Daisuke Ikeda, Petra Galuščáková, Douglas Oard

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

In recent years, a large amount of multimedia data consisting of images and text have been generated in libraries through the digitization of physical materials into data for their preservation. When they are archived, appropriate cataloging metadata are assigned to them by librarians. Automatic annotations are helpful for reducing the cost of manual annotations. To this end, we propose a mapping system that links images and the associated text to entries on Wikipedia as a replacement for annotation by targeting images and associated text from photo-sharing sites. The uploaded images are accompanied by descriptive labels of contents of the sites that can be indexed for the catalogue. However, because users freely tag images with labels, these user-assigned labels are often ambiguous. The label “albatross”, for example, may refer to a type of bird or aircraft. If the ambiguities are resolved, we can use Wikipedia entries for cataloging as an alternative to ontologies. To formalize this, we propose a task called image label disambiguation where, given an image and assigned target labels to be disambiguated, an appropriate Wikipedia page is selected for the given labels. We propose a hybrid approach for this task that makes use of both user tags as textual information and features of images generated through image recognition. To evaluate the proposed task, we develop a freely available test collection containing 450 images and 2,280 ambiguous labels. The proposed method outperformed prevalent text-based approaches in terms of the mean reciprocal rank, attaining a value of over 0.6 on both our collection and the ImageCLEF collection.

Original languageEnglish
Title of host publicationDigital Libraries at the Crossroads of Digital Information for the Future - 21st International Conference on Asia-Pacific Digital Libraries, ICADL 2019, Proceedings
EditorsAdam Jatowt, Akira Maeda, Sue Yeon Syn
PublisherSpringer
Pages167-180
Number of pages14
ISBN (Print)9783030340575
DOIs
Publication statusPublished - Jan 1 2019
Event21st International Conference on Asia-Pacific Digital Libraries, ICADL 2019 - Kuala Lumpur, Malaysia
Duration: Nov 4 2019Nov 7 2019

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume11853 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference21st International Conference on Asia-Pacific Digital Libraries, ICADL 2019
CountryMalaysia
CityKuala Lumpur
Period11/4/1911/7/19

All Science Journal Classification (ASJC) codes

  • Theoretical Computer Science
  • Computer Science(all)

Fingerprint Dive into the research topics of 'Towards Automatic Cataloging of Image and Textual Collections with Wikipedia'. Together they form a unique fingerprint.

Cite this