<?xml version="1.0" encoding="UTF-8"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/"
         xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
         xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/
         http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-10-01T15:54:28Z</responseDate>
  <request verb="GetRecord" identifier="oai:dcpapers.dublincore.org:952649937" metadataPrefix="oai_dc">https://dcpapers.dublincore.org/oai</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:dcpapers.dublincore.org:952649937</identifier>
        <datestamp>2026-09-10</datestamp>
        <setSpec>dcmi-2026</setSpec>
        <setSpec>openaire</setSpec>
      </header>
      <metadata>
    <oai_dc:dc
        xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/"
        xmlns:dc="http://purl.org/dc/elements/1.1/"
        xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
        xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/
        http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
      <dc:title>From MARC to Linked Open Data: AI-Driven Entity Extraction from Hebrew Manuscript Metadata Using Distant Supervision</dc:title>
      <dc:creator>Goldberg, Alexander</dc:creator>
      <dc:creator>Prebor, Gila</dc:creator>
      <dc:creator>Elmalech, Avshalom</dc:creator>
      <dc:description>Cultural heritage institutions preserve invaluable provenance information in Machine-Readable Cataloging (MARC) records, yet much of this knowledge remains trapped in unstructured note fields, inaccessible to computational analysis. Transforming these legacy catalogs into Linked Open Data (LOD) requires extracting structured person-role relationships—identifying authors, scribes, owners, and censors—from cataloger narratives. This study presents an AI-driven system that uses distant supervision
from MARC metadata itself to automatically generate training data, eliminating the prohibitive cost of manual annotation for specialized cultural heritage domains. By exploiting the dual structure of catalog records, where structured fields provide authoritative labels and unstructured notes provide context, we achieve 85.70% F1 for person extraction and 100% role classification accuracy, outperforming general Hebrew NER models by +55.55% F1. Our approach demonstrates how existing metadata can be leveraged to train AI systems that align with the values of cultural heritage preservation: accuracy, provenance tracking, and semantic enrichment. The extracted entities populate ontology instances based on CIDOC-CRM and IFLA-LRM, enabling computational analysis of scribal networks and manuscript circulation at scale.</dc:description>
      <dc:publisher>Dublin Core Metadata Initiative</dc:publisher>
      <dc:date>2026-09-10</dc:date>
      <dc:type>info:eu-repo/semantics/conferenceObject</dc:type>
      <dc:type>Text</dc:type>
      <dc:format>application/pdf</dc:format>
      <dc:format>text/html</dc:format>
      <dc:identifier>https://doi.org/10.23106/dcmi.952649937</dc:identifier>
      <dc:identifier>https://dcpapers.dublincore.org/article/952649937</dc:identifier>
      <dc:source>Dublin Core Metadata Initiative Conference Proceedings</dc:source>
      <dc:language>eng</dc:language>
      <dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
      <dc:rights>https://creativecommons.org/licenses/by/4.0/</dc:rights>
    </oai_dc:dc>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>