1Introduction

This research proposal aims to develop an artificial intelligence system for the automated dating of undated Hebrew manuscripts produced after 1540. While existing projects have primarily focused on manuscripts up to 1540, a substantial gap remains in the study of later periods.

The foundation of this research is the premise that MARC metadata from Hebrew manuscripts contains distinctive patterns and features that can be effectively analyzed using AI algorithms to determine approximate dating and typological classification. We hypothesize that post-1540 Hebrew manuscripts exhibit systematic variations in their physical and codicological characteristics that correlate with their temporal and geographical origins. The integration of AI analysis of MARC metadata with traditional paleographic knowledge is expected to achieve significantly higher accuracy in manuscript dating than either method alone.

2Methodology

We use MARC metadata from the KTIV platform as the primary dataset. This metadata includes physical characteristics, scribal features, and provenance details. Our analysis focuses on:

  • Supervised learning (e.g., support vector machines) for predicting manuscript dates
  • Unsupervised learning (e.g., clustering) for discovering typological groupings
  • Statistical analysis and visualization to validate temporal and regional patterns

A preliminary dataset of 1000 dated manuscripts from 1600-1700 was used for initial testing.

3Research questions

  1. How can AI techniques enhance the analysis of MARC metadata for dating post-1540 manuscripts?
  2. How can unsupervised learning identify new typological features in manuscript metadata?
  3. How can AI algorithms incorporate historical and cultural nuances for accurate dating?

4Preliminary Results

Our initial analysis revealed:

  • A surge in Ashkenazic manuscript production around 1640, with subsequent decline
  • Italian scripts maintaining high production levels into the late 17th century
  • Distinct subject correlations across script types (e.g., Karaite texts with Karaite script)
  • Clear regional patterns, e.g., Yemenite scripts found exclusively in Yemen
  • Strong statistical correlation (r > 0.75) between certain MARC fields and manuscript dates

These findings confirm the potential of AI models to identify scribal traditions and temporal trends through catalog metadata alone.

5Contribution and Significance

This project pioneers the application of AI to post-1540 Hebrew manuscripts by leveraging structured MARC metadata. Building on previous applications of AI in manuscript analysis [2], it bridges a crucial gap in Hebrew manuscript studies and offers a replicable framework for large-scale, data-driven dating and typological analysis. By integrating AI methodologies with traditional paleographic scholarship, the project generates innovative tools for cataloging undated manuscripts, reveals insights into regional scribal practices and their historical evolution, and contributes significantly to both Jewish studies and the digital humanities. It also provides practical benefits for libraries and archives and demonstrates how legacy metadata standards such as MARC can be creatively reimagined to support advanced computational analysis.

6Future Work

Future phases will include expanding the dataset, refining the machine learning models, and developing a user-facing tool for manuscript dating and typology. A key goal is to integrate historical context into the AI models to improve interpretability and scholarly relevance.