Full Paper

When Metadata Is Not Enough: Repository Interoperability in a National Library Open-Source Discovery Platform

  • Dwi Fajar Saputra
  • Taufik Asmiyanto
  • Nina Mayesti
Available
Access and licence
Open access CC BY 4.0
Download PDF
Contents

Abstract

Web scale discovery services (WSDS) depend on the interoperability of institutional repositories to provide reliable access to scholarly content. In practice, repository interoperability is often inferred solely from metadata registration, while operational sustainability receives limited attention. This study examines repository interoperability readiness within the Indonesia One Search (IOS) ecosystem using a multidimensional framework that integrates metadata completeness, harvesting sustainability, and metadata capacity. Drawing on registry and harvesting data from IOS, the study adopts a quantitative and exploratory design supported by visual analytical techniques, including density plots, scatter plots with regression analysis, boxplots, and heatmaps. The findings reveal a weak and non-linear relationship between metadata completeness and harvesting sustainability, indicating that well-documented metadata does not necessarily result in sustained interoperability. Most repositories are concentrated in low harvesting sustainability categories, reflecting widespread inactive or discontinuous integration. Journal site demonstrate higher metadata capacity and are more likely to achieve well-integrated status, whereas repositories managing datasets and electronic theses and dissertations (ETDs) exhibit lower capacity and greater vulnerability to operational stagnation. The study concludes that repository interoperability should be conceptualised as a dynamic and graduated condition rather than a binary state. Assessing metadata completeness in isolation is insufficient, sustained harvesting and metadata capacity must be evaluated jointly. This integrative assessment provides empirical evidence to support improved governance and long-term sustainability of national web-scale discovery services.

1 Introduction

The development of web-scale discovery services (WSDS) has been a key element in the integration of scientific information sources across repositories at national and global scales. Thru interoperability protocols such as the Open Archives Initiative Protocol for Metadata Harvesting (OAI-PMH), WSDS enables the aggregation of metadata from various heterogeneous systems to support unified indexing and information retrieval. In this context, metadata is no longer understood solely as a descriptive tool, but as an operational infrastructure that determines the success of integration, the stability of the index, and the sustainability of discovery services (Baglioni et al., 2025; Knoth et al., 2023)[3, 14].

However, according to Asok et al., the evaluation of repository system quality generally still focuses on metadata completeness, particularly in relation to the principles of open science and increasing content visibility (2024)[1]. Some studies even show a positive relationship between metadata completeness and user engagement in specific contexts, such as institutional repositories (Rasuli et al., 2025)[22]. However, this approach tends to be static and doesn't adequately consider operational dynamics in large-scale aggregation environments. Experience with international aggregation services shows that descriptively complete metadata cannot always be sustainably integrated due to platform differences, variations in institutional technical capacity, and inconsistencies in maintaining OAI-PMH endpoints (Knoth et al., 2023)[14].

This condition indicates a gap between the registration-based assessment of metadata quality and the reality of dynamic, long-term interoperability. Interoperability doesn't stop at the repository registration stage; rather, it depends on the sustainability of harvesting activities and the continuous maintenance of technical readiness. This perspective is becoming increasingly relevant in the context of national aggregation, where the heterogeneity of institutions and repository management systems poses its own structural challenges (Baglioni et al., 2025; Knoth et al., 2025)[3, 15].

Indonesia One Search (IOS), as a national WSDS service, faces similar challenges in integrating thousands of systems with diverse characteristics and levels of readiness. Although the IOS registration system has provided an initial foundation for metadata integration, there are strong indications of stagnation in harvesting activities and disparities in technical readiness among repositories. To date, empirical research comprehensively evaluating the interoperability conditions of systems in WSDS by linking metadata completeness, operational sustainability, and technical readiness is still limited. If we relate this to IOS as the WSDS of the national library, data from reporting.onesearch.id (2025) shows a decrease in data entries from 2022-2023, with the number of entries dropping by 55.4%[23]. Furthermore, no data is available for the years 2024-2025.

This study contributes to the literature by shifting repository interoperability evaluation from a static, registration-based perspective toward a dynamic, sustainability-oriented assessment grounded in real harvesting operations. Based on this gap, this study aims to evaluate the interoperability of repositories within the Indonesian One Search ecosystem using national aggregated data. Specifically, this research analyzes the completeness level of registration metadata, maps the sustainability of harvesting activities as an indicator of long-term interoperability, and classifies repository readiness based on descriptive and operational dimensions. Conceptually, the research is built upon the framework that interoperability is the result of a multidimensional interaction between metadata quality, technical maintenance sustainability, and institutional capacity, so complete metadata does not automatically guaranty stable integration without consistent operational support.

To test the framework, this study employs a quantitative-exploratory approach by analyzing the Indonesian One Search registration and harvesting dataset. This approach allows for an empirical evaluation of the relationship between metadata quality and long-term interoperability, while also providing an evidence-based foundation for developing more adaptive metadata policies and national WSDS management.

2 Literature Review

2.1 Quality and Completeness of Metadata in Digital Repositories

The quality of metadata is a fundamental factor in the success of information retrieval systems and web-scale discovery services. High-quality metadata allows information sources to be effectively indexed, discovered, and utilized within a distributed digital repository ecosystem. One of the most widely studied dimensions of metadata quality is metadata completeness, which refers to the extent to which the required metadata elements are available and well-filled. Within the framework of classical metadata quality, completeness is seen as a minimum prerequisite for a record to function optimally within an information system (Bruce & Hillmann, 2004)[7].

In recent decades, metadata completeness has been increasingly used as a quantitative indicator to assess system quality, primarily because it is relatively easy to measure and can be applied across metadata schemas. Asok et al. (2024) emphasize that complete metadata is an essential foundation for open science practices, particularly in supporting the visibility, interoperability, and reuse of research data[1]. Rasuli et al. also showed that metadata completeness contributes to increased visibility and user engagement, for example, in institutional repositories and library collection automation (2025)[22].

Various approaches have been developed to systematically measure the completeness of metadata. Ochoa and Duval (2009) introduced the automated evaluation of metadata quality using element-presence-based metrics[20], while Király and Büchler (2019 developed a weighted approach that considers the relative importance of each metadata element[13]. Another approach considers the hierarchical structure of metadata, where an element is considered complete only if all its sub-elements are filled (Margaritopoulos et al., 2012)[18]. Although these approaches provide a strong quantitative basis for assessing the descriptive quality of metadata, most still view metadata as a static entity evaluated at a single point in time.

Beside completeness, other quality dimensions such as accuracy and consistency are also widely discussed in the literature. However, several studies indicate that incompleteness, format inconsistency, and filling errors remain common issues in institutional repositories, even in stable systems (Çolak & Eroğlu, 2025)[8]. This finding indicates that the quality of metadata depends not only on the schema or guidelines used, but also on the practices of ongoing metadata management and maintenance.

2.2 Repository Interoperability and the Sustainability of Metadata Harvesting

In a large-scale aggregation environment, the quality of metadata cannot be separated from the issues of interoperability and operational sustainability. Web-scale discovery services rely on the ability of repository systems to provide metadata that can be consistently harvested thru interoperability protocols such as OAI-PMH. In this context, interoperability is not only determined by the completeness of metadata elements, but also by the stability of the endpoint, format compliance, and the continuity of system maintenance.

Global aggregators like CORE indicate that the heterogeneity of repository system platforms, variations in institutional technical capacity, and differences in metadata management practices are major challenges in cross-system metadata integration (Knoth et al., 2023)[14]. Manocci's study highlights that many repositories listed in national or international registries are not always sustainably harvestable due to URL changes, server failures, or a lack of technical maintenance (2022)[17]. This condition indicates that interoperability is a dynamic process susceptible to degradation over time.

Sustainable harvesting is an important indicator in assessing the interoperability readiness of repositories. Outdated metadata or repositories that are no longer responsive to harvesting requests can cause information to become obsolete, out of sync, or even disappear from the aggregator's index. This phenomenon is often associated with the concept of metadata decay, which is the decline in the quality and relevance of metadata due to a lack of long-term maintenance. Although the dimension of timeliness has long been recognized as part of metadata quality, this aspect is relatively rarely operationalized systematically in repository evaluations (Kumar & Harinarayana, 2024)[16].

Some recent studies are beginning to emphasize the importance of combining metadata quality evaluation with operational sustainability indicators. Baglioni et al. (2025) highlight the need for a registry-based interoperability approach that not only records the existence of repositories but also validates the technical readiness and sustainability of metadata integration[3]. A similar approach is also seen in studies on repository registries and national aggregation networks, which emphasize the importance of technical coordination and sustainable maintenance to ensure the stability of discovery services (Knoth et al., 2025)[15].

3 Methodology

This study employs a quantitative-exploratory approach to evaluate the interoperability of repositories within the national web-scale discovery services (WSDS) ecosystem. The quantitative approach was chosen because it allows for objective measurement of metadata characteristics and interoperability performance on a large scale, while the exploratory nature is used to identify patterns, disparities, and trends that have not been extensively mapped in national-scale repository interoperability studies (Creswell & Creswell, 2018; Kumar & Harinarayana, 2024)[9, 16].

The research object includes repositories listed in Indonesia One Search as a national discovery service. Research data is sourced from registration metadata and harvesting records managed by the aggregator system. The use of aggregated data is considered methodologically relevant because interoperability within the WSDS cannot be understood at the individual repository level, but rather as a systemic phenomenon emerging from the interaction between repositories, interoperability protocols, and aggregator infrastructure (Knoth et al., 2023; Baglioni et al., 2025)[3, 14].

Conceptually, the interoperability of repository systems in this study is treated as a multidimensional construct encompassing descriptive, technical, and operational aspects. This multidimensional approach aligns with the literature on metadata quality and interoperability evaluation, which emphasizes that metadata completeness alone is insufficient to represent long-term integration readiness, especially in large-scale aggregation environments (Ochoa & Duval, 2009; Margaritopoulos et al., 2012; Kumar & Harinarayana, 2024)[16, 18, 20].

In this study, metadata capacity is conceptualised as an integrative indicator reflecting the structural, technical, and operational readiness of a repository to sustain metadata-based interoperability over time. Unlike metadata completeness, which focuses on the presence of descriptive elements at the registration level, metadata capacity captures the repository's ability to maintain interoperable metadata practices through stable harvesting, technical compliance, and administrative continuity. This construct is used to bridge descriptive metadata quality and long-term operational sustainability within the WSDS environment.

Data analysis was conducted by combining metadata completeness measurement, statistical distribution analysis, temporal harvesting classification, and comparative analysis across repository types. The use of presence-based scoring for metadata completeness metrics is a common technique in metadata evaluation studies to generate quantitative indicators that are replicable and comparable across systems (Ochoa & Duval, 2009; Király & Büchler, 2019)[13, 20]. Distribution analysis and comparison were used to identify structural variations in interoperability readiness across repository groups, as recommended in the institutional repository evaluation study (Çolak & Eroğlu, 2025)[8].

The temporal dimension of interoperability is analyzed thru the use of data harvesting to assess the sustainability of metadata integration over time. This temporal approach is used to capture the phenomena of interoperability degradation and metadata decay, which have been identified as important issues in global repository aggregation and service discovery studies (Mannocci et al., 2022; Knoth et al., 2023)[14, 17]. Thus, interoperability is not treated as a static condition, but rather as a dynamic process that depends on the sustainability of technical and operational maintenance.

Thru this methodological approach, the research does not aim to establish a causal relationship, but rather to present an empirical mapping based on national aggregate data regarding the interoperability conditions of repositories. This approach aligns with the practice of large-scale information systems evaluation, which places descriptive, classificatory, and trend analysis as the basis for policy decision-making and system development (Creswell & Creswell, 2018; Baglioni et al., 2025)[3, 9].

4 Discussion

4.1 Mapping of IOS Registration Field Metadata Categories

A screenshot of the Indonesia One Search registration dataset opened as a spreadsheet, with one repository per row and columns including date_created, date_modified, status, reason, content_type, title, publisher, oai_url, home_url, cover_url, metadataPrefix and others; the visible rows list Indonesian journals and repositories with their managing institutions and harvesting status values.
Figure 1. The Indonesia One Search registration dataset.

The Indonesian One Search (IOS) registration dataset contains metadata fields that not only function as a description of the repository's identity but also track the technical readiness and operational status of its integration into the national retrieval system. Due to the heterogeneous nature of the field characters, category mapping is necessary to ensure that each attribute can be read consistently within the interoperability evaluation framework. In this study, the entire field is understood thru four functional categories: (1) administrative registration metadata, (2) interoperability technical metadata, (3) operational harvesting metadata, and (4) system administration metadata. The four components form a sequence that represents the repository integration cycle: from identity at registration, technical access for metadata exchange, activity trail of harvesting and indexing, to workflow control and change auditing.

Table 1. Administrative Registration Metadata
FieldDefinitionFunction
idUnique identifier of an entity in the IOS systemRecord identification
titleName of the repository, journal, or collection as registered in IOSEntity identity
publisherName of the managing or responsible institutionService accountability
descriptionDescription of the repositoryAdditional contextual information
library_idInstitutional identifier within the IOS networkInstitutional relationship and linkage
Table 2. Technical Interoperability Metadata
FieldDefinitionFunction
oai_urlOAI-PMH endpoint of the repository or journalHarvesting access
metadataPrefixMetadata format used (e.g., dc, oai_dc, mods, etc.)Interoperability standard
setSpecSet identifier within the OAI-PMH structureHarvesting filter and segmentation
update_frequencyUpdate frequency selected by the repository managerSynchronization scheduling
content_typeType of system (journal, institutional repository, ETD, SLiMS UCS, others)Platform identification
repository_software_idRepository or publishing system used (OJS, EPrints, SLiMS, DSpace, etc.)Platform interoperability
home_urlMain homepage of the repository or journalEntity existence validation
cover_urlCover image or visual representation of the repository (optional)Visual identity and aesthetics
Table 3. Operational Harvesting Metadata
FieldDefinisiFunction
harvest_dateDate of the most recent successful harvesting processSustainability indicator
harvest_statusStatus of the harvesting result (e.g., success, fail, pending)OAI communication quality
harvest_userSystem or reviewer responsible for the harvesting processAudit trail
harvest_urlSource URL used for harvestingEndpoint validation
dedup_dateDate when metadata deduplication was performedPost-harvest processing
dedup_statusStatus of the deduplication processMetadata cleaning indicator
dedup_userUser or process responsible for deduplicationAudit trail
solrupdate_dateDate when metadata were successfully indexed into SolrIndexing success
solrupdate_statusStatus of the indexing processDiscovery validation
solrupdate_userSystem or process responsible for Solr indexingAudit trail
Table 4. System Administration Metadata
FieldDefinitionFunction
date_createdDate when the entity was created in the IOS systemRegistration initiation
date_modifiedDate when the metadata were last modifiedMaintenance history
date_queuedDate when the entity entered the review queueWorkflow management
statusRegistration status (e.g., approved, pending, rejected)Reviewer validation
reasonNotes provided by the reviewer explaining the decisionStandards compliance assessment
notesAdditional remarks or technical notesSupplementary technical comments
queuedBoolean flag indicating whether the entity is in the review queueQueue position indicator
created_byAdministrator or user who created the entityAudit
last_modified_byUser or system responsible for the latest modificationAudit
queued_byUser or system that placed the entity into the review queueAudit

4.2 Data Preprocessing

Data preprocessing in this study is a fundamental step to ensure that the Indonesian One Search (IOS) dataset is in a suitable condition for analysis. The raw dataset obtained from the IOS registration and harvesting system exhibits variations in format, the presence of null values, and heterogeneity in metadata structure, which could potentially affect the accuracy of quantitative analysis. Pre-processing focuses on four main aspects: data cleaning, format consistency, semantic completeness, and mapping of analytical indicators. This series of steps ensures that each metadata field can be interpreted consistently within the framework of repository interoperability evaluation. With systematically processed datasets, further analysis can more accurately reflect the institution's metadata conditions within the Indonesian One Search ecosystem, while also supporting the validity and comparability of research results.

A top-to-bottom flowchart of the research workflow. Data Acquisition (IOS Registry and Harvesting Logs) feeds Data Preprocessing, which branches into Data Cleaning (Strings and Duplicates), Format Standardization (Schema and Elements), Semantic Harmonization (Field Mapping), and Indicator Mapping; these map to Administrative Registry, Technical Interoperability, Operational Harvesting, and System Administration Factors. The flow continues to Metadata Completeness Scoring, then Dimension-Specific Scores, then Harvesting Sustainability Classification, which branches into Active, Stagnant, Dormant, Never Harvested, Comparative Analysis (Repository Type), and Metadata Decay Index, and finally converges on Interpretation and Discussion.
Figure 2. Research Workflow for Evaluating Repository Interoperability in Indonesia One Search

4.3 Analysis of Metadata Completeness: Distribution, Dimensions, and Systemic Patterns

A histogram titled Distribution of Repository Metadata Health Scores, plotting the number of repositories against the metadata health score per repository (0 to 100 percent). The tallest bars are in the 25 to 40 percent range (35.2 percent and 23.9 percent of repositories), with a secondary cluster near 0 percent (14.8 percent) and only a small share above 75 percent.
Figure 3. Distribution of Repository Metadata Health Scores

The distribution of health metadata scores indicates a concentration of repositories in the mid-to-low score range, with a small percentage of repositories at a very high level of completeness. This uneven distribution pattern reflects what is referred to in the literature as metadata quality skewness, which is the condition where most repositories only meet minimum metadata requirements, while more mature metadata practices are only found in a limited group of institutions (Park & Tosaka, 2015)[21].

Some studies confirm that this type of distribution pattern is common in open and decentralized institutional repository ecosystems. Stvilia et al. (2017) explain that without continuous audit and feedback mechanisms, metadata tends to evolve unevenly, depending on the local capacity of the managing institution[24]. In the national WSDS context, this means that the existence of repositories with high metadata scores does not automatically improve the overall quality of the system, as the majority of repositories remain at a minimum level of contribution.

This distribution also strengthens the argument that interoperability assessment cannot be based solely on average values, but must consider the spread and concentration of scores. This kind of distributional approach has been recommended in the evaluation of large-scale information systems to avoid conclusions biased toward high-capacity institutions (Batini, Cappiello & Francalanci, 2016)[4]. Thus, the histogram of iOS metadata scores not only describes the technical condition of the metadata but also reveals structural disparities in the interoperability readiness of repositories.

A boxplot titled Metadata Completeness by Dimension, showing completeness per repository (percent) for four metadata dimensions. Basic metadata has the highest and most stable completeness (mean 79.5 percent), while Administrative (28.7 percent), Technical (12.9 percent), and Harvesting (25.5 percent) dimensions show much lower medians with greater spread.
Figure 4. Metadata Completeness by Dimension

The boxplot visualization shows a sharp difference between the basic, administrative, technical, and harvesting metadata dimensions. The basic metadata dimension shows a relatively high and stable level of completeness, while the technical and harvesting dimensions exhibit a low median with greater variation. This pattern indicates that metadata practices in repositories are more focused on resource description rather than supporting operational interoperability.

This phenomenon aligns with findings in metadata management studies that distinguish between descriptive sufficiency and functional interoperability. Gilliland (2016) emphasizes that descriptive metadata is often prioritized because it is directly related to the visibility and legitimacy of collections, while technical and operational metadata is considered a less institutionally focused supporting layer[11]. As a result, although the repository appears "complete" descriptively, its capacity to actively participate in the aggregation system remains limited.

Additionally, the low completeness in the technical and harvesting dimensions can also be interpreted as an indicator of weak integration sustainability. According to Higgins (2018), long-term interoperability in digital systems is highly dependent on process metadata, including the recording of harvesting and update activities[12]. When this type of metadata is not managed consistently, aggregation systems risk functional degradation even if descriptive metadata remains available.

The differences between dimensions shown by the boxplot also support the multidimensional evaluation approach in this study. As stated by Zeng and Qin (2016), the quality of metadata cannot be reduced to a single indicator, but must be read as a combination of interconnected dimensions[26]. This visualization strengthens the argument that the completeness of administrative metadata is not sufficient to represent the interoperability readiness of repositories within the national WSDS.

These two visualizations together show that the quality of repository metadata within the Indonesian OneSearch ecosystem is uneven and highly dependent on the dimensions being evaluated. While basic metadata is relatively stable and complete, the technical and operational metadata that underpins long-term interoperability still shows structural weaknesses. This finding reinforces the view that WSDS interoperability cannot be evaluated solely thru the completeness of descriptive metadata, but rather requires a multidimensional approach that considers score distribution and the operational dynamics of the repository.

4.4 Dynamics of Repository Harvesting Activity: Patterns, Challenges, and Sustainability

The distribution of sustainability harvesting classes shows that most repositories in the Indonesian OneSearch ecosystem fall into the "never harvested" category, while the proportion of repositories actively harvested in the last year is very small. This pattern indicates that the integration of repositories into the national WSDS is often administrative, without sustained operational connectivity. In other words, the existence of a repository within a registration system does not automatically transform into an active metadata flow within the discovery index.

The dominance of the "never harvested" and "dormant" categories reflects a phenomenon known in repository aggregation literature as structural harvesting failure, which is a systemic and recurring integration failure, not just a temporary technical glitch. Mannocci, Baglioni, and Manghi (2022) show that at a large aggregation scale, many repositories become inaccessible due to endpoint changes, lack of system maintenance, or the absence of periodic monitoring mechanisms[17]. In the context of iOS, this pattern indicates a gap between the registration process and the sustainable integration of metadata.

From an interoperability perspective, these findings confirm that harvesting sustainability is a more sensitive indicator than the completeness of registration metadata in assessing repository readiness. Repositories can have complete registration metadata, but still not contribute to WSDS if harvesting activity is not sustained. This condition strengthens the argument that interoperability needs to be understood as a dynamic process that depends on technical and operational continuity, rather than as a static status at a single point in time (Higgins, 2018)[12].

A 100 percent stacked bar chart titled Harvesting sustainability by repository type, with one bar per type (Book, Dataset, ETD, Journal, Other) and four sustainability classes: Active (≤ 1 year), Stagnant (1–3 years), Dormant (> 3 years), and Never harvested. The Never harvested class dominates every type and is largest for Dataset and Journal; ETD shows the largest Dormant share.
Figure 5. Harvesting sustainability by repository type

If analyzed by repository type, there is a noticeable difference in the sustainability patterns of harvesting. The journal repository shows a relatively better proportion of harvesting activity compared to ETD repositories, datasets, and other types. This pattern reflects that repositories with regular publication cycles and a strong visibility orientation tend to have greater institutional incentives to maintain their technical interoperability. Conversely, ETD and dataset repositories are dominated by the never harvested and dormant categories. This indicates that although these repositories play a crucial role in the knowledge and open science ecosystem, managing their interoperability is often not a long-term priority. Repository management literature indicates that repositories with irregular content flows and fragmented management responsibilities are more prone to experiencing a decline in metadata integration quality (Park & Tosaka, 2015)[21].

The low sustainability of harvesting on dataset repositories also underscores the complexity of research data interoperability. Batini, Cappiello, and Francalanci (2016) emphasize that dataset interoperability demands higher metadata investment and technical maintenance compared to text-based repositories[4]. Without adequate policy and resource support, this type of repository tends to fail to maintain metadata integration within the national aggregation system.

Both visualizations together show that the main interoperability challenge for repositories in IOS does not lie in the registration stage, but rather in the post-registration phase. Low harvesting sustainability indicates the risk of metadata decay, which is a condition where metadata remains structurally present but loses its operational functionality within the discovery system (Gilliland, 2016)[11]. In the long term, this condition can lead to institutional representation imbalances and reduce the reliability of WSDS as a national retrieval service.

The variation in sustainability across repository types also confirms that interoperability is contextual and influenced by domain characteristics, institutional practices, and technical governance. Therefore, the national WSDS evaluation needs to move beyond a uniform approach and consider monitoring and intervention mechanisms tailored to the type of repository. This finding reinforces the need for sustainable harvesting indicators as an integral part of repository interoperability evaluation, not merely a supplement to registration metadata analysis.

4.5 Integrating Metadata Completeness, Harvesting Sustainability, and Interoperability Readiness

A two-dimensional density heatmap titled Density of Repositories by Metadata Completeness and Harvesting Sustainability, with harvesting sustainability score (percent) on the vertical axis and basic metadata completeness (percent) on the horizontal axis; colour encodes count. The highest counts are concentrated in the bottom band where harvesting sustainability is near zero, across all levels of metadata completeness.
Figure 6. Density of Repositories by Metadata Completeness and Harvesting Sustainability

This density visualization shows that most repositories within the Indonesian OneSearch ecosystem are concentrated in a combination of low to medium metadata completeness with very low harvesting sustainability. The highest density is in areas with a harvesting sustainability value approaching zero, although the completeness level of basic metadata varies. This pattern indicates that the completeness of registration metadata is not automatically directly proportional to the sustainability of repository integration into aggregation systems, a phenomenon that has been widely reported in large-scale repository interoperability studies (Park & Tosaka, 2015; Mannocci, Baglioni & Manghi, 2022)[17, 21].

The low repository density in areas with a combination of high metadata completeness and high harvesting sustainability indicates that ideal interoperability conditions are relatively rare. This finding supports the view that descriptive metadata is often prioritized for initial legitimacy and visibility purposes, while the technical and operational metadata that underpins sustainable harvesting receives less institutional attention (Gilliland, 2016; Zeng & Qin, 2016)[11, 26]. Thus, structurally complete metadata does not necessarily reflect a repository's readiness to actively participate in the web-scale discovery services ecosystem.

The distribution of repositories with varying levels of metadata completeness but still at a low harvesting level also indicates a structural decoupling between metadata quality and the operational performance of interoperability. Global repository aggregation literature indicates that the sustainability of harvesting is more influenced by technical and governance factors, such as the stability of the OAI-PMH endpoint, repository software maintenance, and human resource sustainability, than by the mere completeness of registration metadata (Knoth et al., 2023; Higgins, 2018)[12, 14].

From a system evaluation perspective, this density pattern strengthens the argument that repository interoperability is a dynamic process susceptible to degradation if not accompanied by continuous monitoring and maintenance. The condition where metadata remains available but is no longer actively harvested reflects the phenomenon of metadata decay, which directly impacts the timeliness and reliability of national discovery indexes (Higgins, 2018; Mannocci et al., 2022)[12, 17]. Therefore, this visualization confirms the need to integrate metadata completeness indicators with sustainable harvesting indicators in assessing repository interoperability readiness.

A scatterplot titled Relationship between Metadata Completeness and Harvesting Sustainability, with harvesting sustainability score (percent) on the vertical axis and basic metadata completeness (percent) on the horizontal axis, points coloured by repository type (Book, Dataset, ETD, Journal, Other). A dashed regression line slopes only very slightly upward and stays near the bottom of the plot, indicating a weak positive relationship; points are spread across all completeness levels including high-completeness repositories with low harvesting.
Figure 7. Relationship between Metadata Completeness and Harvesting Sustainability

This scatterplot visualization shows that the relationship between metadata completeness and harvesting sustainability is weak and non-linear. The slightly upward sloping regression line indicates a positive trend, but its very low slope confirms that increasing the completeness of basic metadata only contributes marginally to the sustainability of harvesting activities. In other words, repositories with more complete metadata do not consistently show a higher level of sustainable harvesting.

The wide distribution of points across various levels of metadata completeness, including repositories with high completeness but low harvesting values, suggests that metadata completeness is not the primary predictor of operational interoperability. This finding aligns with the literature emphasizing that the quality of descriptive metadata often serves as an initial prerequisite, but is not sufficient to guaranty long-term technical integration success in aggregation and discovery systems (Park & Tosaka, 2015; Zeng & Qin, 2016)[21, 26].

The variation in dot color based on repository type shows that the relationship between metadata completeness and harvesting sustainability is also contextual and domain-specific. Journal repositories tend to appear at a higher harvesting rate compared to other repository types, although they don't always have the highest metadata completeness. This pattern indicates that institutional factors—such as publication rhythm, visibility demands, and management support—have a greater influence on harvesting sustainability than metadata completeness alone. Studies on global repository aggregation show that repositories with regular content update cycles are better able to maintain technical interoperability, regardless of variations in the quality of their descriptive metadata (Knoth et al., 2023)[14].

Conversely, the dataset and ETD repositories show a more inconsistent distribution, with many points at low harvesting levels despite varying metadata completeness. This condition reinforces the finding that the interoperability of this type of repository is highly dependent on the sustainability of technical maintenance and data management policies, not just on filling in registration metadata (Higgins, 2018)[12]. In the national WSDS context, this confirms that the interoperability improvement strategy cannot focus solely on enhancing metadata completeness, but needs to be directed toward strengthening operational harvesting mechanisms.

Conceptually, this visualization supports the argument that repository interoperability is the result of a multidimensional interaction between metadata, technology, and institutional practices. The weak relationship between metadata completeness and harvesting sustainability indicates a structural decoupling between descriptive readiness and operational performance. Therefore, a comprehensive evaluation of WSDS interoperability needs to integrate sustainable harvesting indicators and institutional context as key components of repository readiness assessment.

A boxplot titled Repository metadata capacity by repository type, with the metadata capacity index (percent) on the vertical axis for Book, Dataset, ETD, Journal, and Other. Journal and ETD show the highest median capacity, Book is intermediate, and Dataset and the Other category show lower medians with wide spread.
Figure 8. Repository metadata capacity by repository type

This boxplot visualization clearly shows differences in metadata capacity across repository types, confirming that interoperability readiness is not uniform within the Indonesian OneSearch ecosystem. Journal and ETD repositories show a relatively higher median metadata capacity compared to book, dataset, and other categories of repositories. This pattern indicates that repositories with a strong academic orientation and a more structured content flow tend to have more mature metadata management practices.

The high metadata capacity in journal repositories can be attributed to the demands of visibility, standardization, and the sustainability of scientific publications. Repository management literature indicates that journal repositories generally have stricter metadata policies and more stable integration with global aggregation and indexing systems, thus consistently supporting higher metadata capacity (Knoth et al., 2023; Park & Tosaka, 2015)[14, 21]. This explains why journal repositories, despite not always having the highest metadata completeness, still demonstrate better interoperability readiness.

ETD repositories occupy a middle position with considerable variation, reflecting the heterogeneity of management practices at the institutional level. Some ETD repositories exhibit metadata capabilities approaching those of journal repositories, while others are at a lower level. This variation aligns with the finding that ETD management is highly dependent on institutional policies and the continuity of technical support, resulting in uneven quality of the generated metadata (Higgins, 2018; Gilliland, 2016)[11, 12].

Conversely, the "other" dataset and category repositories show a lower median metadata capacity with a wide range of values. This condition indicates that although dataset repositories have significant potential in the context of open science, their metadata practices and interoperability are not yet fully established. Batini, Cappiello, and Francalanci (2016) emphasize that research data interoperability requires complex and sustained metadata investment; without adequate policy and resource support, dataset repositories are likely to experience metadata capacity limitations as reflected in this visualization[4].

The differences in metadata capacity between these repository types strengthen the argument that repository interoperability needs to be understood as a systemic capability, not merely the result of basic metadata completeness or simple harvesting activities. Metadata capacity serves as an integrative indicator, capturing the structural, technical, and operational readiness of the repository more comprehensively. Thus, this visualization bridges the previous completeness-sustainability relationship analysis toward the interoperability readiness classification of repositories in the next stage.

A heatmap titled Interoperability readiness by repository type, with repository types (Book, Dataset, ETD, Journal, Other) on the vertical axis and four readiness classes on the horizontal axis: A. Ready and sustained, B. Ready but not sustained, C. Partially ready, and D. Not ready; each cell shows the share of repositories, shaded darker for higher share. Most repositories fall into B. Ready but not sustained or D. Not ready, with Journal showing the largest D share (86.3 percent) and Dataset concentrated in B (80.0 percent).
Figure 9. Interoperability readiness by repository type

This heatmap shows that the interoperability readiness of repositories within the Indonesian OneSearch ecosystem varies significantly across repository types and tends to be concentrated at medium to low levels of readiness. Most repositories fall into the "Ready but not active" category, while the proportion of repositories that are truly in the "Well integrated" category is relatively limited and dominated by journal repositories. This pattern shows that although many repositories have met basic administrative and technical prerequisites, only a small fraction have been able to consistently maintain operational integration.

Journal repositories hold the most prominent position in the Well-integrated class, indicating that the combination of metadata capacity, harvesting sustainability, and technical governance is more mature compared to other repository types. International literature confirms that journal repositories generally have more stable interoperability standards because they are tied to the global scientific publishing ecosystem, including indexers, aggregators, and open access initiatives (Björk, 2017; Knoth et al., 2023)[5, 14]. Thus, the interoperability readiness of journal repositories is not only influenced by metadata, but also by external pressures and institutional incentives.

Conversely, the dataset and ETD repositories are dominated by the Ready but not active and Partially ready classes. This reflects that although the repository already has metadata elements and basic infrastructure, the sustainability of integration and operational maintenance remains a major challenge. Studies on research data management show that dataset interoperability requires more complex coordination of metadata, policies, and technical resources compared to text-based repositories, which is why many dataset repositories fail to achieve full integration (Whyte & Wilson, 2017; Borgman, 2015)[6, 25].

The "Not ready" category that still appears for some repository types, especially the "other" category, indicates structural gaps in interoperability readiness. The presence of repositories in this class indicates that the registration of repositories in the national WSDS has not been fully followed by continuous validation and mentoring mechanisms. The literature on discovery system evaluation emphasizes that without continuous interoperability governance, aggregation systems risk becoming mere lists of repositories rather than reliable retrieval infrastructures (Deodato, 2015; NISO, 2014)[10, 19].

Overall, this visualization confirms that repository interoperability readiness is tiered and highly influenced by domain context and institutional practices. This heatmap also reinforces the findings from the previous visualization that interoperability cannot be reduced to a single indicator, but is rather the result of integration between metadata capacity, harvesting sustainability, and system governance. Therefore, the interoperability readiness classification becomes an important instrument for the strategic and policy-oriented evaluation of the national WSDS.

5 Conclusion

This study demonstrates that repository interoperability within the Indonesia One Search ecosystem cannot be adequately assessed through metadata completeness alone, as harvesting sustainability and metadata capacity play decisive roles in determining long-term integration. The findings reveal structural disparities across repository types, where journal repositories tend to achieve higher interoperability readiness, while dataset and ETD repositories remain vulnerable to operational stagnation and metadata decay. By integrating metadata completeness and harvesting dynamics, this study contributes a multidimensional and evidence-based framework that advances the evaluation of web-scale discovery services beyond static metadata assessments. The results offer practical insights for improving repository governance and monitoring mechanisms in national discovery infrastructures. Future research should incorporate longitudinal analysis and institutional policy variables to further refine interoperability readiness models and support sustainable repository ecosystems.

6 Notes

  1. Papers will not be published without the agreement of referees, selected for their thematic specialisation, who will review the articles on a double-blind basis.

7 Authorship statement

Dwi Fajar Saputra: Conceptualization (lead); writing – original draft (lead); formal analysis (lead); writing – review and editing (equal). Taufik Asmiyanto: review and editing (equal). Nina Mayesti: review and editing (equal).

References

  1. [1] K. Asok, S. S. Dandpat, D. K. Gupta, and P. Shrivastava, Common metadata framework for research data repository: necessity to support open science. in Journal of Library Metadata, vol. 24, no. 1, pp. 133-145, 2024.
  2. [2] Introduction to Metadata. Getty Research Institute, Los Angeles, 2016.
  3. [3] M. Baglioni, G. Pavone, A. Mannocci, and P. Manghi, Towards the interoperability of scholarly repository registries. in International Journal on Digital Libraries, vol. 26, 2025. Article 2
  4. [4] C. Batini, C. Cappiello, and C. Francalanci, Data and Information Quality: Dimensions, Principles and Techniques. Springer, Cham, 2016.
  5. [5] B.-C. Björk, Open access to scientific articles: a review of benefits and challenges. in Information Services & Use, vol. 37, no. 4, pp. 395-408, 2017.
  6. [6] C. L. Borgman, Big Data, Little Data, No Data: Scholarship in the Networked World. MIT Press, Cambridge, MA, 2015.
  7. [7] T. R. Bruce and D. I. Hillmann, The Continuum of Metadata Quality: Defining, Expressing, Exploiting. ALA Editions, Chicago, 2004.
  8. [8] F. A. Çolak and Ş. Eroğlu, Evaluating metadata quality in institutional academic repositories of Turkish research universities. in Online Information Review, vol. 49, no. 7, pp. 1335-1352, 2025.
  9. [9] J. W. Creswell and J. D. Creswell, Research Design: Qualitative, Quantitative, and Mixed Methods Approaches. Sage, Thousand Oaks, CA, 2018.
  10. [10] J. Deodato, Evaluating web-scale discovery: A step-by-step guide. in Information Technology and Libraries, vol. 34, no. 2, pp. 19-75, 2015. https://doi.org/10.6017/ital.v34i2.5745.
  11. [11] A. J. Gilliland, Setting the stage. in Introduction to Metadata, Getty Research Institute, pp. 1-19, 2016.
  12. [12] S. Higgins, Digital Curation: The Emergence of a New Discipline. Facet Publishing, London, 2018.
  13. [13] P. Király and M. Büchler, Measuring completeness as a metadata quality metric in Europeana. in Proceedings of the IEEE International Conference on Big Data, IEEE, pp. 2711-2720, 2019.
  14. [14] P. Knoth, D. Herrmannova, M. Cancellieri, L. Anastasiou, N. Pontika, and et al., CORE: a global aggregation service for open access papers. in Scientific Data, vol. 10, 2023. 366
  15. [15] P. Knoth, P. Walk, M. Cancellieri, M. Upshall, and et al., USRN Discovery Pilot: increasing the discoverability of open access content through a national network. in Proceedings of the 20th International Conference on Open Repositories, 2025.
  16. [16] V. C. Kumar and N. S. Harinarayana, Exploring dimensions of metadata quality assessment: a scoping review. in Journal of Librarianship and Information Science, 2024.
  17. [17] A. Mannocci, M. Baglioni, and P. Manghi, "Knock Knock! Who's There?" a study on scholarly repositories' availability. in Theory and Practice of Digital Libraries (TPDL 2022), Springer, pp. 306-312, 2022.
  18. [18] T. Margaritopoulos, M. Margaritopoulos, I. Mavridis, and A. Manitsaris, Quantifying and measuring metadata completeness. in Journal of the American Society for Information Science and Technology, vol. 63, no. 4, pp. 724-737, 2012.
  19. [19] NISO, Open Discovery Initiative: promoting transparency in discovery. National Information Standards Organization, Baltimore, 2014.
  20. [20] X. Ochoa and E. Duval, Automatic evaluation of metadata quality in digital repositories. in International Journal on Digital Libraries, vol. 10, no. 2-3, pp. 67-91, 2009.
  21. [21] J. Park and Y. Tosaka, Metadata quality control in digital repositories and collections: criteria, semantics, and mechanisms. in Cataloging & Classification Quarterly, vol. 53, no. 5-6, pp. 595-613, 2015.
  22. [22] B. Rasuli, M. Boock, J. Schöpfel, and B. Van Wyk, The link between dissertation metadata completeness and user engagement in an institutional repository. in Scientometrics, vol. 130, no. 5, pp. 2875-2899, 2025.
  23. [23] Reporting Indonesia OneSearch, Global summary report: All institutions. 2025. https://reporting.onesearch.id/index.php/global/all/summary.
  24. [24] B. Stvilia, L. Gasser, M. B. Twidale, and L. C. Smith, A framework for information quality assessment. in Journal of the Association for Information Science and Technology, vol. 68, no. 1, pp. 12-25, 2017.
  25. [25] A. Whyte and A. Wilson, How to Appraise and Select Research Data for Curation. Digital Curation Centre, Edinburgh, 2017.
  26. [26] M. L. Zeng and J. Qin, Metadata. Facet Publishing, London, 2016.

Article details

Available
Section
Full Papers
DOI
10.23106/dcmi.952637150
License
CC BY 4.0 · open access

Described in Dublin Core

This article's metadata, in the vocabulary these proceedings are about.

dcterms:title
When Metadata Is Not Enough: Repository Interoperability in a National Library Open-Source Discovery Platform
dcterms:creator
Saputra, Dwi Fajar
Asmiyanto, Taufik
Mayesti, Nina
dcterms:available
2026-08-01
dcterms:identifier
doi:10.23106/dcmi.952637150
dcterms:publisher
Dublin Core Metadata Initiative
dcterms:type
Text
dcterms:language
en
dcterms:rights
CC BY 4.0