Poster

Finding Identities : Machine Learning Metadata Analysis on LGBTQ+ Knowledge Representation in the University of the Philippines Diliman Academic Libraries

  • Jessie Rose M. Bagunu 1 ORCID
  • Miriam Charmigrace Q. Salcedo 1 ORCID
  • Michael D. Amandy 1 ORCID
  • Maria Maura S. Tinao 2 ORCID
  • 1 Library, School of Library and Information Studies, University of the Philippines Diliman, PH
  • 2 Faculty, School of Library and Information Studies, University of the Philippines Diliman, PH
Available
Access and licence
Open access CC BY 4.0
Download PDF
Contents

Abstract

This study investigates the gap between the rich conceptual depth of LGBTQ+ knowledge representation and the limited metadata used to describe identities within the Philippine academic landscape. A mixed-methods approach was employed and focused on a corpus of over 1,400 bibliographic records across five library units at the University of the Philippines Diliman. The study concludes that the future of library metadata must evolve from static, rigid filing systems into dynamic and socially responsive maps of human identity. There is an urgent need to move beyond exact-match keyword searches and similarity-based discovery be embrace which allows users to explore the conceptual relationships between diverse works.

1 Introduction

Metadata is the primary mediator of visibility and accessibility of repositories in academic libraries. However, it is never a neutral technical stance as it is a reflection of the values, biases, rules and cultural assumptions from the time from which it was created. This study investigates the gap between the rich conceptual depth of LGBTQ+ knowledge representation and the limited metadata used to describe identities within the Philippine academic landscape. In many knowledge repositories, the standardized language used to organize research functions as a mechanism of silence were adopted from classification and descriptors that fail to mirror the vibrant and evolving reality of queer identities. When metadata remains overly generalized or Western-centric, it goes beyond making research material harder to find it effectively erases the identity of the work and the community it represents, creating a barrier to the true discovery and representation of identities.

3 Methodology

This research employed a mixed-methods approach to bridge the gap between traditional technical cataloging and contemporary user-centered information retrieval. The study focused on a corpus of over 1,400 bibliographic records across five library units at the University of the Philippines Diliman.

Central to the methodology was a three-stage computational pipeline designed to analyze and enrich these records. The first stage, Preprocessing, utilized Rapid Automatic Keyword Extraction (RAKE) and Term Frequency-Inverse Document Frequency (TF-IDF) to generate high-impact tags directly from titles and abstracts, capturing the natural language of the authors. The second stage, Exploration, applied K-Means clustering to group these records based on their semantic content, visually mapping the disconnect between existing "library silos" and the actual thematic intersections of the research. The final stage, Bridge-building, employed Cosine

Similarity to develop a semantic-based search prototype, enabling a system that suggests related materials based on conceptual proximity rather than shared classification codes. This quantitative pipeline was complemented by qualitative inquiry, including interviews with librarians and user feedback, ensuring the findings remained grounded in the practical realities of library service

4 Findings and Discussions

Through this deep computational analysis, the research revealed a profound disconnect between the richness of academic content and the poverty of its metadata. The analysis uncovered a pervasive silence where the vast majority of records lacked the specific descriptors necessary for modern discovery. While the physical collections contain a wealth of information on human rights, local politics, and intersectional health, these nuanced themes are frequently collapsed under the single, monolithic subject heading of "Gender."

By applying machine learning techniques, the study was able to uncover latent conceptual structures hidden within the library’s silos. This process demonstrated that AI can effectively reveal thematic connections across disparate departments, transforming the collection into an interconnected and visible digital knowledge system that reflects the actual complexity of the research.

5 Conclusion

The study concludes that the future of library metadata must evolve from static, rigid filing systems into dynamic and socially responsive maps of human identity. To achieve this, a hybrid approach is essential—one that harnesses the scalable power of machine learning to enrich records while maintaining the ethical oversight and cultural empathy of human librarians. There is an urgent need to move beyond exact-match keyword searches and embrace similarity-based discovery that allows users to explore the conceptual relationships between diverse works. By integrating localized vocabularies and community-informed descriptors into the metadata pipeline, knowledge repositories can transform their catalogs from mere inventories into instruments of social inclusion. Finally, by reimagining how we find identities, we ensure that the stories of marginalized communities are not just safely stored away, but are actively found, recognized, and celebrated.

6 About the Authors

Jessie Rose M. Bagunu is a College Librarian at the University Library, University of the Philippines - Diliman and currently is the Head Librarian at the University of the Philippines School of Library and Information Studies Library or #UPSLISLib. Her research interest includes information needs and behavior, personal learning networks, metadata and classification as well as on the well-being of senior librarians and members of the elderly population, which she actively supports and advocates on their behalf.

Miriam Charmigrace Q. Salcedo is the Reference Instruction Librarian at the University of the Philippines Diliman’s School of Library and Information Studies (SLIS) Library. Her work centers on bridging the gap between historical preservation and modern accessibility. Through her research in social media, digital democratization, and user engagement, she leads initiatives that transform library services. Committed to community outreach, she is a passionate advocate for equitable information access, ensuring that valuable resources remain available and accessible to the public.

Michael D. Amandy is a library staff member at the University of the Philippines School of Library and Information Studies (UP SLIS) Library and an LIS student at Manuel S. Enverga University Foundation. Passionate about information management and queer studies, he is currently working toward becoming a licensed professional librarian in the Philippines.

Maria Maura S. Tinao has been with the University of the Philippines since 1993. She is a full-time faculty at the School of Library and Information Studies (SLIS) at the University of the Philippines Diliman. A certified IBM Data Science and Data Analyst Professional, her expertise lies in data science, machine learning, predictive analytics, and data analysis. Her research spans breast cancer detection, sarcopenia risk prediction, and AI adoption among the aging population, demonstrating her commitment to projects that bridge technology and meaningful societal impact.

References

  1. [1] M. Adler, Cruising the Library: Perversities in the Organization of Knowledge. Fordham University Press, 2017.
  2. [2] M. J. Bates, The invisible substrate of information science. in J. Am. Soc. Inf. Sci., vol. 50, no. 12, pp. 1048, 1999.
  3. [3] S. Berman, Prejudices and Antipathies: A Tract on the LC Subject Heads Concerning People. Scarecrow Press, 1971.
  4. [4] A. Crystal and J. Greenberg, Usability of a metadata creation application for resource authors. in Library and Information Science, vol. 27, pp. 177-189, 2005.
  5. [5] E. Drabinski, Queering the Catalog: Queer Theory and the Politics of Correction. in The Library Quarterly, vol. 83, no. 2, pp. 94-111, 2013.
  6. [6] J. Greenberg, Metadata and Digital Information. in Encyclopedia of Library and Information Science, Marcel Dekker, Inc., New York, pp. 3610-3623, 2010. https://doi.org/10.1081/E-ELIS3-120044415.
  7. [7] J. B. C. Inciong, Representation of Filipiniana LGBT books in selected university libraries in Metro Manila. University of the Philippines Diliman, 2014. Unpublished Undergraduate Thesis
  8. [8] T. Kunesh and J. Romines, Changing the subject: The Homosaurus in Emory University's library catalogue. in PINES Bibliographic Services Specialist, 2025.
  9. [9] A. Ogletree, R. Koskela, K. Jeffery, A. Ball, D. Dublin, J. Greenberg, and F. Berman, RDA Overview and Metadata Activities. in DC-2015: Metadata and Ubiquitous Access to Culture, Science, and Digital Humanities: Proceedings of the International Conference on Dublin Core and Metadata Applications, São Paulo, Brazil, 2015.
  10. [10] The Queer Metadata Collective, K. Adolpho, A. Bailund, E. Beck, J. Bradshaw, E. Butler, B. Bárcenas, M. Caelin, R. Carpenter, A. Day, T. Day, D. Dixon, A. Dover, S. Frizzell, G. Goodrich, B. L. Hendrickson, T. Keller, C. Misorski, D. Murphy, and C. Yragui, Best Practices for Queer Metadata (1.0). Zenodo, 2024. https://doi.org/10.5281/zenodo.12580531.
  11. [11] B., et al.) UBC Researchers (Watson, The Power of Names: Exploring Queer Cataloguing Conventions in Libraries. UBC School of Information, 2024.

Article details

Available
Section
Posters
DOI
10.23106/dcmi.952690206
License
CC BY 4.0 · open access

Described in Dublin Core

This article's metadata, in the vocabulary these proceedings are about.

dcterms:title
Finding Identities : Machine Learning Metadata Analysis on LGBTQ+ Knowledge Representation in the University of the Philippines Diliman Academic Libraries
dcterms:creator
Bagunu, Jessie Rose M.
Salcedo, Miriam Charmigrace Q.
Amandy, Michael D.
Tinao, Maria Maura S.
dcterms:available
2026-08-01
dcterms:identifier
doi:10.23106/dcmi.952690206
dcterms:publisher
Dublin Core Metadata Initiative
dcterms:type
Text
dcterms:language
en
dcterms:rights
CC BY 4.0