Research Libraries UK

Inclusive Collections, Inclusive Libraries – Machine Learning Approaches to Gender Bias in Archival Curation

Loading Events

Inclusive Collections, Inclusive Libraries is an RLUK programme of events that aims to foster conversation around decolonisation and inclusive practice in collecting, describing, presenting, and engaging with content in research library collections. It seeks to raise awareness about the opportunities and challenges of dealing with, contextualising, and engaging with offensive collections while also identifying and sharing examples of

Machine Learning Approaches to Gender Bias in Archival Curation

6 February 2025, 16:00 – 17:00 (GMT), 18:00 – 19:00 (SAST/EET), 17:00 – 18:00 (CET), 11:00 – 12:00 (EST), 08:00 – 09:00 (PST)

In this presentation Lucy Havens will report on research combining Machine Learning (ML) and human-centered research methods to identify gender biased language in archival catalogue metadata descriptions. Though the ML community has achieved significant advances in language modeling, state-of-the-art models still have social biases encoded within them, the implications of which are under-explored. ML researchers and practitioners have focused more on minimising and removing bias than on understanding and managing bias. Adopting pre-trained ML models thus risks harm to marginalised communities that information and heritage professionals have been working to better represent. Motivated by information and heritage professionals’ desire to manage bias, as well as the scalability and efficiency that ML offers, Lucy has been investigating a new use case for ML: identifying bias. This work began during her Ph.D. research at the University of Edinburgh, where she created the first text classification models to identify gender biased language in the University’s archival catalog metadata descriptions. Now, thanks to a Research and Innovation Grant from The National Archives, she is building upon that work with Ian Johnson (Head of Special Collections & Archives) and the Newcastle University Special Collections team. In this talk, she will provide an overview of the models and mixed-methods approach to their development and evaluation, and discuss insights this research has offered on the capabilities and limitations of ML for description and critical cataloging workflows.

Share This Story, Choose Your Platform!

Go to Top