Jump to content

Research:Measuring the Gender Gap: Attribute-based Class Completeness Estimation

From Meta, a Wikimedia project coordination wiki
Created
06:05, 1 November 2022 (UTC)
Contact
Gianluca Demartini
Duration:  2023-07 – 2025-06
Grant ID: G-RS-2303-12081
This page documents a completed research project.


Introduction

[edit]

The problem. Successful crowdsourcing projects like Wikipedia and Wikidata naturally grow and evolve over time. This happens while having editors focussing on certain parts of the project instead of others. While the ability for editors to decide what to contribute to comes with the advantage of flexibility, it may result in biased content where, for example, one gender is better represented than others. An example of this is the number of male astronauts as compared to the number of female astronauts (73 out of 574, https://en.wikipedia.org/wiki/List_of_female_astronauts https://en.wikipedia.org/wiki/List_of_space_travelers_by_name) in Wikipedia.

The solution. There are possible viable approaches to address this issue. For example, the editor community may decide to stop adding new male astronauts to the project to allow for content about female astronauts to catch up. Alternatively, the community may decide to represent the real distribution in the profession. In any case, this remains a community decision.

Research Contribution. Rather than deciding how to deal with gender unbalanced content in Wikimedia projects, the aim of this research is to automatically identify underrepresented classes by quantifying and measuring the expected size of a class in order to empower the community in taking decisions and setting editorial priorities. This is possible by making use of the edit history for a Wikimedia project.

Literature Review

[edit]

Gender in Wikipedia

Despite its broad reach and influence of knowledge sharing, Wikipedia faces persistent challenges related to the gender gap (Cabrera et al. 2018; Falenska and Çetinoğlu 2021; Hill and Shaw 2013; Miquel-Ribé and Laniado 2021; Redi et al. 2021). Previous research such as those by Farzan et al. (Farzan et al. 2016) has already looked at the gender gap, highlighting the importance of addressing such disparities across Wikimedia projects (Konieczny and Klein 2018; Redi et al. 2021). To name a few, Pellissier Tanon and Suchanek (Pellissier Tanon and Suchanek 2019) proposed a system to access the data through a SPARQL endpoint and index changes in the Wikidata graph and enable users to trace the contributions of human or automated contributors. Bourli and Pitoura (Bourli and Pitoura 2020) proposed measures for identifying bias in a Wikidata dataset and a debiasing approach based on projections on the gender subspace. Abián et al. (Abián et al. 2022) proposed methods to identify content gaps in crowdsourced knowledge graphs and found that Wikidata editors usually tend not to work on under-represented entities. While many attempts to address this issue exist, the key difference in the approach we propose to take is that we do not attempt to address the issue, but rather, to apply effective methods and develop accurate instruments for the different Wikipedia user groups to be empowered in making data-driven editorial decisions and able to address the issue by themselves.

Class Cardinality Estimation

Previous research (Lueggen et al. 2019) has looked at how to use statistical estimators to estimate class cardinality in Wikidata. They used the knowledge graph edit history as evidence for the estimators. In this paper, we extend this approach by looking at attribute-specific cardinality estimations (e.g., How many female astronauts should be there? Do we have them all?) and beyond the Wikidata project. In the area of crowdsourced databases, the problem of answering queries under the open-world assumption has been studied in the past. Researchers have encountered the problem that popular entities are reported by crowd members more frequently than “tail” (i.e., unpopular) entities, thus making it difficult to complete the answer set (and, in our case, to estimate the class cardinality). The approach followed by Trushkowsky et al. (Trushkowsky et al. 2015) has looked at using statistical estimators to understand how far from the complete set the incrementally constructed query answer set currently is. We apply these methods to understand how complete Wikipedia is.

Abián, D., Meroño-Peñuela, A., and Simperl, E. 2022. “An Analysis of Content Gaps Versus User Needs in the Wikidata Knowledge Graph,” in The Semantic Web – ISWC 2022, U. Sattler, A. Hogan, M. Keet, V. Presutti, J. P. A. Almeida, H. Takeda, P. Monnin, G. Pirrò, and C. d’Amato (eds.), Cham: Springer International Publishing, pp. 354–374. (https://doi.org/10.1007/978-3-031- 19433-7_21).

Bourli, S., and Pitoura, E. 2020. “Bias in Knowledge Graph Embeddings,” in 2020 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM), , December, pp. 6–10. (https://doi.org/10.1109/ASONAM49781.2020.9381459).

Cabrera, B., Ross, B., Dado, M., and Heisel, M. 2018. “The Gender Gap in Wikipedia Talk Pages,” Proceedings of the International AAAI Conference on Web and Social Media (12:1). (https://doi.org/10.1609/icwsm.v12i1.15053).

Falenska, A., and Çetinoğlu, Ö. 2021. “Assessing Gender Bias in Wikipedia: Inequalities in Article Titles,” in Proceedings of the 3rd Workshop on Gender Bias in Natural Language Processing, M. Costa-jussa, H. Gonen, C. Hardmeier, and K. Webster (eds.), Online: Association for Computational Linguistics, August, pp. 75–85. (https://doi.org/10.18653/v1/2021.gebnlp-1.9).

Farzan, R., Savage, S., and Saviaga, C. F. 2016. “Bring on Board New Enthusiasts! A Case Study of Impact of Wikipedia Art + Feminism Edit-A-Thon Events on Newcomers,” in Social Informatics, E. Spiro and Y.-Y. Ahn (eds.), Cham: Springer International Publishing, pp. 24–40. (https://doi.org/10.1007/978-3-319-47880-7_2).

Hill, B. M., and Shaw, A. 2013. “The Wikipedia Gender Gap Revisited: Characterizing Survey Response Bias with Propensity Score Estimation,” PLoS ONE (8:6), p. e65782. (https://doi.org/10.1371/journal.pone.0065782).

Konieczny, P., and Klein, M. 2018. “Gender Gap through Time and Space: A Journey through Wikipedia Biographies via the Wikidata Human Gender Indicator,” New Media & Society (20:12), SAGE Publications, pp. 4608–4633. (https://doi.org/10.1177/1461444818779080).

Lueggen, M., Difallah, D., Sarasua, C., Demartini, G., and Cudré-Mauroux, P. 2019. “Non-Parametric Class Completeness Estimators for Collaborative Knowledge Graphs—The Case of Wikidata,” in The Semantic Web – ISWC 2019, C. Ghidini, O. Hartig, M. Maleshkova, V. Svátek, I. Cruz, A. Hogan, J. Song, M. Lefrançois, and F. Gandon (eds.), Cham: Springer International Publishing, pp. 453–469. (https://doi.org/10.1007/978-3-030-30793-6_26).

Miquel-Ribé, M., and Laniado, D. 2021. “The Wikipedia Diversity Observatory: Helping Communities to Bridge Content Gaps through Interactive Interfaces,” Journal of Internet Services and Applications (12:1), p. 10. (https://doi.org/10.1186/s13174-021-00141-y).

Pellissier Tanon, T., and Suchanek, F. 2019. “Querying the Edit History of Wikidata,” in The Semantic Web: ESWC 2019 Satellite Events, P. Hitzler, S. Kirrane, O. Hartig, V. de Boer, M.-E. Vidal, M. Maleshkova, S. Schlobach, K. Hammar, N. Lasierra, S. Stadtmüller, K. Hose, and R. Verborgh (eds.), Cham: Springer International Publishing, pp. 161–166. (https://doi.org/10.1007/978-3- 030-32327-1_32).

Redi, M., Gerlach, M., Johnson, I., Morgan, J., and Zia, L. 2021. A Taxonomy of Knowledge Gaps for Wikimedia Projects (Second Draft), arXiv. (https://doi.org/10.48550/arXiv.2008.12314).

Trushkowsky, B., Kraska, T., Franklin, M. J., Sarkar, P., and Ramachandran, V. 2015. “Crowdsourcing Enumeration Queries: Estimators and Interfaces,” IEEE Transactions on Knowledge and Data Engineering (27:7), pp. 1796–1809. (https://doi.org/10.1109/TKDE.2014.2339857).

Methods and Data

[edit]

The research conducted for this project explored three complementary directions:

Study 1. Data-driven work. In this work-package we focussed on adapting statistical methods to estimate the completeness of Wikipedia for different types of persons. The aim was to estimate the level of Wikipedia coverage and completeness across genders. Our results indicate high level of estimated completeness across different sub-classes of the class person.

Our method can estimate the completeness of a class of entities. Hence can be used to answer questions such as “Does the knowledge base have a complete list of all female astronauts?”. Our techniques are derived from species estimation and data management and are applied to the case of collaborative editing. We make use of entities observed in a project’s edit history as a proxy for observations in a capture/recapture study setup. This allows us to use estimators for species population (e.g., Jackknife Estimators [1]) to predict class cardinality. A full report is available here: https://doi.org/10.48550/arXiv.2401.08993 This work was also presented at the Wiki workshop 2024 and later published in the proceedings of the 35th Australasian Conference on Information Systems. Full paper can be viewed here: https://aisel.aisnet.org/acis2024/99/

Study 2. We have been interviewing Wikipedia editors who focus on gender balance to understand the strategies they use and use this to inform the design of data tools that could aid their work.

Study Design. We have conducted semi-structured interviews with Wikipedia editors to understand their decision process in balancing the gender gap when contributing content to Wikipedia. In total, we planned to interview 25 to 30 editors. For each participant, we asked the following semi-structured questions to understand their behaviours in line with our objectives of the study: (1) Do you think about gender balance when you edit Wikipedia? Do you find gender an important dimension that you think about when editing Wikipedia? Why is, in your opinion, important to focus on gender balance? (2) Based on your editorial work for Wikipedia, what’s your usual process when considering gender balance while editing? Would you share an example? (3) How do you make decisions on what content to contribute and what to focus on? And why? (4) Could you comment on any difficulties or challenges you experienced when working towards balancing gender in the Wikimedia? And how do you address them (5) What sources of information do you use when balancing gender in Wikipedia? (6) What data and information would you like to be able to access to help you make decisions on gender balancing contributions to Wikipedia? In which format, modality, or interface would you like to access this data? For example, if you are provided with a dashboard of gender distribution data in Wikipedia, would you use it? How?

This interview protocol allows participants to reveal their decision process in balancing the gender gap, allowing us to understand the intentions behind the editing process, as well as the difficulties or challenges they experienced.

Ethical considerations. We obtained human research ethics approval from the University of Queensland Human Research Ethics committee before recruiting Wikipedia editor participants. We provided all participants with the participant information sheet and consent form and asked them to provide either written or verbal consent. The target participants are all Wikipedia editors. Any Wikipedia editors could participate to this project and no screening was required.

Recruitment. We recruited interviewees to the groups of Wikimedia editors recommended by other researchers, as well as the Wikipedia user talk pages, snowball sampling, and face-to-face interactions at Wikimedia-related events. The target participants do not belong to any special ethnic, age or gender groups.

Data collection. The interviews with Wikipedia editors are conducted using the semi-structured protocol online using Zoom, currently spanning from 2023 to 2024 over the course of 14 months. All the data collected from the participants are anonymised and are not identifiable or re-identifiable. The overall interview for each participant took 70 - 90 minutes. All interviews were conducted in English and audio recorded and then transcribed.

Data analysis. As for analysing the user insights collected from the interviews, we used NVIVO 12 (For the use of Nvivo 12, please view https://lumivero.com/products/nvivo/) for verbal protocol analysis. The verbal protocols have been transcribed and provided to two independent coders for analysis to reduce bias in our analysis. We follow the procedure outlined by Gioia et al., to show the `dynamic relationships among the emergent concepts that describe or explain the phenomenon of interest and one that makes clear all relevant data-to-theory connections’. We started the analysis in an inductive approach and then transited into a more abductive approach, to ensure `data and existing theory are now considered in tandem’.

Study 3. We have built a dashboard to visualise gender data distributions and statistics to support editors in their editorial choices.

This work has been accepted for publication at the ACM Web Conference 2025 (WWW 2025) [2].

[1] Heltshe, J.F., Forrester, N.E.: Estimating species richness using the jackknife procedure. Biometrics pp. 1–11 (1983)

[2] Yahya Yunus, Tianwa Chen, and Gianluca Demartini. 2025. Exploring Wikipedia Gender Diversity Over Time — The Wikipedia Gender Dashboard (WGD). In Companion Proceedings of the ACM Web Conference 2025 (WWW Companion ’25), April 28-May 2, 2025, Sydney, NSW, Australia. ACM, New York, NY, USA, 4 pages. https://doi.org/10.1145/3701716.3715175

Policy, Ethics and Human Subjects Research The work makes use of Wikipedia edit history, thus will not disrupt the current work of Wikipedia editors. We have obtained ethics approval from our institution for this type of analysis as well as for the interviews conducted with editors.

Results

[edit]

Research Publications that originated from this research project:

(1) Hrishikesh Patel, Tianwa Chen, Ivano Bongiovanni, and Gianluca Demartini (2024). Estimating Gender Completeness in Wikipedia The 35th Australasian Conference on Information Systems (ACIS 2024) https://aisel.aisnet.org/acis2024/99/

(2) Yahya Yunus, Tianwa Chen, and Gianluca Demartini. Exploring Wikipedia Gender Diversity Over Time - The Wikipedia Gender Dashboard (WGD) In: The 2025 ACM Web Conference (TheWebConf 2025) - Demo track. Sydney, Australia, April 2025. (DOI)

Research Presentations

(1) Estimating Gender Completeness in Wikipedia., by Tianwa Chen. Presented at The 35th Australasian Conference on Information Systems (ACIS 2024), Wellington, New Zealand.

(2) Measuring the Gender Gap, by Tianwa Chen. Presented at Wikimedia Research/Showcase, March 19, 2025 https://www.mediawiki.org/wiki/Wikimedia_Research/Showcase Presentation slides can be viewed from https://commons.wikimedia.org/wiki/File:Measuring_the_Gender_Gap_on_Wikimedia_Research_Showcase_March_2025_slides.pdf

(3) Exploring Wikipedia Gender Diversity Over Time, by Tianwa Chen. Presented at Wikimedia Australia Community Meeting, July, 2025

(4) Exploring Wikipedia Gender Diversity Over Time - The Wikipedia Gender Dashboard (WGD), by Tianwa Chen. The 2025 ACM Web Conference (TheWebConf 2025) - Demo track. Sydney, Australia, May 2025.

Conclusion

[edit]

Research Summary

The project Measuring the Gender Gap: Attribute-based Class Completeness Estimation was formally completed in June 2025.

This project investigated the framework for estimating gender representation gaps in Wikipedia by assessing class completeness across various attributes. The research led to the publication of two peer-reviewed papers at The Web Conference (WWW) and Australasian Conference on Information Systems (ACIS), identifying gender disparities in Wikipedia content.

To complement this quantitative work, we initiated a qualitative phase involving interviews with Wikipedia editors and community members to better understand how gender gaps are perceived, addressed, and acted upon in real editorial practice. As of June 2025, we have conducted interviews with 27 participants across diverse roles and language communities within the Wikimedia ecosystem. These interviews explored editors’ awareness of gender imbalances, data needs for supporting inclusive editing, and preferences for potential tools such as gender-distribution dashboards.

Next Steps

Moving forward, we see different possible direction for future research. We will focus on two key areas:

(1) Endangered Articles and Underdeveloped Content: Identifying and analyzing articles, particularly biographies and topical pages about underrepresented genders, that are at risk of being overlooked, poorly sourced, or tagged for deletion.

(2) Barriers to Entry for New Editors: Investigating the challenges faced by new or potential contributors, especially those from gender-diverse or marginalized groups—through additional interviews and focus groups, with the aim of informing design strategies and policy recommendations that support more inclusive participation.

These next steps will contribute to a broader understanding of how platform design, editorial norms, and data availability influence bias and equity in collaborative knowledge production.