Jump to content

Research:The state of science and Wikimedia: Who is doing what, and who is funding it?

From Meta, a Wikimedia project coordination wiki

This page documents a research project in progress.
Information may be incomplete and change as the project progresses.
Please contact the project lead before formally citing or reusing results from this page.


Research Report

[edit]

Submission PDF on OpenReview

A partially revised version of this proposal was resubmitted.

Project description

[edit]

This research examines Wikimedia-related publications to see who did them, what was done, and most importantly who has funded this sort of work in the past. Our goal is to understand what has been done and identify which sponsors or contributors supported those activities, both as a service to the community and as the basis from which to grow future projects and collaborations. We use the Scholia (Nielsen et al., 2017) dataset, employing bibliometric analysis, to enhance our existing efforts to identify relevant projects, people, and opportunities for the Wikimedia Research and Science communities. This work is mainly for the community, and to lay the framework for more coordinated action.

Introduction

[edit]

Wikimedia projects are among the world's most valuable and popular sources of science information (Eveleth, 2013; Ford, 2020), including science-related articles in Wikipedia (Economist, 2021; Heilman et al., 2011; Gherkin,2010; Penev et al., 2011), the reuse of general reference Wikidata datasets for off-wiki research (Mehdi et al., 2017; Arroyo-Machado etal., 2020; Bragazzi et al., 2017; Cao et al., 2020;Falk & Hagsten, 2022; Nielsen, 2007; Rasberry etal., 2022; Rasberry & Mietchen, 2024), and use of Wikimedia Commons media for illustrations in global media of all sources (Erickson et al.,2018). Despite this, the Wikimedia Movement lacks a narrative of its successes in the sciences, a profile of the hundreds of past projects, an accounting of the tens of millions of dollars of funding which scientific Wikimedia projects have solicited outside of Wikimedia Foundation donations, and the overall documentation infrastructure which would enable engagement between science and Wikimedia and coordination between science related affiliates. Who is doing what, exactly? How can we help? Even while there is an overall lack of coordination, there exist many initiatives across academic (Buttliere et al., 2024; Jemielniak &Aibar, 2016; Shafee et al., 2017), educational (Ackerly & Michelitch, 2022; Friesen & Hopkins, 2008; Konieczny, 2016; Konieczny, 2012; Lim,2009; Mkrtchyan, 2021), scientific (Buttliere, 2014; Severo, 2019; Teplitskiy et al., 2017; Shafeeet al., 2018; Waagmeester et al., 2020), and professional (Davenport, 2015; Duncan, 2020)domains, so many that it is difficult to actually know all of who is doing what, or even what isbeing done actually, even for the most well connected of Wikimedians. Having a better understanding of who is doing what, and especially who is funding that work, also helps us better organize and makes it easier to take up projects together. There are many people working on similar topics but often they do not know each other in order to coordinate and collaborate on relevant grants.

Background

[edit]

Wikimedia projects are among the world's most valuable and popular sources of science information (Eveleth, 2013; Ford, 2020), including science-related articles in Wikipedia (Economist, 2021; Heilman et al., 2011; Gherkin, 2010; Penev et al., 2011), the reuse of general reference Wikidata datasets for off-wiki research (Mehdi et al., 2017; Arroyo-Machado et al., 2020; Bragazzi et al., 2017; Cao et al., 2020; Falk & Hagsten, 2022; Nielsen, 2007; Rasberry et al., 2022; Rasberry & Mietchen, 2024), and use of Wikimedia Commons media for illustrations in global media of all sources (Erickson et al., 2018). Despite this, the Wikimedia Movement lacks a narrative of its successes in the sciences, a profile of the hundreds of past projects, and a good understanding of how all that research was funded, especially funding which scientific Wikimedia projects have solicited outside of Wikimedia Foundation donations, and the overall documentation infrastructure which would enable engagement between science and Wikimedia and coordination between science related affiliates toward coordinated actions.

Who is doing what, exactly? How can we help?

[edit]

Even while there is an overall lack of coordination, there exist many initiatives across academic (Buttliere et al., 2024; Jemielniak & Aibar, 2016; Shafee et al., 2017), educational (Ackerly & Michelitch, 2022; Friesen & Hopkins, 2008; Konieczny, 2016; Konieczny, 2012; Lim, 2009; Mkrtchyan, 2021), scientific (Buttliere, 2014; Severo, 2019; Teplitskiy et al., 2017; Shafee et al., 2018; Waagmeester et al., 2020), and professional (Davenport, 2015; Duncan, 2020) domains, so many that it is difficult to actually know all of who is doing what, or even what is being done actually, even for the most well connected of Wikimedians. Having a better understanding of who is doing what, and especially who is funding that work, also helps us better organize and makes it easier to take up projects together. There are many people working on similar topics but often they do not know each other in order to coordinate and collaborate on relevant grants.

Project Goal

[edit]

The goal of this grant is to understand what the Wikimedia Research and Science Communities have done, who did it, and importantly, who funded it; as a part of our longer term efforts to replicate and build on these successes and engage large/ state funders with Wikimedia.

[edit]

Our team has been working in a coordinated manner for the last years on getting Wikimedia taken more seriously as a science communication platform. Academics have the choice to choose what they do and the expertise that Wikimedia projects want and need, but there have been few coordinated efforts to engage academics with Wikimedia. The academics that do engage often suffer both as a part of the Wikimedia community because they cannot be full time for Wikimedia, but also in the academic hierarchy because they are spending time organizing the Wikimedia community rather than publishing traditional academic output that is more visible there.

Our theory of change is that if academics can get professional recognition for developing open resources in the Wikimedia platform, then they would contribute more, and the scholarly reputation of Wikimedia would improve (Buttliere, Vetter, & Ross, 2024). This project identified many highly engaged networks of Wikimedia community members who contribute, but also found that the community is so large, varied, and undocumented that no one has a usable and actionable narrative of the scientific Wikimedia community. Thus this project is aimed at getting a really good understanding of what is going on, who is doing what, who is funding that work, and where we can reasonably look to create synergies.

One outcome of this work from the last year was an early version of the Research Persons database, also in partnership with Kinneret Gordon and the Wikimedia Research Community more generally. This was a quite good success, the project has by now over 140 signatories, but it is not as good as it could be, and it is certainly not exhaustive or as detailed as it could be. For instance, it could be better linked with ORCID, or be more explicit about opportunities where people can collaborate.

This project is intended to be a quite explicit search for high impact and interest projects, and especially at who or what entities are funding that research. This both serves the purpose of finding impactful actors, but also funders with an interest and record of funding good work.

Initiatives to understand who is doing what.

[edit]

An example of this work is Flavia Varella and her team, who are also collaborators of ours (but notably not on the research persons database), who recently did a 3 year community and strategy grant in a quite similar direction, but for educational initiatives (Varella & Figuredo, 2023). Whereas Varella et al. focused on Wikimedia’s educational initiatives, which has also led to a much stronger team in this area, and especially in Latin America, we intend to focus on Science initiatives. They found 13,000 educational initiatives in general. We believe a systematic and detailed study within science is warranted and will bring benefits similar to the way that the education community has benefitted from such. Here we want to go systematically through scientific publications about Wikimedia.

Look at who is funding this work.

[edit]

To our understanding, there has been little work done to explicitly look at who is funding Wikimedia related research. This also makes our work somewhat novel and potentially important to help fund future work to build the community or get Wikimedians research jobs.

Another example of this work upon which we expect to expand is the European research projects related to Wikimedia during 2021. The goal is to develop some understanding also for the Foundation so that it can more effectively fundraise on behalf of the community.

The need to get contributions taken seriously.

[edit]

One thing that all of this past work made clear is that the Wikimedia community could get more professional credit for this work and it would benefit the community. This also helps Wikimedians of all types meet their goals in terms of professional development.

In order to gain respect for Wikimedia contributions, and to organize individuals to apply more systematically for funds external to the Wikimedia Foundation, the field needs to establish a shared understanding and citable narrative which demonstrates what successes the community has achieved, and what funders have already recognized these successes with grants and sponsorship.

Methods

[edit]

This project is essentially a bibliometric review of papers, with the ultimate goal of developing a communal understanding of what is being done, who is doing it, and who is funding that work.

The core of the grant is to collect publications which study or are about Wikimedia-related topics, and then look at their bibliometric data, in particular the ‘sponsor’ variable as made available in Scholia, looking both for ways to improve the variable and at who is funding this work in general.

This data will be used as a basis not only to understand what is currently happening, but also from which to grow new projects by linking partners for grantmaking opportunities at the funders that we expect to identify.

Beyond creating these databases of projects, people, and funders, we expect the project to be a key update as to what exactly is going on with Wikimedia and Science/ Research, and a key part of additional projects to develop the community into a force for WikiScience.

The nature of the database can be explored at this link. The goal is to dig into the Scholia dataset of about 15,000 papers that include Wikimedia/ Wikipedia/ Wikidata or other closely related subjects as a topic and to see where there are clusters of projects, who is doing it, and who is funding it. This is a bibliometric analysis in the style of Turki et al. (2024) of the keywords and annotations in Wikidata on any amount of papers, but looking in particular at who is funding the work for the most important 200 works. Check in the acknowledgements for who or what grantmaker is funding it in the style of Hegde, Garg, Murray-Rust, & Mietchen (2022).

Key research questions this project answers

[edit]

This project addresses itself to many relevant questions for the research and community more generally:

  • What keywords and in what areas is the work? What exactly is being done?
  • What journals are publishing Wikimedia work?
  • Who are the main authors in different areas? Are there particular groups?
  • Who is funding this Wikimedia work?
  • What are the main universities?
  • How to bring these people together?

Looking systematically at the scientific output around Wikimedia will be of benefit to many areas of the community, especially in relation to future works to bring people together. Below we outline some of the specific questions and how we intend to tackle them while analyzing these data. The project will focus on academic papers and scholarly content published about Wikimedia, as identified by the Scholia tool. Scholia currently has a database with several thousands of papers in it.

The goal is to categorize what is being done, identify who is active, and what they are doing. The goal is to also bring people together so that in the future we can survey them as to what they want to do and recruit them for other projects.

While organizations like OpenAlex offer an overarching and broad modeling of these different communities, for instance with authorship networks, we would like to use these existing metrics and really do the work of reading those important papers and bringing those individuals into some community.

Analytical Questions

[edit]

The goal of the work is to identify funders which have engaged with Wikimedia before which will also lead us to inquire about the following questions:

What are the most cited papers in the area and what are the growing subareas in this field? Given the huge amount of papers about Wikimedia, we will first sort the papers by citation metrics, so that we can focus on understanding the most traditionally impactful works. A summary paper would make the full dataset available and have likely the top 10 papers (per subfield) in a table. Such data can also be used to develop syllabi.

Who are the most prolific authors? This is not a main analysis, but running a ‘quick and dirty’ network analysis will allow us to identify the key partners and people in this network. Using this list or knowledge, the research community could selectively activate people and bring them together for future action. The goal would be to identify key actors in that network as levers for change.

What journals are publishing Wikimedia work? This is a bibliometric analysis of which journals (or other venues, such as conferences) the papers are published in. This analysis will be useful for people in the area to know where they should best be targeting their papers and again useful to the community to know where to e.g., put in a special issue on some topic.

In what topics or areas is the most work being done? A keyword analysis in terms of frequency and co-occurrences. This will show the most common areas that works are being published in and the most common keywords. Already OpenAlex and Scholia have some rough analysis, but knowing who is doing what allows us to contact particularly those people.

Who is funding the work? The Key and most interesting analysis is who exactly is funding this work. Scholia has a ‘Sponsor’ variable that indicates the sponsor of the work, similar to e.g., the Web of Science’s funder variables. The idea will be to analyze and fill in the variable.

Creating a disambiguated list of funders that have sponsored Wikimedia work, focusing on the last decade. The first analysis will be to simply populate and create a list of funders or sponsors that have ever funded a Wikimedia related project.

How many of the papers about Wikimedia report any funding? The first analysis of this key sponsor data will be to see what % of the papers report any sort of sponsor for their work. While there are likely to be a large number of papers that have a sponsor even if it is not reported, this analysis will at least give a baseline to work from in the future when we ask people more explicitly to include their sponsors.

Who are the main sponsors? Probably the most important piece of data will be to look at the list of top funders, first in the number of papers but also in the dollar amounts if we can manage it. The number of papers funded should be relatively straightforward, but the sponsor variable is a text variable, most often with a project number, rather than a specific dollar amount.

What percentage of the works are Wikimedia sponsoring? One of the main goals of this project is to help Wikimedians seek external funds. To this end, we intend to look briefly at how many papers Wikimedia is funding vs. external funders.

Looking at the top 100 papers specifically. The sponsor variable should capture all of the data that publishers or authors indicate in their related grants section, but still we believe looking at the top top papers in particular will be useful. Thus, we intend to look at the top 100 papers in terms of citations, in more detail than the others, both to understand the nature of the relationships between sponsors and the Wikimedia research they support, and to identify trends.

Expected outputs

[edit]

This research grant is expected to result in a number of outputs as described below:

Database of science contributors.

[edit]

While the main goal is looking at who is funding this work, we also hope to be able to bring the research community together. To this end, we intend to develop and extend the existing Research Persons Database, particularly with information from Wikidata. Most probably, we will send an email to the corresponding authors on these papers, and invite them to get involved. While many of these emails may not be functional any more, we do expect to reach a significant number of authors this way, particularly of recent papers.  

[edit]

Ultimately, our long term goal is to facilitate getting grants and building networks within the Wikimedia community. This research project builds on recent successes to push this work in a more systematic and needed direction to get full coverage of the community.

A database of funders that have already engaged with Wikimedia Science.

[edit]

The ultimate goal will be to create some list of large scale funders that have already funded or engaged with Wikimedia. We expect this for instance to be national science funders, private charities, or etc. For instance, if we see that a particular national government has already funded much WikiWork or with a particular person, it can make sense for the community to bring them on board in more systematic ways.

Another benefit of this work will be that it will allow future Wikimedians or the community in general to look at those particular funders for particular projects.

Improvements to Scholia.

[edit]

Given that we are using the Scholia dataset and in some ways trying to improve it, and that two of our collaborators are key players in Scholia, it makes sense to add our contributions to their system. For instance, we will intend to contribute to their disambiguation protocol and their funding statement indexer how we can.

Presentations about the project at WikiVenues.

[edit]

We believe our project will be relevant for a number of Wikimedia Communities including the WikiData and Research conference, the WikiCite Conferences, and the WikidataCon series of events. This would be aside from more mainstream Wikimedia and Academic Conferences like Wikimanias, WikiConNorth America, the Metascience conference, etc.

Paper(s) presenting the results.

[edit]

One key outcome of this work will be a paper presenting what the Academic and Wikimedia Research communities are working on and calling for scientists to join us in these efforts. This project is a part of a longer term program of research focused on getting scientists engaged with Wikimedia.

[edit]

The ultimate goal of this project is to help the research community understand who is funding what types of work and also to help those in the community get funding in the future.


Risks

[edit]

We have identified the following potential risks in relation to this project.

  • People access the database and use it for malign purposes.
  • The Wiki community is not interested in organizing (into a hub structure).
  • The projects we organize fail.

Mixing incentives can cause problems, especially when people start doing it for work and looking for ways to get more for doing less. Still, we believe that overall, the risks of not doing this project and organizing ourselves into a community are far greater than organizing. The most significant risk is simply that the work takes much more time than originally anticipated.

Adding: GPDR management with data related to researchers. Using ORCID and professional contacts. Create a procedure to remove their names from lists.

Community impact plan

[edit]

The goal of our proposal is to understand what is actually being done in the field, and how we can help develop this action.

This project will help the community in a number of ways. In particular, we will have quite a comprehensive list of e.g., scientific authors that are doing some Wikimedia related Science related work, what is being done, and who is funding such work.

Understanding what is going on.

There are a wild number of projects, affiliates, user groups, and other organizations that are Wiki affiliated. So many that it makes it hard for many people to find their places or even know where to begin. Having an overview will help the whole community know what it wants.

Developing some organized directions to go in.

Having an understanding of what exactly is being done among all of the disparate groups and fields also allows us to understand what is important to the research community.

Helping people find collaborators

One of the biggest problems that we have experienced in the work so far is finding the right people at the right time to take advantage of a grant opportunity. Wikimedia is in a unique position of having people all over the world and with diverse backgrounds, which is a huge benefit for complex grants.

Fostering projects and bringing resources to the community. The ultimate goal is to help people get resources for the Wikimedia-related work they want to do. A major hurdle in this endeavor is finding the grants themselves and more importantly project partners with the proper expertise to do the parts that the team is not expert in, and to meet the grant criteria.

Evaluation

[edit]

Number of papers and projects found.

[edit]

The proximal goal of the project is to have a pretty complete overview of the Wikimedia Research Landscape. The goal is to have quite an exhaustive overview and categorize and understand these actions such that we have an overview of what is happening in the community. To this extent, the more projects and people we find, the better.

Number of champions identified.

[edit]

These initiatives are not actors in themselves, so the key will be to find particular people associated with these initiatives that we can potentially engage with when and where relevant. The ultimate goal will be to have a database of people who are interested in working on Wikimedia, especially for future granting opportunities that are international.

White paper presenting our understanding of who is funding what.

[edit]

The main outcome of our grant will be a white paper for other Wikimedia researchers and affiliates on who is funding this type of work. This paper will be submitted to a journal, but is expected to be primarily of interest to the Wikimedia (Research) community, so potentially it ends up in something like the WikiJournal, which is also an initiative we are keen to support.

Grants and further projects generated.

[edit]

Ultimately, the goal of these activities is to bring more resources to the community, thus helping people do the work they want to do without financial or time commitment struggles. To this end, the whole point of this project and trying to develop a community in general is to support grants in the community. By better identifying opportunities and the appropriate people to engage with on that project, we will help everyone do the work they want to do with the appropriate resources needed.

Budget

[edit]

The budget is here, but in general we plan to use most of the funds to pay for people to work on the project. This is justified because the work is explicitly for the community, it is community building rather than career building with the ultimate goal of helping professionalize contributing to Wikimedia in general (Buttliere, Vetter, & Ross, 2024).

The figures in the budget are gross gross, meaning before currency exchange, social insurances, and all sorts of taxes. Net pay will be more like 70% of the quoted figures. This money will be distributed among the authors corresponding with their workload, with most of it expected to go to the first authors (Buttliere & Vetter) who are expected to do the most work.

These are gross gross estimates also of realistic wages for the time of these professionals.

  • Selecting and downloading the data from Scholia, Web of Science, Open Alex, for papers that include ‘Wikipedia’ in the title or abstract:
    • ~1 week each system. $2,000 per system 6,000 USD
  • Scanning/ Collating the data. Collating authors and journals. Collating and analyzing funders. Developing reproducible R code for this.
    • 10 minutes per paper to find, skim, code the data: up to 500 papers = 5,000 minutes = 88 hours = 11 work days = 2 boring weeks 6,000USD
  • Reading/ mining high relevance papers in preparation for a group paper outlining the history and different areas of Wikimedia.
    • Authors reading ~10 papers: 1 paper = half day: 5 days per author (30 work days) 6 weeks: 12,000USD
  • Impacting the community: Reaching out to authors of high relevance. Putting these data into e.g., the Research persons database. Using the funder data to help Wikimedians get projects.
    • Regular meetings, open science meetings, conference submission/ presentation:  2 weeks: 5000USD.
  • Reporting/ writing the paper: Writing the report for the grant, preparing the paper for publication. Goal would be to invite the key network levers to join a paper about science for Wikimedia.
    • Writing the report for the WikiResearch project. 1 week: Writing the paper for a journal. 6 weeks: Going through the Review process. 2 weeks 9 weeks 15,000USD
  • 15% for the university/ fiscal sponsor: 5,500USD

In total we are estimating spending approximately 23 full time weeks, half a year, on the project across the six collaborators, and we are asking for 49,450USD in total for the project. We believe this project will have many downstream benefits for the community.

Results and open data

[edit]

Our project looked at papers published in Open Alex (Priem, Piwowar, & Orr, 2022), one of the largest open access bibliometric databases, for papers related to Wikimedia between 2015 and 2024 (the last completed year before the research project began). This is also one of the few open access bibliometrics resources that also makes available data about who funded the work. In an initial analysis comparing other bibliometrics databases, for instance Web of Science and Scholia, we found OpenAlex to be the best fit for purpose, it had the most papers, it had the most complete funding data, and it was open access.

In order to identify papers in these bibliometric databases, we used the following search within titles and abstracts: “Wikipedia OR Wikimedia OR Wikidata OR Wikisource OR Wikinews OR Wikibooks OR Wikispecies OR Wikivoyage OR Wikinews OR Wikisource OR Wikiquote OR Wiktionary OR Wikiversity OR Wikimania OR MediaWiki OR Wikifunctions”

We found 9,137 records from Open Alex, which also was the largest sample across the databases, perhaps because it has more preprints and conference records. All of the analyses reported here were done in R, and included the use of stringr, tm, MASS, psych, SnowballC, wordcloud, RColorBrewer, psych, quanteda, dplyr, stringr, tidyr and jsonlite. All data and the analysis script to create these results or to run additional analyzes are available here. We would also encourage you to directly use Open Alex to look at these, or to download your own set of data, which can be more particular for your use case.

All of our datasets and analysis scripts are hosted on https://osf.io/6e9as - Who funds research on Wikimedia projects? Initial datasets.

Starting with 9,137 records, we removed duplicates based on the title, removing 191 papers. One important caveat to remember here and across all of our analyses are that we are only in this case removing exact duplicates of the title, meaning that if the title was changed even if it is the same project, those projects will still be in the dataset. This leaves us with 8,946 papers with unique title records. Given that some analyses compare how the number of publications change over time, we next removed 207 papers from 2025, because the data were collected part way through the year, making the results misleading unless one is careful. Thus we have 8,739 papers which we will analyze further. How many Authors are working on Wikimedia related work?

One initial question we had was how big actually the Wikimedia Research space is, which can also be represented by how many ‘unique’ authors exist within the Wikimedia Research Space since 2015. By splitting the authorship lists and having the computer count how often each unique string occurs, we can estimate not only how many authors there are, but also how many papers each author wrote. This analysis finds 18,454 unique authors on those 8,739 papers.

Again it is important to keep in mind that we are only counting exact matches, so even if there is a middle initial or not, those names will not be counted as one. A small test of the uniqueness of the names shows that the data are actually quite clean. When we looked at the top authors (who are also the most likely to have multiple names in there), we found that most were in the top choice. For instance, if you look for one of our co-authors Matthew Vetter, 12 of his 16 (75%) are assigned to Matthew A Vetter, and 4 are assigned to Matthew Vetter. This is similar across a number of examples, and it indicates that while the counts of how many papers might be slightly unreliable at the very top 1% of contributors, this doesn’t matter that much to the conclusions because the ranking is not exactly crucial - we simply want to know who is in the top, what they are working on, and etc, which this data accomplishes.

At approximately 20,000 unique authors, it is clear that there is a substantial community of researchers who have written something that centers Wikimedia or its sister projects. But even at 20,000 authors, with an estimated 20 million active authors over the last decade, we are talking about .1% of authors i.e., 1 in 1,000. This is still a small fraction of researchers overall. There is great room for improvement, especially given the very interdisciplinary and international nature of Wikimedia. Average citations in the database

In order to understand whether working on Wikimedia related activities is academically valuable as a topic or not, we decided to look at the average citations. The papers are cited 11.16 times on average. The most cited paper in the dataset is cited 6,113 times; which is the ‘SQuAD: 100,000+ Questions for Machine Comprehension of Text’ paper. This paper from Stanford presents a dataset of questions and answers that are available in corresponding texts, which is often used as a baseline in machine learning and comprehension tests (Rajpurkar, Zhang, Lopyrev, & Liang, 2016). Only one other paper is cited over 3,000 times, Lehmann et al., (2015) paper presents Dbpedia, another large scale database often used in machine learning. Five papers are cited more than 1,000 times, all of which are about machine learning and understanding, and 136 of the 8,739 papers are cited more than 100 times.

Looking at the normalized percentile value puts this number into context and gives an indication of how cited work on Wikimedia is to an overall, standard, average. The average paper about Wikipedia scores at the 44.05th percentile, with the median being at the 52.57th percentile. These are both significantly different from an overall average of 50%, the expected average in the population of papers. The heavy left bunch and still ok amount in the top scores are indicated in Figure 1.

Figure 1 normalized citation percentile.png

Looking further, in our dataset, 3,508 of the 8,739 papers have 0 citations, which is approximately 40% of the papers in the dataset. 7,154 of the papers are cited less than 10 times, meaning that only 1,585 of the papers are cited more than 10 times, with 11 being the average. To put this into a percentage, 19.2% of the papers are cited more than 10 times. This data indicates that the average paper about Wikimedia does worse than the average paper, but there is a cluster of highly cited papers as well.

This speaks to a need to increase the activity and especially citations within the community.

Table *: Number of papers per citation count up to 10 citations on the paper. 7,154 of 8,739 papers overall (81.86%) have less than 10 citations. Average is 11.16.

Citations per paper, tabled

More Than Half of the Papers are Closed Access

[edit]

One potential explanation for the relatively fewer citations is that so many of them are closed access. Of the 8,739 papers, 4,775 (54.64%) of the papers were closed access, 978 (10.50%) were hybrid open access, 814 (9.31%) were bronze open access, 540 were green open access, 1,160 were gold open access and 472 were diamond open access. These effects might be explained through multiple factors, for instance the adoption rate of open access from 2015 to 2025. While this study cannot answer all of these questions in general, it is somewhat ironic that papers about an open project are more than half the time closed access. It would be a worthy goal for the community of researchers working on Wikimedia to have at least 50% open access publications over the next ten years compared to the last.

Most Wikimedia Research is Published in English

[edit]

Of the 8,739 papers published, 7,640 (87.4%) of them were published in English. Spanish is the second most common language, with 244 papers in Spanish. This is a huge difference (> 31x) and suggests that Wikimedia Research, or at least research being published and recorded in this database, is primarily anglophone. This might limit our conclusions and utility for e.g., Latin American funders.

The main languages of Wikimedia Research

Where is this work being done?

[edit]

While the work is primarily in English, the country distribution is much better. It is important to recognize that these numbers are derived from authorships data and not aggregated at the author level. That means, if there are 10 papers written by the same author in Germany, they are counted in these data 10 times, since DE shows up in the authorship list 10 times. USA, China, Germany, and India, are the largest countries, and the only ones having more than 1,000 each. The long list of countries with authors working on Wikimedia is a good indication of the worldwide nature of Wikimedia. Still, as shown above, less than 1% of scientific authors have written something about Wikimedia since 2015.

Wikimedia research papers published per country.

What institutes have been involved?

[edit]

It is one thing to ask where in the world are these projects being produced, but another to identify the specific institutes. By looking at the author’s affiliations we can identify ‘where the work was done’. Once again we need to keep in mind the caveat that this is the number of times the university shows up in the authorship data at all. Thus, this is not unique authors at e.g., Carnegie Mellon, but rather saying that Carnegie Mellon had 195 authorships affiliated with it. It is notable and interesting that Google has 157 authorships about Wikimedia projects while the Wikimedia Foundation has 95 authorships about Wikimedia projects. In terms of conclusions that could be drawn from this data, this is where Wikimedia based researchers should probably be looking at jobs, for instance. This is also where the Foundation or community can look to create connections which will link the names or scale initiatives. If the goal is to start some international consortium or Center for Wikimedia Research, these are the places where it most likely will be easiest to initiate.

Authorships per institute

What Open Alex defined topics are being studied?

[edit]

Another way to look at these data is to ask what are the main topics that are being studied by Wikimedia researchers? This variable is applied by Open Alex. The largest topic was the role of ‘Wikis in Education and Collaboration’, with 3,084 papers labeled with the topic over the 10 years. The next topics are all essentially computer science, indicating the important computer science background that Wikimedia enables. 890 papers were written about ‘Topic Modeling’, 712 papers about ‘Natural Language Processing Techniques’, 116 papers are labeled with ‘Advanced Text Analysis Techniques’; the rest are smaller from there. It is noteworthy that 169 papers were not labeled with any topics.

Main display topics

Wikidata Concepts or Keywords of the Papers

[edit]

A more detailed way to consider the topic could be to look at the ‘concepts’ in Open Alex, which are more like keywords than one general descriptor. These concepts appear to be linked to Wikidata concepts. At least, the concepts listed here have Wikidata links to concepts. Table * contains the top 100 Open Alex concepts in the dataset. Notably, the average paper contains 12.52 concepts, with a maximum of 42 concepts assigned to a single paper, so the numbers are higher.

Min. 1st Qu. Median Mean 3rd Qu. Max.

  1.00    7.00   13.00   12.52   17.00   42.00

There are 109,434 individual concepts in the dataset (i.e., the sum of all keyword occurrences). Computer Science has 6,997 uses, double the next World Wide Web with 3,372 uses. Artificial Intelligence has 3,238 uses and Philosophy has 2,839 uses. 159 keywords are used more than 100 times. 892 concepts are used more than 10 times.

There are 5,865 unique concepts mentioned in the dataset. 2,734 of them are used only a single time. There is a lot of analysis that can be done here, but for now our main focus is to understand which funders are doing what types of work and we won’t go farther.

Main Wikidata topics assigned to the papers.

Journal Analysis - where is WikiResearch published?

[edit]

Across the 8,937 papers, we found 3,079 unique journals, again attesting to the large dispersion and many fields where Wikimedia works. 2,755 papers have no noted journal, which is the most popular category and likely is related to preprints. Another obvious result is that many of the largest outlets are either preprint repositories or conference proceedings, with specific conferences appearing to have a Wikimedia theme, for instance, ‘Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing’ which has 82 papers associated with it. This effect can be interpreted as either the result of so much of the work being focused on computer science, or as evidence of the difficulty of getting academic work about Wikimedia published. If papers are getting stuck between conference and publication, this is another reason to look and understand which journals are publishing on the topic. One notable take-aways not available by looking at the most popular journals are that the Nature group appears to be more open to publishing Wikimedia related content than Science, with Nature having 9 papers and others at subjournals, but with Science publishing only 3 papers and PNAS having 5 papers in the dataset.

Another take away here might be to put together a series of special editions related to Wikimedia in some of these journals. In any case, as was demonstrated before, this acts as a guide to Wikimedia related authors on where they should be putting their papers. Using our dataset, you can also match papers and keywords to journals, to get a further indication of for instance which terms to use in submissions.

Twenty- five of the papers report themselves as having a deleted journal, which is different from it being blank in general. Looking at these 25, there does not appear to be any pattern of them being from any particular journal, and many of the texts are still available at the DOI link. One thing we did notice is that all of the papers were originally hosted on CrossRef, maybe this has something to do with it. At this point 25 out of 8,739 is not our main concern, so we put it aside in the interest of time and attention.

Main journals where WikiResearch is published

Looking at this list, it is pretty clear that there is not a mainstream high impact venue through which many Wikimedia papers are published. PloS ONE has published 55 papers mentioning Wikipedia in the title or abstract, but this is not quite an A class journal. Journal of the Association for Information Science and Technology is quite a good journal but still not really of general interest. One notable missing set of journals are WikiJournals, which one might think are more open to publishing Wiki related things, but which are not really being engaged with by the Wiki Research community. This is likely something which should be developed.

Funder analysis - Who is funding Wiki Research?

[edit]

Our main goal was to look at who is funding Wikimedia Research papers, also to gain some strategic intelligence as to where we might encourage Wiki Researchers to look for funding in the future. The first and most notable outcome of the study is the large number of papers which do not have any reported funder; 8,091 (92.6%) of the papers do not have any reported or recognized funder. We looked at a random sample of 10 papers that are gold open access and do not indicate a funder and we found that 7 of the 10 truly had no funding data available, two had funding data that we think probably should have been picked up, and one paper that was quite ambiguous so it makes sense that there is nothing picked up. While a 20% miss rate is regrettable, there was no opportunity to check across all of the 8,091 papers that had no funding reported and try to correct them.

These suggests that either many papers are being done without any explicit funding, which matches with the volunteer nature of the Wikimedia projects, or that there is not a culture of indicating funders for work, most likely in general. The idea that there is or at least was not a culture of indicating who is funding the work is also made obvious by looking at how many papers the Wikimedia Foundation Funded. That is, there are 7 papers in the dataset for which WMF is named as a funder; we did not look at all of the research projects that WMF funded, but we guess more than 8 papers came out of those projects. Additionally, and in looking at those papers, we found one example where the WMF is only acknowledged by the authors for providing access to the data, rather than for explicit funding.

The 648 papers which have a recorded funder were cited more often than the average paper in the dataset, having on average 17.93 citations (overall average 11.16), and being in the 72nd percentile, again remember that the average paper was in the 42nd percentile. Still, none of the top 11 cited papers report any funding, with the 12th reporting funding from the National Science Foundation in the USA. Only 7 of the top 50 report funding in a way that is recognized or catalogued by OpenAlex. In order to understand how many papers each funder funded, we could not simply count how often each funder occurred, because some papers had multiple funders, and sometimes the same funder was mentioned multiple times on the same paper, if they funded more than one person on the project. In this case we had to split the funders if multiple funders existed, and also remove multiple instances of funders on the same papers.

Of the 648 papers that report a funder, after splitting and counting how often each unique name occurred, we found 313 unique funders, an average of 2.07 projects per funder on average. While this suggests that funders are coming back on average, 200 of 313 have only a single project that they funded, and 47 have only 2 projects that they funded. 39 funding agencies had funded more than 5 papers. Table * indicates those funders that have funded more than 10 papers. With 117 papers funded, the National Natural Science Foundation of China (NNSFC) is the largest funder in the dataset, followed by National Science Foundation (NSF) from the US in second, with 105 papers. The National Science Foundation of the USA was named on 105 papers, with Horizon 2020 (an EU program) being named on 35 papers. The largest private foundation was the Alfred P. Sloan foundation funding 10 papers. The Wikimedia Foundation was named as a funder on only 7 papers between 2014 and 2024. Table * makes available those funders which funded more than ten papers, along with how many times they were mentioned at all and also their most commonly funded keywords. The full dataset of 313 funders, along with their most commonly used keywords, can be found at this link.

Aside from the number of papers published, another way one could look at the data is to see how many times a funder is mentioned at all, since a funder can be named multiple times on the same paper. That is, if there are 5 authors on a paper and all of them are funded by a particular organization (and they all acknowledge that), then that funder name will be in the dataset 5 times. In this sense we are measuring the number of authors that are sponsored, rather than the number of papers sponsored. The notable increase for the NNSFC in the number of mentions relative to the number of papers suggests that the NNSFC is either funding more people per project, or is more consistently being added for each author, as compared to one blanket statement for the team. One notable finding from the data is that there are not many private foundations (except Alfred P Sloan Foundation) that have recognized funding within the Wikimedia research landscape. This seems like a major area of potential growth. At a very minimum, the Wikimedia Research team and community can do more to make sure that we are giving credit to our funders. Funder by keyword analysis

To go beyond simply who is funding Wikimedia Research, we wanted to understand who is funding what sorts of studies. In order to accomplish this, we created one row in the dataset for each paper by funder pair (creating 2 lines if there were 2 funders on the project), gathered all of the papers for each funder, and counted how often each keyword occurred. From this data we can get a sense of how much funding each funder put into each topic; or more generally, who is funding what type of work. In Table *, we present only the top ten concepts, and only for those funders that had at least ten projects.

One can see that most of the funded work is related to computer science, topic modeling, and AI. This is interesting and in contrast to the main topics on the papers, which were mostly labeled as Wikimedia in Education. This might also indicate that some work is getting funded and others not, for instance Wiki in the classroom might be less likely to get funded.

While the top lists are quite similar, looking across them can also reveal some neat findings. For instance, One thing that can be seen is that all China, the USA, and other top funders have Artificial Intelligence as top keywords, but the European Union’s main funding instruction, Horizon 2020 does not. There has been much discussion about the need for Europe to catch up in the AI race, which can be seen in the data.

This keyword data is highly interesting data and rich for additional reuse, especially if one is applying to a particular funding body. One can create clusters, word clouds, and networks using these data. Another way to use this process would be to look at particular journals, or institutions, at the types of work they are publishing. Another way to look at these data would be to go by the funder, see which authors are being funded, and to speak to them about their experience and recommendations.

Top funders and the main keywords they are funding.

National Natural Science Foundation of China Analysis

[edit]

One unexpected finding which raised particular questions was the fact that the National Natural Science Foundation of China (NNSFC) is the most mentioned funder in the dataset. We looked more closely at these papers to see what and why. We found that NNSFC was mentioned on 117 unique papers. They are almost exclusively focused on computer science. The most cited paper by NNSFC is 'A correlation comparison between Altmetric Attention Scores and citations for six PLOS journals’ cited 109 times, which also puts it in the top 1% of papers overall. In fact that paper is not centered on Wikimedia. Number 2 paper is ‘SEthesaurus: WordNet in Software Engineering’ with 52 citations. These papers are cited in, Journal of the Association for Information Science and Technology and Journal of Information Science respectively. The highest relevance score papers include, ‘Assessing the quality of information on wikipedia: A deep-learning approach’, ‘A domain knowledge graph construction method based on Wikipedia’, ‘Transforming Wikipedia Into Augmented Data for Query-Focused Summarization’, ‘Concept over time: the combination of probabilistic topic model with wikipedia knowledge’, ‘An efficient approach for measuring semantic relatedness using Wikipedia bidirectional links’.

These are papers related to computer science, using wikipedia data, most often to do topic modeling or query answering based on similar modeling. It is also interesting that the papers often indicate multiple authors funded by the NNSFC, which also accounts for the larger increase in mentions relative to e.g., the NSF or other funders. This effect in general could again be a difference in the expectations, where they are better about identifying their funders.

Top Authors in the WikiSpace

[edit]

We also wanted to look at the top authors within the Wikimedia Research area. Overall we found 18,454 authors, which also provides some context for the potential size of the Wikimedia Research Community. Of these, 298 or 1.6% of authors had written 5 or more papers. Only 48 (.2%) of the 18,454 had written more than 10 papers, which we suggest is some barrier for ‘making a career out of it’.

In the overall sample, 83% had written only 1 paper. 93.6% have written 1 or 2 papers. Only 6% have written 3 or more papers. This suggests that while there is a large community of Wikimedia Researchers, with many people having written 1 paper, very few were writing many papers. 5 papers between preprints, conference publications, and journal publications is not that many, but only 1.2% of WikiResearchers had done it.

In the process of analyzing these data, we looked at how ‘clean’ these data are, for instance if there are two versions of the same person’s name in the dataset. While these data are quite clean, they are not perfect, and we found several repeated profiles for especially the authors with many papers. Among our own authors and in the names we checked, we found a 75% to 80% hit rate, meaning if the author had 15 papers, 3 were logged under different profiles than the main one. This is encouraging for the data, especially because how many papers the person has written is not the main interest of our analysis, and if that person comes up in any funder or keyword analysis, both versions are likely to come up in cases where it is of interest. Because there are 18,454 unique author names and only a few hundred are in the dataset more than twice, we decided not to pursue further cleaning of the names.

When looking at the data in general, it should be understood that these are the people that should be forming the core of the WikiScience initiative. Many of them are by now in significant positions. Markus Strohmaier runs GESIS, one of the larger research organizations in Germany. Ulrike Cress is also the head of a Leibniz Center for Knowledge Media in Tubingen. Dariusz Jemielniak is the Vice-President of the Polish Academy of Sciences. If the research community is going to establish itself, bringing these people into the fold as mentors, sponsors, or etc would be of great benefit most likely.

Top Wikimedia Research Authors

Discussion

[edit]

This project makes a number of suggestions for the Wikimedia Community including e.g., where to send delegations, where to publish their work, and where they might be able to look for funding. Beyond the simple or example uses that we have demonstrated here, we have demonstrated and highlighted how the analyses could be combined to answer far more specific and interesting questions for your own particular case. For instance, one can look at who is funded in particular, at which institutes, and for doing what type of work. In this way one can very quickly create a short list of projects, funders, or institutes which make sense for further inquiry.

In general, and for our specific research question, the most consequential finding is that the majority of papers report no explicit funding, even when looking by hand at the paper itself. This is definitely something that can be improved upon within the Wikimedia Research Community which might also increase the chances of it being funded in the future. One next question is whether it is true that more than half of the papers are done without any funding what-so-ever, and or how we can improve the rate at which funding is reported.

Please take these data and use them; data also added to Scholia These data have a large number of uses, and it should be clear to the reader that they can do analyses like these not only using our data but in fact OpenAlex or any bibliometric database that logs funders. The analyses here should be seen as an overview and introduction, and a demonstration of what could be done using more specific keywords in your context. For instance, we showed which keywords funding bodies were funding. You can go far beyond this, to see for instance what actually are the papers, what types of projects are the projects in general, which ones of them were highly cited, who wrote those papers, where they are located, and of course with this information also email them to ask questions.

Another direction from which to go with the data might be to start at the country level, to see who is funding what in your country or region. This can be done for instance by selecting all of those papers done at an institution within that region (Table *), and then simply looking at the funding agencies, topics, authors, or institutions which are associated within that subsample. This process can essentially be done for any entity.

If you are a young scholar looking to work within a particular area, you could again select those papers which contain your keyword of interest, and look at which institutions, who is doing them, and also who is funding them. Another direction would be to see which journals are publishing work like yours. Or another who are the other people that you should probably know about at least within the Wikiworld, if not your topic or country. Again, we can only emphasize that this has tried to be an introduction to the data, rather than an exhaustive study of the possibilities, because only you know your situation. Limitations

These data are not elucidating fundamental truths, and we do not claim to have identified any strong conclusions based on these data alone. The data are to help us understand and answer questions like, which funders might be interested in my work, or which journals might be interested in publishing it? That being said, we can still show some limitations which could reasonably affect our conclusions.

First is that we are only using academic papers, and those academic papers that are available in Open Alex. Thus, we are probably leaving out many initiatives and stories, because they are not in the dataset. Additionally, and as was pointed out in an earlier round of review, There is a potential that we are missing academic papers written in different languages due to them not being indexed. This paper is meant to be an example and summary at least for English language papers, though we recognize the need and encourage researchers to look within their own preferred bibliographic databases.

Another limitation is the way that our dataset often relies on strictly similar snippets of text, as was shown in our author analysis. If the same paper has two slightly different titles, or an author has used two different versions of their name, the records will not be exact. We did go through the funders in a manual way, looking them up and filling in some information about them and rectifying discrepancies where we felt that the data are wrong but even this is not a sure proof method. There are more than 8000 papers and 18,000 authors, it is simply not possible to go through all of the names and rectify them within the amount of time that we had for this project. The best approach to solve this problem in your own situation is just to identify the key people, institutes, or funder that you want to look at more and then to manually look for other abbreviations. For instance if you are interested in different versions of the lead author’s name, one can easily look at their first name, and then see among all of the Brett s in the database, are there any others.

A final and perhaps most crucial limitation that we will examine here is the difficult state of the funder variable, which was in some ways the key of our analysis. We found 313 funders, which is really a lot and I would guess the longest list of at least science funders aside from perhaps WMF itself. Still, less than 10% had really indicated a funder in such a clear way that it is logged as a part of the largest open access bibliometric database in the world. We did look and found approximately 20% of our sample could have funders that are not recognized by the process that made the funder variable. Still, this would be far less than half of the papers having any reported funder at all. This can be a symptom of how Wiki Research works, on volunteer time similar to normal editing, or because people are simply not reporting their funders. People not reporting their funders could be the case, because also we found that the Wikimedia Foundation is only recognized 7 times in the dataset, and one of the times they were not thanked for funding, but for access to the data.

Overall, while our method seems to have done a quite ok job of identifying when there is funding on science papers, it is first limited to those who report their funding, and second only science funding. We are probably missing many grants for instance to Wikimedians-in-residence or libraries that do not write papers. One thing that is positive about our sample is that it is probably unique from the general list that WMF keeps https://wikimediafoundation.org/who-we-are/annualreport/2022-annual-report/donors/. Our list is different not only because these are funders that funded science papers, but also because they did not give to WMF itself, but most often rather to a person in the community. In this sense we think the dataset can offer an interesting addition to WikiResearch and community knowledge about funding.

Conclusion

[edit]

There is a large community, more than 15,000 unique authors, who have written a paper where a Wiki project is centered enough to be in the title or abstract of the paper. These authors are across a wide range of countries and disciplines. There do appear to be some signs that the work is not being taken as seriously as maybe it should, with very high profile journals not having many papers, and many papers being preprints or in conference proceedings. We found more than 300 funders from many areas of the world, but perhaps the more important line is that more than 90% didn’t have a funder indicated in the dataset or at all. This is definitely an area for improvement for the community, which could also help the community get more funds in the future. This list of funders can be especially valuable for the research community not only because they have all funded science related activity, but also because they gave to research projects rather than WMF itself. Let us all in the future put just a few minutes into better thanking our funders and with that we thank the Wikimedia Research Fund for making this work possible.

Dissemination

[edit]

Wiki Workshop 2026

  • Brett Buttliere, Matthew A. Vetter, Lane Rasberry, Iolanda Pensa, Daniel Mietchen, Susanna Mkrtchyan. Who is Funding Research Related to Wikimedia? A Bibliometric Study of Funders, Journals, and Authors.
Abstract: ::https://wikiworkshop.org/2026/paper/wikiworkshop_2026_13_who_is_funding_research_related_to_wikimedia_a_bibliometric_study_of_funders_journals_and_authors

Video Presentation: https://www.youtube.com/watch?v=2Wsix72Wy6Y

  • White paper currently being finalized to be published to OSF.


References

[edit]
  • Ackerly B, Michelitch K (2022) Wikipedia and Political Science: Addressing Systematic Biases with Student Initiatives. PS: Political Science & Politics 55 (2): 429‑433. https:// doi.org/10.1017/S1049096521001463
  • American Society for Cell Biology. (2012). San Francisco Declaration On Research Assessment (DORA).
  • Arroyo-Machado, W., Torres-Salinas, D., Herrera-Viedma, E., & Romero-Frías, E. (2020). Science through Wikipedia: A novel representation of open knowledge through co-citation networks. PloS one, 15(2), e0228713.
  • Bragazzi NL, Watad A, Brigo F, Adawi M, Amital H, Shoenfeld Y (2017) Public health awareness of autoimmune diseases after the death of a celebrity. Clinical Rheumatology 36 (8): 1911‑1917. https://doi.org/10.1007/s10067-016-3513-5
  • Buttliere, B. T. (2014). Using science and psychology to improve the dissemination and evaluation of scientific work. Frontiers in computational neuroscience, 8, 82. https://doi.org/10.3389/fncom.2014.00082
  • Buttliere, B., Vetter, M., Ross, S., (2024). Developing Wikimedia Impact Metrics as a Sociotechnical Solution for Encouraging Funder/ Academic Engagement. Wikimedia Research. https://w.wiki/BYix
  • Cao Y, Mehta H, Norcross A, Taniguchi M, Lindsey J (2020) Analysis of Wikipedia pageviews to identify popular chemicals. Reporters, Markers, Dyes, Nanoparticles, and Molecular Probes for Biomedical Applications XII. [ISBN
  • Ciocirdel GD, Varga M (2016) Election Prediction Based on Wikipedia Pageviews.Centrum Wiskunde & Informatica. URL: https://event.cwi.nl/lsde/2016/papers/group02.pdf
  • Clark J, Faris R, Heacock Jones R (2017) Analyzing Accessibility of Wikipedia Projects Around the World. Berkman Klein Center. https://doi.org/10.2139/ssrn.2951312
  • Coalition for Advancing Research Assessment (CoARA). The agreement, 2022. URL https://coara.eu/agreement/the-agreement-full-text/. Unofficial Report.
  • Davenport M (2015) Working With Wikipedia. American Chemical Society. URL: https://cen.acs.org/articles/93/i36/Working-Wikipedia.html
  • Digital Botanical Gardens Initiative Consortium. (2022). The Digital Botanical Gardens Initiative. Manubot. https://www.dbgi.org/
  • DigComp (2022). Digital Competence Framework for Citizens (DigComp). EU Science Hub. URL https://joint-research-centre.ec.europa.eu/scientific-activities-z/education-and-training/digital-transformation-education/digital-competence-framework-citizens-digcomp_en
  • DORA. San francisco declaration on research assessment. Technical report, DORA, 2012. URL https://https://sfdora.org/
  • Duncan A (2020) Towards an activist research: Is Wikipedia the problem or the solution? Art Libraries Journal 45 (4): 155‑161. https://doi.org/10.1017/alj.2020.24
  • Economist (2021) Wikipedia is 20, and its reputation has never been higher. The Economist. URL: https://www.economist.com/international/2021/01/09/wikipedia-is-20-and-its-reputation-has-never-been-higher
  • Erickson K, Perez FR, Perez JR (2018) What is the Commons Worth?: Estimating the Value of Wikimedia Imagery by Observing Downstream Use. Proceedings of the 14th International Symposium on Open Collaboration. [ISBN 978-1-4503-5936-8]. https://doi.org/10.1145/3233391.3233533
  • Eveleth R (2013) How Much is Wikipedia Worth? Smithsonian Institution. URL: https://www.smithsonianmag.com/smart-news/how-much-is-wikipedia-worth-704865/
  • Falk MT, Hagsten E (2022) Digital indicators of interest in natural world heritage sites. Journal of Environmental Management 324 https://doi.org/10.1016/j.jenvman.2022.116250
  • Ford H (2020) Rise of the Underdog. In: Reagle J, Koerner J (Eds) Wikipedia @ 20. URL: https://wikipedia20.mitpress.mit.edu/pub/fcgjp9ul/release/2 [ISBN 978-0-262-53817-6].
  • Friesen N, Hopkins J (2008) Wikiversity; or education meets the free culture movement: An ethnographic investigation. First Monday https://doi.org/10.5210/fm.v13i10.2234
  • Gerken J (2010) How Courts Use Wikipedia. The Journal of Appellate Practice and Process 11 (1): 191‑227. URL: https://lawrepository.ualr.edu/appellatepracticeprocess/vol11/iss1/8
  • Glammons (2024). Resilient, sustainable and participatory practices: Towards the GLAMs of the commons – https://glammons.eu/
  • GSRMI (2018). The Transformative Potential of Research in Museums https://www.leibniz-forschungsmuseen.de/gsrm-2022
  • Hegde, S., Garg, A., Murray-Rust, P., & Mietchen, D. (2022). Mining the literature for ethics statements: A step towards standardizing research ethics. Research Ideas and Outcomes, 8, e94685.
  • Heilman, J.M., Kemmann, E., Bonert, M., Chatterjee, A., Ragar, B., Beards, G.M., ..., Laurent, M.R. (2011). Wikipedia: A key tool for global public health promotion. Journal of Medical Internet Research, 13(1), e14. Retrieved from http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3221335/.
  • Jemielniak, D., & Aibar, E. (2016). Bridging the gap between Wikipedia and academia. Journal of the Association for Information Science and Technology, 67(7), 1773-1776.
  • Konieczny, P. (2016). Teaching with Wikipedia in a 21 st -century classroom: Perceptions of Wikipedia and its educational benefits. Journal of the Association for Information Science and Technology, 67(7), 1523–1534. https://doi.org/10.1002/asi.23616.
  • Konieczny, P. (2012). Wikis and Wikipedia as a teaching tool: Five years later. First Monday, 17(9). Retrieved from http://firstmonday. org/htbin/cgiwrap/bin/ojs/index.php/fm/article/viewArticle/3583/3313.
  • Lim, S. (2009). How and why do college students use Wikipedia? Journal of the American Society for Information Science and Technology, 60(11), 2189–2202.
  • McGranaghan, E., Klein, S., Cameron, A., Young, E., Schonfeld, S., Higginson, A., Ringuette, R., Halford, A., Bard, C., Narock, A., & Thompson, B., (2021). The need for a Space Data Knowledge Commons. https://knowledgestructure.pubpub.org/pub/space-knowledge-commons/release/2
  • Mehdi, Mohamad; Okoli, Chitu; Mesgari, Mostafa; Nielsen, Finn Årup; Lanamäki, Arto (March 2017). "Excavating the mother lode of human-generated text: A systematic review of research that uses the Wikipedia corpus". Information Processing & Management. 53 (2): 505–529. https://doi.org/10.1016/j.ipm.2016.07.003
  • Mkrtchyan, S. M., (2021). Education through Wikipedia. Mathematical Problems of Computer Science, 55, 62-58. https://doi.org/10.51408/1963-0074
  • Nielsen, F. Å. (2007). Scientific citations in Wikipedia. arXiv preprint arXiv:0705.2106.
  • Nielsen, F.Å., Mietchen, D., Willighagen, E. (2017). Scholia, Scientometrics and Wikidata. In: Blomqvist, E., Hose, K., Paulheim, H., Ławrynowicz, A., Ciravegna, F., Hartig, O. (eds) The Semantic Web: ESWC 2017 Satellite Events. ESWC 2017. Lecture Notes in Computer Science(), vol 10577. Springer, Cham. https://doi.org/10.1007/978-3-319-70407-4_36
  • Nosek, B. A., Alter, G., Banks, G. C., Borsboom, D., Bowman, S., Breckler, S., ... & DeHaven, A. C. (2016). Transparency and openness promotion (TOP) guidelines. https://osf.io/9f6gx/
  • Penev, L., Hagedorn, G., Mietchen, D., Georgiev, T., Stoev, P., Sautter, G., ... & Erwin, T. (2011). Interlinking journal and wiki publications through joint citation: Working examples from ZooKeys and Plazi on Species-ID. ZooKeys, (90), 1.https://doi.org/10.3897/zookeys.90.1369
  • Racheva, V., (2012). Sofia Zoo and Bulgarian Wikipedians/Sofia Zoo Powered by Wikimedia. https://meta.wikimedia.org/wiki/Grants:PEG/Sofia_Zoo_and_Bulgarian_Wikipedians/Sofia_Zoo_Powered_by_Wikimedia/Report
  • Rasberry, L., & Mietchen, D., (2024).
  • Readership of Wikipedia. ARPHA Preprints. https://preprints.arphahub.com/article/139375/
  • Rasberry, L., Tibbs, S., Hoos, W., Westermann, A., Keefer, J., Baskauf, S. J., ... & Mietchen, D. (2022). WikiProject clinical trials for Wikidata. medRxiv, 2022-04.https://doi.org/10.1101/2022.04.01.22273328
  • ReCreating Europe (2021) GLAM Definition https://recreating.eu/stakeholders/wp5-glam/,
  • Severo, M. (2019). Can Wikipedia serve as a citizen science tool? Building knowledge between amateurs and institutions. Wikipedia@ 20.https://wikipedia20.mitpress.mit.edu/pub/u9tt7i19/release/3
  • Teplitskiy, M., Lu, G., & Duede, E. (2017). Amplifying the impact of open access: Wikipedia and the diffusion of science. Journal of the Association for Information Science and Technology, 68(9), 2116-2127.
  • Shafee, T., Schenone, F., Sumter, M., Whalley, W.B., Pössel, M., Alexander, I., … & Häggström, M. (2018). The aims and scope of WikiJournal of Science. WikiJournal of Science, 1(1), 1. doi: 10.15347/wjs/2018.001
  • Shafee, T., Mietchen, D., & Su, A. I. (2017). Academics can help shape Wikipedia. Science, 357(6351), 557-558. https://doi.org/10.1126/science.aao0462
  • Turki, H.; Taieb, M. A. H.; Aouicha, M. B.; Rasberry, L.; Mietchen, D. (2024). Ten Years of Wikidata: A Bibliometric Study (PDF). Wikidata Workshop 2023. Proceedings of the Wikidata Workshop 2023. Athens, Greece. https://ceur-ws.org/Vol-3640/paper13.pdf
  • Waagmeester, A., Stupp, G., Burgstaller-Muehlbacher, S., Good, B.M., .. Su, A.I. (2020) Science Forum: Wikidata as a knowledge graph for the life scienceseLife 9:e52614.
  • Wikimedian in Residence Exchange Network. https://meta.wikimedia.org/wiki/Wikimedians_in_Residence_Exchange_Network
  • Wikimedia Foundation (2011). Editor Survey Report. Retrieved from http://upload.wikimedia.org/wikipedia/commons/7/76/Editor_Survey_ Report_-_April_2011.pdf. San Francisco: Wikimedia Foundation.
  • Wikimedia Foundation. (2023). Wikipedia statistics. Wikipedia. Retrieved from https://en.wikipedia.org/wiki/Wikipedia

Community feedback

[edit]

Please add any feedback or endorsements to the grant discussion page.

Give feedback on the discussion page