Strategie/Mehrgenerations-/Künstliche Intelligenz für Redakteure
Dieses Strategiepapier wurde im April 2025 veröffentlicht und von Chris Albon und Leila Zia von der Wikimedia-Foundation verfasst. Es ist auch auf Figshare und Wikimedia Commons verfügbar.
Zusammenfassung
Die Gemeinschaft der Freiwilligen ist der wichtigste und bestimmende Erfolgsfaktor von Wikipedia – einem Vorbild für weltweite enzyklopädische Wissensverwaltung. Seit über einem Jahrzehnt entwickeln und nutzen die Community und die Wikimedia-Foundation (WMF) künstliche Intelligenz (KI) für Bots, Produkte und Funktionen.[1] Die jüngsten Fortschritte im Bereich der KI und die zunehmenden Herausforderungen bei der Moderation eines komplexen Wissensökosystems haben uns veranlasst, eine Strategie zu entwickeln, um das Potenzial der KI weiter auszuschöpfen und ihre Risiken für die Projekte zu minimieren. Diese Strategie konzentriert sich darauf, KI zur Unterstützung der Arbeit von Redakteuren in Bereichen einzusetzen, in denen KI den größten Einfluss auf die Projekte haben kann: Automatisierung von Routineaufgaben, die kein menschliches Urteilsvermögen oder Diskussionen erfordern, Einarbeitung und Betreuung neuer Redakteure, Zeitersparnis für Redakteure, die sich auf lokales und spezialisiertes enzyklopädisches Wissen in verschiedenen Sprachen konzentrieren, und Vorrang für Schutz der Inhaltsintegrität – alles unter Beibehaltung der menschlichen Handlungsfreiheit. Indem die Arbeit der Redakteure in diesen Bereichen gezielt und substanziell durch KI unterstützt wird, zielt diese Strategie darauf ab, Wikipedias Position als vertrauenswürdige, von der Community getragene Wissensquelle für kommende Generationen zu erhalten und zu stärken.
Projektrahmen
Diese Strategie wird die allgemeine Richtung der Entwicklung, des Hosting und der Nutzung von KI für Produkt, Infrastruktur und Forschung der WMF vorgeben, zum direkten Nutzen der Bearbeiter der Wikimedia-Projekte. Wir konzentrieren uns dabei insbesondere auf Wikimedia-Projekte, die für den Austausch enzyklopädischen Wissens als wichtig erachtet werden.
Mit der oben genannten Definition möchten wir ausdrücklich darauf hinweisen, dass folgende Punkte nicht Gegenstand dieser Strategie sind: 1) Die KI-Förderung der WMF sowie die Aktivitäten ihrer Partner und Freiwilligen im Bereich KI; 2) KI-Projekte der WMF außerhalb der Wikimedia-Projekte und der von ihr betriebenen Technologieplattformen; 3) Strategische Empfehlungen für die Interaktion der WMF mit Technologieunternehmen, die Wikimedia-Inhalte zur Entwicklung eigener KI nutzen; 4) die Nutzung von KI für organisationsinterne Anwendungen der WMF.
Ausblick
Diese Strategie zielt darauf ab, die Arbeit der Organisation zwischen dem 1. Juli 2025 und dem 30. Juni 2028 innerhalb des oben genannten Rahmens zu lenken.
Aktualisierungsrhythmus
Der Forschungs- und Entwicklungsraum in der KI ist sehr dynamisch. Wir empfehlen, diese Strategie einmal im Jahr zu überprüfen und ggf. zu aktualisieren. Bei großen Durchbrüchen oder Veränderungen werden wir diese Strategie auch abseits des Jahreszyklus überdenken.
Ziel
KI bietet Möglichkeiten für viele Aspekte der Wikipedia-Nutzung. In dieser Strategie empfehlen wir Entwicklung und Hosting von KI-gestützten Technologien besonders für:
- Einarbeitung neuer Autoren.
- Motivation der vorhandenen Autoren, weiterhin zur Enzyklopädie beizutragen, indem ihre Arbeitsbelastung reduziert und ihre Beiträge auf Gebieten unterstützt werden, auf denen nur sie beitragen können.
- Die Position der Wikipedia als zuverlässigster Quelle für enzyklopädisches Wissen in weiteren Sprachen zu stärken.
- Dier Stiftung in der Entwicklung und Nutzung von KI mit menschlichem Ansatz führend zu etablieren, indem diejenigen Werkzeuge priorisiert werden, die Redakteure ermächtigen und nicht ersetzen, die die Werte der Gemeinschaft schützen, und die den Wissenszugang erleichtern.
Kernfragen
- Welche verschiedenen Strategien könnte die WMF im Bereich Artikelbearbeitung (Schreiben und Moderieren der Inhalte) in Bezug auf KI ausprobieren? Was sind die Kompromisse, die bei den verschiedenen Strategien eingegangen werden müssen?
- Welche der Strategien sollten wir verfolgen?
- Wie wird sich in dieser Strategie die Entwicklung, das Hosting und die Nutzung von KI mit unseren Grundwerten in Einklang bringen lassen?
- Wie werden in dieser Strategie Menschen mit KI interagieren?
- Wie sollten wir in dieser Strategie KI bei der Erstellung von Inhalten einsetzen? Wie sollten wir KI bei der Moderation von Inhalten einsetzen? Wie sollten wir die KI-Nutzung zwischen der Verbesserung bestehender Inhalte und der Erzeugung neuer Inhalte priorisieren?
- Welche Investitionen werden für den Erfolg dieser Strategie nötig sein? Welche Investitionen sollten abgeändert oder gestoppt werden?
Aktueller Status
Da sich das Internet sich weiter verändert und die KI-Nutzung zunimmt, erwarten wir, dass das Wissensökosystem zunehmend mit minderwertigen Inhalten, Fehlinformationen und Desinformationen kontaminiert wird. Wir hoffen, dass sich die Menschen weiterhin um hochwertiges, überprüfbares Wissen bemühen und nach Wahrheit streben. Wir wetten, dass sie sich dabei auf echte Menschen als Schiedsrichter des Wissens verlassen wollen. Wir glauben, dass Wikipedia das Rückgrat der Wahrheit sein kann, an das sich die Menschen wenden wollen, entweder bei den Wikimedia-Projekten oder bei der Nachnutzung durch Dritte.
Wikipedias Modell der kollaborativen Wissenserstellung hat seine Fähigkeit, verifizierbares und neutrales enzyklopädisches Wissen zu erstellen, unter Beweis gestellt. Die Wikpedia-Community und WMF haben seit langem KI zur Unterstützung der Arbeit der Freiwilligen eingesetzt, wobei der Mensch stets im Mittelpunkt sand. Wir nutzen unter anderem KI, um die Wikipedianer*innen dabei zu unterstützen, auf allen Wikipediaseiten Vandalismus zu erkennen, Inhalte für Leser zu übersetzen, Artikelqualität einzuschätzen, die Lesbarkeit von Artikeln zu messen und den Freiwilligen Bearbeitungsvorschläge zu machen. Wir befolgen dabei die Werte von Wikipedia in Bezug auf Community-Governance, Transparenz, die Förderung der Menschenrechte, Open Source und weitere Aspekte. Das heißt, wir haben KI in geringem Umfang in die Bearbeitungsprozesse integriert, wenn sich Gelegenheiten dazu geboten haben. Wir haben jedoch keine gezielten Anstrengungen unternommen, die Editiermöglichkeiten der Freiwilligen mithilfe von KI zu verbessern, da wir beschlossen haben, diesem Aspekt keinen Vorrang vor anderen Möglichkeiten einzuräumen.
Die neuesten Fortschritte in der Künstlichen Intelligenz haben zu neuen Möglichkeiten bei der Erstellung und der Nutzung von Inhalten geführt. Große Sprachmodelle (LLMs), die dazu in der Lage sind, Texte in natürlicher Sprache zusammenzufassen und zu generieren, eigenen sich besonders für Wikipedias Schwerpunkt auf schriftlich fixiertes Wissen. Das langfristige Potenzial dieser Techologien, qualitativ hochwertige, skalierbare Benutzererfahrungen zu schaffen, ist von großer Bedeutung und erfordert eine sorgfältige Erwägung. Gleichzeitig stellen dieselben Techologien Risiken für die Wikimediaprojekte und die Arbeitsabläufe der Wikipedianer*innen dar. Zum Beispiel können weltweit Menschen und Regierungen nun innerhalb kürzester Zeit tausende wikipediaartige Artikel erzeuen, die auf den ersten Blick valide erscheinen und nur schwer als Wikipedia-Hoaxes erkennbar sind. Während es extrem einfach und kostengünstig geworden ist, neue Inhalte zu erzeugen, ist die Überprüfung dieser Inhalte langsam und ressourcenintensiv geblieben. Die Überprüfbarkeit ist jedoch das Rückgrat der enzyklopädischen Arbeit. Wikipedianer*innen benötigen erhebliche Unterstützung der WMF, um angesichts dieser Chancen und Herausforderungen das Beste aus dem zu machen, was KI den Wikimediaprojekten zu bieten hat.
Lösungsmöglichkeiten
Bei der Entwicklung dieser Strategie haben wir viele Möglichkeiten untersucht und viele Kompromisse in Betracht gezogen. Unser Ziel war, den besten Weg für die Integration von KI in unsere Arbeit zu finden während unsere Werte eingehalten werden und wir zudem sicherstellen, dass die Wikipedianer*innen und Communities der Mittelpunkt des Projekts bleiben.
We first explored the option to support editors with AI incidentally. This is the status quo option that is closest to how we currently do research, product and feature development. We invest some resources on AI but not a lot. Following this path means that we will not respond to a changing internet[2] at the peril of the projects. As shared earlier, given how recent advancements in AI have made content generation easy, and given that the cost of verifiability remains high for editors, the editors will be at significant risk of overload and burnout. Maintaining the status quo risks leaving Wikipedia’s user experience behind the expectations of modern internet users and even more so with the next generation. While Wikipedia as a platform has traditionally evolved at a slower pace, the broader internet landscape continues to advance rapidly, setting new standards for usability, mobile-first design, interactivity, and accessibility. Without adapting to meet these expectations, we risk alienating both current and future users and diminishing Wikipedia’s value and relevance.
The second possible strategy would be to invest in AI knowledge generation over human knowledge generation. Advances in AI make it clear that in the coming years there will be increasing attempts at using AI for knowledge creation or curation by directly summarizing primary, secondary, and tertiary sources. There are obvious appeals to this approach for companies, such as efficiency and scalability. But there are also downsides, such as limited human oversight, vulnerability to biases, hallucinations, potential for misinformation and disinformation spread, limited local context, significant invisible human labor,[3] and weak ability to handle nuanced topics. Adopting this strategy has a further and more important risk. Since volunteers are the core and unique element of success for the Wikimedia projects this strategy can discourage existing volunteers, to the peril of the projects.
The strategies discussed above will not enable us to reach our goals for a multi-generational Wikipedia. Therefore, we recommend a third strategy: make a significant and targeted investment in supporting editors with AI. While companies are leaning away from human-created knowledge, we should lean in on the collective power of editors and use AI to assist them. Humans supported by AI will be more effective at generating knowledge than humans or AI alone. In addition, we propose using AI in areas where AI is uniquely positioned to support editors and advance the goals of the Wikimedia Movement. This targeted approach is important because it allows us to achieve the highest impact within the realities of our budget and resources.
Prioritized strategy and tradeoffs
Our prioritized strategy is to invest in AI to support editors in areas where AI can have a unique advantage over other technologies to solve problems of impact and to prioritize editors’ agency in interacting with AI. More specifically, we recommend investing in AI to support the editors as follows:
Prioritize AI assisted workflows in support of moderators and patrollers. The recent advancements in AI, particularly generative AI, have made content generation significantly easier and that introduces significant risk for the editors and the projects as validating content remains costly. Thousands of hard-to-detect Wikipedia hoaxes and other forms of misinformation and disinformation can be produced in minutes.[4] Therefore we should prioritize using AI to support knowledge integrity and increase the moderator's capacity to maintain Wikipedia’s quality. Overloading editors with the task of managing an influx of AI-assisted content risks burnout and compromises Wikipedia’s quality and existence. This focus on workflows for moderators and patrollers ensures Wikipedia remains a trusted source, allowing editors to do their work effectively.
Create more time for editing, human judgment, discussion, and consensus building. Editors spend a significant amount of time before they can edit Wikipedia. Part of this time is invested in finding the information they need for their editing, discussion, or decision making. AI excels at handling tasks such as information retrieval, translation, and pattern detection. By automating these repetitive tasks, AI frees up editors’ time to focus on areas of encyclopedic work that require human expertise: editing, discussions, consensus building, and making judgment calls in complex situations where the stakes are high and the impact is significant.
Create more time for editors to share local perspectives or context on a topic. Editors of less represented languages face pressure to create more content in their local languages. Automating the translation and adaptation of common topics[5] allows editors to enrich the encyclopedic knowledge with cultural and local knowledge and nuances that AI models cannot provide. This allows editors to invest more time in creating content that strengthens Wikipedia as a diverse, global encyclopedia.
Engage new generations of editors with guided mentorship, workflows and assistance. Editors drive knowledge curation and governance. For the projects to be multigenerational, new editors must encounter editing workflows that fit their expectations and must find effective ways to get help. AI offers opportunities for generating valuable types of suggested edits that make sense for a new generation. And generative AI in particular offers a promising solution for automated mentoring and guidance of newcomers. AI can provide personalized support, from retrieving information and understanding policies to giving feedback on edits, helping newcomers feel confident and capable.
How we will implement this strategy
Our implementation of this strategy is shaped by WMF’s vision, mission, guiding principles, privacy policy, human rights policy, 2030 Movement Strategy and Multigenerational pillars. Below we highlight the core tenets drawn or informed by these sources that should define how we implement this strategy.
- We adopt a human-centered approach. We empower and engage humans, and we prioritize human agency.
- We prioritize using open source AI technologies or open weights, and we develop only open source AI.[6]
- Our use of AI will allow editors to focus more on what they want to accomplish, not how to technically achieve it.
- We coordinate with Wikimedia affiliates and we invest in the distributed network of people, institutions, and organizations to contribute to this strategy.
- We prioritize transparency.
- We prioritize multilinguality in nuanced ways.
- We continue to offer a space where humans can share in the sum of all encyclopedic knowledge without the fear of persecution or censorship.
Trade-offs
In order to arrive at the above prioritized strategy, we faced trade-offs and we had to make choices. We will share more about them below. Note that implementing this strategy will require more trade-offs and choices to be made by us and other decision makers in WMF. We currently capture draft implementation trade-offs in the Appendix.
Content generation vs. content integrity. Our resources are limited and we cannot significantly support editors with AI for both content generation and content integrity at the same time. We made a decision to first prioritize using AI to support editors to assure content integrity. By doing so we want to assure that moderators and patrollers are adequately supported for any surge of new content on the projects. Our reasoning is that new encyclopedic knowledge can only be added to Wikipedia at a throughput that is defined by the capacity of existing editors to moderate that content. If we invest heavily in new content generation before moderation, the new content will overwhelm the capacity of the editors to moderate them. This balance might well shift in time as the needs of moderation vs. new content shifts.
Open source models vs. open weight models.[7] We commit to building open source models for AI. However, we have to acknowledge our resources are too limited to develop our own open source foundational model which would require thousands of new servers[8] and hundreds of thousands of hours of work by machine learning engineers and researchers. For this reason we have made a choice to use open weight models when necessary to build features to support editors. We hope open source foundational models able to compete on best practice evaluations are released in the future.
Use AI in many places vs. use AI for specific areas of impact. Our resources, even considering the collective resources in the community and the broader free knowledge ecosystem, are not sufficient (i.e. expertise, funds for infrastructure, etc.) for planning, developing, tuning, and using AI for myriad different applications without focus. We have made a choice through this strategy to limit the applications of impact by focusing on four main areas. We acknowledge that there is hype and excitement around AI. We expect that we will be asked to apply AI to more and more applications, which may create friction when balancing new proposals with our focused approach. This tension can create frustration among those advocating for other areas of impact and may require us to regularly revisit our prioritization. It can also slow down the work on this strategy as we may need to regularly revisit our prioritization.
Acknowledgements
This strategy brief is possible thanks to the contributions and input of numerous people who engaged with us between June 2024-February 2025. We recognize and thank these individuals below.
Throughout the process, Selena Deckelmann and Marshall Miller supported us in a variety of ways including re-scoping of the work to specifically focus on editors as well as providing extensive feedback, particularly to the early stages of strategy. Nadee Gunasena partnered with Selena and us to create spaces and opportunities for us to engage and gather input from different groups. Miriam Redi provided frequent feedback to the early stages of our thinking and work. These conversations had varying dimensions: from the importance of prioritizing “open and free” to prioritizing a sustainable symbiosis between Wikipedia and generative AI. We also would like to thank Isaac Johnson for supporting us early on in arriving to a more nuanced understanding of generative AI for Wikipedia and multilingualism; and for proposing the framework of using AI where AI is more uniquely positioned (compared to other social or technical solutions that may not scale) to support editors (e.g., Mentorship).
In July-August 2024, we had a few sessions with WMF’s senior leadership to learn about their perspectives and priorities. These sessions were important for us because we wanted to have organizational alignment about the strategy we were developing and aligning with the leadership was one important aspect of it. We thank (in order of last name) Lane Becker, Nadee Gunasena, Maryana Iskander, Stephen LaPorte, Lisa Seitz Gruwell, Amy Tsay, Denny Vrandečić, and Yael Weissburg for deeply engaging with our questions and sharing their thoughts and perspectives freely.
In August, we held a session with some of the affiliates and volunteers during Wikimania 2024. We thank those who joined us in that conversation, sharing their perspectives, and providing valuable feedback to our thinking. One of the learnings we had from the session was that multiple affiliates looked forward to having clarity on “how” we implement the AI strategy. We dedicated one subsection in this strategy brief to this topic inspired by those conversations.
And lastly, we would like to thank Pablo Aragón, Adam Baso, Suman Cherukuwada, Rita Ho, Caroline Myrick, and Santhosh Thottingal for their questions, comments or input that helped improve this work.
Footnotes
- ↑ Siehe ClueBot NG, einer der ersten KI-gestützten Community-Bots, und die von der WMF entwickelten und gehosteten KI-Modelle.
- ↑ Special:MyLanguage/Strategy/Multigenerational
- ↑ Humans in the AI loop: the data labelers behind some of the most powerful LLMs' training datasets
- ↑ See Asaf Bartov's presentation in CEE 2024 for examples.
- ↑ Examples of such common topics include but is not limited to the list of articles every Wikipedia should have
- ↑ Note that the code for the major open source LLM technologies is currently not open. For some of these AI models the weights are open.
- ↑ Open source models provide access to the training data and code, while open weights only provide the trained parameters (weights), often in Safetensors format. These weights can be hosted on Wikimedia’s infrastructure using open source software libraries
- ↑ For comparison, according to one source Meta has 600,000 GPUs for AI, while the Wikimedia Foundation currently has fewer than 20.