Talk:Artificial intelligence/2026 Wiki AI
Add topicWelcome to this pre-conf event!
This is a place to discuss constructive uses of LLMs and other AI models as editing tools on Wikipedia.
- See also
- Waikiki
- Wikipedia:WikiProject_AI_Tools on English Wikipedia
- Using generative AI on English Wikiversity
- discussions at recent hackathons about building good tools.
Proposals and suggestions
[edit]Share ideas and links to tools and proposals you'd like to see discussed, below:
- Playground
- Wikiomnia Library: access to resources
- ... add yours!
AI Wikimedia and project discussions
[edit]See Requests for comment/Artificial intelligence policy, the broader movement appears to be generally against almost all use of AI to generate content (though opinions do vary and it is far from a monolith), and I'd imagine there's a hefty amount of opposition and scepticism about other uses. Have a discussion, float some ideas, sure, but efforts behind the scenes that don't align with the movement would be very controversial and frankly subversive. Kowal2701 (talk) 09:10, 10 June 2026 (UTC)
- Apologies, I didn't realise this is en:WP:WPAIT. Nevertheless, en:WP:AINB#Fuzheado is concerning Kowal2701 (talk) 10:05, 10 June 2026 (UTC)
- Hi @Kowal2701, the pre-conference is very much a public discussion. We'll look at the tools that are already being used by editors and we'll discuss both positive and negative impacts of AI. Andrew (fuzheado) is one of the speakers who will be sharing his experience. If you're at Wikimania we'd be happy to see you at the pre-conference. Alaexis (talk) 15:18, 10 June 2026 (UTC)
- Thanks, apologies but I'm not attending Wikimania. Discussion is good, I've got no quarrels with that, I was just alarmed at what looks like advocacy against important PAGs. On enwiki a ban on LLM-generated content was adopted with near-unanimous consensus, wherein some people only supported due to concerns about article quality, and others supported due to ideological concerns/concerns outside of that (I need to put this into an essay at some point). I understand why Andrew might be interested in improving LLM-generated content, but quality concerns are only part of it. Kowal2701 (talk) 18:01, 10 June 2026 (UTC)
- I've responded on en.wp but will reiterate here: We will be discussing a wide range of ways AI can have beneficial applications in the Wikimedia community that don't entail one-shot LLM article generation or source fabrication. These include creating tools for linting, checking, insights, metrics, patrolling, multimedia browsing, and anti-vandal fighting, just to name a few. The state of the art with AI has come a long way, with agentic coding for implementing deterministic workflows (versus the stochastic parroting of LLMs) that can add to our Toolforge capabilities. For those who cannot make it to Wikimania, we will be recording it for later viewing and discussion. To be clear, I am in favor of restricting mass use of LLM tools of the 2023 ChatGPT vintage, as they peppered Wikipedia with low-quality and unchecked contributions from inexperienced users. But we should also not be so pedantic that we leave no room for evolving our perspectives on the rapid refinement of our technological tools. - Fuzheado (talk) 13:58, 11 June 2026 (UTC)
- Thanks, apologies but I'm not attending Wikimania. Discussion is good, I've got no quarrels with that, I was just alarmed at what looks like advocacy against important PAGs. On enwiki a ban on LLM-generated content was adopted with near-unanimous consensus, wherein some people only supported due to concerns about article quality, and others supported due to ideological concerns/concerns outside of that (I need to put this into an essay at some point). I understand why Andrew might be interested in improving LLM-generated content, but quality concerns are only part of it. Kowal2701 (talk) 18:01, 10 June 2026 (UTC)
- Argumentum ad populum. schiste (talk) 15:41, 10 June 2026 (UTC)
- this is a consensus-driven project, argumentum ad populum has done us pretty well thus far! Kowal2701 (talk) 18:07, 10 June 2026 (UTC)
Commons RfC
[edit]There is a proposal to change the Wikimedia Commons scope to require watermarking of some types of AI-generated/modified media: Commons:Commons:Requests_for_comment/Policy_update_for_AI_content. Consider voting or commenting. -- Jtneill - Talk 10:39, 24 June 2026 (UTC)
Will it be also online?
[edit]I have an online ticket for Wikimania, cannot come to Paris. Will the preconference also be online? I hope so and would strongly recommend it. We need such tools and development should be coordinated. You should declare wether a separate online registration is necessary and howto do it. Thanks Wortulo (talk) 13:45, 19 June 2026 (UTC)
- Hi Wortulo at least the tools session will be online! We will see about other sessions once the program is finalized in a week. We will also have one or two remote lightning talks during that session. –SJ talk 01:31, 21 June 2026 (UTC)
- Great. Let us know then if we need to to more to take part online. Wortulo (talk) 04:28, 21 June 2026 (UTC)
- Hi @Sj and @Alaexis! Do you have any news whether some or all of the pre-conference will be available online for those of us not able to attend Wikimania in person? Thanks! Diegodlh (talk) 22:35, 19 July 2026 (UTC)
- pour moment je suis pas en ligne pour y participer, comment nous allons faire Bwasomo (talk) 07:59, 21 July 2026 (UTC)
- Unfortunately only offline, what a pity Wortulo (talk) 08:37, 21 July 2026 (UTC)
- @Bwasomo@Wortulo@Diegodlh
- Wiki AI - Fact-checking at scale: Source Verifier demo & panel Tuesday, 21 July · 2:30 – 3:30pm CET Video call link: https://meet.google.com/zpv-rsyw-kyk
- Wiki AI demos - AIlog - fighting AI slop, WikiShield - fighting vandalism Tuesday, 21 July · 4:45 – 5:30pm CET Video call link: https://meet.google.com/hmk-hqzc-xqm
- Alaexis (talk) 09:20, 21 July 2026 (UTC)
- pour moment je suis pas en ligne pour y participer, comment nous allons faire Bwasomo (talk) 07:59, 21 July 2026 (UTC)
- Hi @Sj and @Alaexis! Do you have any news whether some or all of the pre-conference will be available online for those of us not able to attend Wikimania in person? Thanks! Diegodlh (talk) 22:35, 19 July 2026 (UTC)
- Great. Let us know then if we need to to more to take part online. Wortulo (talk) 04:28, 21 June 2026 (UTC)
Wiki tools gallery ideas
[edit]A gallery with the most used tools and datasets could be helpful for orienting people, and a longer list including the smaller projects could be an appendix. Starting some lists here for ideas:
Tools
[edit]Current tools
- Patrolling Workbench: SP42 (by Schiste, LuisVilla, and others)
- Citation verification: Source Verifier
- Citation-needed checkers: BAWUG
- Citation finders: IA Europe, CNfirmed
- Controllable sentence simplification (Italian, 2026)
- SmolLM2-1.7b-wiki - SmolLM grounded w/ WP
- Six Degrees of Wikipedia
- DNS over Wikipedia (via infobox links)
- Xikipedia (simple articles as social feed)
- Pywikibot, once and future
- Python wiktionary parser
- Wikiquote tool
- Commons Category Diffusor
- WikiVault for En->Ko translation
Older tools
- wikipedia2vec
- Citation Needed, WikiQuickie
- Propositionizer (proposition retrieval from text, 2023)
- WMF tool experiments (202x)
- Wikinews feed*
Datasets
[edit]Current datasets (HF)
- Wikimedia pages markdown (2025, OpenLLM-France. WP,WB,WN,WQ,WS,WV,WVoy,WT, 15 euro languages)
- Cosmos 1.0 (innovation)
- Structured Wikipedia (WME)
- Wikipedia Lists of Lists graph (May 2026, warning)
- CulturaX + NL.Wikipedia data (Dutch, 2024) for training
- Wikipedia multilingual embedding (Cohere, 2023)
- aligned paragraphs-conversations
- dbpedia (much more than just W)
- All Wikidata triples (1/2026)
- Wikidata truthy rdf (all core facts, 2025)
- Wikidata 5M KG (5M entity knowledge graph, 2025)
- Wikidata multilingual labels + descriptions
- Wikidata query logs (250k queries, 2025)
- Wikicommons CC0 (11TB, 2025)
- Commons maps (450 historical maps, 2025)
- Wikiteam archive (600k non-WM wikis)
Older datasets
- WikiNeural Named Entity Recognition (2021)* BERT topic model (2023)
- Wikidata simple questions (2017)
- Kaikki / Wiktionary data (2026)
- Wiktionary pronunciation (2022)
Linear schedule
[edit]When will setup start // do we know when morning refreshments will be ready? How long should we invite people to stay in the evening? –SJ talk
9:30 Welcome – SJ Klein
9:40 Keynote – Mark Graham (Internet Archive)
10:15 Productive Coexistence (workshop) – Hay Kranen
11:00 Coffee break
11:30 Wikigame - spot the AI errors
11:45 Building Wiki Tools with AI – Andrew Lih
13:00 Lunch
14:30 Fact-checking at scale: Source Verifier demo & panel, Elena Simperl, Kevin Payravi, Alex Ostrovsky Hybrid
15:10 From Hope to Mistrust: AI hypotheses at the Futures Lab, Eva Martin
15:40 "What Would It Take?" – Netha Hussain
16:15 Towards a sensible AI policy on Wikimedia Commons, Abzeronow
16:30 Coffee break
16:45 AIlog - fighting AI slop, Fermiboson Hybrid
17:00 WikiShield - fighting vandalism, LuniZunie (remote) Hybrid
17:20 AI for visualising data, Galder Gonzalez
17:35 Small Wikis lightning talks – Micheal Kaluba & Dev Jadiya
18:00 Builder tables – Netha Hussain, Houcemeddine Turki, Britta Gustafson, Andrew Lih - until 19:00
AI slop as the header image
[edit]Hi @Fuzheado.
You uploaded File:Wikimania-2026-wiki-ai-preconf-graphic-an-title.png to Commons with the description Uploaded a work by Google Gemini AI from Google Gemini AI with UploadWizard and the used that image as the advertising header for this event.
Am I to assume this was done with a sense of irony, to illustrate the dangers of generating AI slop and uploading it to Wikimedia projects? If so, that is not at all clear. Otherwise, if this was genuinely created and uploaded to advertise the event, I find the choice really concerning.
The text around the Wikipedia logo appears to read something along the lines of WIKIIIPEDIA WIKIPEDIA · RIIPEDIA EHIRYPE?EDIA ORC?TREDIUS with the globe itself filled with malformed characters.
The Foundation's Visual identity guidelines asks that the public do not alter the Wikipedia or related project identities nor any of their elements. This slop image clearly breaks the Wikipedia identity in exactly that way.
It is particularly inappropriate for a Wikimedia event about AI to use, as the principal promotional image, an obviously defective AI-generated slopped version of Wikipedia's own branding.
Could you not have mocked up something simple using one of the many free graphic-design packages available and released the result under a free licence?
I politely request you remove this image - and request their deletion from Commons - and replace it with something more suitable.
Many thanks, qcne (talk) 20:25, 18 July 2026 (UTC)
- Also not a fan of that header image, and had noted that earlier in an adjacent space. It uses AI as output, with the expected associated issues including quality, rather than using AI as a tool to help a person build a tool, or as part of a tool that helps a person with a task (where the decisions and output remain in the hands of the person), which were the main use cases actually discussed. Dreamyshade (talk) 00:20, 22 July 2026 (UTC)
- I wouldn't call it irony, but it seems accurate: showcasing both the affordances and problems of current AI systems, which we have to understand and grapply with, not trying to hide imperfections and defects. The actual use cases being developed and deployed in the wikiverse mainly help with tool building and review, as Dreamyshade noted. –SJ talk 01:38, 10 August 2026 (UTC)
Using AI to detect AI-generated news content farms ?
[edit]Hi,
as a french investigative journalist, I've spent two years tracking AI-generated ‘news’ (with my own eyes, not AI detection tools), and identified more than 15,000 of them (in french, plus about 1,500 in English, 200+ in German, 130 in Spanish, and dozens in other languages).
More than 1 900 of those 15 000 AI-generated 'news' sites are mentionned as sources on fr.wikipedia.org (most of them expired domain names which has since been bought by SEO professionnals who fills them with AI-generated news, but also junk news sites which ask their journalists to mass-publish AI-generated articles), and I'm working with volunteers to help them clean such sources on Wikipedia.
That said, I've identified a number of serial patterns which could help figure out whether a news site is AI-generated :
- how many articles published/author/day, at which hours, do they publish on week-ends ?
- do their authors have names+surnames (or only a pseudonym), photos, active profiles on social networks ?
- are their legal terms refer to a company, and if so is it a media, a SEO professionnal, and what other news sites are link to it ?
- do their advertising (ADSENSE DIRECT) iD refers to other news sites ?
(...)
I tend to think that AI could help us pre-check those serial patterns, and help volunteers figure out wether sources are AI-generated of human made, but fail to find people in France to help us develop such a framework. Do you know people or organizations who could help, or do you think people from the WikiProject AI Tools could be interested in ? Regards, Manhack (talk) 15:01, 21 July 2026 (UTC)
- Hi @Manhack, apologies for the radio silence. This is indeed very worrying and I'm sure that both deterministic algos and LLMs can help us. Just to make sure I understood the shape of the problem, you believe that there are more such fake news websites and you'd like to build a system/process that detects them, right? Alaexis (talk) 20:49, 27 July 2026 (UTC)
Were any of the sessions recorded?
[edit]There are very many recordings at https://www.youtube.com/@TheWikimediaFoundation/streams but I can't find any from this workshop there, or on Meta. Did they get recorded? ~2026-40742-08 (talk) 09:16, 23 July 2026 (UTC)
- Unfortunately not, you're welcome to review the etherpad and to ask questions here. Alaexis (talk) 20:50, 27 July 2026 (UTC)