Talk:Community Resources and Partnerships/India General Support Project/Evaluation of AI generated Wikipedia content
Add topicFollow up on your project application
[edit]Dear Pavanaja,
Thank you for submitting your proposal. To help the committee review your proposal effectively, we would appreciate your clarifications on the following points:
- Has the Kannada Wikipedia community discussed or formally approved the idea of testing AI generated content on its platform? Please share links to community discussions or endorsements that show consent for such experimentation.
- Could you explain why this needs to be a 12 month long funded project to create only around 200 articles? Would a smaller, time-bound pilot study (with limited scope and lower cost) be more appropriate at this stage?
- The project proposes engaging 20 undergraduate students with limited prior Wikipedia experience. How do you plan to train, mentor and retain these contributors? especially given the technical complexity of working with AI generated text?
- Have you considered running the same project with the existing experienced Kannada Wikimedia editors to ensure quality control and continuity after the project ends?
- How do you plan to bring in technical oversight or collaboration to ensure accuracy and ethical use of AI in generating the content?
- Please clarify which AI tools, models, or datasets you plan to use for generating and evaluating content, and how you will ensure ethical, transparent and responsible use of these tools.
- What evaluation framework will you use to measure quality, accuracy, and potential harm of AI generated articles? Could you also provide details on the team’s technical expertise?
- The proposal mentions rural women participants. Given challenges like limited access to devices and internet connectivity, what steps will you take to ensure safe participation and equitable access for them?
- How does this current proposal build upon or improve past efforts of yours?
Please share your responses by 25th October so we can move forward with the next stage of deliberation. Looking forward to your responses.
Regards, Praveen (on behalf of South Asia Regional Funds Committee) PDas (WMF) (talk) 21:26, 11 October 2025 (UTC)
Answers to Follow up on your project application
[edit]Thank you @Praveen Das for reviewing the application and for the followup questions. Here I am trying to answer your queries,
1. Has the Kannada Wikipedia community discussed or formally approved the idea of testing AI generated content on its platform? Please share links to community discussions or endorsements that show consent for such experimentation.
Yes. The project idea originated with me creating an AI generated article (with improvements and corrections) for Kannada Wikipedia by me. Then I gave a presentation to our community about “Creating Kannada Wikipedia Article using Generative AI”. Later the project idea was discussed in the community meeting. The project idea was approved by the community in this meeting. I also wrote a blog titled “Creating Kannada Wikipedia Article using Generative AI” available in Diff. I have been to Bogota, Colombia as a participant and presenter in the EduWiki Conference 2025. There, during some informal meetings, over lunch and dinner, I had mentioned this project idea. It was well appreciated by the people then. In fact Netha Hussain mentioned about a project she has got the grant to work with titled “Between Prompt and Publish: Community Perceptions and Practices Related to AI-Generated Wikipedia Content”. The project that I am proposing is in similar lines but doing much more.
2. Could you explain why this needs to be a 12 month long funded project to create only around 200 articles? Would a smaller, time-bound pilot study (with limited scope and lower cost) be more appropriate at this stage?
The project isn’t just about article creation — it’s about training, evaluation, and empowerment. We are building:
- Digital and AI literacy among rural women,
- Evaluation rubrics and methodologies to measure the usefulness and limitations of AI, and
- Bridging gender gap in Kannada Wikimedia.
- Digitally empowering and training rural women students in AI.
A full 12-month period is necessary because the participating students are new to Wikipedia and will require time for training, mentoring, drafting, and quality evaluation of AI-assisted content.
Moreover, since the students have regular academic classes, their availability for the project will be part-time. This means the project cannot run as an intensive short-term pilot; instead, it needs to be distributed over a year to allow steady progress without affecting their studies. This duration ensures sustainable learning, ethical experimentation, and measurable outcomes rather than rushed content creation.
3. The project proposes engaging 20 undergraduate students with limited prior Wikipedia experience. How do you plan to train, mentor and retain these contributors? especially given the technical complexity of working with AI generated text?
We will conduct hands-on workshops covering basic Wikipedia editing, referencing, neutral writing, and responsible AI use. These students will be guided by mentors — experienced Kannada Wikipedians — for continuous guidance. Training materials will be in Kannada, and all sessions will be done in the college’s computer lab with internet access. Retention will be encouraged through certificates, recognition, and continued mentorship, motivating students to stay active beyond the project. We’ll also try to integrate this into their academic activities to give it lasting value.
4. Have you considered running the same project with the existing experienced Kannada Wikimedia editors to ensure quality control and continuity after the project ends?
Experienced editors are very much part of the plan — they will serve as mentors, reviewers, and evaluators. However, involving rural women students brings fresh voices, local perspectives, and helps reduce the gender gap on Kannada Wikipedia. By combining experienced editors’ knowledge with new contributors’ enthusiasm, we ensure quality control and community growth together. This mixed model makes the project both sustainable and inclusive.
5. How do you plan to bring in technical oversight or collaboration to ensure accuracy and ethical use of AI in generating the content?
The project follows a “human-in-the-loop” approach — AI is used only for drafting, never for automatic publication. Every AI-generated draft will be fact-checked, edited, and referenced by students and reviewed by mentors before it reaches Wikipedia. All AI prompts, model details, and outputs will be documented for transparency. We will also consult with experts from Wikimedia AI initiatives for ethical and technical guidance. This ensures that the project stays responsible and transparent.
6. Please clarify which AI tools, models, or datasets you plan to use for generating and evaluating content, and how you will ensure ethical, transparent and responsible use of these tools.
We plan to experiment with two kinds of models — one open-source and one commercial large language model that supports Kannada text generation. Each AI tool used will be clearly documented with model name, version, and access method. We will ensure:
- AI output is always reviewed by humans,
- Any unverifiable or hallucinated information is removed,
- All sources are checked for reliability, and
The process is fully transparent in reports and article talk pages. This approach meets Wikimedia’s standards for ethical AI experimentation.
7. What evaluation framework will you use to measure quality, accuracy, and potential harm of AI generated articles? Could you also provide details on the team’s technical expertise?
We are developing a quantitative rubric with key criteria such as: Accuracy and factual correctness, Completeness of coverage, Neutrality and tone, references and citations, language quality, article structure and potential cultural or factual harm. Each AI draft and final version will be scored independently by reviewers. The results will show how much value human editors add over AI drafts. The project lead and mentors have strong backgrounds in Wikipedia editing, Kannada writing, and evaluation, supported by local faculty and technical advisors.
Evaluation rubrics:
| Criterion | Weight | AI Draft (0-5) | Final (0-5) | Notes / Evidence |
|---|---|---|---|---|
| Content coverage & depth | 20 | |||
| Verifiability & citations | 20 | |||
| Language quality (Kannada grammar, clarity, neutrality) | 15 | |||
| Infobox presence & completeness | 10 | |||
| Structure & wiki markup (sections, templates, categories) | 10 | |||
| Lead quality (concise, summary of key facts) | 5 | |||
| Internal links & categories | 5 | |||
| Images/media & licensing | 5 | |||
| Topic notability & scope (sourcing justifies topic) | 5 | |||
| Maintenance issues (tags avoided, policy compliance) | 5 |
Evaluation metrics:
| Field | Value (AI Draft) | Value (Final) | Notes |
|---|---|---|---|
| Characters (approx.) | |||
| Citations (...) count | |||
| Internal links (...)) | |||
| External links (http/https) | |||
| Infobox present? (0/1) | |||
| Infobox fields count |
8. The proposal mentions rural women participants. Given challenges like limited access to devices and internet connectivity, what steps will you take to ensure safe participation and equitable access for them?
The project will be conducted within the college computer lab, which provides reliable computers, internet, and a safe environment. We will schedule sessions during non-class hours to ensure accessibility. Offline materials and Kannada-language guides will help those with limited connectivity at home. We will also ensure consent, digital safety, and privacy training, and appoint a faculty coordinator to support participants. This guarantees equal opportunity and a safe learning space for all students.
9. How does this current proposal build upon or improve past efforts of yours?
This project builds upon my previous work in Kannada Wikipedia outreach, digital literacy, and AI evaluation. Earlier initiatives focused mainly on editing and translation — this one adds a new dimension: quantitative evaluation of AI’s role in knowledge creation. It brings together capacity building, gender inclusion, and AI research in one model. The framework developed here — evaluation rubrics, workflows, and training materials — will be reusable for other Indian-language Wikipedias, making this project a model for future AI–Wiki collaborations.
-Once again, thanks - Pavanaja (talk) 04:30, 24 October 2025 (UTC)
Your Project Application is Not Funded
[edit]Dear @Pavanaja,
After careful deliberation, the committee decided not to fund the project proposal "Evaluation of AI Generated Wikipedia Content" in its current form. This was difficult decision for us, and we want to be clear about the reasoning and provide a constructive path forward.
The core issue is that AI generated content on Wikipedia is highly sensitive across many communities, and the committee felt there hasn't been sufficient broad community consultation as many language communities actively discourage this type of content.
The committee strongly recommends initiating via the India Rapid project with a phased approach. Begin with an open community consultation on the Kannada Wikipedia village pump and Meta-Wiki discussion pages, documenting feedback and concerns in a transparent manner. Then run a small pilot with a few experienced Kannada editors creating small pool of articles to test your evaluation rubric while documenting what works and what doesn't. This would validate your methodology and build community trust before scaling to the capacity building component with rural women students.
The pathway forward just requires broader community engagement and starting small with experienced Wikimedians before scaling as a proof of concept. Wishing you all the best!
Thanks,
On behalf of the South Asia Regional Funds Committee PDas (WMF) (talk) 17:39, 22 November 2025 (UTC)