Jump to content

OpenSpeaks/Archives/July 2025 – June 2026

From Meta, a Wikimedia project coordination wiki

Overview

[edit]

This project ran between June 2025 and June 2026 as the first phase of a three-phase project with a focus on documenting oral history/knowledge in local languages, archiving them and citing them as a source of information.

Oral history of the majority of the world's peoples in their respective languages is not widely recognised as a form of knowledge, reducing their use to merely representational. We challenge this status quo, particularly within the Wikimedia movement, and work towards practical ways to bring oral history as a source of knowledge.

Language community members will be a part of this project, bringing their stories and knowledge into Wikimedia projects. In over ten low-resourced South Asian languages, high-quality, accessible audiovisual media will be published as an outcome. Additionally, we will create open source tools, workflows and open educational resources (OER) to help train both Wikimedia and language speaker communities. We will also collaborate with GLAM (galleries, libraries, archives and museums) institutions to ensure oral history is considered a reliable source of information within Wikimedia projects. As a result, the oral history media will be citable, translatable, and usable across Wikimedia platforms.

Why this project?

[edit]

Despite being multilingual and diverse in certain aspects, Wikimedia projects lack knowledge from the majority world and their speaker communities. For example, Wikipedia articles often cover fictional languages in detail, while many living languages and their speakers remain invisible.

We identify three major barriers within Wikimedia projects:

  1. Content gaps: Indigenous and other low-resource languages are poorly documented or entirely missing.
  2. Lack of tech tools: Open, affordable, cross-platform, offline-friendly tools for community-led language documentation do not exist or are locked behind proprietary systems.
  3. Citation bias: Oral histories and Indigenous knowledge face barriers to being accepted as sources in Wikimedia projects.

This project tackles these barriers by:

  1. Publishing accessible, multilingual oral histories that can be cited.
  2. Developing the missing tools and workflows for documentation.
  3. Training and mentoring archivist-Wikimedians and language experts.
  4. Partnering with GLAM institutions to ensure the material is discoverable and reusable.

Our Approaches

[edit]
Theory of Change
The Wikimedia movement can serve low-resourced language communities by embedding community knowledge in their own languages, by providing tools for local assertion, and by ensuring accessibility for wider audiences.
High-quality, accessible language media
Selected recordings (oral histories, descriptive videos) from over ten languages will be edited, subtitled, and published. Subtitles will be bi/multilingual (a local dominant language + English) for accessibility and translation support.
GLAM partnerships and citations
Collaborating with GLAM institutions to ensure oral histories are integrated into catalogues and made citable for Wikipedia, Wikidata, and beyond.
Tools, workflows, and OER
Identifying and co-developing missing tools for language documentation, releasing them as open source, and documenting them for archivists.
Community and capacity building
Training archivist-Wikimedians through workshops (e.g. WikiConference India, Celtic Knot), mentorship, and campaigns like Wiki Loves Languages.

Strategies and Activities

[edit]
Content creation
Process and publish oral history recordings in 10 languages with subtitles and metadata, ensuring they can be cited in Wikimedia projects.
Tools & OER
Develop and document open-source workflows for audiovisual language media, so archivists can replicate and extend the model.
Training & community building
Host in-person and remote workshops, develop training curricula, and create a peer-mentor network among language experts.
Campaigns
Co-lead Wiki Loves Languages to encourage communities in enriching Wikipedia and Wikimedia projects about languages and their speakers.
GLAM collaborations
Work with partner GLAM institutions to catalogue and cite oral histories, enhancing their reach and credibility.

Phases of Work

[edit]
Phase 1 (ongoing; July 2025–June 2026)
Community building, training, media processing, tool development, OER creation, GLAM partnerships.
Phase 2
Expand self-paced learning modules (e.g. WikiLearn), grow community-led content documentation, deepen local collaborations.
Phase 3
Showcase tangible documentation of oral cultures, broaden GLAM partnerships, and strengthen train-the-trainer and peer-learning networks.

Languages and people

[edit]

The project is guided by OpenSpeaks Fellows, native speakers of the focus languages. The focus languages are divided into three clusters: Nepal, northern India, and eastern-southeastern India. Apart from their larger advocacy role, the Fellows will contribute in three important ways:

  • Reviewing and subtitling media, in consultation with other community members
  • Ensuring community ownership and consent
  • Acting as mentors for new archivist-Wikimedians

OpenSpeaks Fellows

[edit]

The first phase of OpenSpeaks Fellowship was awarded to seven community members who will be co-leading the media identification, subtitling, translation and publication. The Fellows are:

  • Kimmi Pal is a communication designer and a Marcha-Rongpo speaker from Uttarakhand, India, who will coordinate the Rongpo archive.
  • Surendra Singh Pangtey is a noted author, former administrator, and Johari-language speaker based in Dehradun, Uttarakhand.
  • Arun Gour is an environmental activist from Bangsil, Uttarakhand, where he founded the Devalsari Paryavaran Sanrakshan Awam Tekniki Vikas Samiti, and is a speaker of the Jaunpuri dialect of Garhwali.
  • Opino Gomango is a researcher and activist of the Sora-cluster languages from Palakhemundi, Odisha, India.
  • Nenavath Mohan is a speaker of the Lambani language and is based in Telangana, India.
  • Sanjib Chaudhary is an author and Saptariya/Eastern Tharu-language activist from Nepal.
  • Uday Raj Aaley is a language researcher, lexicographer, writer, and activist from Nepal.

Activities conducted

[edit]
2025
  • July:
    • Discussed with project plan with archivists.
    • Discussed with Indic Mediawiki User Group collaboration with tools development.
    • Early funding, resource person and other logistical coordination.
  • August:
    • Decision to call all language leads OpenSpeaks Fellow.
    • Organised workshop (mentioned above) in Dehradun.
    • Developed prototype tools that directly address technical gaps identified in the pilot last year, to create and edit subtitles, inspect media properties and compress files for sharing/editing, batch calculate total media duration inside folders for project planning/budgeting.
    • Shared prototypes and discussed with Indic Mediawiki User Group to collaborate for tools development.
    • In-person field translation conducted for Johari in Dehradun, India. One output video used in multiple Wikipedia articles.
    • To plan for subtitling and translation, met in person with Arun Gour, OpenSpeaks Fellow for Jaunpuri, who flagged the need for a tutorial to understand subtitling.
    • Kimmi Pal, Fellow for Rongpo, finished first draft of subtitles for interviews of Bimla and K.S. Bharwal.
  • September:
    • Seven OpenSpeaks Fellows confirmed participation.
    • Created three prototype tools addressing all technical gaps identified in the pilot:
      • OpenSpeaks Subtitler: key tool; detects pauses, generates dummy subtitles and helps subtitle offline.
      • Media Optimizer: displays essential metadata of audio/video files and generates a command-line prompt to compress media for sharing with collaborators.
      • Additional tools prototyped for media folder organisation and transcript word counting.
    • Successful follow-up meeting with Indic MediaWiki Developers User Group; key members expressed interest in collaborating on tool-building.
    • Presented at Celtic Knot (virtual, 23 September), initiating a broader movement conversation on oral citation practices.
    • First phase of language documentation begun in Gorum/Parengi and Juang languages by Opino Gomango in Mysore, India.
    • Kimmi Pal (Rongpo) finished subtitling most of the Rongpo recordings; videos edited and subtitles integrated, scheduled for publication on Wikimedia Commons in mid-October.
  • October:
    • OpenSpeaks/Tools page created on Meta-Wiki, documenting all tools in development.
    • Five prototype tools created (yet to be fully tested before publishing but will be published eventually) as webapps:
      • Commons Metadata Generator]]: generates wikicode for files created during a language documentation project.
      • Media Folder Analyzer: analyses audio/video duration for project budgeting.
      • [Multimedia Folder Organizer: organises, tags, categorises and renames media files in a production folder.
      • Transcript Word Counter: counts transcript words to support billing and project planning.
      • Print Subtitles: prints subtitles for translators to correct offline.
    • First batch of subtitled oral history videos from multiple languages prepared for upload to Wikimedia Commons.
    • Selected for participation in Wikimedia Futures Lab, to be held in Frankfurt from 30 January–1 February 2026.
    • In conversation with the Songhay language diaspora community to support mentorship and the Wikipedia incubation process for the Songhay language.
  • November:
    • Invited to Deutsche Welle (DW) Akademie's invite-only "The next chapter: Journalism in the age of AI" gathering (Chiang Mai, Thailand, 25 November).
    • Presented OpenSpeaks Archives at FosterLang: "Strategies to Increase Equitable Forms of Exchange and Partnerships in Support of Linguistic Capital and Language Justice", a working group organised by Linguapax International.
    • Indic MediaWiki User Group collaboration still being confirmed; contingency plan to engage a paid developer if needed, focusing on essential tool feedback from active users.
  • December:
    • Spoke at WikiConference Kerala 2025: "Digital Tools and Strategy for Indigenous Languages" (recorded talk).
    • Significant progress with media translation across multiple languages; processed media uploaded to Commons and embedded into Wikipedia and Wikidata entries.
    • Two new Wikimedia Commons templates created for community use:
    • Verbal agreement for a GLAM collaboration with a New Delhi, India-based institution related to language data, capacity building, educational resources, and tools.
    • New GLAM partnership finalised with an European public archive to permanently archive all videos in the OpenSpeaks archive; outcomes include:
      • assigning a DOI to each video, enabling citation on Wikipedia, Wikidata, and other Wikimedia projects.
      • Field linguist training for Wikimedians interested in language documentation, to be delivered by their staff.
      • OER to be created for how language archivists can contribute data to their archive and subsequently to Wikimedia projects.
      • OpenSpeaks to act as a bridge between Wikimedian-archivists and them, rather than as a gatekeeper; formal agreement to be signed soon.
    • Webapp created for adding rich metadata to Wikimedia Commons.
    • Invited to and recorded for a podcast by Radio Taiwan International (Taiwanese public broadcaster).
2026
  • January:
    • Moderated a panel on community activism at the AI Impact Summit 2026 pre-Summit event.
    • Tech lead finalised for building tools; prototyping and UI design in progress.
    • Submitted a paper to Wiki Workshop 2026, co-authored with two OpenSpeaks Fellows, Opino Gomango and Kimmi Pal.
    • Acquisition of Eastern Tharu media archive from Sanjib Chaudhary completed; currently being edited for publication, retaining his copyright.
    • Media in three languages processed, subtitled, and published on Commons; used in Wikipedia and Wikidata.

Tools

[edit]
OpenSpeaks Subtitler
Webapp for creating audio/video subtitles both offline and online.
Screenshot of a prototype of OpenSpeaks Subtitler in action
Media Metadata Viewer & Compress Helper
Quickly inspect media properties and compress files for sharing/editing with collaborators.
Media Duration Calculator
Batch calculates total media duration of audio and video files inside folders for project planning/budgeting.
Multimedia Organization Tool
Organise, categorise, tag, and batch-rename multimedia files (video, audio, image) inside a folder using structured naming conventions for production workflows.

Open Educational Resources

[edit]

Awareness & capacity building

[edit]

Impact/outcome

[edit]
New pages (articles, entries) created

Wikipedia

[edit]
English
  1. w:Johari (dialect)
  2. w:Surendra Singh Pangtey
  3. Parenga (title was created earlier as redirect without any content)
Odia
  1. ଜୁରାୟ ଭାଷା
  2. ରଙ୍ଗପା
  3. ପାରେଙ୍ଗା
Hindi
  1. जौनपुरी बोली (गढ़वाल)
  2. पारेंगा

Wikidata

[edit]
  1. Oral history documentation of Gejmehac (Kusunda) language (Q139373278)
  2. Meghavath Sathish (Q139289734)
  3. Nenavath Mohan (Q139289570)
  4. Ram Prasad (Q139250121)
  5. Surendra Dutt (Q139250105)
  6. Maari Jaban Maari Birsa (Q139200369)
  7. Dinabandhu Gamango (Q138978470)
  8. Achhai Chaudhary (Q138949609)
  9. Category:Kshetrabasi Juanga (Q138835954)
  10. Kshetrabasi Juanga (Q138835636)
  11. K.S. Badwal (Q138769667)
  12. Bimla Badwal (Q138769659)
  13. Manjula Bhuyan (Q138762809)
  14. Category:Shauka people (Q138756949)
  15. Parenga (Q138637146)
  16. Namad Dalbehera (Q136825803)
  17. Johari (Q135781451)
  18. Arun Prasad (Q135778497)
  19. Ruum Sakam (Q135485904)

Impact of created multimedia

[edit]

As of 5 May 2026,[1] nearly 900 pages across 127 Wikimedia projects, including about 100 Wikipedia language editions, have been enhanced by files uploaded through this project. Files related to this grant alone enhanced 475 pages on 99 wikis (approximately 8 pages per file on average). Coverage spans Wikipedias in relatively smaller South Asian languages (Santali, Odia, Assamese, Maithili), African languages (Igbo, Swahili, Hausa, Malagasy), as well as major European and East Asian ones.

Overall metric Value
Wikimedia projects reached (all OpenSpeaks) 127
Wikimedia projects reached (grant category only) 99
Wikipedia language editions reached 97
Pages enhanced (all OpenSpeaks) 23,875
Pages using grant-category files 475
Distinct languages documented 14
Countries covered 3 (India, Nepal, Sri Lanka)
April 2026 views (all OpenSpeaks, full month) 215,491
March 2026 views (all OpenSpeaks) 152,571
March 2026 views (grant category only) 86,264
Files with DOI (citable) All in progress via ELAR; also with LAC
Metric Figure Source
Total files in grant category 116 PetScan / Commons category (live)
Files with at least one use 69 GLAMorgan
Files actively viewed 60 GLAMorgan
Total media files across all OpenSpeaks categories 71,812 GLAM Dashboard
Distinct media actively used 96 / 206 GLAM Dashboard

Usage rate: 69 of 116 grant-category files (59%) are in active use; 60 are actively viewed. Across all OpenSpeaks categories, 47% of files are actively used. This is a mid-grant analysis, with many files still to be fully processed, subtitled, and interlinked.

Reach (Wikimedia projects)

[edit]
Metric Figure Source
Total Wikimedia projects reached (all OpenSpeaks) 108 GLAM Dashboard
Wikimedia projects reached (grant-category files) 99 wikis GLAMorgan
Total pages enhanced (all OpenSpeaks) 897 GLAM Dashboard
Pages using grant-category uploaded files 475 GLAMorgan

Wikipedia language editions where files are used in article namespace

[edit]

Afrikaans (af), Angika (anp), Arabic (ar), Egyptian Arabic (arz), Assamese (as), Asturian (ast), Awadhi (awa), Azerbaijani (azb), Belarusian (be), Bhojpuri (bh), Bengali (bn), Breton (br), Catalan (ca), Cebuano (ceb), Czech (cs), Danish (da), German (de), Zazaki (diq), Greek (el), English (en), Esperanto (eo), Spanish (es), Basque (eu), Persian (fa), Finnish (fi), French (fr), Irish (ga), Galician (gl), Gujarati (gu), Hausa (ha), Fiji Hindi (hif), Hindi (hi), Croatian (hr), Hungarian (hu), Armenian (hy), Indonesian (id), Igbo (ig), Ilocano (ilo), Icelandic (is), Italian (it), Japanese (ja), Javanese (jv), Georgian (ka), Koro (kge), Kazakh (kk), Kannada (kn), Korean (ko), Kashmiri (ks), Komi (kv), Latin (la), Lingua Franca Nova (lfn), Lithuanian (lt), Latvian (lv), Maithili (mai), Malagasy (mg), Macedonian (mk), Malayalam (ml), Marathi (mr), Malay (ms), Mazanderani (mzn), Nepali (ne), Nepal Bhasa/Newari (new), Norwegian Nynorsk (nn), Norwegian Bokmål (no), Occitan (oc), Odia (or), Punjabi (pa), Polish (pl), Piedmontese (pms), Western Punjabi (pnb), Pashto (ps), Portuguese (pt), Romani (rmy), Russian (ru), Santali (sat), Sanskrit (sa), Sindhi (sd), Serbo-Croatian (sh), Sinhala (si), Simple English (simple), Slovak (sk), Swedish (sv), Swahili (sw), Tamil (ta), Tagalog (tl), Tulu (tcy), Telugu (te), Thai (th), Turkish (tr), Tatar (tt), Ukrainian (uk), Urdu (ur), Uzbek (uz), Vietnamese (vi), Min Nan Chinese (zh-min-nan), and Chinese (zh).

Other Wikimedia projects enriched are Wikimedia Commons (host), Wikidata, Meta-Wiki, Wikisource, and Wikiversity (six non-Wikipedia projects in total).

Views

[edit]

Monthly and cumulative[2]

[edit]
Period Views Source
March 2026 (file views, grant category) 86,264 GLAMorgan
March 2026 (all OpenSpeaks media) 152,571 GLAM Dashboard
April 2026 (partial, all OpenSpeaks media) 45,571 GLAM Dashboard

Languages and dialects covered[3]

[edit]
# Language/dialect ISO Cluster Files in category
1 Marcha-Rongpo rnp Northern India 9
2 Johari-Kumaoni kfy Northern India 8
3 Jaunpuri-Garhwali gbm Northern India 2
4 Jaunsari jns Northern India 4
5 Bangani him Northern India 2
6 Sora srb Eastern/SE India 12
7 Juray juy Eastern/SE India 9
8 Juang jun Eastern/SE India 7
9 Gorum/Parenga pcj Eastern/SE India 4
10 Lambadi lmn Eastern/SE India 3
11 Raji rji Nepal 4
12 Nepali ne Nepal 1
13 Sri Lanka Malay sci Sri Lanka 2
14 Eastern (Saptaria) Tharu thq Nepal/Northern India 5

File type breakdown

[edit]
Type Count (of 100) Notes
Video (.webm) ~23 Primary oral history and interview content; highest aggregate fileviews per file
Audio (.wav) ~27 Pronunciation samples and short clips
Image (.jpg/.png) ~46 Speaker portraits, workshop photos, cultural documentation, tool screenshots (most are for archiving and not to use in Wikimedia projects per se)
Document (.pdf) 4 Conference presentations (Celtic Knot 2025, WikiConference Kerala 2025, FOSS Meetup Bangalore, citing guide); low fileviews but strategically important for citation and capacity building

Note that all the files that are unused are uploaded for archiving (multiple audio files with the pronunciation of the same word) or educational purposes (slide decks, video recordings, etc.)

Testimonials

[edit]

This page includes interviews and other testimonials with collaborators. They are taken from audio interviews and are gently edited for grammar and conciseness if and where needed. The quotes can be used outside of Meta-wiki as long as the interviewee(s) and the project (OpenSpeaks Archives) are attributed.

Kshetrabasi Juanga

[edit]
Kshetrabasi Juanga is an educator, writer, and a native speaker of the Juanga, one of the focus languages
On audio-visual language documentation

"Imagine someone from the next generation doing their PhD on tribal languages and cultures. When they do, they will find a great deal of information waiting for them."

On the value for young speakers

"Today's children are tomorrow's young people, and if they one day watch videos in Juang, they will feel that some effort has been made to preserve what is theirs."

On documentation as a record of language change over time

"Suppose a recording is made today, in 2026. When someone listens to it decades from now, they will be able to understand how the language was gradually changing, how a culture was being challenged, and how it survived."

On the resilience of Indigenous language and culture

"I think of it like a bitumen road. In some places, the surface has cracked, the bitumen risen, and green grass has pushed through. I see our language like that green grass — surviving despite countless vehicles crushing it every day. No matter how globalised the world becomes, Indigenous language and culture are surviving, and will continue to survive."

On audio-visual documentation for the Juang language

"This is helpful and a welcome step. Juang is shifting. But while globalisation is bringing more languages to people through technology, our language and culture risk becoming stagnant. Documentation is essential, and that is why this work matters."

On the value of subtitles and captions specifically

"If a Juang person living in a city comes across this documentation, it will feel like history to them. They will find self-respect and feel proud."

On what oral traditions carry that writing cannot

"Such traditions carry meaning that cannot easily be written down… That trust, that unspoken moral code, is something a society outside our community can learn from. The subtitles help them understand."

On the universal value of this documentation

"To me, documenting these traditions will create an invaluable resource for the entire humankind."

Opino Gomango

[edit]
Opino Gomango is a researcher of the Sora-Juray languages, and coordinated documentation of Sora, Juray, Juang and Gorum languages
What do you think is the use of publishing detailed metadata with audio/video from a speaker community standpoint?

"Earlier, what used to happen was that if a linguist wrote any book after researching a language, the linguist's publication—be it a book or a journal—would be cited as a reference. But the knowledge of those who are speaking in their own language, the knowledge that was recorded, whether audio or video, could not be used as a reference. So we are trying to give the person who speaks that language as a reference. Not only a linguist or researcher, but even the native speaker should also get the reference; they should also have their name referenced."

if and how adding multimedia to Wikipedia/Wikimdia entries help?

"[Wikipedia] is open-source, it is free, and will be useful for everyone. Wikimedia is free—it is free for everyone. So everyone can access it. This is very good for everyone, including tribal people and others, as they can see everything and look ahead. So it is the best."

"Whenever we were documenting Juang/Parengi, we showed that video to the interviewees—'your video has come, see'—doing this made them very happy."

How will this documentation help the native speaker community?

"Previously, people used to think that if we studied English or Hindi, we would get a job. Now, the people who have jobs think that if the language is lost, then we are also lost."

How will subtitling audio/video help?

"Through those subtitles, we can understand — yes, this is what it means, this is how it is said. Subtitles are very important for many different people."

How will our work help non-native speakers?

"Language is not just for us. For people all over the world who want to know the language, who want to know the culture—it is very beneficial that this is all subtitled."

Taukeer Alam

[edit]
Taukeer Alam is a conservationist, author, and community organiser from the Van Gujjar community and a native speaker of the Van Gujjar, one of the focus languages of the OpenSpeaks Archives pilot

"I'd definitely like to endorse. First, whatever community or citizen data or knowledge is documented, whatever is being done is done in different forms. That is my objective, and this is our team's objective. We want it documented, transcribed, translated, and submitted somewhere. It should be done as quickly as possible. Whatever knowledge the community holds, if it must be documented in its true form, it should be documented quickly. That's why I want to endorse it: the community's real knowledge. We can document it as quickly as possible with collaboration. That's why I endorse this."

Full interview:

Panigrahi, Subhashish (2026-04-17). "Vulnerable language speakers fear that their unique and sensitive knowledge will be exploited or abused through AI". Global Voices. Retrieved 2026-04-17. 

Sanjib Chaudhary

[edit]
Sanjib Chaudhary is a Nepalese activist and writer, and a native speaker of Eastern (Saptariya) Tharu, one of the focus languages of the OpenSpeaks Archives

"As part of this project, I am interviewing Tharu traditional healers and documenting their practices in Eastern (Saptariya) Tharu. When we write, we lose so much: tone, rhythm, body language, and the emotion in the healer’s voice.

Many people also do not have the time or habit to read long articles; they skim. With audio and video, they can hear the healer’s voice directly instead of reading my description of them. That creates trust and credibility — people see that these are real community members speaking about their own practices, not someone speaking on their behalf."

Full interview:

Panigrahi, Subhashish (2026-05-10). "Archiving Nepal's medicinal plants and Indigenous knowledge for future generations". Global Voices. Retrieved 2026-05-12. 

Publications

[edit]

Footnotes

[edit]
  1. Excludes any 404 errors which are common for live analysis.
  2. Wikipedia editions such as English are high-traffic (12% media are in English Wikipedia), whereas the majority of contributed media is in other language editions.
  3. The number of files are not a direct indication of overall progress. In some cases, the captioning is over but final review/file export/conversion might be pending.