OpenSpeaks/Archives
|
OpenSpeaks Archives is a public digital archive supporting community-led documentation of low-resourced languages. It contributes to Wikimedia projects, communities, and the open knowledge movement. So far, we have documented 20 South Asian languages, and improved nearly 1000 pages across 100+ Wikipedia and Wikimedia projects. Framework • Tools • Style Guide • Media • Workshops • People |
-
A native speaker shares the barriers to native-language education in his near-extinct language, Gorum
-
A Marcha-Rongpo-speaking couple living in a city remembers their childhoods in different villages
-
A Jaunpuri-Garhwali speaker discussing access to public information and social welfare
-
-
A joke narrated by author Surendra Singh Pangtey in Johari-Kumaoni dialect
-
Overview
OpenSpeaks Archives addresses three broader gaps: content gaps about Indigenous and other low-resourced languages, a lack of practical tools for community-led audiovisual documentation, and citation bias against oral history as a source of knowledge. Working with community language archivists, GLAM partners, and Wikimedians, we create citable, accessible media to embed and use as sources on Wikimedia projects and open educational resources, and technological tools to support language documenters.
Our 2024–2025 pilot archived five tongues, and our first phase (2025–2026) aimed to document 13 tongues from three South Asian countries. In our second phase, we are expanding into a network focused on language documentation and archiving.
Why it matters
Despite being multilingual and diverse, many living languages and their speaker communities remain largely invisible on Wikimedia projects (content) and communities (participation). This is when even fictional are richly documented. Oral histories from the majorityworld communities are rarely treated as citable knowledge within Wikimedia policies and workflows, limiting how far community-held knowledge can travel. OpenSpeaks Archives challenges this status quo by demonstrating how oral history can be recorded ethically, adhering to community protocols, consent, and verifiability, and aligning with emerging movement work on oral citations. The project aims to make it normal, not exceptional, to cite low-resourced language oral history on Wikipedia, Wikidata, and sister projects.
What we do
The project works across three main areas:
Educational resources
- OpenSpeaks Oral Knowledge Framework: a practical framework for documenting oral history using FAIR–CARE principles, peer-reviewed in a short version at Wiki Workshop 2026.[1]
- OpenSpeaks Captioning: guidance on creating accessible, multilingual subtitles and captions for audio and video used in Wikimedia projects.
- OpenSpeaks on Wikiversity: toolkit covering planning, recording, consent, accessibility, and publishing workflows for language documentation.
These resources are to encourage community archivists, Wikimedians, and GLAM practitioners to learn at their own pace and then adapt the workflows to their languages and contexts.
Tools and infrastructure
Several open-source prototype tools are being developed as part of the project, documented at OpenSpeaks/Tools:
- OpenSpeaks Subtitler: key webapp for for community subtitle creators/translators for subtitling offline and online; integrated into Wikimedia Commons to pull audio/video.
- Media Metadata Viewer & Compress Helper: helps inspect media properties quickly and generate compression commands for sharing and editing.
- Media Duration Calculator: batch-calculates total duration of media in folders to support budgeting and project planning.
- Multimedia Folder Organizer and related utilities: support structured naming, tagging, and transcript word counting, so production workflows can be replicated and scaled.
Communities and languages
Our primary focus is to gradually build community strength so that they can document their languages themselves.
OpenSpeaks Fellows, who are native speakers and community organisers, guide the work in three regional clusters. For our 2024–2025 pilot, we focused on one language from Nepal and four tongues from India. During our first implementation phase, we focus on four clusters: Nepal (two languages), northern India (five tongues), eastern–southeastern India (five languages), and Sri Lanka (one language). They review and subtitle media, ensure consent and community ownership, and mentor new archivist‑Wikimedians in their languages.
| Cluster | Language (ISO 639 code) | Fellow/Coordinator | Interviewees | ||
|---|---|---|---|---|---|
| Northern India | Marcha (dial. Rongpo - rnp) |
Kimmi Pal (Fellow), Bhawani Pal (advisor) | Bimla and K.S. Badwal | ||
| Johari (dial. Kumaoni - kfy) |
Surendra Singh Pangtey, Bhuppi Pangtey (reviewer) | Surendra Singh Pangtey | |||
| Jaunpuri (dial. Garhwali - gbm) |
Arun Gour (Fellow) | Sampati, Bhagwandi, Suchita | |||
| Jaunsari (jns) |
Arun Gour (Fellow) | Deepak Joshi | |||
| Bangani |
Jaiprakash Chauhan | ||||
| Eastern-Southeastern India | Sora (srb) |
Opino Gomango (Fellow) | Ramani Dalbehera, Namad Dalbehera, Opino Gomango | ||
| Juray (juy) |
Opino Gomango (Fellow) | Dinabandhu Gomango, Manjula Bhuyan, Srinivas Gomango | |||
| Juang (jun) |
Opino Gomango (coordinator) | ||||
| Gorum/Parengi (pcj) |
Opino Gomango (coordinator) | ||||
| Lambadi (lmn) |
Nenavath Mohan (Fellow) | Nenavath Mohan and Meghavath Sathish | |||
| Nepal | Saptariya Tharu (thq) |
Sanjib Chaudhary (Fellow) | |||
| Raji (rji) |
Uday Raj Aaley (Fellow) | ||||
| Sri Lanka | Sri Lankan Malay (sci) |
Sajehan Buckman |
-
Uday Raj Aaley, researcher and OpenSpeaks Fellow, speaking with Kaliprasad Raji, a Raji-language elder during documentation in a public place in the latter's village
-
Fellow Opino Gomango interviewing Kshetrabasi Gomango to document Juang language
-
In community-based language documentation like the Ho-language documentation here, often a group of community members reviews what is being recorded. This process informed our Oral History Framework.
Publications
- Panigrahi, Subhashish (2026-04-17). "Vulnerable language speakers fear that their unique and sensitive knowledge will be exploited or abused through AI". Global Voices. Global Voices. Retrieved 2026-04-17.
- (Panigrahi, Subhashish), (Gomango, O.), (Pal, K.) (2026), OpenSpeaks Archives: Citing Low-Resourced Language Oral History Multimedia, Wikimedia Foundation
- Panigrahi, Subhashish (2026-03-31). "OpenSpeaks Archives: Language Documentation Field diary (July 2025–March 2026)". Diff. Retrieved 2026-04-05.
Tools
- Subtitler: offline-capable caption/subtitle editor and translator (can fetch and upload/edit Commons' TimedText subtitles)
- Tome: audio, video, images, and document metadata editor (can also upload media and metadata to Commons)
- Bento: audio and video file/folder organiser, analyser and compressor
Open Educational Resources
- Oral Knowledge Framework for documenting oral knowledge using FAIR-CARE principles
- OpenSpeaks Text Style Guide for captioning/subtitling/transcription
Awareness & capacity building
- Larsen, Solana (2026-06-18). "Roundtable Recap: How to digitise a physical archive to work strategically with AI?". Open Knowledge Blog - Content for a fair. Retrieved 2026-07-22. (recording)
- "Subhashish Panigrahi, OpenSpeaks". Notes from a Maintainer. 2026-05-26. Retrieved 2026-05-27.
- "Indigenous Data Sovereignty: Centering Justice, Care, and Co-Ownership in the Digital Economy". RightsCon. IT for Change, Samdhana Institute, and ETC Group. (upcoming panel)
- Prishchepova, Valentina (2026-04-17). "Oberseminar Program @ HDSM: SoSe 2026". Resistance AI? Pespectives from Digital Humanities. HDSM TU Darmstadt. Retrieved 2026-04-19. (lecture by Subhashish Panigrahi titled, "Oral history is knowledge: lessons from low-resourced language documentation")
- "How we're building a language archive using open source and open licensing". FOSS United. FOSS United. 2026-03-28. Retrieved 2026-03-29. (deck)
- "Explorer Spotlight". National Geographic Society and National Centre for Biological Sciences (NCBS). 2026-01-29. Retrieved 2026-03-16.
- "Cultural Rights, Innovation, and Development in the AI Moment -Towards a Public Domain Framing". UNESCO civil society network on AI ethics and policy. (virtual, 10 March 2026)
- "Wikimedia Futures Lab". Wikimedia Deutschland. 2026-01-29. Retrieved 2026-03-05. (screened of Gyani Maiya, discussion on citation of oral history for low-resourced languages; co-built alpha prototype of WikiVoice, a platform for reliable, citable oral history together with Tochi Precious and Biyanto R.)
- "Indigenous languages and Small Language Models: Creating Open Source Protocols for Community Toolkits". Official Pre-Summit Event of the AI Impact Summit 2026. Design Beku. 2026-01-21.
- "Digital Tools and Strategy for Indigenous Languages" (PDF). WikiConference Kerala 2025. 2025-12-21. Retrieved 2025-12-24. (video)
- "Speaking in Our Voices: Preserving Local Languages through Wikimedia Projects". Wiki in Africa. 2025-12-02. Retrieved 2025-12-02.
- "Strategies to Increase Equitable Forms of Exchange and Partnerships in Support of Linguistic Capital and Language Justice". Fosterlang. FosterLang, Linguapax International. 2025-10-29. Retrieved 2025-12-02.
- "When Can We Cite Low-Resourced Language Oral Histories in Wikimedia Projects?", Celtic Knot Conference 2025 (virtual, 23 September 2025)
- Language Documentation for Open Knowledge Workshop, Dehradun, India (15 August 2025)
- "OpenSpeaks Archives: Language digital archive for Wikimedia projects ", Wikimania 2025 Nairobi (virtual, 8 August 2025)
- "Oral Knowledge in the Digital Commons: Reflections on Documentation and Citational Practice". Future of the Commons Collective (virtual, 28 July 2025)
In media
- Panigrahi, Subhashish (2026-05-10). "Archiving Nepal's medicinal plants and Indigenous knowledge for future generations". Global Voices. Retrieved 2026-05-11.
- Panigrahi, Subhashish (2026-04-17). "Language documentation needs community rights, consent, and recognition: Interview with Van Gujjari writer Taukeer Alam". Global Voices. Retrieved 2026-05-11.
- Shyn, Oleksandr (2025-12-02). "The Untranslatables Project: Johar, with Subhashish Panigrahi". Radio Taiwan International. Retrieved 2025-12-02.
Friends and collaborators
People involved
References
- ↑ Panigrahi, Subhashish; Gomango, Opino; Pal, Kimmi (2026-03-25). OpenSpeaks Archives: Citing Low-Resourced Language Oral History Multimedia. Wiki Workshop 2026. Online: Wikimedia Foundation.
Cite as
- "OpenSpeaks Archives". Endangered Languages Archive. DOI: 2196/c6ab6125-379b-46b2-b756-ee5e1b0d744e. www.elararchive.org.. Retrieved 2026-03-25.
