Grants:Programs/Wikimedia Community Fund/Rapid Fund/UrduLitGraph:Expanding an Open Dataset for Computational Analysis of Urdu Literature (ID: 23911736)
This is an automatically generated Meta-Wiki page. The page was copied from Fluxx, the web service of Wikimedia Foundation Funds, where the user has submitted their application. Please do not make any changes to this page because all changes will be removed after the next update. Use the discussion page for your feedback. The page was created by CR-FluxxBot.
Applicant Details
[edit]- Main Wikimedia username. (required)
Hm849c (talk • accounts • contributions • edit count)
- Organization
N/A
- If you are a group or organization leader, board member, president, executive director, or staff member at any Wikimedia group, affiliate, or Wikimedia Foundation, you are required to self-identify and present all roles. (required)
N/A
- Describe all relevant roles with the name of the group or organization and description of the role. (required)
Main Proposal
[edit]- 1. Please state the title of your proposal. This will also be the Meta-Wiki page title.
UrduLitGraph: Expanding an Open Dataset for Computational Analysis of Urdu Literature
- 2. and 3. Proposed start and end dates for the proposal.
2026-09-30 - 2026-12-31
- 4. Where will this proposal be implemented? (required)
Pakistan
- 5. Are your activities part of a Wikimedia movement campaign, project, or event? If so, please select the relevant project or campaign. (required)
Not applicable
- 6. What is the change you are trying to bring? What are the main challenges or problems you are trying to solve? Describe this change or challenges, as well as main approaches to achieve it. (required)
Urdu literature is a rich cultural, historical, and intellectual resource, but it remains underrepresented in computational research and open knowledge ecosystems. A major challenge is the lack of structured, machine-readable datasets for Urdu novels. Many Urdu literary works are available only in printed form or in formats that are not suitable for computational analysis. This limits research in Urdu natural language processing, digital humanities, literary studies, authorship analysis, and educational reuse.
In our prior research, published on arXiv, we developed a graph-based pipeline for analyzing Urdu novels using character interaction networks. That work used a dataset of 52 novels and demonstrated that graph representations of literary texts can support computational analysis such as authorship modeling and narrative structure analysis. However, the study also identified a clear limitation: the available Urdu literary dataset is still too small for broader research and reuse.
This proposal aims to address that gap through Phase 1 of UrduLitGraph, a larger effort to expand and structure Urdu literary data for open research. The project will collect, digitize, process, and structure 50–80 additional Urdu novels into character interaction graphs, character annotations, metadata, and reusable documentation.
The main change we want to bring is to make Urdu literature more accessible for computational and open knowledge work. Instead of keeping Urdu literary material locked in printed books or unstructured text, this project will create structured derived datasets that can be used by students, researchers, digital humanities scholars, and open knowledge contributors.
Our approach has four main strategies:
Digitization and processing of Urdu literary works We will acquire selected Urdu novels, scan them where necessary, and use OCR to create machine-readable text for further processing. Transformation into structured graph-based data We will extract and normalize character entities, construct character interaction graphs, and prepare node-level features and annotations. Open documentation and reusable pipeline development We will document the processing steps so that the work can be extended in later phases and reused by others working on low-resource languages. Open knowledge alignment Where legally permissible, we will release derived data such as metadata, graphs, annotations, and documentation. We will also explore how relevant outputs can support Wikimedia platforms such as Wikisource, Wikidata, or Wikipedia, especially through metadata, bibliographic information, and structured knowledge about Urdu literary works.
This approach is based on our prior implementation and research experience, so the project is not starting from zero. Phase 1 will build directly on an existing pipeline and dataset, while improving scale, documentation, and potential Wikimedia relevance.
- 7. What are the planned activities? (required) Please provide a list of main activities. You can also add a link to the public page for your project where details about your project can be found. Alternatively, you can upload a timeline document. When the activities include partnerships, include details about your partners and planned partnerships.
The project will be implemented over approximately 4–6 months through the following activities:
1. Book selection and acquisition
We will identify 80–120 Urdu novels that are suitable for inclusion in the expanded dataset. Priority will be given to works that are important for Urdu literary studies, underrepresented in computational datasets, and useful for future research. We will purchase printed books where digital versions are not available and explore collaboration with authors, publishers, or libraries where possible.
2. Scanning and digitization
Printed books will be scanned using a dedicated scanner or consistent scanning setup to improve OCR quality. Scanning will be organized carefully to ensure readable page images and consistent formatting.
3. OCR processing
We will use a hybrid OCR approach. Open-source tools such as Tesseract will be tested and used where possible. For selected difficult or high-priority subsets, paid OCR services may be used to improve accuracy. OCR outputs will be checked and organized for downstream processing.
4. Text cleaning and post-processing
OCR text will be cleaned to remove common OCR noise, page artifacts, formatting issues, and recognition errors. This step is especially important for Urdu because OCR errors can affect character name extraction and graph construction.
5. Character extraction and normalization
Character names will be extracted from the processed texts. Variations of names, honorifics, spelling differences, and OCR-based distortions will be normalized. This will help ensure that the same character is not incorrectly treated as multiple separate entities.
6. Character annotation
Characters will be annotated with relevant attributes where possible, such as gender, narrative role, and other useful features. Annotation will be done carefully and documented so that the dataset remains reusable and transparent.
7. Graph construction
Character interaction graphs will be constructed using co-occurrence-based methods. Nodes will represent characters, while edges will represent interactions or co-occurrence relationships within the narrative. Graph outputs will be prepared in reusable formats.
8. Dataset structuring
The project outputs will be organized into structured folders and formats, including graph files, node features, annotations, metadata, and documentation.
9. Validation and quality checks
A sample of outputs will be manually reviewed to check OCR quality, entity normalization, graph construction, and metadata consistency. Where errors are found, the pipeline and outputs will be corrected.
10. Documentation and open release
We will prepare documentation describing the dataset structure, processing workflow, limitations, and reuse instructions. Derived data such as graphs, annotations, metadata, and processing code will be released openly where legally permissible. Full text will only be shared when it is legally allowed, such as public domain works or works with permission.
11. Wikimedia alignment
We will explore how the dataset can support Wikimedia projects. Possible contributions include improving bibliographic metadata, creating or improving relevant Wikidata items, supporting Urdu literature pages on Wikipedia, and identifying public domain texts that may be suitable for Wikisource.
- 8. Describe your team. Please provide their roles, Wikimedia Usernames and other details. (required) Include more details of the team, including their roles, usernames, Wikimedia group, and whether they are salaried, volunteers, consultants/contractors, etc. Team members involved in the grant application need to be aware of their involvement in the project.
The project will be led by Hassan Mujtaba, whose work focuses on computational analysis of Urdu literature, graph-based modeling, and low-resource language datasets. Hassan has prior research experience through an arXiv-published study on Urdu literary character interaction networks using a dataset of 52 novels and will oversee the overall implementation, research design, quality assurance, and reporting for this project.
The core project team includes:
1. Project Lead / Research Coordinator – Hassan Mujtaba Responsible for project planning, dataset design, graph construction, quality control, documentation, community engagement, and final reporting. Wikimedia Username: Hm849c Status: Grant-supported lead
2. Research Collaborator – Hamza Naveed Responsible for supporting dataset development, annotation review, validation activities, research coordination, and documentation of project outputs. Status: Grant-supported Collaborator
In addition to the core team, the project may engage short-term hired assistance or contractors for specific operational tasks where needed, subject to the approved budget. These may include:
3. Digitization and OCR Assistant Responsible for scanning books, organizing page images, running OCR tools, and preparing OCR outputs for cleaning.
4. Additional Support Roles
Depending on project needs, additional assistance may be provided for character annotation, entity normalization, data validation, book selection, literary review, and technical guidance related to the processing pipeline, dataset release, and long-term sustainability of project outputs.
All team members involved in the grant will be informed about their roles before submission. No confidential personal information will be included in the public proposal.
- 9. Who are the target participants and from which community? How will you engage participants before and during the activities? How will you follow up with participants after the activities? (required)
The target participants and beneficiaries include:
1. Urdu NLP researchers and students 2. Digital humanities researchers working on South Asian literature 3. Urdu literature scholars and educators 4. Wikimedia contributors interested in Urdu literature, Wikisource, Wikidata, and Wikipedia 5. Open knowledge communities working on low-resource languages 6. Experienced Urdu literary readers, critics, and authors who have deep knowledge of Urdu literature and can provide valuable insights into book selection, character interpretation, and literary context
Before the activities, we will engage relevant participants by sharing the project idea with Urdu language, Wikimedia, research, and literary communities. We will request feedback on book selection, dataset structure, metadata fields, annotation approaches, and possible Wikimedia integration. In particular, we will consult experienced Urdu readers and authors to help identify significant novels, assess the suitability of selected works, and provide guidance on literary and cultural considerations.
During the project, engagement will happen through progress updates, documentation, and selected review requests. Community members may help identify important Urdu novels, suggest metadata improvements, review sample outputs, advise on Wikimedia relevance, and provide feedback on the overall processing workflow. Urdu literary readers and authors will be invited to review aspects of the character extraction, normalization, annotation, and graph construction process to ensure that the outputs remain meaningful and useful from a literary perspective.
After the project, we will follow up by sharing the released dataset, documentation, GitHub repository, and final report. We will invite feedback for future phases, including expansion to 200+ novels and deeper integration with Wikimedia projects. Feedback from researchers, Wikimedia contributors, Urdu literary readers, and authors will help guide improvements to future dataset releases and project development.
- 10. Does your project involve work with children or youth? (required)
No
- 10.1. Please provide a link to your Youth Safety Policy. (required) If the proposal indicates direct contact with children or youth, you are required to outline compliance with international and local laws for working with children and youth, and provide a youth safety policy aligned with these laws. Read more here.
N/A
- 11. How did you discuss the idea of your project with your community members and/or any relevant groups? Please describe steps taken and provide links to any on-wiki community discussion(s) about the proposal. (required) You need to inform the community and/or group, discuss the project with them, and involve them in planning this proposal. You also need to align the activities with other projects happening in the planned area of implementation to ensure collaboration within the community.
This proposal is based on work we have already carried out on a smaller scale through our previous research on graph-based computational analysis of Urdu novels. During that work, we encountered several challenges related to the availability, quality, and structure of Urdu literary data. These challenges highlighted the need for a larger and more systematically organized dataset for Urdu literature.
To better understand the problem and its broader impact, we discussed our findings and ideas with a range of stakeholders, including Urdu readers, literature professors, authors, and researchers. These conversations confirmed that the lack of accessible and structured digital resources is a significant barrier not only for literary research but also for the preservation, study, and wider appreciation of Urdu literature.
We also discussed the issue with computer scientists and researchers interested in natural language processing and digital humanities. Many expressed that the scarcity of high-quality Urdu datasets limits their ability to develop tools, conduct experiments, and contribute to research involving the Urdu language. As a result, Urdu remains underrepresented in many computational and open knowledge initiatives compared to higher-resource languages.
The feedback we received helped shape the goals of this project. Community members emphasized the importance of creating reusable resources that can benefit multiple groups, including readers, writers, literary scholars, students, and computer scientists. This project aims to address a shared need by expanding the availability of structured Urdu literary data and making it more useful for research, education, and open knowledge initiatives.
We will continue engaging relevant communities throughout the project by sharing progress updates, seeking feedback on dataset design and documentation, and exploring opportunities for future collaboration with Wikimedia contributors and Urdu language communities.
- 12. Does your proposal aim to work to bridge any of the content knowledge gaps (Knowledge Inequity)? Select one option that most apply to your work. (required)
Language
- 13. Does your proposal include any of these areas or thematic focus? Select one option that most applies to your work. (required)
Culture, heritage or GLAM
- 14. Will your work focus on involving participants from any underrepresented communities? Select one option that most apply to your work. (required)
Linguistic / Language
- 15. In what ways do you think your proposal most contributes to the Movement Strategy 2030 recommendations. Select one that most applies. (required)
Innovate in Free Knowledge
Learning and metrics
[edit]- 17. What do you hope to learn from your work in this project or proposal? (required)
The project will help us learn how to build a scalable and legally responsible pipeline for expanding structured Urdu literary datasets. The main learning questions are:
1. How accurately can Urdu literary texts be digitized using a hybrid OCR approach that combines open-source tools and selected paid OCR services? 2. What are the most common OCR and normalization challenges in Urdu novels, especially for character names, honorifics, spelling variations, and narrative references? 3. How much manual review is needed to produce reliable character interaction graphs from OCR-based Urdu text? 4. Which metadata and annotation fields are most useful for researchers, students, and Wikimedia contributors? 5. What forms of derived data can be released openly while respecting copyright restrictions on full literary texts? 6. How can this dataset support future Wikimedia-related work, such as Wikidata bibliographic items, Urdu literature pages, or possible public domain Wikisource contributions? 7. What improvements are needed before scaling the project from 80–120 additional novels to 200+ novels?
- 18. What are your Wikimedia project targets in numbers (metrics)? (required)
| Other Metrics | Target | Optional description |
|---|---|---|
| Number of participants | 10 | These may include researchers, students, Wikimedia contributors, Urdu literature scholars, and community reviewers who provide feedback, review sample outputs, or participate in project discussions. |
| Number of editors | 3 | These will be people who may create or improve Wikimedia-related content as part of the project, such as Wikidata items, Wikipedia references, or documentation pages related to Urdu literary works |
| Number of organizers | 4 | This includes the project lead, digitization/OCR assistant, annotation assistant, and any technical or Wikimedia advisor involved in implementation. |
| Wikimedia project | Number of content created or improved |
|---|---|
| Wikipedia | 3 |
| Wikimedia Commons | |
| Wikidata | 30 |
| Wiktionary | |
| Wikisource | 3 |
| Wikimedia Incubator | |
| Translatewiki | |
| MediaWiki | |
| Wikiquote | |
| Wikivoyage | |
| Wikibooks | |
| Wikiversity | |
| Wikinews | |
| Wikispecies | |
| Wikifunctions or Abstract Wikipedia |
- Optional description for content contributions.
N/A
- 19. Do you have any other project targets in numbers (metrics)? (optional)
Yes
| Main Open Metrics | Description | Target |
|---|---|---|
| Urdu Novels Processed | Number of Urdu novels selected, digitized, cleaned, and processed through the project workflow. | 120 |
| Character Interaction Graphs Created | Number of validated character interaction graphs generated from processed Urdu novels. | 120 |
| Metadata and Entity Lists Prepared | Number of novels for which structured metadata and character entity lists are completed. | 120 |
| Node-Level Annotations Completed | Number of processed novels with documented node-level character annotations included in the dataset. | 120 |
| Documentation and Dataset Release | Completion of a documented processing pipeline, public repository or dataset release page, and a future expansion plan for scaling to 200+ novels. | N/A |
- 20. What tools would you use to measure each metrics? Please refer to the guide for a list of tools. You can also write that you are not sure and need support. (required)
We will use the following tools and methods to measure project targets:
1. GitHub repository tracking To track dataset files, graph outputs, code, documentation, issues, and release versions. 2. Project spreadsheet To track books acquired, scanned pages, OCR status, cleaning status, character extraction status, graph construction status, annotation status, and validation status. 3. OCR logs and page counts To measure the number of pages processed and compare OCR quality across open-source and paid OCR tools. 4. Manual validation sheets To record review results for character extraction, normalization, graph structure, and metadata consistency. 5. Wikimedia contribution tools Wikimedia project histories, user contribution pages, Wikidata item histories, and relevant dashboards if used. 6. Final project report To summarize completed outputs, lessons learned, limitations, metrics, and recommendations for the next phase.
Financial proposal
[edit]- 21. Please upload your budget for this proposal or indicate the link to it. (required)
- 22. and 22.1. What is the amount you are requesting for this proposal? Please provide the amount in your local currency. (required)
1250000 PKR
- 22.2. Convert the amount requested into USD using the Oanda converter. This is done only to help you assess the USD equivalent of the requested amount. Your request should be between 500 - 5,000 USD.
4487.5 USD
- By submitting your proposal/funding request you confirm that you have read and agree to the Application Privacy Statement, WMF Friendly Space Policy, and the Universal Code of Conduct.
Yes
Endorsements and Feedback
[edit]Please add endorsements and feedback to the grant discussion page only. Endorsements added here will be removed automatically.
Community members are invited to share meaningful feedback on the proposal and include reasons why they endorse the proposal. Consider the following:
- Stating why the proposal is important for the communities involved and why they think the strategies chosen will achieve the results that are expected.
- Highlighting any aspects they think are particularly well developed: for instance, the strategies and activities proposed, the levels of community engagement, outreach to underrepresented groups, addressing knowledge gaps, partnerships, the overall budget and learning and evaluation section of the proposal, etc.
- Highlighting if the proposal focuses on any interesting research, learning or innovation, etc. Also if it builds on learning from past proposals developed by the individual or organization, or other Wikimedia communities.
- Analyzing if the proposal is going to contribute in any way to important developments around specific Wikimedia projects or Movement Strategy.
- Analysing if the proposal is coherent in terms of the objectives, strategies, budget, and expected results (metrics).