Jump to content

Grants talk:Programs/Wikimedia Community Fund/Rapid Fund/Making DelintBot able to reliably fix more lint categories (ID: 23811013)

Add topic
From Meta, a Wikimedia project coordination wiki

Review responses, questions and action items for your application

[edit]

Hello @Redmin,

Thank you for your application. We have completed the initial review of your application. Here are the combined comments from the review team (i.e. Wikimedia Bangladesh, SA Regional Funds Committee and SA (ex India) Programme Officer).

  • To summarise: Section 6 requires your response.


We look forward to your response by 17 May 2026. Please let us know if you require more time to respond. We can adapt.

  • A quick note, as your application has technology components, it is also undergoing a parallel Product and Technology review and you may receive additional questions from that process.


Thank you.

Regards, Jacqueline

--


Criteria Inputs/ Comments/ Questions
1) Is the proposal clear in terms of the change it wants to make? Does the proposal indicate value/impact potential? Is the proposal viable? YES

Detailed explanation

  • The proposal is very clear about the problem. The planned improvements are well-described and technically viable. The Phabricator tag and existing code show this is a natural evolution of ongoing volunteer work.
  • The proposal clearly explains the technical problem and shows good potential impact for improving Wikimedia content maintenance and reducing lint errors.
  • https://xtools.wmcloud.org/ec/meta.wikimedia.org/DelintBot and https://xtools.wmcloud.org/ec/meta.wikimedia.org/RedminBot show sizable number of global edit counts in Bengali Wikimedia projects which serve as proof of work.
  • Innovation and learning: The project demonstrates a good level of technical innovation, particularly in its approach to automated lint fixing using validation and visual regression methods. However, opportunities for collaboration and learning within the wider community are limited, as the project is largely designed as an individual effort. Incorporating more structured plans for community engagement, knowledge sharing, or onboarding contributors would strengthen its broader impact.
2) Does the applicant have the experience (technical / organizing) capacity to implement this project? Does the applicant or team have experience on target Wikimedia projects? YES

Detailed explanation

  • Strong technical background: active since 2020, multiple bots (RedminBot, LexemeBot), MediaWiki extensions/skins and public code on GitLab. Has already secured bot approvals on Bengali projects and Wikidata. Single-person team is the main limitation, but the applicant’s track record supports successful delivery.
  • The applicant appears to have sufficient technical experience and understanding of Wikimedia projects to complete the work successfully.
  • Yes, technically strong, but structurally fragile due to being a single-person project. The proposal is entirely dependent on a single individual, which introduces risks related to continuity, scalability, and long-term maintenance.
3) Does this proposal have support from Wikimedia community members?  Has there been sufficient engagement of community members through the endorsement and feedback process? TO SOME EXTENT

Detailed explanation

  • The applicant links to bot approval discussions on Bengali Wikibooks, Bengali Wiktionary, and Wikidata, which demonstrates meaningful community engagement on the wikis where the bot already runs. However, the proposal mentions expanding to more Wiktionary communities without evidence of outreach already begun to those communities. Community support is solid where the bot is deployed but remains untested for the expansion scope.
  • There is visible community interest, though additional endorsements and broader engagement would strengthen the proposal further.
4) Does the proposed budget adequately reflect the investment needed to achieve the proposed goals? TO SOME EXTENT

Detailed explanation

  • The total amount of 2,348 USD is reasonable. However, the 33 units × 8,500 BDT for programming needs clearer justification (what does one unit represent? hours/days per task?). Other line items are minimal. With better breakdown linking units to specific deliverables (visual testing, concurrency, etc.), it would be fully adequate.
  • The proposed budget seems reasonable and aligned with the project goals and activities.
  • It is unclear why he cannot continue his bot development and running in volunteer capacity. The number of units mentioned in the budget is unclear. Is it 33 hours or 33 modules?
  • While the overall budget is within a reasonable range, the programming cost appears relatively high given the lack of clarity in how it is calculated. The proposal does not define what constitutes a “unit” or how the 33 units relate to specific tasks or deliverables. Without a clear breakdown of effort, it is difficult to assess whether the cost is justified. A more detailed explanation linking time, tasks, and expected outputs would strengthen the budget’s credibility.
5) Does the applicant show clarity in what they hope to learn from their work given the change they are hoping to achieve? TO SOME EXTENT

Detailed explanation

  • The applicant demonstrates good learning orientation through iterative improvement (simple regex to robust validation), documentation plans, and openness to community feedback. Clear focus on making the bot more reliable and scalable. However, the proposal would benefit from time-bound milestones within the grant period.
  • The proposal clearly describes expected learning outcomes related to improving automated lint fixing processes.
  • The project addresses a problem at significant scale and proposes technically sophisticated solutions. However, the proposed metrics do not yet match this level of ambition. While reducing lint errors is an appropriate core outcome, the proposal would benefit from more specific, time-bound, and quantifiable targets (e.g., percentage reduction per wiki, number of wikis supported, or expected error reductions within the grant period). Strengthening the structure with clearer milestones, measurable targets, and better alignment between problem, approach, and expected outcomes would significantly improve its quality.
6) Questions or feedback to the applicant based on your review We need a response to these questions

a) Programming cost & unit definition: Could you define what one unit represents in the programming line item (e.g., hours, days, or task blocks), and provide an approximate breakdown of how the 33 units are allocated across your six planned deliverables?

b) Timeline: Could you provide time-bound targets for the grant period for example, what error reduction percentage you expect on Wikidata and the Bengali wikis within 3 months vs. 6 months? How many errors do you hope to fix?

c) Community engagement: Have you begun any outreach to the additional Wiktionary communities you plan to approach during the project? If so, what has the initial response been? Could you provide the list of projects in which you would run these bots? Would they include languages other than Bengali?

d) Core team: This project depends entirely on one person. How will you mitigate continuity risks? Do you have any plans to involve other developers or communities (e.g., through documentation, onboarding, or collaboration), to reduce reliance on a single maintainer?

e) Budget justification: Some items, such as documentation and deployment, appear under-costed relative to their complexity. Could you provide more detailed assumptions behind these costs?

f) Quality control & risk mitigation: Given that earlier deployments to Wikidata produced some incorrect edits, how will you measure and control error rates (e.g., revert rates, false positives) when scaling to more wikis? What specific thresholds or monitoring will you use to catch and halt problematic edits when scaling to new wikis?

g) Sustainability: While you mention long-term maintenance, could you elaborate on how the project will remain usable and maintained if you are unavailable?

JChen (WMF) (talk) 01:24, 12 May 2026 (UTC)Reply

Hello, @JChen (WMF). Thanks to all of you for taking the time to review this, and for the questions. My answers follow:
a) Yes, my sincere apologies for not explaining it before; I had not realized from checking out other applications that it was necessary. Each unit is a task block that includes development, debugging and deployment of the newly developed/updated code to verify the code works as expected on a wiki. Depending on the specific task, these take between two and three hours. This is an approximate breakdown of how the units are allocated across the primary deliverables:
  • “Implement fixes for anything not currently being fixed…”: 16 units
  • “Making the bot rely on visual regression testing…”: 5 units
  • “Write proper logic for generating edit summaries….”: 2 units
  • “Inform administrators exactly what needs to be changed…”: 1.5 units
  • “Using Wikimedia's REST APIs…”: 1.5 unit
  • “Making the bot perform edits concurrently…”: 1 unit
And this is a breakdown of how the additional units are allocated to secondary deliverables (documented on Phabricator):
  • Other improvements to how the bot operates (described in the tasks already filed on Phabricator) and refactoring code as needed as the codebase evolves over the course of the grant period: 4 units
  • Miscellaneous bug fixes as required (take this bug, for example): 2 units
b) Yes but please note that because widely used templates are often protected from editing by non-admins, exact predictions are hard to make as there is no guarantee that an administrator would be available to make the edit proposed by the bot and fix the errors and as such, these estimates are very conservative:
On Bengali Wikibooks:
Over the first three months:
~2850 of the about 5000 total lint errors (assuming all of the more than 2000 dark mode-related errors, 800 of all 996 of the errors from obsolete tags being used, 50 of the more than 150 errors from bogus file options and all 47 of the ‘miscellaneous issues’ can be fixed) or an error reduction of approximately 57% (achievable thanks to the fact that most templates with lint errors on this wiki are not protected)
Over the first six months:
~2860 of the about 5000 total lint errors (assuming all of the more than 2000 dark mode-related errors, 800 of all 996 of the errors from obsolete tags being used, 50 of the more than 150 errors from bogus file options, all 47 of the ‘miscellaneous issues’, all 12 of the errors from the ‘paragraph wrapping bug workaround’ and the sole error from ‘unclosed quote in heading’ can be fixed) or an error reduction of approximately 57.2% (achievable thanks to the fact that most templates with lint errors on this wiki are not protected)
Bengali Wikibooks is a very small wiki with only a little more than 25 thousand pages in total; it may thus be better to look at Malagasy Wiktionary than another small-ish wiki like Bengali Wiktionary since it is the largest Wiktionary the bot is currently approved on (and the third largest Wiktionary overall) with more than 6 million pages in total:
Over the first three months:
~109350 of the about 675000 total lint errors (assuming 109047 or about 70% of the nearly 156000 of the dark mode-related errors, 10 of the less than 20 errors from bogus file options and 300 of the 332 errors from the use of obsolete tags can be fixed) or an error reduction of approximately 16.2% (achievable thanks to the fact that most templates with lint errors on this wiki are not protected)
Over the first six months:
~119350 of the about 675000 total lint errors (assuming 109047 or about 70% of the nearly 156000 of the dark mode-related errors, 10 of the less than 20 errors from bogus file options, 10000 of the 19942 errors from misnested tags and 300 of the 332 errors from the use of obsolete tags can be fixed) or an error reduction of approximately 17.68% (achievable thanks to the fact that most templates with lint errors on this wiki are not protected)
On Wikidata, which has more active admins who could respond to the edit requests:
Over the first three months:
~168900 of the about 278900 total lint errors (assuming 148724 or 80% of the 185905 dark mode-related errors can be fixed, 200 of the 260 errors from bogus file options, 20000 of the 22156 errors from the use of obsolete tags can be fixed) or an error reduction of approximately 60.55%
Over the first six months:
~171400 of the about 278900 total lint errors (assuming 148724 or 80% of the 185905 dark mode-related errors can be fixed, 200 of the 260 errors from bogus file options, 2500 of the 2910 errors from misnested tags, 20000 of the 22156 errors from the use of obsolete tags and all 2 of the errors from unclosed quotes in headings can be fixed) or an error reduction of approximately 61.45%
Based on these, I would be targeting an error reduction of at least 50% on Bengali Wikibooks and 55% on Wikidata within the first three months and 57% on Bengali Wikibooks and 61% on Wikidata over the grant period. I would also keep track of the error reduction rate on other Wiktionary editions where the bot would receive approval (I expand more on that below). The number of additional errors fixed in the second 'half' of the grant period is lower than those fixed in the first because these are complex issues that are less prevalent on the wikis.
Please note that all of the total lint error counts excludes hidden lint categories as is recommend practice.
c) Yes, I have already reached out to twelve more Wiktionary communities, and the bot has since been approved on ten of them. The Spanish Wiktionary community decided not to use the bot; instead opting to fix some of the errors manually and by replacing custom signatures that produced most of the errors on that wiki with default signatures provided by MediaWiki. The discussion on the Czech Wiktionary has not ended yet but a contributor there has noted the problem caused by the bot's current inability to verify whether lint errors were fixed before saving the edit, which can lead to it saving edits where only improper Wikitext syntax that do not lead to any lint errors are fixed when the bot tries to fix dark mode-related errors in some cases. This would be addressed by the proposed use of Wikimedia's REST APIs to verify the merits of an edit, which I have planned for this exact reason, as well as general improvement of regular expressions used to find the styles that have to be fixed. Another volunteer expressed their concern about the semantic correctness of a specific replacement the bot carries out when dealing with chemical formulae, which can be easily fixed as I have noted in that discussion (but have not implemented as the community has not yet expressed whether it is important enough to them).
To be clear though, adding support for a new Wiktionary is currently as simple as changing three lines of code at most, and I hope to reach out to most (, if not all,) of the more than 150 active Wiktionary communities if this is approved but these are the results of primary outreach. I am happy to reach out to more of them at this time if you would like that.
d) This project is unique in that the planned edits would make sure the bot does not need to be run again on the wikis where it is run in the future once the project is completed so that in itself reduces any continuity risk. That being said, I welcome contributions from anyone and have guided a newcomer to the Wikimedia movement through this process in the past (and am currently trying to guide another through the same process). The documentation that is planned to be written would certainly help potential contributors as well since that would explain how to make changes to the code, and what piece of code has what purpose.
e) As for documentation, writing it is a less complex task as unlike most software, this is a bot, and so the documentation would not have to explain, for example, how to use its GUI, which would have made the task more complex. Instead, I would be able to focus on explaining how to make changes to the code, set up and run the bot on Toolforge with the target audience being anyone with a Wikimedia developer account without any assumption about their prior experience.
As for deployment of the visual regression testing system on Toolforge, I would only carry it out once I have worked on and tested it on my own computer (the cost for which is accounted for within the programming cost), and would be focused exclusively on Toolforge-specific deployment necessities rather than development of the system itself.
f) The first layer of quality assurance lies in running tests on all new code before they are ever deployed to check whether they would produce unexpected Wikitext, which I keep increasing in number to watch out for edge-cases. If, however, that fails to catch some error, that would be caught by the proposed visual regression testing mechanism if the edit would change the appearance of the page. If the appearance is not changed but the edit would introduce new lint errors or fail to fix existing lint errors, that would be caught by the proposed use of Wikimedia’s REST APIs. If even that fails to catch any error, such an error would not cause the visual appearance of the page to change but would produce incorrect Wikitext that does not cause lint errors and I would look out for reverts (I have set up cross-wiki notifications and emails for edit reverts on the bot’s account). If I notice a revert then I would stop the bot from making any further edits, fix any bug, go back and fix any pages where the erroneous edit(s) w[ere] made and only then restart the bot. I believe this would be adequate in capturing any error but am all ears if you think there is room for improvement.
g) All of the code would continue to be hosted on Wikimedia GitLab under the “toolforge-repos” namespace and the bot itself would continue to be hosted on Toolforge, which would allow any Wikimedia community member to adopt the project in the event of my unavailability in accordance with Toolforge’s well-established policy for dealing with abandoned tools as they could be given access to both the original Git repository and Toolforge tool account. Another way for someone to continue to maintain it would be by forking the repository, making changes to the code if needed (which is unlikely) and using Toolforge to make their own tool account and run the code there. The documentation which is planned to be written would explain how to do this, as is standard practice for established open source projects where the goal is to help anyone contribute to the code.
I hope this answers all of your questions. Thanks again for your review and feedback, and please let me know if there are any further questions, concerns or suggestions.
Best regards, Redmin (talk) 14:23, 15 May 2026 (UTC)Reply
Hello again, @JChen (WMF), I thought I would note regarding the answer to question c that the bot just received approval to operate on MediaWiki.org yesterday. (Not bothering you with an email since it does not change my answer too much but still noting it since that wiki is very different in structure to a Wiktionary.) Best regards, Redmin (talk) 16:55, 19 May 2026 (UTC)Reply

An update on your application - next steps

[edit]

Hello @Redmin,

Thank you for your responses. We are in touch via email to set up a time for a conversation to discuss next steps.

Thank you.

Regards, Jacqueline JChen (WMF) (talk) 04:04, 20 May 2026 (UTC)Reply

Your application has been approved

[edit]

Hello @Redmin

Thank you for your application and for taking time to respond to the review team's questions. Congratulations, your application has been approved in the amount of BDT 319,550 from 15 August 2026 to 30 June 2027.

  • I extended the implementation timeline. If you complete your project ahead of time, please feel free to submit your final report early via Fluxx.
  • I also added a 10% contingency to your project budget for unforeseen costs that may occur during the implementation of your project
  • I am also acknowledging your responses to questions consolidated in section 6 in the table above.  
  • Here’s also further feedback about your project from the Product and Technology team:
    • The review committee wanted to ensure that the applicant is aware of the mwbot-rs bot library in Rust since it offers mwbot and parsoid crates and mwbot-rs already includes a linter bot that uses the Parsoid API to fix lints and other wikitext issues. It might be worthwhile considering building on existing delinting solutions there. To be clear, this is not a blocker or a requirement, but more an encouragement to consider reuse and enhance existing solutions rather than build new ones from scratch.
  • As this is your first application, I will be reaching out to you to discuss the next steps post award.


Additional resources which may be useful


We thank you for your participation in the grant application process and hope to continue to journey with you as you embark on this project. Good luck!

Regards, Jacqueline on behalf of review team JChen (WMF) (talk) 03:44, 17 July 2026 (UTC)Reply

Hello @JChen (WMF),
Thank you very much for the review, approval and adjustments. I appreciate the kind suggestion from the Product and Technology team to use existing technology, but given that a substantial amount of code has already been written for the bot in Python and the amount of resources that would be needed to rewrite it in Rust, I am choosing not to use the mentioned Rust libraries.
Best regards, Redmin (talk) 16:44, 17 July 2026 (UTC)Reply