Talk:Community Wishlist/W518
Add topicAbsolutely not
[edit]No, no, and hell no. ChompyTheGogoat (talk) 23:48, 6 March 2026 (UTC)
- No thanks. I don't want Big Brother watching me. * Pppery * it has begun 19:58, 7 March 2026 (UTC)
- I just don't think what's proposed here is needed or would be very useful: users that comment things like that already get warned and maybe blocked if adequate. I don't see how this would be a real big problem on Wikimedia projects currently that this would help address effectively. Prototyperspective (talk) 22:40, 8 March 2026 (UTC)
- @ChompyTheGogoat, Pppery, and Prototyperspective: I feel like the Wikipedia community fails to protect users. Evidence of this includes 15 years of complaints that women get continuous harassment in Wikipedia, me having first-hand awareness that LGBTQ+ people experience disproportionate harassment and also that Wikimedia LGBT+ is a de-facto channel for receiving these complaints, because our other infrastructure fails to meet the need, and that the Wikimedia Foundation has spent many millions of dollars trying to be more welcoming to diverse demographics. I think we have great success in recruiting all kinds of people to Wikimedia projects, but we fail to keep them because of toxicity.
- We have never had a third-party evaluation of our systems for responding to reports of harassment. If we did, I have no doubt that any evaluation would immediately identify points of failure, and that an evaluation would also note our Movement aversion to talking about those points of failure as evidenced by a lack of conversation about it.
- I am not sure what the solution to hate speech is, but I know the solution is not to assume it does not exist, or to assume without checking that our processes for detecting and responding to it are sufficient for needs.
- I was a recent attendee at WikiConference 2025 which had a gunman come. He was tackled on stage without shooting. This follows a lot of other violent incidents, including 2014 malicious disruption at WikiConference NYC of that year. it is just my own theory, but I see a pattern by the people who do violence and harassment that they start by bringing hate conversation into the Wikimedia platform and get tolerance there. I think if we ever had a sociologist track who has done crazy violent disruptions of Wikimedia programs, then we we see a pattern of these people finding tolerance of their hate speech in Wikimedia platforms.
- Perhaps AI is not the solution but from another perspective, the humans we have with the resources available to them are currently also insufficient, and have been insufficient every year for many years. We have no plan to change.
- I appreciate the people who comment but also, I wish your comments were higher value. I hear you complaining but I am not hearing your ideas for a solution, or even whether you are willing to react to my claim that a problem exists. Bluerasberry (talk) 17:21, 20 March 2026 (UTC)
- "the humans we have with the resources available to them are currently also insufficient": The critical mass that exists on English Wikipedia and in a relatively small number of other places exists because the volunteer editors are empowered to be part of the solution. Take people out of the process, or cloak the evidence-gathering and analysis in LLM-based technical obscurity, and you may see a short term gain, but then stagnation will set in as people become further removed from the issues and have less agency over the process. Values and value judgments have to be regularly examined and renewed, by people who believe in the process. How else can we claim and hold the moral high ground?
- "The tool that I am imagining is one in which a user's edits can be entered, and the tool makes an objective evaluation based on a public training dataset about whether the text is hate speech.": How would the parameters of this value judgment be conformed to the prevailing community views? Right now, it is imperfectly debated in a usually-public forum. Some people enjoy that, and I acknowledge that others would just as soon be rid of it. But without that debate, aren't we left with nothing but a bare choice: accept or reject the LLM output? I don't think we're ready for that, because it is almost inconceivable that the technology is ready for that.
- "I have the idea that other platforms - the major social media platforms - very well manage hate speech with automated tools": Do they? They have no ethos that we would recognize as such, and avoid accountability for individual cases as if only the aggregate statistics matter.
- "presume that Nazi speech is not possible to detect reliably": Maybe it is, maybe it isn't. Relative to ostensibly normal speech, it's probably easy to detect. But what should we expect with regard to statements consistent with other broadly nationalist, authoritarian, traditionalist or ethnocentrist movements, that have not yet engaged in genocide or wars of conquest? At least with humans running the show, we focus on the merits of the argument and evidence before us, and draw a line, subject to revision as the consensus shifts. That flexibility and locus of discussion is what breaks the taboo of calling out objectionable statements—in contrast to preordained prohibitions or unintelligible LLM correlations.
- TheFeds 03:05, 18 April 2026 (UTC)
- Other platforms ABSOLUTELY do not manage it well with existing tools. Facebook has massive bias and will ignore things up to and including overt threats against oppressed demographics, while overreacting to comments that are negative towards men, white folks, etc. Reddit leaves almost all of it up to mods on individual subs. Any tool will inevitably reflect the bias of the programmers, not community consensus. ChompyTheGogoat (talk) 05:37, 19 April 2026 (UTC)
- I just don't think what's proposed here is needed or would be very useful: users that comment things like that already get warned and maybe blocked if adequate. I don't see how this would be a real big problem on Wikimedia projects currently that this would help address effectively. Prototyperspective (talk) 22:40, 8 March 2026 (UTC)
- This is how the community dies: delegate decision of what can and cannot be said to a machine out of the community's control Ita140188 (talk) 14:12, 8 June 2026 (UTC)
- Personally, I don't think it is a pure bad idea. In my imagination, it will be a feature to detect hate speech or other uncivil comments in the discussion, and the tool will flags the comment to sysops (or patrollers) to review if there is any enforcement needed, or if it is just a false positive. It may possibly be useful for sysops as it can help to figure out the hostile comments (especially if the discussion is not actively monitored) and the sysops can take action on time. Thanks. SCP-2000 16:18, 18 June 2026 (UTC)
Update from WMF
[edit]Thank you for submitting your wish! We are working with internal stakeholders to figure out next steps, such as feasibility and prioritization. We’ll keep you posted about it. MikeZ-WMF (talk) 14:51, 10 March 2026 (UTC)
- The team responsible for this is focused on other priorities, but they may consider it in the future for the next planning year, so we’ll mark this as a long-term opportunity for now, and revisit in due time.
- To read about what the team is currently focused on, see the Product & Technology OKRs. MikeZ-WMF (talk) 16:03, 20 March 2026 (UTC)
- @MikeZ-WMF: I sincerely hope that you do come back and post an update in due time. Bluerasberry (talk) 17:21, 20 March 2026 (UTC)
- Hi @Bluerasberry - just to close the loop, relevant efforts to address negative comments has been prioritized by our Product Safety and Integrity team. They will follow up with more details. MikeZ-WMF (talk) 14:32, 5 June 2026 (UTC)
- @Bluerasberry We have actually begun work in the last couple of months on a project related to using AI-backed tools to detect different kinds of content that violates community policies:
- https://www.mediawiki.org/wiki/Product_Safety_and_Integrity/Detecting_abusive_content
- As that says, for right now we're focused on suppressible content (e.g. doxxing), but we plan to explore other abusive content as well. We expect that both we and the community will want a fairly high bar of precision for this kind of thing, so part of this will be figuring out what categories of abusive content are most amenable to this sort of approach.
- This work is also what we're referencing in the draft Annual Plan for next FY, where it says " Platform interventions based on centrally managed automation, that precisely and usefully implements community policies, to detect and contain bad-faith activity earlier in the process.". EMill-WMF (talk) 21:00, 5 June 2026 (UTC)
- We've been discussing this some more internally. Neither "Prioritized" nor "Declined" really appropriately reflects the state here. We aren't currently focused on detecting the specific kind of content being asked for in this wish and don't know if we will do so, and we want to be straightforward about that. But we are doing some clearly closely related foundational work, which if it's successful would help us (and community members) understand the feasibility and trade-offs that that would go into what's being specifically asked for here.
- It may be that we need a new designation entirely for situations like this, where we're doing work that is obviously related, and may even be influenced by the wish, but is not the exact wish. That's something we're considering. In the meantime, we're going to move this back to "Under Review" (this may take a couple days, since I don't have the permissions to do that and it's a Saturday). Our apologies for the rapid-fire status changes here, and please do feel welcome to engage on the related project page as well. EMill-WMF (talk) 21:29, 6 June 2026 (UTC)
- I would suggest "we're doing work that is obviously related, and may even be influenced by the wish, but is not the exact wish" is not a meaningfully different state than any other wish which is eligible to receive votes, other than perhaps notifying those who commented on a wish that this new project exists. Wishes are specific things that the community desires technical expertise in implementing and some cousin project that is achieving a WMF objective is not the same thing. Sometimes the editors of Wikipedia really do understand important things that should exist but don't and the Wishlist should be our way to help make those things happen. Best, Barkeep49 (talk) 23:34, 6 June 2026 (UTC)
- Hi @Bluerasberry - just to close the loop, relevant efforts to address negative comments has been prioritized by our Product Safety and Integrity team. They will follow up with more details. MikeZ-WMF (talk) 14:32, 5 June 2026 (UTC)
- @MikeZ-WMF: I sincerely hope that you do come back and post an update in due time. Bluerasberry (talk) 17:21, 20 March 2026 (UTC)
What constitutes as "hate speech"
[edit]I think that's a bigger question about your request. What makes LLM language == hate speech? Is it certain words and phrases? We are still trying to detect LLM language on a basic level and it's not 100%. I can a big can of worms opening with this as it's too broad. The Grid (talk) 15:02, 15 April 2026 (UTC)
- If I understand the proposal right, it's not saying LLM language == hate speech. The tool is what would be AI and it would be detecting patterns of behavior in manual edits made by users. Certain words and phrases can objectively be shown to come from Nazi circles so there wouldn't be much of a problem in that area. No filter on Wikipedia is reliable, but they are very useful for detecting long-term abuse and then having human review for the actual blocking etc.
- Whatever about the sociological claims, it would be bad for PR if someone started hiding crypto-fascist content in articles especially considering the Anti-Defence League have researched antisemitism on Wikipedia in the past. Mullafacation (talk) 14:26, 16 April 2026 (UTC)
Objective(?)
[edit]“ the tool makes an objective evaluation based on a public training dataset about whether the text is hate speech ”
I strongly doubt that speech could be evaluated objectively, but using a "public training dataset" is definitely not a method of objective evaluation, as the contents of such dataset reflect the subjective opinions of its creators.
Consider the following: If someone proclaims that they oppose gay marriage, is that hate speech? If yes, does supporting an organization that opposes gay marriage constitute hate speech? I being Catholic hate speech?
Is supporting post-WWII deportations of Germans and of other ethnic minorities hate speech? But what if a person who says this lives in a country which happens to criminalize opposing these measures.
Is saying "there are only men and women, and sex is determined biologically" an anti-trans dogwhistle? But what if I happen to live in a country where the constitution happens to say so?