Jump to content

Oulu Löyly/Documentation/May the Source Be With You

From Meta, a Wikimedia project coordination wiki

Process

[edit]

In the age of the source some enthusiastic people with similar problem statements connected in a group and gathered around a perceived problem, which in general was that we see that Cultural heritage Institutions are sharing less.

We used a method of What? - So what? - Now what?

We deepened and clarified the problem statement under What? We had a lot of talk and discussions how we have noticed this and how this problem was existing before AI but has been greatly amplified.

After this first discussion, we were ready to write our mission statement.

Mission statement

[edit]

What is the source of common worries and anxieties around the technologies that now surrounf the preservation inititatives that result in reduced sharing of cultural heritage data?

  • What is happening to data?
  • Who/what is using it and how?
  • How is technophopia and big tech governing the narrative?
    • What is true from it?
    • How does it effect the culture of sharing?

Continued discussions

[edit]

Then we moved on to the So what? going deeper on the effects of this.

Finally, we spent considerable time trying to figure out solutions, also reflecting on some of the weaknesses/limits of our proposals.

Cleaned up view of all our notes and ideas

[edit]

AI Transparency Demands

[edit]

From our solutions, we chose to describe some demands about AI transparency even more.

What?

[edit]
  • A "Cookie for bots and LLMs"
  • Bots to identify themselves
  • Builds reports of what has been scraped and preferably also for what (or what promt)
  • Data provenance should be included
  • Sourse attribution would be included

Why?

[edit]
  • Misuse of opt-out materiariels would be treasable
  • Reports for archives for measurements of use (more use = better success)
  • Archives can trace the use of theis data and use the it as success stories
  • Archives would understand their value and the value of sharing: also for AI
  • The bias will stop growing because data from oroginal sourses are found and used

Also needed

[edit]
  • Effective opt-out mechanism for trust & peace of mind when publishing more sensitive materials not suitable for AI

Regarding reuse that encourage sharing (Wikimedia specific)

[edit]

There are some technical improvements that may increase the incentives for people and institutions to share material to Wikimedia. Here is a few examples:

  • Notification: Your file was used - makes a user aware that an image shared on Wikimedia Commons is actually being used in a page/article in our ecosystem. Yes, some tools make it possible to retrieve this kinds of stats in bulk, but having it as a notification would give a signal that it was just put into use.
  • Receiving/Sending Pingbacks/Trackbacks/Refbacks - a mature technology for letting someone know that you are referring to them on your website. Was very common among blogs earlier, but kind of went out of style with social media platforms wanting to create walled gardens.
  • Better media attributions. Currently, you need to click-through an image to get the attribution for it. If the attribution was next to the media, like is commons in newspapers, that could encourage institutions to share more. This is is a prototype user script (but it could perhaps be improved by the new attribution API).