Microsoft to Credit AI Training Data Contributors - New Initiative

Microsoft Investigates Data Provenance in Generative AI
Microsoft has initiated a research undertaking focused on determining the extent to which specific training examples impact the outputs generated by artificial intelligence models, encompassing text, imagery, and various other media formats.
This initiative was revealed through a job posting originally published in December and recently gaining renewed attention on LinkedIn.
Project Goals and Methodology
The research, seeking a research intern, aims to demonstrate the feasibility of training models to allow for the “efficient and useful estimation” of the influence individual data points – such as photographs and literary works – have on the model’s generated content.
The job listing highlights the current lack of transparency in neural network architectures regarding the origins of their outputs, stating, “Current neural network architectures are opaque in terms of providing sources for their generations, and there are […] good reasons to change this.”
It further suggests that establishing data provenance could create incentives, recognition, and potential compensation for individuals contributing valuable data to future AI models.
Copyright and Intellectual Property Concerns
The development of AI-driven generators for text, code, images, video, and music has triggered numerous intellectual property lawsuits against AI companies.
A common practice among these companies is training their models on extensive datasets sourced from public websites, some of which contain copyrighted material.
Many companies defend this practice under the fair use doctrine, a position contested by creators including artists, programmers, and authors.
Microsoft is currently facing at least two legal challenges related to copyright infringement.
Legal Challenges and Data Dignity
In December, The New York Times filed a lawsuit against Microsoft and OpenAI, alleging copyright violations stemming from the use of millions of its articles in training their AI models.
Additionally, several software developers have initiated legal action against Microsoft, asserting that its GitHub Copilot AI coding assistant was illegally trained using their copyrighted code.
This new research effort, described as “training-time provenance,” reportedly involves Jaron Lanier, a prominent technologist and scientist at Microsoft Research.
Lanier previously articulated the concept of “data dignity” in a 2023 op-ed for The New Yorker, advocating for connecting “digital stuff” with “the humans who want to be known for having made it.”
He proposed that a data-dignity system would identify and acknowledge key contributors when an AI model generates valuable output, potentially leading to compensation for their contributions.
Existing Initiatives and Challenges
Several companies are already exploring similar approaches.
Bria, an AI model developer that recently secured $40 million in funding, claims to “programmatically” compensate data owners based on their “overall influence.”
Adobe and Shutterstock also provide payouts to dataset contributors, though the specific amounts are often undisclosed.
However, few large AI labs have implemented individual contributor payout programs, instead focusing on offering copyright holders the option to “opt out” of training datasets.
These opt-out processes can be complex and typically only apply to future models, not those already trained.
Potential Outcomes and Skepticism
It is possible that Microsoft’s project will remain a proof of concept.
OpenAI announced a similar technology in May, intended to allow creators to control the inclusion of their work in training data, but this tool has yet to be released and has not been prioritized internally.
Some speculate that Microsoft’s initiative may be an attempt to mitigate ethical concerns or preempt potential regulatory or legal challenges.
Contrasting Approaches to Copyright
Microsoft’s investigation into data provenance is noteworthy given the stances taken by other AI labs regarding fair use.
Google and OpenAI have advocated for weakening copyright protections related to AI development, with OpenAI explicitly urging the U.S. government to codify fair use for model training.
This would, they argue, alleviate restrictions on developers.
Microsoft did not respond to a request for comment regarding this matter.