Can text and data mining exceptions and synthetic data training mitigate copyright-related concerns in generative AI?

8Citations
Citations of this article
22Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Rapidly emerging generative artificial intelligence (GenAI) models stand at the epicentre of current public discourse. They demonstrate impressive abilities to generate various types of data promptly and cost-effectively. However, AI developers need to train their systems on massive volumes of data which is usually copyrighted. Therefore, the growth of copyright-related concerns in the field of GenAI comes as no surprise. The study introduces two solutions which could mitigate the tension between copyright holders and AI developers, one legal (text and data mining (TDM) exceptions of the CDSM Directive) and one technical (synthetic data), highlighting the promises and challenges of both. First, the article will discuss the capability of TDM exceptions to facilitate the fundamental right to information and the freedom of research in the context of AI development. Next, the paper will analyse how providers of GenAI models can leverage synthetic data to comply with copyright law while training their systems and what risks might be associated with this approach. The findings of this study will indicate what issues, in both legal and technical spheres, should be addressed to ensure a balance of powers in the digital environment and effective functionality of the EU AI sector.

Cite

CITATION STYLE

APA

Manteghi, M. (2024). Can text and data mining exceptions and synthetic data training mitigate copyright-related concerns in generative AI? Law, Innovation and Technology, 16(2), 663–686. https://doi.org/10.1080/17579961.2024.2392928

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free