It takes a lot of data to train AI models — data that may come from third parties. Even if this information can be readily accessed on public websites, it may be unlawful to use it for AI training. Data licenses may not permit AI-related use cases if they aren’t spelled out in the agreement. There are even restrictions on the use of data that users enter into a pretrained generative AI tool and the output generated by AI platforms.

Data ownership becomes murky when used for AI, fueling major legal battles over creator rights, copyright infringement, and fair use. Organizations that develop or use AI models should not underestimate the complexity of data licensing and control.

Why AI Companies Are Moving to a License-First Approach

AI data licensing has emerged as a critical, rapidly growing market. The market for AI training data alone reached $2.68 billion in 2024 and is expected to grow to $11.16 billion by 2030. Increasingly, formal licensing agreements are replacing scraping practices to avoid copyright infringement lawsuits. While some firms argue that training is transformative and thus exempt from licensing, high-profile lawsuits are pushing the industry toward a license-first approach.

Licensing reduces risk by providing explicit permission to use data for training. As a bonus, licensed data is typically cleaner and better curated than raw data scraped from the internet, leading to more trustworthy AI models. For content owners, licensing provides a way to monetize their archives and provides a legal framework for ethical usage and control.

However, AI data licensing can be fraught with pitfalls, particularly as the technology continues to evolve. Often, disputes arise over the scope of authorized use.

The Critical Distinction Between Ownership and Control

Generally, data is simply information that cannot be owned in the property sense. The laws of most countries classify some types of data as intellectual property, which includes creative works, copyrighted software and databases, trademarks, and patentable inventions. Data that falls outside those categories may not be legally owned, but it may not be free to use either.

For example, individuals have the right to control their personal data, with a few limited exceptions. Companies have the right to control confidential and proprietary business information. In addition, many open datasets contain material that is restrictively licensed or ShareAlike, meaning any new works or derivatives created from it must be released under the same license.

At the same time, AI vendors have an interest in controlling the data they use to train and improve their AI models. AI contracts must spell out the boundaries and obligations around each of these interests.

In AI licensing agreements, ownership refers to the legal right to sell, license, or be compensated for the data. Control refers to the operational and technical capability to manage, use, and restrict access to that data. Ownership does not guarantee control, and control does not guarantee ownership.

Understanding the Three Main Types of AI Data

Data ownership and control can vary widely depending on the type of data and the way it’s used. As a result, AI contracts typically categorize data into three primary categories.

Training Data

Training data is further divided into two buckets: first-party data and third-party data. First-party data includes the proprietary information an organization collects directly from its own customers, users, or systems through its owned channels. Third-party data is purchased or licensed from content owners or aggregators. It also includes outputs from other AI models.

AI vendors own first-party data and are generally free to use it for model training — if customers have consented to such use. Customers may forbid vendors from using their sensitive or proprietary information for AI training. However, vendors often negotiate for the right to use aggregated or de-identified data, arguing that it no longer constitutes protected customer data.

Third-party data is much trickier. To mitigate risks, vendors should use data that is in the public domain, licensed from creators, or scraped under clear terms of service. They should also review licenses carefully and remove unlicensed third-party material before training begins.

Inputs

Input data refers to any information that users feed into an AI system to enable it to analyze or generate outputs. It includes raw data, prompts, and data files. It may also include specialized datasets that customers used to train a pre-existing model and live data feeds from sensors or devices acting as input to make automated decisions.

Generally, customers retain control over their inputs with private, enterprise-class AI platforms. The vendor is granted a license to use the data only for the purpose of providing the service to that customer. Standard enterprise agreements may also forbid the vendor from using customer inputs to train their AI models. However, the use of customer inputs to improve AI models can be a source of disagreement.

Customers generally do not retain full control over input data in free or public AI platforms, which often use inputs to train models and may share data with third parties. While users own their input, they usually forfeit control over how it’s used. To maintain control, users must manually opt out or use enterprise versions.

Outputs

Outputs refer to the responses, predictions, or creative works generated by an AI system based on a user’s input. Public AI platforms generally grant users broad rights to use, sell, and commercialize outputs. However, these rights are limited by legal realities — specifically, AI-generated content cannot be copyrighted unless it’s heavily edited by the user.

Enterprise AI platforms generally assign customers all right, title, and interest in outputs, treating the output as customer data. However, these rights are often subject to contractual limitations.

Some vendors grant the user a license to the output rather than full ownership, particularly if outputs incorporate elements of the vendor’s proprietary model. This can create risks for customers if outputs reveal sensitive or proprietary information. Users should check enterprise agreements to see if the vendor retains the right to use output data.

Reducing AI Data Risk

Although the concept appears straightforward, applying it can be complex. AI contracts should spell out all these details to avoid future litigation.

For example, if a contract doesn’t explicitly prohibit it, content owners should assume the vendor is using their data to train the vendor's model for other clients. Anti-training clauses are critical for customers to prevent their trade secrets or privileged information from becoming part of the vendor’s model. Agreements should also address what happens to the model when a license expires, given that removing data from a trained model can be virtually impossible.

As companies turn to synthetic data to avoid copyright issues, licensing agreements should include terms specifically covering the generation and reuse of synthetic content. They should also specify who owns the derivatives, fine-tuned models, or embeddings generated from the licensed data. Agreements should define whether data can be used for technologies “now known or later developed” to avoid legal challenges as AI evolves.

Organizations should not assume that existing license agreements are sufficient to cover AI. They should engage counsel with expertise in both AI technology and data licensing to ensure that contracts meet their objectives.

AI Is Creating a Virtual Quagmire of Legal Issues

Many organizations are adopting AI without a clear understanding of the risks surrounding data ownership, licensing, and control. Purdue Global Law School’s online Executive Juris Doctor (EJD) program includes a Law and Technology focus to prepare business leaders and professionals to address this risk.

For students interested in becoming a licensed attorney, Purdue Global Law School’s online Juris Doctor (JD) program provides a solid foundation in technology-related issues. Graduates of the JD program are academically eligible upon graduation to sit for the California, Connecticut, or Washington bar or, with an approved petition, the Indiana bar.

Online individual law courses are also available for those not seeking a degree.

Request more information to find the path that works for you.

About the Author

Purdue Global Law School

Established in 1998, Purdue Global Law School (formerly Concord Law School) is Purdue University's fully online law school for working adults.