Guides

    How to read a GPAI training-data summary

    General-purpose AI providers have published training-content summaries since August 2025. Six questions that separate a useful one from a decorative one.

    Versely Team8 min read

    Since 2 August 2025, providers of general-purpose AI models have had two obligations that produce public documents: a policy for complying with EU copyright law, and a sufficiently detailed summary of the content used to train the model. As of 2 August 2026 the Commission's enforcement powers over those obligations are live. Which means the documents have existed for a year, are addressed to you, and almost nobody in a creative pipeline has read one.

    They are worth reading, with realistic expectations. A training-content summary is not a dataset manifest, not a licence chain, and not a legal clearance. It is a first-party disclosure, in a broadly comparable format, about where a model's training material came from. That is a narrow artefact — and it is still the only one of its kind you get.

    What the summary is and is not

    It is It is not
    A public, first-party statement about training content A file-level list of what was ingested
    Broadly comparable across providers Standardised down to the field
    Evidence of what a provider is willing to say on the record Evidence that the training was lawful
    A useful input to a procurement decision A warranty, an indemnity or a clearance
    Paired with a stated copyright policy Proof that the policy was followed

    Hold that right-hand column firmly, because the temptation runs the other way. A detailed summary reads like reassurance. It is disclosure, not adjudication. The legality of training on copyrighted material is genuinely unsettled: US district courts have gone in both directions, the first appellate word on AI-training fair use is still pending before the Third Circuit after argument in June 2026, and the largest sum paid out so far in that space settled a claim about how a library was acquired rather than establishing anything about training itself. No summary resolves any of that.

    What a summary does resolve is what a provider will commit to publicly — and how specific they were willing to be.

    Six questions to ask of any summary

    Run these in order. The first two sort most documents into useful or decorative within a couple of minutes.

    1. Are named sources named?

    The tell is specificity. A summary that identifies major datasets, named corpora, licensed catalogues and specific crawl sources is doing the job. A summary whose main category is "publicly available data from the internet" has described the internet. Both are technically summaries; only one changes a decision.

    2. Are licensed and proprietary sources separated from scraped ones?

    The most commercially relevant split in any training corpus is between material the provider had a deal for and material it collected. A summary that keeps those categories distinct — and says which is which — is far more useful than one that lists everything in a single undifferentiated block. This is also the field that maps most directly to how buyers are starting to choose models, which is the argument in licensed training data as a buying criterion.

    3. Is web collection described concretely?

    Look for the crawler's identity, the period of collection, and how reservations of rights were handled. "We respect robots.txt" is a different statement from naming the user agent, the opt-out mechanism honoured, and the date range. The second is checkable by a rightsholder; the first is not.

    4. Is synthetic and model-generated data broken out?

    Increasingly material, and increasingly common. If a substantial share of training content was itself generated, that belongs in the summary, and its absence in a modern model's document is a question worth asking.

    5. Does the copyright policy do anything?

    The paired copyright policy is the other half of the disclosure. A policy that restates the applicable law is decorative. A policy that names an actual reservation-of-rights procedure, gives a contact route for rightsholders, and describes what happens when a reservation is found is operational. The difference is visible in one read.

    6. Which model does it cover, and when was it published?

    Summaries attach to models, and providers ship models continuously. A document dated to a previous generation tells you about a previous generation. Check that the summary you are reading covers the model you would actually call — the same version discipline you would apply when comparing options on the model catalogue or a head-to-head on compare.

    Useful versus decorative, in practice

    Two paraphrased shapes, to calibrate.

    Decorative. "The model was trained on a large and diverse corpus of publicly available data, licensed data, and data created by our teams. We respect applicable law and the rights of creators." Four sentences, no named source, no date range, no crawler, no opt-out mechanism, no split between licensed and scraped. Nothing in it could be checked by anyone.

    Useful. Named public datasets with versions. Named licensed partners or categories of licensor. A crawl described by user agent and collection window, with the reservation-of-rights mechanism identified. Synthetic data broken out as its own category. A copyright policy with a rightsholder contact route and a stated process. None of that proves the training was lawful. All of it is falsifiable, which is the point.

    What a good summary still cannot tell you

    Four questions people try to answer from these documents and cannot.

    Whether your output is protectable. That turns on human authorship, not on training data. In the US, purely AI-generated work is unregistrable, and the Supreme Court declined to take up the question in March 2026, leaving the appellate position in place. Prompts alone — however detailed or iterated — do not make you the author, though human selection, arrangement and modification can be protected. Owning the file under a provider's terms is a contract fact; copyright in the file is a statutory one, and they come apart constantly.

    Whether you are indemnified. Indemnities live in your agreement, not in a summary, and where they exist they are generally enterprise or API tier and conditioned on using the provider's safety filters. Read the contract you are actually on.

    Whether a specific output infringes. A clean corpus does not stop a model from producing a recognisable protected character on request. That exposure is user-side, live in litigation, and unresolved.

    Whether you can use the model's weights the way you plan to. That is a licence question, separate again — open-weight video licences and what you can ship covers the version of it that bites hardest for video work.

    For the broader contract-and-clearance picture around commissioned work, legal and licensing for AI content in business is the wider frame this fits into.

    Where this actually changes a decision

    Three places, concretely.

    1. Client briefs with a provenance clause. More agency contracts now ask you to state which models produced the deliverable and what is publicly disclosed about their training. A summary is the citable answer, and reading it beforehand means you are not writing that clause from a marketing page.
    2. Model shortlisting for regulated categories. Health, finance, kids' content and anything near a rightsholder-heavy vertical increasingly gets model choice reviewed. A provider that publishes a specific, checkable summary is easier to defend than one that publishes a paragraph.
    3. Sequencing your own tests. Once the shortlist is set on paper, the remaining question is output quality, and the fastest way to settle it is to run one prompt across several named models in a single request through the agent and compare the results side by side. Paper first, then pixels — it stops you burning credits on models you would have rejected on the document.

    The summaries will not tell you a model is safe. They will tell you which providers were willing to be specific, and after a year of these documents existing, that variance is informative on its own.

    FAQ

    Where do I find a model's training-content summary?

    Providers publish them alongside their model documentation, typically on the same pages as model cards, terms and policy documents. If a general-purpose model in your stack has no locatable summary, that absence is itself worth raising with the vendor.

    Does a detailed summary mean the model was trained lawfully?

    No. It is disclosure, not adjudication. The legality question is contested and, at appellate level in the US, still pending. A summary tells you what a provider says it used, which is a different thing from whether it was entitled to use it.

    Does this obligation apply to image and video models?

    The general-purpose AI model obligations attach based on how a model is characterised under the Act rather than on output modality. Check whether the specific provider publishes a summary rather than assuming from media type, and treat the presence or absence as a data point in procurement.

    If a model has a great summary, can I skip the disclosure on my output?

    No — unrelated duties. Training-content disclosure sits on the model provider. Labelling synthetic content for your audience sits on whoever publishes it, and a well-documented model changes nothing about that.