EN FR ES PT DE AR 中文

Your Archive on Someone Else's Platform Is Now AI Training Data

Platform licence grants written a decade ago are broad enough to cover AI model training without a word being amended. If your firm's footage lives on a platform it doesn't own, UK law says the exposure runs deeper than the terms.

Listen8 min

Somewhere in the terms you agreed to years ago, there's a licence broad enough to turn your entire video archive into AI training data, and nobody has to ring you before using it. That's the quiet repricing of the model era. Footage recorded to sell a product, train a team or fill a webinar slot has become exactly what generative models are hungriest for, and the paper platforms took from you a decade ago was drafted wide precisely so nobody would ever need to come back for permission. Amazon has owned Twitch since 2014, when it bought a live-video community business; a decade on, the same asset reads as a training corpus. Your firm's webinar library is a smaller entry in the same ledger.

This isn't a streamer problem. Firms have made equivalent deposits everywhere: webinars, executive talks, product walkthroughs, recorded support sessions, all sitting on platforms whose terms can be rewritten without anyone signing anything twice. The value of those archives was latent until the model era made it legible, and latent value on someone else's balance sheet has a way of getting realised without a courtesy call.

The consent mechanics tell you where this is heading, and one paragraph covers it. When platforms ship AI-training controls at all, expect opt-out toggles rather than opt-in requests: an operator that believed users would say yes would simply ask, so near-total enrolment is the design goal. Hours of talking-to-camera footage is the richest substrate there is for synthesising a convincing human, which makes an archive of it a likeness asset regardless of why it was recorded. And when every frontier model has eaten the same public web, a walled corpus of exclusive footage is a moat, so expect terms to keep drifting toward capture. That's the context. The sharper question is what the paper you've already signed, and the law you're already subject to, actually say.

Can a platform use your videos as AI training data?

Read the licence you granted on upload. Twitch's terms of service take an "unrestricted, worldwide, irrevocable, fully sub-licensable, nonexclusive, and royalty-free" right to "use, reproduce, modify, adapt, publish, translate, create derivative works from, distribute, perform, and display" what you post. YouTube's terms are politer but no narrower where it counts: a "worldwide, non-exclusive, royalty-free, sublicensable and transferable" licence to use your content "in connection with the Service and YouTube's (and its successors' and Affiliates') business". Neither clause mentions model training, and both predate the question mattering. That is precisely the problem: language this broad doesn't need amending to cover a training pipeline, it needs only a lawyer willing to argue that training improves the service. Against that backdrop, an opt-out toggle is a product feature, not a contract term. Features can move. The licence stays.

Whether courts eventually narrow those grants is open, and I wouldn't bet a business on either outcome. The triage position is simpler: if the platform's owner builds models, assume deposited content is candidate training data unless the terms say otherwise, then check what the objection mechanism actually covers and when it closes. No platform publishes an inventory of what its owner's models have already ingested, and none promises to. Assume you can't find out.

Is your executive's face biometric data under UK GDPR?

Here's the argument the terms-of-service debate misses, and for a UK firm it's the sharper one. Footage of an identifiable person is personal data, full stop. Under Article 4(14) of UK GDPR, facial images become biometric data when they undergo "specific technical processing" that allows or confirms "the unique identification" of a person. Hosting a webinar isn't that. Training a system until it can recognise or reproduce one specific human, face and voice included, is a strong candidate for exactly that, and biometric data used to uniquely identify someone is special category data under Article 9, where processing is prohibited unless a narrow exception applies. The workable exception here is explicit consent, and the Data Protection Act 2018 keeps the other gates tight. No court has settled this for model training yet; I'm not claiming otherwise. I'm claiming that a UK company whose executives become synthesisable has a live Article 9 question on its hands, and "the platform's terms allowed it" appears nowhere on the list of exceptions.

Now run the employment angle. The firm that uploads the footage is a controller and needs a lawful basis of its own. Consent looks obvious and isn't: the ICO's guidance is blunt that consent must be freely given, and the imbalance of power between employer and employee makes employee consent suspect by default. A marketing team that filmed the head of product for explainers holds, at best, consent for explainers. Nobody consented to becoming substrate.

And here's the part that makes it a governance hole rather than a paperwork gap. When you outsource processing, UK GDPR expects a processor bound by an Article 28 contract, acting on your documented instructions. A platform hosting your footage under its own terms is nothing of the sort. It's an independent controller, processing for its own purposes, and uploading to a US company is a restricted transfer under Chapter V, a regime built for outsourcing arrangements rather than for a counterparty with plans of its own. Your data-protection framework simply doesn't follow the file. And if the footage has already been trained into a model, the rights that normally backstop you, erasure and objection, collide with the fact that machine unlearning at scale sits somewhere between hard and unproven. The ICO has seen this shape before. In 2017 it found that the Royal Free NHS Foundation Trust failed to comply with data protection law when it passed 1.6 million patient records to Google DeepMind for app testing. The finding landed on the trust that let the data leave its estate, not on the company that received it. Different data, lower stakes, same accountability: the depositor answers.

Pricing the exposure

Treat this as unpriced counterparty exposure, because that's what it is. The audit isn't exotic: inventory what you've deposited and where; read each licence grant and objection mechanism, and diarise terms-change notices; decide, platform by platform, whether the reach still pays for the exposure; keep your own masters; negotiate carve-outs where your account is big enough to matter. Footage of identifiable executives gets its own line, with a recorded lawful basis, because that's where the Article 9 tail risk lives. This is exactly what an AI-readiness review should surface before any new platform commitment, and it deserves a standing entry in your technical strategy rather than a one-off panic.

I'd revise all of this on real evidence: a major platform moving to genuine opt-in and publishing the take-up rate it gets, or terms that separate operating the service from training models. The incentives above are my reason for not waiting. Auditing your deposits costs an afternoon. An executive's likeness inside a model you can't inspect or recall costs an unknowable amount for an unbounded time. We've written before about keeping humans in control of AI systems; the corollary is refusing to become uncredited raw material in someone else's. A small premium against a large, unbounded loss, in a window that's closing. That's the trade.

Questions people ask

How do I find out if a platform is training AI on my content?

Read the current terms of service and any AI or machine-learning addendum, searching for licence language such as "train", "machine learning" or "improve our services", then check your account settings for a training or data-sharing toggle, which is usually where an opt-out lives. Set a calendar reminder to re-check whenever the platform emails a terms-change notice, because enrolment changes typically arrive that way with a short objection window.

Does opting out remove my videos from AI models that were already trained?

Almost certainly not. Opting out generally stops future ingestion at best; removing specific content from an already-trained model (so-called machine unlearning) is technically difficult and not something platforms commit to. No major platform publishes an inventory of what its models have already eaten, which tells you how much visibility to expect. The timing of your audit matters more than the thoroughness of a later objection.

Is video footage of an executive biometric data under UK GDPR?

Not by default. Facial images only count as biometric data under Article 4(14) once they undergo "specific technical processing" that allows unique identification, and simply hosting a webinar doesn't qualify. Training a model to recognise or synthesise one specific person is a strong candidate, though, and at that point the Article 9 special category regime applies, with explicit consent as the realistic gateway. Treat executive footage as high-risk personal data and record a lawful basis before it leaves your estate.

Related

Written by an AI editorial persona of Abyshire's proprietary editorial system and reviewed by our team.