AI Was Trained on Your Data Without Asking. Here Is What You Can Take Back
AI Was Trained on Your Data Without Asking. Here Is What You Can Take Back
You did not volunteer for this. A photo you posted in 2014, a forum reply, a comment under a news story, a resume you uploaded one time and forgot about. None of it was meant to teach a machine anything. That is exactly where a lot of it ended up.
The AI tools millions of people now use every day did not learn from nothing. They learned from the open internet, and the open internet is full of you.
How your data ended up inside a model
The method is called web scraping. Automated crawlers move across the internet and copy everything they can reach, then that material gets bundled into the giant training sets behind large AI models. There is no consent screen in this process. No notification. Nobody asks whether the names, faces, and personal details swept up in the haul belong to people who agreed to take part.
Plenty of that data was technically public, but public was never the same as fair game. A picture shared with friends, a profile meant for a small community, an old record sitting on a people-search site. All of it is reachable by a crawler, and reachable is all that matters to one.
In a May 2026 briefing titled Unlawful by Design, Amnesty International described the bulk collection of training data through web scraping as a mass invasion of privacy by design, and called on governments to prohibit standalone generative AI systems built that way. On the regulatory side, the European Union now requires makers of general-purpose AI to publish a summary of the datasets used to train their models, a sign that the era of collecting quietly and explaining nothing is ending.
What you can do, and what it will not fix
There are real steps, and there are limits to be honest about.
What helps:
- Use the opt-out where a platform offers one. Some now let you exclude your public posts from future AI training, though the controls are often buried and easy to miss.
- Send deletion or objection requests to companies that hold your data. In several regions they are legally required to respond.
- Remove yourself from data broker and people-search sites. Those are the easiest, richest sources a crawler can grab, so clearing them shrinks what gets scraped on the next pass.
What it will not fix: you cannot pull your information back out of a model that has already trained on it. There is no delete button for a finished model. That is why the realistic goal is forward-looking. You are not undoing the past. You are reducing what is available to be taken next time.
See what is still out there with your name on it
Our free scan checks the data broker sites that feed scrapers and aggregators, so you can see your exposure before deciding what to do about it.
Run my free scanIf you decide to chase this down yourself, it helps to know what the numbers look like. When Consumer Reports tested the cleanup process, automated and do-it-yourself removal cleared only about 27 percent of listings, while removal handled by actual people reached about 70 percent. Machines are excellent at finding your data. They are far worse at getting it taken down and keeping it down.
Frequently asked questions
Can I make an AI company delete my data from its model?
Not in any clean way. You can request deletion of the data they store about you, and in some regions they must respond, but a model that already trained on your information cannot be neatly reversed. The achievable goal is limiting future collection, not erasing the past.
Does setting my accounts to private stop the scraping?
It helps going forward, but it does nothing about data already collected, and nothing about the broker sites that copied your details years ago. Those copies sit on separate servers you do not control.
Where is most of my data being pulled from?
A large share comes from people-search and data broker sites that compile public records, purchase history, and online activity into one tidy profile. A neatly assembled profile is the single easiest thing for a crawler to grab.
Shrink your footprint before the next training run
Start with a free scan to see where you show up, then let a real person handle the removals that machines keep missing.
Run my free scan Start free removal