It’s clear that companies are currently unable to make chatbots like ChatGPT comply with EU law, when processing data about individuals. If a system cannot produce accurate and transparent results, it cannot be used to generate data about individuals. The technology has to follow the legal requirements, not the other way around.
They’re trained on far more than reddit. But it’s not a training data problem, it’s a wrong tool problem. It’s called “generative AI” for a reason: it generates text, same way a Markov chain does. You want it to tell you something, it’ll tell you. You want factual data, don’t ask a storyteller.