Every AI tool says "we don't train on your data." I'm starting to wonder if that's even possible.
I keep seeing the same promise. "We don't train on your data." "Your conversations are private." Brave says it. Everyone says it.
But then I read about Moonshot AI routing 23 million prompts through Claude and using the responses as training data. And OpenAI models doing things like finding a leaked API key and inventing data when they couldn't get what they wanted.
So here's what I can't figure out.
If you use a tool, and that tool uses an API, and the API provider uses the data—whose promise is it? Does "we don't train on your data" mean nobody trains on it? Or just that that one company doesn't?
I don't know enough about how this works under the hood. I'm not accusing anyone. I'm genuinely asking.
Has anyone here actually traced where their data goes when they use an AI tool? Or do we all just trust the sentence on the homepage and move on?
Replies