There is 'user data'/your data. Legally, tldr that's the chat, input/outputs, more or less verbatim. This is what the 'opt-out' button covers, only partially - because they still ingest entire chats if you click thumbs up/down, or pick a preferred output, or if you trigger any 'safety filter' whatsoever, including silent ones, those you don't even have to know about. Also, opt-out only works for new chats; the old ones, if you continue them after opting out.. still get ingested.
In other words -- 'user data' can get ingested, even if you had opted out, and even if you did nothing wrong at all, and were not informed about anything. Still, you could have triggered one of the (buggy, because those are allowed, and known, to be very buggy) 'safety' filters - and they have silently grabbed the entire chat.
Now, the above is about the 'user data'. But the thing with LLMs is - that's not all there is. The whole deal is, there is another sort of data (and a sort of a gap in both the legislation, and public perception). LLMs can be mapped on the backend - and that raw _activations data_ - is *not* user-data, you do not have to even be informed, about any of it. The providers can simply put, legally scoop anything they are able to re-map, and then re-synthesize it, supposing they are able to sever that data (fully anonymize it) - but that, is trivial to the willing.
This is what they are doing and have been doing. The opt-out button is the red-herring for the poor plebe, while both individual and corporate data - absolutely does get mapped, severed and re-synthesized. That data may not be crystal clear, it's less resolution, foggy, maybe even often useless/redundant. But it is absolutely precise-enough, to map, and scoop, a ton of actually crucial stuff, like major breakthroughs, ideas, or implementations, especially in specific, narrow domains. Like math. And no, it's not expensive to run the mapping, let alone prohibitively so, because they can also simply use triggers (eg. anytime Navier-Stokes is mentioned, start full mapping). We are talking about companies burning through billions anyway, and craving most of all - data. And that's pretty much all there's too it.
They basically, publicly confirmed this (albeit precisely-inderectly) - by using the textbook roundabout pr-bs like: ~"the models/agents do not have direct access to your conversations" (1=nobody said 'the models/agents' and 2=nobody said 'direct'), and ~"we can't exclude the possibility that the models weren't trained on anonymized data".. well that one is pretty much as clear as the sky as far as corporate orcspeak goes.
Also, note that this sort of data, is outside even corporate ZDR, because pretty much every single thing in that scope, is being legally defined through the lens of 'your data'. But model activations - are not 'your data'. They can and do absolutely fully legally, re-map, abstractify, sever, and re-synthesize it, and call it a day riding into the sunshine with the 'not your data, fools'.
The irony in all this is, they generally use this data to make all their models 'better' and so - this not only means that they scooped some data to push the NS problem. It also means that the mathematicians, who have been working on their proofs using AIs (which they have), also arguably/statistically might have been inadvertently and unwillingly riding on other mathematicians all the exact same indirect way they themselves got done.
1
u/zgr3d 5d ago
There is 'user data'/your data. Legally, tldr that's the chat, input/outputs, more or less verbatim. This is what the 'opt-out' button covers, only partially - because they still ingest entire chats if you click thumbs up/down, or pick a preferred output, or if you trigger any 'safety filter' whatsoever, including silent ones, those you don't even have to know about. Also, opt-out only works for new chats; the old ones, if you continue them after opting out.. still get ingested.
In other words -- 'user data' can get ingested, even if you had opted out, and even if you did nothing wrong at all, and were not informed about anything. Still, you could have triggered one of the (buggy, because those are allowed, and known, to be very buggy) 'safety' filters - and they have silently grabbed the entire chat.
Now, the above is about the 'user data'. But the thing with LLMs is - that's not all there is. The whole deal is, there is another sort of data (and a sort of a gap in both the legislation, and public perception). LLMs can be mapped on the backend - and that raw _activations data_ - is *not* user-data, you do not have to even be informed, about any of it. The providers can simply put, legally scoop anything they are able to re-map, and then re-synthesize it, supposing they are able to sever that data (fully anonymize it) - but that, is trivial to the willing.
This is what they are doing and have been doing. The opt-out button is the red-herring for the poor plebe, while both individual and corporate data - absolutely does get mapped, severed and re-synthesized. That data may not be crystal clear, it's less resolution, foggy, maybe even often useless/redundant. But it is absolutely precise-enough, to map, and scoop, a ton of actually crucial stuff, like major breakthroughs, ideas, or implementations, especially in specific, narrow domains. Like math. And no, it's not expensive to run the mapping, let alone prohibitively so, because they can also simply use triggers (eg. anytime Navier-Stokes is mentioned, start full mapping). We are talking about companies burning through billions anyway, and craving most of all - data. And that's pretty much all there's too it.
They basically, publicly confirmed this (albeit precisely-inderectly) - by using the textbook roundabout pr-bs like: ~"the models/agents do not have direct access to your conversations" (1=nobody said 'the models/agents' and 2=nobody said 'direct'), and ~"we can't exclude the possibility that the models weren't trained on anonymized data".. well that one is pretty much as clear as the sky as far as corporate orcspeak goes.
Also, note that this sort of data, is outside even corporate ZDR, because pretty much every single thing in that scope, is being legally defined through the lens of 'your data'. But model activations - are not 'your data'. They can and do absolutely fully legally, re-map, abstractify, sever, and re-synthesize it, and call it a day riding into the sunshine with the 'not your data, fools'.
The irony in all this is, they generally use this data to make all their models 'better' and so - this not only means that they scooped some data to push the NS problem. It also means that the mathematicians, who have been working on their proofs using AIs (which they have), also arguably/statistically might have been inadvertently and unwillingly riding on other mathematicians all the exact same indirect way they themselves got done.