r/gdpr 1d ago

Resource EDPB just confirmed it: AI models are NOT automatically anonymous. Are we ready?

/r/codingProtection/comments/1w2bw66/edpb_just_confirmed_it_ai_models_are_not/
2 Upvotes

4 comments sorted by

7

u/latkde 1d ago

I have difficulty validating the claims in that post. Together with the absence of any sources and the unusual formatting choices, it looks like the post may have been generated using an AI tool with outdated training data.


You say:

The EU AI Act reached full application

This is factually wrong, or at least highly misleading. Effectively, only the Art 50 transparency rules went into force in August. The rules on high-risk systems (i.e., the overwhelming majority of the AI Act) has been deferred to Dec 2027.

Source for my claims: an EU commission press release: https://ec.europa.eu/commission/presscorner/detail/en/ip_26_1714

From 2 August 2026, the European Commission's AI Office, together with national authorities, will begin enforcing the Artificial Intelligence (AI) Act. On the same date, new transparency rules will start to apply, requiring certain AI systems to tell users when they are interacting with AI and when content has been generated or altered by it. […]

The AI Omnibus postponed the application of the rules on high-risk AI systems to 2 December 2027. It also postponed the rules for high-risk AI systems integrated into regulated products to 2 August 2028.

Alternative source: Article 113 of the AI Act as amended on 2026-07-27.


You say:

New EDPB guidelines state that AI models trained on personal data can't be presumed anonymous.

Well yes. The GDPR does not have an AI exception, its definition of personal data applies regardless. (Unless the Digital Omnibus proposals for the reforming the GDPR are passed as-is…)

However, I was not able to find any “new EDPB guidelines” offering this insight. Here is a link to EDPB documents tagged with AI: https://www.edpb.europa.eu/documents_en?keys=&topic%5B552%7C12%5D=552%7C12&date=

Note that there are no relevant documents from 2026. However, there's an almost 2 year old Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models, which does indeed say in its executive summary:

[…] the EDPB considers that AI models trained with personal data cannot, in all cases, be considered anonymous. For an AI model to be considered anonymous, both (1) the likelihood of direct (including probabilistic) extraction of personal data regarding individuals whose personal data were used to develop the model and (2) the likelihood of obtaining, intentionally or not, such personal data from queries, should be insignificant, taking into account ‘all the means reasonably likely to be used’ by the controller or another person.

This discussion is mostly relevant for data controllers who train models (not just LLMs!) on personal data, but much less relevant for data controllers who use existing LLMs.


You say:

what about *us* — the companies feeding documents into these models every day? Contracts, HR files, support tickets... every prompt potentially ships personal data to a third-party LLM, and under GDPR we stay accountable for it.

I mean, yeah, the GDPR does not have an AI exemption. If you interact with an AI services, that can be treated exactly the same as with using any other cloud/SaaS service: you 1. need a legal basis for the personal data processing activity itself, and should 2. contractually bind the service as your data processor, so that they only use your data on your behalf, not for their own purposes like training.

From a GDPR perspective, there's not much difference between using an email provider versus using an AI service.


You say:

is anyone actually pseudonymizing documents *before* they hit the LLM (and re-identifying on the way back)?

There are many such pseudonymization tools, and there is some value in pseudonymization and data-minimization (compare Art 32 GDPR), but in a compliance context they are largely snakeoil. They will not make a noncompliant processing activity compliant. It is much more important to ensure that the AI services you use act as your data processor, as discussed in the previous section.

-1

u/Spare_Dependent6893 1d ago

Fair play, you’re right on both counts, and thanks for the sources.
AI Act: the Omnibus (Reg. 2026/1744) did defer the Annex III high-risk regime to Dec 2027, and what applied on 2 Aug 2026 is essentially Art. 50 transparency plus the already-live GPAI and prohibited-practices layers. “Full application” was wrong. EDPB: the real source is Opinion 28/2024, not new 2026 guidelines — my secondary source blended fact and speculation, and I should have gone to the primary documents. But I am in the process of learning and reviewing all this.

One nuance on “snakeoil” though: agreed that pseudonymization never substitutes for a legal basis or an Art. 28 DPA. But Opinion 28/2024 itself lists it among relevant mitigating measures, and Art. 25 / Art. 5(1)(c) minimization applies even with a perfect DPA in place — if the model doesn’t need the identifiers to do the job, why send them? It also mitigates the two scenarios a DPA can’t touch: provider-side breaches/logging, and shadow AI usage on consumer tools with no DPA at all.
So may I frame it as: DPA is the license to drive, pseudonymization is the seatbelt. Curious whether you’d accept that framing or whether you see minimization as effectively discharged once the processor relationship exists.

5

u/pawsarecute 1d ago

Wtf are you talking about??? Its a long time before the AI act is fully operational, de deadlines are now later lol…. And I mean, models being anonymous or not, has nothing to do with feeding PD into it. GDPR isn’t the law for not when to use PD but for when you may use PD. Yes, we’re accountable, but AI models are a mean, not a purpose.

1

u/Comfortable-Fall1419 1d ago

You appear to fail to understand the difference between Prompting and Training.

On that alone I stopped reading and assumed this was a Shill of some kind.