Skip to content
All articles
Data, privacy and GDPR

Can an AI model be anonymous? The EDPB answered, and the answer is not yes

Opinion 28/2024 states that AI models trained with personal data “cannot, in all cases, be considered anonymous”, and sets a two-part test that has to be met before anyone claims otherwise.

By Bastion, Cluj-NapocaPublished 7 min read

"The model is anonymous" is an attractive sentence for a vendor and a dangerous one for a buyer. The European Data Protection Board addressed it directly in Opinion 28/2024, published on 18 December 2024, and the answer is narrower than the marketing.

The finding

The Opinion states that claims of an AI model's anonymity “should be assessed by competent SAs on a case-by-case basis, since the EDPB considers that AI models trained with personal data cannot, in all cases, be considered anonymous”.

It repeats the conclusion in the body of the document: “the EDPB considers that AI models trained on personal data cannot, in all cases, be considered anonymous. Instead, the determination of whether an AI model is anonymous should be assessed, based on specific criteria, on a case-by-case basis.”

Note what that does and does not say. It is not a ruling that models are never anonymous. It is a refusal to accept anonymity as a property of the category rather than a finding about a specific model.

The two-part test

for an AI model to be considered anonymous, using reasonable means, both (i) the likelihood of direct (including probabilistic) extraction of personal data regarding individuals whose personal data were used to train the model; as well as (ii) the likelihood of obtaining, intentionally or not, such personal data from queries, should be insignificant for any data subject

EDPB Opinion 28/2024

Both conditions, not either. Extraction from the model, and accidental disclosure through ordinary queries. And the standard is insignificant “for any data subject”, not on average.

The assessment must consider “all the means reasonably likely to be used” by the controller or another person, and the Opinion adds that supervisory authorities should consider the risk of identification by “different types of 'other persons', including unintended third parties accessing the AI model”.

Why the intuitive test fails

The common argument is that the model holds no names, therefore no personal data. Anonymisation has never worked that way. A description can identify a person with no name in it at all: the chief executive of the only software company in a town of eight thousand people, who joined in March 2018.

The Opinion also notes that the underlying research is live: “research on training data extraction is particularly dynamic. It shows that it is possible, in some cases, to use means reasonably likely to extract personal data from some AI models, or simply to accidentally obtain personal data through interactions with an AI model”.

Anonymity is a conclusion supported by documentation about a specific model. It is not an adjective a supplier can apply to a product line.

What a company should do before the question arises

  • Identify what personal data is actually present in prompts, documents and any fine-tuning set
  • Decide whether it is necessary for the task, and remove what is not
  • Record the legal basis for processing it
  • Record who processes it, where, and for how long it is retained
  • Keep the technical controls, and the reasoning behind them, in writing

The Opinion expects exactly this kind of record. It says supervisory authorities “should review the documentation provided by the controller to demonstrate the anonymity of the model”. The burden sits with whoever makes the claim.

Where local processing helps, and where it does not

Running the model inside the organisation gives more control over who can query it and who can reach it, which is directly relevant to the "other persons" limb of the test. Fewer parties with access is a smaller identification surface.

It does not convert personal data into anonymous data. A model that could regurgitate a training example in a data centre can do the same in your server room. Privacy work starts with understanding the data, and no deployment choice substitutes for it.

With Bastion

What Bastion changes

Bastion is a private AI system delivered as one sealed appliance that runs inside your building. One monthly fee covers the hardware, the model, the hardened operating system and support, and nothing your team types leaves the building.

Bastion ships an open-weight model that the customer does not train on their own data by default, and it runs where the customer controls who can query it.

Neither fact makes a model anonymous in the sense Opinion 28/2024 uses, and we will not say otherwise. What they do is narrow the set of people who could ever attempt extraction.

On the record

  1. 1