Tainted Training Data and AI: A Case-by-Case Approach

0

1. When unlawful data enters the AI supply chain

Artificial intelligence systems are insatiable consumers of data, much of which is obtained through large-scale web scraping. Whenever these datasets contain personal data, the lawfulness of their collection and seubsequent processing becomes a central issue under Regulation (EU) 2016/679 (hereinafter ‘GDPR’).

Yet one question is becoming increasingly relevant for AI governance and remains only partially settled: what happens when personal data used to train an AI system was unlawfully collected at an earlier stage? More specifically, do the legal consequences of that initial violation extend to all subsequent uses of the model?

The issue is far from theoretical. AI systems rarely remain in the hands of a single actor. Today’s AI ecosystem is characterised by increasingly fragmented supply chains, where one or more entities collect data, others develop the model, and others deploy it in downstream products and services (for an overview of some of the actors involved, see OECD, 2025, “Competition in artificial intelligence infrastructure”, OECD Roundtables on Competition Policy Papers, No. 330, OECD Publishing, Paris).

Against this background, a blanket affirmative answer would produce significant consequences, as any AI model trained, even partially, on data of uncertain provenance could become unusable throughout its entire lifecycle. At the same time, the opposite solution, which would allow downstream third-parties and deployers to ignore the unlawful origin of training data would undermine the core principles of the GDPR, particularly lawfulness, accountability and purpose limitation.

This question has recently been addressed in Austrian case law, albeit outside the specific context of artificial intelligence. Nevertheless, the legal principles emerging from that debate may have significant repercussions for AI ecosystems, where data, models and responsibilities are distributed across multiple actors. This blog post therefore analyses the Austrian discussion and considers how its findings may inform the governance of AI models trained on unlawfully collected personal data.

2. The Austrian Precedent: From a Rigid Cascade to a Contextual Assessment

In proceedings initiated by noyb (none of your business), the Austrian Data Protection Authority (Datenschutzbehörde, hereinafter ‘DSB’) issued a decision on 24 March 2023 (case no. D124.3816 2023-0.193.268) finding that the credit reference agency CRIF GmbH had violated the GDPR by processing personal data (names, addresses, dates of birth) obtained from the address publisher AZ Direct (part of the Bertelsmann Group). AZ Direct was legally authorised to share this data only for advertising purposes; CRIF used it to calculate creditworthiness scores and sell them to third parties, without any legal basis and without the data subjects’ knowledge or consent.

Building on a line of reasoning already established by the Austrian Administrative High Court (VwGH ruling Ra 2017/04/0034; Ra 2019/04/0054), the DSB adopted a strict ‘cascading unlawfulness’ approach: the unlawful collection by one controller, here, AZ Direct’s transmission of data beyond its authorised purpose, rendered inadmissible any subsequent processing by the recipient. Under this reading, a third party acquiring tainted data could never process it lawfully, regardless of any independent legal basis it might invoke.

The Austrian Federal Administrative Court (Bundesverwaltungsgericht, hereinafter ‘BVwG’), ruling in the subsequent appeals (cases W605 2270910-1 and W605 2271598-1, concerning the respective appeals by the complainant and CRIF against the DSB decision; see also BVwG W176 2259543-1 concerning AZ Direct’s appeal on the purpose limitation issue), declined to follow the DSB’s ‘cascade’ logic. The Court instead applied Articles 5 and 6 GDPR directly to each controller’s conduct in isolation. It found both AZ Direct and CRIF independently in violation: AZ Direct for transferring data beyond its authorised purpose (Art. 5(1)(b) GDPR), and CRIF for processing without any valid legal basis (Art. 6 GDPR), without needing to derive the second violation from the first. The ‘continued effect of the error’ was therefore not the operative ground for the BVwG’s findings (for further information on the Austrian cases see Münch and Kempermann, 2026).

The significance of this distinction is considerable. By grounding its decision on independent violations of GDPR principles by each party, rather than on a mechanical transmission of unlawfulness, the BVwG implicitly opened the door to a scenario in which a downstream controller could, in principle, process data of uncertain provenance lawfully, provided it has a valid and independent legal basis for doing so, and provided the specific circumstances of the case do not lead to a different conclusion. The court’s reasoning is, in other words, neither a blanket permission nor a blanket prohibition: it is an invitation to contextual analysis.

Applying the same line of thought to AI systems, the original question is not whether unlawfulness automatically propagates through the entire AI lifecycle, but rather under which circumstances subsequent processing remains compatible with GDPR obligations.

3. The EDPB’s Framework: The Due Diligence Approach

The EDPB’s Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models (hereinafter ‘Opinion 28/2024’), adopted on 17 December 2024 following a request by the Irish supervisory authority, addresses this question head-on in its fourth section (paras. 109–135), distinguishing three scenarios depending on whether personal data is retained in the deployed model and whether the deploying controller is the same as or different from the one that conducted the initial (unlawful) processing. The scenario most relevant to the AI supply chain is Scenario 2: a controller unlawfully processes personal data to develop the model, the personal data is retained in the model, and a different controller then deploys it (EDPB, Opinion 28/2024, paras. 124–132).

The EDPB’s answer is emphatically not a blanket cascade. It recalls that each controller bears independent responsibility under Articles 5(1)(a) and 5(2) GDPR to ensure and demonstrate the lawfulness of its own processing activities. Supervisory authorities retain discretion to choose appropriate corrective measures on a case-by-case basis. Crucially, however, the Board does not say the deploying controller can simply ignore the tainted origin of the model it is using: it must conduct “an appropriate assessment, as part of its accountability obligations” to ascertain that the model was not developed through unlawful data processing (EDPB, Opinion 28/2024, para. 129).

The required depth of that assessment scales with risk. The EDPB identifies several non-exhaustive factors: the source of the personal data; whether the development-phase processing was the subject of a finding of infringement by a supervisory authority or a court; the nature and severity of the unlawfulness; and the existence of technical filters or access limitations that might prevent personal data from being accessible or disclosed during deployment (EDPB, Opinion 28/2024, paras. 129–132). The AI Act’s self-declaration of conformity by providers of high-risk AI systems, while potentially relevant context, is expressly noted as insufficient to constitute a conclusive finding of GDPR compliance (EDPB, Opinion 28/2024, para. 131; see also Hohmann & Kollár, 2025).

The German Federal Commissioner for Data Protection and Freedom of Information (BfDI), in its guide on data protection and AI in public authorities published on 22 December 2025, took a position that mirrors both the BVwG’s autonomy-of-assessment approach and the EDPB’s scenario-based framework.

What, then, does a compliant due-diligence process look like? Drawing on the EDPB’s guidance and the broader accountability architecture of the GDPR, a deploying controller should at minimum: verify the provenance of the training dataset and any available documentation regarding the legal bases relied upon during development; review whether any supervisory authority or court has found infringements in relation to that model or dataset; assess whether meaningful mitigation measures (e.g. pseudonymisation, differential privacy, output filters, opt-out mechanisms) have been implemented by the developer; and determine whether, taking all of this into account, its own chosen legal basis for deployment can withstand scrutiny (Pernice, 2025; EDPB, Opinion 28/2024, paras. 96–108). The GEDI/OpenAI case, in which the Italian Garante issued a formal warning in November 2024 (Provvedimento del 27 novembre 2024 [10077129]) precisely because GEDI appeared unable to guarantee a proper legal basis for the processing of sensitive data in its data-sharing agreement with OpenAI, illustrates what happens when this accountability chain breaks down(Cruz, 2025).

4. Conclusions

The Austrian cases and the EDPB’s Opinion 28/2024 converge on a single, clear message: the unlawfulness of initial data processing does not automatically cascade into a permanent and blanket prohibition on any subsequent use of a model trained on that data. The question of lawfulness is, and must be, answered by each controller independently, with reference to its own legal basis, its own accountability obligations, and the specific risk profile of its intended processing activity (for an in-depth analysis on these issues, specifically related to developers and deployers, see Pernice, 2025).

This case-by-case framework is not a soft option. For AI developers, it demands genuine attention to the legal status of training data from the outset, not least because a known infringement finding dramatically narrows the room for downstream controllers to claim compliance in good faith. For deployers, it requires active due diligence rather than passive reliance on contractual warranties. And for regulators, it offers a proportionate tool: one that can impose corrective measures where the risk to data subjects justifies them, without categorically paralysing an entire industry.

The broader implications for AI governance are significant. As AI models increasingly circulate across multiple organisations, the question of who bears responsibility for tainted training data will recur. The answer offered by the Austrian courts and the EDPB is architecturally sound: accountability is individual, proportionate, and contextual. What it requires in practice is a data protection culture that takes provenance seriously at every stage of the AI lifecycle, something the confluence of the GDPR and the AI Act, read together (Hohmann & Kollár, 2025), now firmly demands.

Share this article!

About Author

Eleonora Silva

PhD student in Law, Ethics and Economics for Sustainable Development, University of Milan Statale

Leave A Reply