EDPB Draft Guidelines on Anonymization and AI Web Scraping Spotlight Re-Identification Risks Amid Digital Omnibus Debates

Table of Contents

The European Data Protection Board has opened a public consultation on two closely related draft guidelines that address some of the most persistent friction points in modern data protection practice: when data can truly be considered anonymous, and how organizations may lawfully scrape publicly available information for artificial intelligence training. The consultation runs through 30 October 2026. The documents arrive at a pivotal moment, as EU institutions continue negotiations on the European Commission’s Digital Omnibus package—a set of proposed reforms that could reshape core concepts under the General Data Protection Regulation, including the definition of personal data itself.

In a recent LinkedIn Live discussion with the International Association of Privacy Professionals, EDPB Secretariat Deputy Head Gwendal Le Grand walked through the drafts and the Board’s broader positioning. The conversation underscored both the practical need for updated guidance and the EDPB’s firm stance against weakening the GDPR’s foundational definitions while reforms remain in flux.

The Digital Omnibus Context and the Fight Over “Personal Data”

The new guidelines were prepared in the wake of a joint opinion issued by the EDPB and the European Data Protection Supervisor on the Digital Omnibus proposals. That package has raised the prospect of material changes to the GDPR, with particular attention focused on the definition of personal data. Le Grand made clear that while the Board supported elements of the overall reform effort, it “advised quite strongly against any change to the definition of personal data.”

That resistance is not merely institutional caution. The current definition—any information relating to an identified or identifiable natural person—has been the bedrock of European data protection law for decades. Narrowing it, even with well-intentioned aims of reducing compliance burdens or clarifying edge cases, risks creating durable loopholes. If certain categories of information are deemed non-personal by legislative fiat rather than by rigorous factual assessment, controllers could more easily claim that GDPR obligations simply do not apply. The EDPB’s position reflects a concern that such a shift would undermine the regulation’s protective purpose at precisely the moment when technological capabilities for re-identification are expanding rapidly.

By issuing detailed guidance under the existing legal framework, the Board is effectively reinforcing the current standard while the legislative process continues. The guidelines offer organizations a clearer operational path under today’s rules and simultaneously signal what the EDPB believes should remain non-negotiable even after any Omnibus reforms are finalized.

Updating Anonymization Guidance for a New Technological Reality

The anonymization draft updates guidance last issued in 2014 by the Article 29 Working Party—the EDPB’s predecessor. That earlier opinion predated the GDPR and, more importantly, predated the widespread deployment of advanced machine learning, large-scale data linkage techniques, and generative AI systems capable of inferring sensitive attributes from seemingly innocuous datasets.

The central analytical shift in the new draft is a move away from absolute questions of identifiability toward a contextual, likelihood-based assessment. “Rather than asking if an individual is identified or identifiable in an absolute sense,” Le Grand explained, “the question is rather about the likelihood that the individual will be identified or identifiable by some entity. This may vary from one entity to another, and therefore anonymity has to be assessed from each relevant entity’s perspective.”

This approach recognizes a practical reality that privacy professionals have long observed: data that appears anonymous when held by one organization may become highly identifying when combined with other datasets or computational resources available to a different actor. Modern re-identification risks are no longer limited to traditional linkage attacks using quasi-identifiers. Sophisticated models can reconstruct identities or sensitive attributes through statistical inference, membership inference attacks, or the exploitation of residual patterns left after classic anonymization techniques such as k-anonymity or differential privacy implementations that prove insufficient against adaptive adversaries.

The draft outlines both a contextual approach and a simplified approach for organizations conducting anonymization assessments. The simplified route allows controllers to treat data as personal even when the residual risk of identification is low—an intentional shift toward false positives rather than false negatives. Le Grand noted that while this may lead organizations to apply GDPR protections more broadly than strictly required, it “can provide greater confidence, and it can be complemented with the contextualized approach to refine the finding.”

Critically, the guidance stresses that anonymization is not a one-time exercise. As organizations expand their processing capabilities, acquire new data sources, or deploy more powerful analytical tools, they must reassess whether previously anonymized datasets remain so. What was effectively anonymous in 2018 may no longer meet the standard in 2026. This ongoing duty of reassessment is particularly relevant for companies that retain large historical datasets or that share “anonymized” data with third parties whose technical capabilities differ from their own.

AI Web Scraping: Minimization, Special Categories, and the Limits of Safeguards

The companion draft on AI web scraping sits at the intersection of the GDPR and the EU AI Act. It runs parallel to joint work underway between the EDPB and the European AI Office on guidance covering compliance with both regimes. The document makes clear that organizations collecting data from publicly available sources for model training remain fully subject to GDPR principles, including data minimization, purpose limitation, and transparency obligations.

Controllers must limit collection to what is necessary for the stated training purpose and must provide transparency to individuals where required—even when the data originates from public websites. The practical difficulties are substantial. “Regardless of the safeguards you put in place,” Le Grand observed, “it’s quite difficult to have 100% assurance that you’re not collecting special categories of data.”

The draft draws on the Court of Justice of the European Union’s judgment in Case C-136/17, which addressed search engines’ incidental processing of sensitive data. Applying similar reasoning, the EDPB indicates that residual or incidental collection of special category data during AI training is not automatically unlawful, provided the controller implements measures designed to prevent the dissemination or further processing of that data in a manner that would cause harm. This is a pragmatic acknowledgment rather than a free pass: organizations still bear the burden of demonstrating appropriate technical and organizational measures.

To reduce reliance on questionable personal data altogether, the Board recommends considering synthetic data for model training. Other suggested techniques include syntax-based filtering to exclude certain content types, replacement of real data with synthetic alternatives where feasible, and robust anonymization applied before training begins. These recommendations align with the AI Act’s emphasis on high-quality, representative datasets and sound data governance practices for high-risk systems. Controllers that can demonstrate careful data sourcing and minimization will be better positioned both under the GDPR and when facing conformity assessments or market surveillance under the AI Act.

Re-Identification Risks in the Age of Advanced Analytics

The two drafts together highlight a structural challenge that has only intensified since 2014. Classic anonymization techniques were designed for a world of structured databases and relatively static analytical methods. Today’s environment features foundation models trained on vast corpora, continuous data flows, and adversaries equipped with comparable computational resources. Re-identification is no longer a binary technical question; it is a dynamic risk assessment that must account for the specific means reasonably likely to be used by particular entities.

This has direct implications for cross-border data transfers, secondary research uses, and the growing market for “anonymized” datasets. Organizations that treat anonymization as a checkbox exercise—applying a standard technique and then declaring the data free of GDPR constraints—face elevated regulatory and litigation risk. The EDPB’s contextual framing effectively requires documentation of the threat model, the residual risks, and the reasons why those risks are considered acceptable for each relevant recipient or processing context.

Practical Implications and the Path Forward

For companies developing or deploying AI systems that rely on web-scraped data, the message is unambiguous: public availability does not equal free-for-all processing. Data minimization and purpose limitation apply with full force. Special category data is difficult to exclude completely, and incidental collection must be paired with effective containment measures. Synthetic data and advanced filtering should be evaluated seriously rather than treated as aspirational best practices.

For any organization that relies on anonymized datasets—whether for analytics, research, or commercial products—the updated guidance demands a more rigorous, entity-specific, and iterative approach to risk assessment. Documentation of the contextual analysis will become increasingly important as supervisory authorities begin applying the new standards.

The public consultation period offers a concrete opportunity for stakeholders to shape the final texts. Le Grand emphasized that the EDPB has strengthened its processes for considering external input: “We’ve made a lot of efforts to engage more proactively and be more transparent about our stakeholder engagement and about how we use the input that we receive through the stakeholder consultations. I can assure you that all the input you send us is considered, everything is analyzed and taken into account.”

As negotiations on the Digital Omnibus continue, these draft guidelines serve a dual purpose. They provide immediate operational clarity under the existing GDPR and they reinforce the EDPB’s view that the definition of personal data should remain robust. Organizations that engage thoughtfully with the consultation—and that begin aligning their anonymization and AI data practices with the principles outlined in the drafts—will be better prepared for both the final guidance and whatever legislative adjustments ultimately emerge from the Omnibus process.

The stakes are high. In an environment where the technical capacity to re-identify individuals continues to advance, the legal and practical boundaries of anonymity and lawful scraping will determine how much of the digital economy remains subject to meaningful data protection constraints. The EDPB’s drafts make clear that those boundaries are not expected to loosen.

Written by: 

Online Privacy Compliance Made Easy

Captain Compliance makes it easy to develop, oversee, and expand your privacy program. Book a demo or start a trial now.