DATA PROTECTION | EDPB releases guidelines on personal data scraping for GenAI development

On 7 July 2026, the European Data Protection Board (EDPB) adopted its Guidelines 03/2026 on web scraping in the context of generative AI (the Guidelines). These Guidelines apply to private entities using automated tools to extract personal data from external internet sources to either train or fine‑tune AI models.

The Guidelines provide practical guidance on compliance with the GDPR, focusing on key obligations. As regards transparency, the EDPB emphasizes that the exemption from the obligation to provide a privacy policy under Article 14(5)(b) of GDPR should not be applied routinely, but only when strictly necessary. As regards the principle of data minimisation, the EDPB underlines that controllers should assess the necessity of collecting personal data before scraping begins and, where feasible, prioritise the use of synthetic data. The Guidelines also identify specific technical and organisational measures to support compliance.

Furthermore, the EDPB confirms that consent would generally not be an appropriate legal basis for web scraping. Instead, entities should rely on legitimate interest under Article 6(1)(f) of GDPR as a lawful basis, for which the Guidelines provide detailed guidance on the required balancing test. Notably, the EDPB extends the reasoning of the Court of Justice of the European Union in Case C 136/17 (GC and Others v CNIL) to the incidental collection of special category data during web scraping for AI training. The Guidelines specify that this is permissible only where: (a) the processing is comparable to a search engine’s processing activity; (b) the collection is incidental and residual, not intentional; (c) preventing such collection ex ante is genuinely difficult; and (d) the controller implements measures before, during and after training to prevent dissemination of the data.

The Guidelines are open for public consultation until 30 October 2026.

Newsletter n. 120 – July 206