{"id":13619,"date":"2026-07-27T15:32:17","date_gmt":"2026-07-27T13:32:17","guid":{"rendered":"https:\/\/orsingher.com\/?p=13619"},"modified":"2026-07-27T17:01:09","modified_gmt":"2026-07-27T15:01:09","slug":"data-protection-edpb-releases-guidelines-on-personal-data-scraping-for-genai-development","status":"publish","type":"post","link":"https:\/\/orsingher.com\/it\/data-protection-edpb-releases-guidelines-on-personal-data-scraping-for-genai-development\/","title":{"rendered":"DATA PROTECTION | EDPB releases guidelines on personal data scraping for GenAI development"},"content":{"rendered":"<p>On 7 July 2026, the European Data Protection Board (<em>EDPB<\/em>) adopted its\u00a0<a href=\"https:\/\/www.edpb.europa.eu\/system\/files\/2026-07\/edpb_guidelines_2020603_webscraping_v1_en_0.pdf\" target=\"_blank\" rel=\"noopener\">Guidelines\u00a003\/2026 on web scraping in the context of generative AI<\/a>\u00a0(the <strong><em>Guidelines<\/em><\/strong>). These\u00a0Guidelines\u00a0apply to private entities using automated tools to extract personal data from external internet sources to either train or fine\u2011tune AI models.<\/p>\n<p>The Guidelines\u00a0provide\u00a0practical guidance on compliance with the GDPR,\u00a0focusing on key obligations.\u00a0As regards transparency, the EDPB\u00a0emphasizes\u00a0that the exemption\u00a0from the obligation to provide a privacy policy\u00a0under\u00a0<a href=\"https:\/\/gdpr-info.eu\/art-14-gdpr\/\" target=\"_blank\" rel=\"noopener\">Article 14(5)(b) of GDPR<\/a>\u00a0should not be\u00a0applied routinely, but only when strictly necessary.\u00a0As regards the principle of data\u00a0minimisation, the EDPB underlines that controllers should assess the necessity of collecting personal data before scraping begins and, where\u00a0feasible,\u00a0prioritise\u00a0the use of synthetic data.\u00a0The Guidelines also\u00a0identify\u00a0specific\u00a0technical and\u00a0organisational\u00a0measures to support compliance.<\/p>\n<p>Furthermore, the EDPB confirms that consent would\u00a0generally not\u00a0be\u00a0an\u00a0appropriate\u00a0legal\u00a0basis for web scraping. Instead, entities\u00a0should\u00a0rely on legitimate interest\u00a0under\u00a0<a href=\"https:\/\/gdpr-info.eu\/art-6-gdpr\/\" target=\"_blank\" rel=\"noopener\">Article 6(1)(f) of GDPR<\/a> as a lawful basis, for which the Guidelines provide detailed guidance on the required balancing test. Notably, the EDPB extends the reasoning of the Court of Justice of the European Union in <a href=\"https:\/\/eur-lex.europa.eu\/legal-content\/EN\/TXT\/HTML\/?uri=CELEX:62017CJ0136\" target=\"_blank\" rel=\"noopener\">Case\u00a0C 136\/17<\/a> (GC and Others v CNIL)\u00a0to the incidental collection of special category data during web scraping for AI training. The\u00a0Guidelines\u00a0specify that this is permissible only where: (a) the processing is comparable to a search engine\u2019s processing activity; (b) the collection is incidental and residual, not intentional; (c) preventing such collection ex ante is genuinely difficult; and (d) the controller implements measures before, during and after training to prevent dissemination of the data.<\/p>\n<p>The\u00a0Guidelines\u00a0are open for public consultation until 30 October 2026.<\/p>\n<p>Newsletter n. 120 &#8211; July 206<\/p>\n","protected":false},"excerpt":{"rendered":"<p>On 7 July 2026, the European Data Protection Board (EDPB) adopted its\u00a0Guidelines\u00a003\/2026 on web scraping in the context of generative AI\u00a0(the Guidelines). These\u00a0Guidelines\u00a0apply to private entities using automated tools to extract personal data from external internet sources to either train or fine\u2011tune AI models. The Guidelines\u00a0provide\u00a0practical guidance on compliance with the GDPR,\u00a0focusing on key obligations.\u00a0As [&hellip;]<\/p>\n","protected":false},"author":6,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":"","_members_access_role":[],"_members_access_error":""},"categories":[28],"tags":[],"class_list":["post-13619","post","type-post","status-publish","format-standard","hentry","category-newsletters"],"acf":[],"_links":{"self":[{"href":"https:\/\/orsingher.com\/it\/wp-json\/wp\/v2\/posts\/13619","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/orsingher.com\/it\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/orsingher.com\/it\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/orsingher.com\/it\/wp-json\/wp\/v2\/users\/6"}],"replies":[{"embeddable":true,"href":"https:\/\/orsingher.com\/it\/wp-json\/wp\/v2\/comments?post=13619"}],"version-history":[{"count":2,"href":"https:\/\/orsingher.com\/it\/wp-json\/wp\/v2\/posts\/13619\/revisions"}],"predecessor-version":[{"id":13674,"href":"https:\/\/orsingher.com\/it\/wp-json\/wp\/v2\/posts\/13619\/revisions\/13674"}],"wp:attachment":[{"href":"https:\/\/orsingher.com\/it\/wp-json\/wp\/v2\/media?parent=13619"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/orsingher.com\/it\/wp-json\/wp\/v2\/categories?post=13619"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/orsingher.com\/it\/wp-json\/wp\/v2\/tags?post=13619"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}