Generative AI, web scraping and blockchain

The EDPB's new orientations

Data accessible on the internet isn't necessarily data that can be freely used to train artificial intelligence. That's the central message of Draft guidelines on web scraping in the context of generative AI that the European Data Protection Board (EDPB) adopted at its plenary meeting on 8 July 2026, and submitted for public consultation until 30 October 2026. The text criticises a practice that has become the bedrock of generative AI: web scraping, which is the automated and massive collection of content published on the web to form training datasets.

This article presents the main contributions of this project, its concrete implications for businesses, and the reasons why compliance is as much about protecting investments as it is a legal obligation.

Why GDPR applies to publicly accessible data

The most common misconception regarding scraping can be summarised in one sentence: «this data is public, therefore I can use it». This reasoning is legally flawed, and the EDPB has stated this unambiguously. As soon as the collected information allows for the identification of a person, directly or indirectly, it constitutes personal data, and its collection, storage, organisation, and reuse are all processing operations subject to the GDPR. The public nature of the source does not negate the rights of the individuals concerned.

The practical consequence is immediate: any organisation collecting data from the web to train, refine or enrich an AI system must be able to justify this processing. This applies to large language model developers, but also to any company that builds its own dataset, refines an existing model on scraped data, or uses a service provider to do so. In the latter case, the first question to be decided is that of the roles: who is responsible for the processing, who is the sub-contractor, and are the parties jointly responsible?

From this qualification arise the obligations of each party.

The flag of the European Union fluttering beneath the Cinquantenaire Arch in Brussels, Belgium.

Legitimate interest, the reference legal basis, but subject to conditions

Data scraping

Il n'existe pas de base légale unique et explicite autorisant le scraping à grande échelle. La légalité du scraping dépend fortement du contexte, des données que vous scrapez et de la manière dont vous les utilisez. Cependant, voici plusieurs bases légales et concepts juridiques qui sont souvent discutés et peuvent influencer la légalité du scraping à grande échelle : * **Droit d'auteur (Copyright Law) :** En général, le contenu d'un site web est protégé par le droit d'auteur. La copie de ce contenu sans autorisation peut constituer une violation. Cependant, des exceptions existent, notamment pour la fouille de textes et de données (Text and Data Mining - TDM) à des fins de recherche, ou si le contenu est sous licence ouverte. * **Conditions d'utilisation (Terms of Service / Terms of Use) :** La plupart des sites web ont des conditions d'utilisation qui interdisent explicitement le scraping automatisé ou l'extraction de données. Le non-respect de ces conditions peut entraîner des poursuites civiles pour rupture de contrat. * **Accès non autorisé à un système informatique (Computer Misuse Act / CFAA) :** Dans certains cas, le scraping peut être considéré comme un accès non autorisé à un système informatique, surtout s'il contourne des mesures de sécurité techniques (comme les CAPTCHAs, les blocages d'IP, etc.). Cela peut relever de la criminalité informatique. * **Droit de la base de données (Database Rights) :** Dans l'Union Européenne (et au Royaume-Uni via la législation britannique), les bases de données peuvent être protégées par un droit sui generis, qui interdit l'extraction et la réutilisation de tout ou partie substantielle de la base de données. * **Protection des données personnelles (RGPD / GDPR) :** Si vous scrapez des données personnelles, vous devez respecter le Règlement Général sur la Protection des Données (RGPD) dans l'UE et le Royaume-Uni. Cela inclut d'avoir une base légale pour le traitement (par exemple, le consentement, l'intérêt légitime), de respecter la minimisation des données, et d'assurer la sécurité des données. * **Droit de la concurrence / Dénigrement commercial :** Si le scraping est effectué dans le but de nuire à un concurrent ou d'extraire des informations pour des pratiques déloyales, cela pourrait être l'objet de poursuites relatives au droit de la concurrence ou au dénigrement. * **Intérêt légitime (Legal Basis for Processing under GDPR) :** Dans certains contextes (et sous réserve de l'applicabilité du RGPD), une entreprise peut baser son scraping sur son "intérêt légitime", mais cela nécessite une analyse rigoureuse pour s'assurer que les droits et libertés des personnes dont les données sont traitées ne sont pas outrepassés. **En résumé, pour fonder légalement le scraping à grande échelle, il faudrait souvent s'assurer que :** 1. **Les données sont publiques et exemptes de restrictions claires (comme un "no-scraping" explicite dans les conditions d'utilisation).** 2. **L'extraction ne viole pas le droit d'auteur (par exemple, s'il s'agit d'une analyse statistique de données publiques, ou si l'on se limite à des métadonnées).** 3. **Il n'y a pas de contournement de mesures de sécurité techniques. ** 4. **Si des données personnelles sont collectées, le RGPD est strictement respecté.** 5. **L'utilisation des données scrapees est légale et éthique (pas de concurrence déloyale, pas de diffusion d'informations fausses, etc.).** Il est **crucial de consulter un avocat spécialisé en droit du numérique et en protection des données** avant de se lancer dans du scraping à grande échelle, car les risques juridiques sont considérables.

Consent, the cornerstone of GDPR, is practically inaccessible: it is impossible to obtain the prior agreement of millions of people whose publications are collected indiscriminately. This leaves legitimate interest, which the EDPB recognises as the most realistic basis for this type of processing, while strictly regulating its invocation. On this point, the guidelines extend and clarify, with examples specific to scraping, the opinion that the Committee had dedicated to AI models.

The company must pass a three-stage test. Firstly, it must demonstrate a real and clearly defined interest: developing an AI system for a specific purpose can be one, provided that this purpose is explicit and documented. Secondly, it must establish the necessity of the processing: is the envisaged collection essential, or is there a less intrusive solution to achieve the same objective? Finally, it must weigh its interest against the rights and freedoms of the data subjects, and this balancing exercise must not be to their detriment.

It is in this balancing act that protective measures take on their full importance. The EDPB indicates that concrete safeguards can tip the balance favourably: exclusion of certain categories of data from the outset of collection, enhanced transparency on scraping activities, limitation of volumes collected, and anonymisation or pseudonymisation as early as possible. In other words, the lawfulness of scraping is not decreed, it is built through precautions taken at each step.

Minimisation begins before collection

One of the most tangible contributions of the draft guidelines lies in their interpretation of the minimisation principle. This principle is not limited to sorting data after the event; it applies before, during, and after collection.

Before collection, the company must consider whether scraping is actually necessary, and in particular whether synthetic data could suffice for its needs. It must define precise collection criteria rather than scraping the web indiscriminately. During collection, exclusion filters must discard what should not be captured, and certain sources must be avoided from the outset: sites structurally likely to contain sensitive data, sites intended for minors, and those that clearly oppose scraping. After collection, sorting, deletion, and anonymisation extend the process.

Data minimisation

Sensitive data: no general exemption

Sensitive data

On so-called sensitive data (health, political opinions, sexual orientation, biometric data, among others), the draft guidelines shut the door on any accommodating interpretation. Their processing is in principle prohibited, except for strict exceptions: it requires both a legal basis under Article 6 of the GDPR and an exception provided for by Article 9, paragraph 2. And the EDPB specifies that there is no general exemption for scraping.

However, the Committee acknowledges a practical reality: despite filters, massive collection can accidentally and marginally capture sensitive data. Drawing on the judgment of the Court of Justice of the European Union in GC and others (C-136/17), the EDPB suggests that this incidental or residual collection can be tolerated on a case-by-case basis, under stringent conditions: the controller must act within their responsibilities, competencies, and capabilities, and demonstrate that they have implemented technical and organisational measures to prevent such collection, promptly delete the data concerned when detected, and prevent its dissemination. Therefore, the tolerance does not concern the principle itself, but the duly controlled accident.

Accuracy: reliable, dated, and validated sources

The project also dedicates developments to the principle of accuracy, often overlooked in discussions on scraping. The EDPB recommends three concrete practices: only collect from reliable sources, record the timestamp of collection, and validate the data before using it for training. The logic is simple: the web contains false, outdated, or contradictory information, and a model trained without caution will absorb it. For the company, the stakes are twofold: compliance with the GDPR's principle of accuracy, and the quality of the resulting system itself, as a model fed with dubious data will produce dubious results.

Data, reliable sources

End-to-end transparency and memorisation: the obligations

Transparency

Two further requirements merit attention. The first concerns transparency: when it proves impossible or disproportionate to provide individual information to each person concerned, which is frequent in the case of scraping, the company is not thereby exempted from informing them. It must then make public information about its collection activities, accessible to all.

The second point concerns a risk inherent to AI models: memorisation. A model trained on personal data can, under certain conditions, reproduce this data in its responses. The EDPB expects companies to implement measures against this memorisation and regurgitation, which shifts the compliance obligation from the collection phase alone to the model's behaviour itself, throughout its life cycle.

From constraint to competitive advantage

These requirements could be read as a stack of constraints. This would miss their economic scope. A dataset collected without a solid foundation is a fragile asset: a company might be forced to delete it, revise its project, or retrain its model, and an AI investment could then lose a large part of its value. Conversely, the ability to demonstrate the origin of the data, the legality of its collection, and the protections put in place strengthens the confidence of customers, partners, and regulators. In a market where questions about the provenance of training data are multiplying, this traceability becomes a commercial argument as much as a legal shield.

The right question is therefore no longer «is this data accessible?», but «is its collection and use necessary, proportionate, transparent, and documented?»

Conclusion: the right time to act

The draft guidelines are open for public consultation until 30 October 2026, and the final version will take account of it. Affected businesses therefore have a dual interest in engaging with it now: auditing their data collection practices and governance mechanisms in light of the text, and, for those who wish to, contributing to the consultation to make their operational realities heard.

It should also be noted that the EDPB adopted, during the same plenary session, Guidelines on anonymisation. The two texts are complementary: early anonymisation is precisely among the guarantees that support the lawfulness of scraping. Reading them together is therefore essential for any organisation concerned.

GDPR compliance is not just a constraint: it helps protect the value created by AI projects. The EDPB's guidelines now provide the instructions for doing so.

Foire aux questions (FAQ)

This is a text adopted on 8 July 2026 by the EDPB which clarifies the GDPR obligations applicable to web scraping used for training generative AI systems. It covers, in particular, the legal bases, the principles of minimisation, accuracy and transparency, as well as the rules applicable to sensitive data.

The two texts are currently subject to public consultation until 30 October 2026. The final versions will be published after analysis of the contributions received.

For private companies, legitimate interest (Article 6(1)(f) of the GDPR) is generally the most suitable basis. It requires demonstrating a specific purpose, the necessity of the collection, and the absence of a disproportionate infringement of individuals' rights.

In principle, no: their processing is prohibited unless both a legal basis (Article 6 of the GDPR) and an exception under Article 9(2) are met. Only accidental and marginal collection can be tolerated, on a case-by-case basis, provided that the company has implemented measures to prevent it, quickly delete the data concerned and prevent its dissemination.

Any organisation that trains, fine-tunes or enriches an AI system with data from the web can be affected, not just large model developers. Start-ups, SMEs and research institutions are also targeted.

Contributions can be submitted directly on the EDPB website, via the dedicated pages for public consultations on anonymisation and web scraping, until 30 October 2026.

Governance first

Risk, responsibility, and clarity of decisions before tools.

Enduring capabilities

Internal systems that persist, even when teams change.

IA responsible

Ethical, compliant, and context-appropriate adoption.

Programme approach

From a status report to a real, scalable system anchored in your organisation.