Framing an AI project: In what order should you ask the right questions?

A scoping sequence can be organised into five steps: define the business problem, specify the use cases, identify the legal and regulatory framework, analyse gaps and risks, and then determine the necessary data. However, this progression should not be seen as strictly linear: legal feasibility, risks, and data must be re-examined throughout the project.

Step 1: Why should you start with the business problem, not the tool?

The first question to ask should not be «what technology should we use?» but «what concrete problem are we trying to solve?». This distinction may seem obvious, yet it is the most common pitfall. A project that starts with the desire to adopt AI, without a clearly identified need, produces a system in search of a use, which is costly and rarely adopted.

Framing the problem in business terms requires answering simple yet fundamental questions: what isn't working today, how much does it cost, and what would a measurable improvement look like? If the answer isn't clear, the relevance of the future system will be difficult to demonstrate, and even harder to justify to management or a funder. An honest assessment at this stage may even conclude that an AI is not the best solution: sometimes, a simplified process or better team training resolves the problem at a much lower cost and risk.

AI business problem

Step 2: How to translate a business problem into concrete use cases?

Data scraping

Once the problem is stated, it must be translated into concrete uses. What exactly should the system do? For which users? In which situations? And above all, what uses must be explicitly excluded?

This last question is the most neglected and protective. Defining what a system must not do confines its scope and prevents misuse. It is at this stage that the abstract idea becomes an operational perimeter, that is, something that can be conceived, tested, and monitored. A vague perimeter, conversely, quickly turns a useful tool into a source of risk: one can neither evaluate nor control a system whose limits have never been defined.

Step 3: what legal and regulatory rules govern an AI project?

 

Next comes the question of applicable rules. This is not limited to the GDPR and the AI Act, even though these two texts are central. Depending on the use case, labour law, intellectual property, cybersecurity, consumer law, and especially sector-specific rules for health, finance, insurance, or education also come into play.

The aim of this stage is to identify not only the obligations but also the restrictions and prohibitions governing the intended use. Certain practices are, in fact, strictly prohibited, regardless of any transparency or consent requirements. Identifying these limitations at an early stage prevents investment in a project that is bound to be blocked.

 Questions to ask at this stage

  • What personal data will the system process, and on what legal basis?
  • Is this use case subject to any specific sector-specific regulations?
  • Are there any AI practices prohibited by the AI Act that might apply?
  • What transparency or explainability obligations are imposed?
Data minimisation

Step 4: How can you identify gaps and risks before deployment?

Sensitive data

With the framework established, we must examine the flaws.

Where can the system produce biases, to the detriment of certain groups? How can it be attacked or diverted from its intended use? Are its results sufficiently transparent and controllable by a human? Can it infringe upon the rights of the individuals concerned?

This analysis does not only serve to document risks: it determines the controls to be put in place before deployment. An identified risk upstream is generally less costly to address than the same risk discovered after deployment. The same risk discovered in production results in incidents, loss of confidence, and sometimes non-compliance.

Step 5: How do you determine what data the AI system needs?

The issue of data comes last of all. What data will enable the system to function? Where does it come from? Is it lawfully usable, relevant, reliable and sufficiently representative of the reality that the system will have to deal with?

Placing the data at the end of the process does not mean that it is of secondary importance; on the contrary: a system is only as good as the data that feeds into it. However, the relevance of the data can only be assessed once the problem, the use cases, the legal framework and the risks have been established. Seeking out data before defining the need amounts to collecting first and thinking later, which is precisely the source of many problems relating to quality, representativeness and compliance.

AI data

Why the framing of an AI project cannot be linear

This sequence is relevant because it is based on business needs rather than technology. However, it would be misleading if it were understood as a strictly linear process, in which each stage is completed once and for all.

One aspect, in particular, cannot be treated as simply a choice between two other options: the legal, regulatory and ethical framework.

Positioned at step 3, it might suggest that compliance is examined once, after the use cases have been defined. This is an error. Certain legal questions need to be raised as early as step 1. If the business need involves processing sensitive data, monitoring individuals, or implementing a prohibited practice, its legal feasibility must be established before the use cases are even detailed.

If you only realise in step 3 that a particular use is unlawful, it means you have wasted your time on the previous steps.

The correct view is therefore not that of a legal step inserted into a sequence of processes. The law first acts as a feasibility filter, at an early stage, and then supports the design, testing, deployment and monitoring of the system throughout its life cycle. Where the project involves personal data, this approach reflects the principle of data protection by design and by default, as set out in Article 25 of the GDPR: data protection is not an afterthought to the project; it is an integral part of it from the outset.

Step 5: How do you determine what data the AI system needs?

Framing an AI project is therefore not just a matter of putting the steps in the right order. It also involves identifying, right from the start, which aspects will need to be reviewed at each stage of the life cycle. The five-step sequence provides a guiding principle: start with the problem, not the tool. But its true value lies in what it implicitly emphasises: a sound AI project is not one that simply ticks boxes; it is one where, before investing, we have understood which questions will need to be addressed throughout its entire lifecycle.

The key question for any organisation embarking on an AI project is this: does the legal and regulatory framework come into play right from the definition of the need, or only at the deployment stage? The answer speaks volumes about the robustness of the project.

Frequently asked questions about framing an AI project

The majority of failures in AI projects are not caused by technological shortcomings, but by poor scoping. Organisations that start with the tool rather than the business problem end up building systems with no clearly defined purpose, which are difficult to justify and rarely adopted by teams.

The general sequence can be structured as follows: (1) define the business problem, (2) specify the use cases, (3) identify the legal and regulatory framework, (4) analyse the gaps and risks, and then (5) determine the necessary data.

This order is not strictly linear, however. Depending on the nature of the project, certain legal and regulatory questions must be examined from the feasibility study stage. If the need involves, for example, processing sensitive data, monitoring individuals, or implementing a practice that may be prohibited, the legal analysis may dictate whether the project can proceed even before the use case is fully defined.

This approach prevents investment in a technically unfeasible solution, one that is likely to be legally prohibited or insufficiently controlled.

 

Legal issues must be integrated from stage 1, not solely during a dedicated compliance phase. Certain legal prohibitions, notably those provided for by the AI Act or GDPR, can call into question the feasibility of the project before the use cases have even been clarified. Waiting until stage 3 to examine them exposes the project to the risk of having to restart all or part of it.

When defining the business problem, an honest assessment must consider non-AI alternatives: a streamlined process, better training for teams, or an existing tool. If a measurable improvement can be achieved at lower cost and with less risk without AI, this option deserves serious consideration.

 

Data evaluation occurring last does not mean it is less important. Its quality remains crucial for the system's performance. However, its relevance, legality, and representativeness can only be fully appreciated once the problem, use cases, legal framework, and risks have been clearly defined. Collecting data before these questions have been answered can lead to compliance and quality issues.

The principle of ‘privacy by design’, as set out in Article 25 of the GDPR, requires that the protection of personal data be built into the very design of the system, rather than being added as an afterthought. When applied to an AI project, this principle means that architectural choices, use cases and control measures must be designed with data protection requirements in mind from the very earliest stages of scoping.