In AI projects, data quality determines success or failure earlier than the choice of model. If you apply a language model or a classification model to incomplete, contradictory, or outdated data, you’ll get a result that reflects exactly those shortcomings—only faster and on a larger scale. This article explains how to identify these issues in concrete terms, how much time data preparation realistically requires, and when a pilot project still makes sense. This applies equally to a classification model in production and to a language model designed to respond to internal company knowledge.
Good AI results require that the data source meet six criteria: completeness, accuracy, consistency, timeliness, uniqueness, and accessibility. According to an analysis by superkind.ai from April 2026, data preparation accounts for an average of about 61 percent of the project duration. Those who factor in this effort—rather than underestimating it—will moving forward even with mediocre data.
Data Quality in AI Projects: Six Dimensions
The term “data quality” seems abstract until you break it down into specific criteria. Six dimensions are most commonly encountered in practice and can be assessed individually for each data source.
| Dimension | Exam Question |
|---|---|
| Completeness | Are there any missing values in required fields? |
| Accuracy | Do the figures match reality? |
| Consistency | Do values differ across systems? |
| Timeliness | How old is the latest update? |
| Uniqueness | Are there any duplicates or alternative spellings? |
| Accessibility | Is it technically possible to connect to the source? |
According to an analysis of superkind.ai According to the analysis from April 2026, a significant portion of the required data fields are incomplete in the projects examined, most frequently in terms of completeness and timeliness. According to the same analysis, formal data quality assessments conducted before the start of a project significantly increase the success rate.
How a Data Quality Assessment Works in Practice
An assessment begins with a sample, not with the selection of a tool. Several hundred records are drawn from the target source and checked against the six dimensions—either manually or using a simple script that flags missing fields, duplicates, and outliers. This sample yields an error rate for each dimension—for example, about 8 percent of missing phone numbers or 12 percent of conflicting addresses between two systems, as illustrative values from a typical customer database.
This error rate determines the next step. If it falls within a range that a prototype can handle, the actual modeling work begins. If it is significantly higher, it is worth performing a targeted cleanup of the most important fields before the pilot project—but only those fields that the use case actually needs, not the entire database.
How Much Time Data Preparation Actually Takes
The same analysis by superkind.ai from April 2026 estimates that data preparation accounts for an average of about 61 percent of the total project duration. Anyone planning an AI project with four weeks of model development should factor in ten weeks of data preparation beforehand, not two. In project plans that underestimate this proportion, the deadline for the first usable prototype is usually pushed back; the data work itself can hardly be shortened. For high-risk systems, this effort is now also legally mandated. Article 10 of the AI Regulation requires that training, validation, and test data be relevant and sufficiently representative, and—as far as possible—free of errors and complete. Anyone who reads this requirement only after the prototype has been developed will have to plan the data work twice. The article provides a complete overview of the deadlines through 2028 EU AI Act: Compliance Roadmap Through 2028.
On average, the project duration is spent on data preparation, not on the model.
superkind.ai, April 2026
A frequently cited statistic—that a large proportion of all AI projects fail due to poor data quality—is often mentioned in many presentations and articles. It is based on individual analyses by specific providers, not on a broad, independent study, and should therefore be viewed as an estimate rather than a reliable measurement.
Why a pilot program Still Makes Sense
There is no such thing as perfect data. Every source has gaps, duplicates, or outdated entries—and this applies just as much to an in-house ERP system as it does to a public database. For a pilot project, the most important factor is whether a single, well-maintained source is sufficient for an initial use case.
A pilot project with a single, well-maintained data source is like a data project that never gets completed because it tries to clean up all the data sources in the company first. If you wait for perfect data, you’ll never get started.
In our own projects, we often see the opposite problem. Teams spend months cleaning up a database before the first prototype is even up and running, and in the process lose sight of the actual use case. A narrow use case with a single data source reveals more quickly whether the effort is worth it than a company-wide data project with no clear endpoint.
From the Database to the Pilot Project
A verified, well-defined data source paves the way for the next step. Our „Methodology for an AI Pilot Project in 90 Days“ describes how this leads to the creation of a prototype with real users. Anyone who would like to have their own data situation assessed in advance can do so in a Initial Consultation on AI Consulting An initial, independent assessment.
Classify the data available prior to the prototype
A quick sample check before prototyping can determine whether an existing data source is sufficient for an initial use case. This reveals whether specific fields need to be cleaned up before modeling begins.
Frequently Asked Questions
How do you check data quality?
The quickest way is to perform a spot check across the six dimensions: completeness, accuracy, consistency, timeliness, unambiguity, and accessibility. For a customer database, for example, this means manually checking 100 to 200 records and recording the error rate for each dimension, rather than cleaning the entire source in advance. The results fit on a single page and are sufficient to make the next decision.
Who is responsible?
Data quality is rarely the responsibility of a single role. The business department understands the meaning of the fields, IT understands the technical aspects, and the project team for the AI initiative must bring these two perspectives together. Without a designated person to coordinate this alignment, it remains an issue that, in the end, no one feels responsible for. In smaller teams, a single point of contact with access to both sides is sufficient; in larger organizations, a set format is needed—such as a brief weekly coordination meeting during the pilot phase.
Is a pilot worth it if the data is mediocre?
In most cases, yes—as long as the use case is narrowly defined enough. A pilot with known gaps reveals which errors are truly problematic in practice and which have little impact on the result. This insight is difficult to gain from a table with missing values, but can be obtained from a running prototype after just a few weeks. The key is to keep the scope of the pilot narrow enough so that a single missing data point does not immediately render the results unusable.
Why torck
torck builds AI applications with its own development teams in Maxhütte-Haidhof, Vienna, and Rabat, and begins every project by assessing the actual data situation—not an idealized version of it. This assessment often determines whether a use case will yield a prototype in six weeks or only after a lengthy data project. Anyone who wants to realistically assess their own data set can do so in a Initial Consultation on AI Consulting .
This article refers to laws and regulations to put technical decisions in context. It is not legal advice. Whether and how a rule applies to your company is a question for your legal department or a law firm.