Data Quality as the Foundation of AI: Data First, Then Models

Data Quality as the Foundation of AI: Data First, Then Models

In AI projects, data quality determines success or failure earlier than the choice of model. If you apply a language model or a classification model to incomplete, contradictory, or outdated data, you’ll get a result that reflects exactly those shortcomings—only faster and on a larger scale. This article explains how to identify these issues in concrete terms, how much time data preparation realistically requires, and when a pilot project still makes sense. This applies equally to a classification model in production and to a language model designed to respond to internal company knowledge.

In a nutshell

Good AI results require that the data source meet six criteria: completeness, accuracy, consistency, timeliness, uniqueness, and accessibility. According to an analysis by superkind.ai from April 2026, data preparation accounts for an average of about 61 percent of the project duration. Those who factor in this effort—rather than underestimating it—will moving forward even with mediocre data.

Data Quality in AI Projects: Six Dimensions

The term “data quality” seems abstract until you break it down into specific criteria. Six dimensions are most commonly encountered in practice and can be assessed individually for each data source.

Six Dimensions of Data Quality
Dimension Exam Question
Completeness Are there any missing values in required fields?
Accuracy Do the figures match reality?
Consistency Do values differ across systems?
Timeliness How old is the latest update?
Uniqueness Are there any duplicates or alternative spellings?
Accessibility Is it technically possible to connect to the source?

According to an analysis of superkind.ai According to the analysis from April 2026, a significant portion of the required data fields are incomplete in the projects examined, most frequently in terms of completeness and timeliness. According to the same analysis, formal data quality assessments conducted before the start of a project significantly increase the success rate.

How a Data Quality Assessment Works in Practice

An assessment begins with a sample, not with the selection of a tool. Several hundred records are drawn from the target source and checked against the six dimensions—either manually or using a simple script that flags missing fields, duplicates, and outliers. This sample yields an error rate for each dimension—for example, about 8 percent of missing phone numbers or 12 percent of conflicting addresses between two systems, as illustrative values from a typical customer database.

This error rate determines the next step. If it falls within a range that a prototype can handle, the actual modeling work begins. If it is significantly higher, it is worth performing a targeted cleanup of the most important fields before the pilot project—but only those fields that the use case actually needs, not the entire database.

How Much Time Data Preparation Actually Takes

The same analysis by superkind.ai from April 2026 estimates that data preparation accounts for an average of about 61 percent of the total project duration. Anyone planning an AI project with four weeks of model development should factor in ten weeks of data preparation beforehand, not two. In project plans that underestimate this proportion, the deadline for the first usable prototype is usually pushed back; the data work itself can hardly be shortened. For high-risk systems, this effort is now also legally mandated. Article 10 of the AI Regulation requires that training, validation, and test data be relevant and sufficiently representative, and—as far as possible—free of errors and complete. Anyone who reads this requirement only after the prototype has been developed will have to plan the data work twice. The article provides a complete overview of the deadlines through 2028 EU AI Act: Compliance Roadmap Through 2028.

61 %

On average, the project duration is spent on data preparation, not on the model.

superkind.ai, April 2026

A frequently cited statistic—that a large proportion of all AI projects fail due to poor data quality—is often mentioned in many presentations and articles. It is based on individual analyses by specific providers, not on a broad, independent study, and should therefore be viewed as an estimate rather than a reliable measurement.

Why a pilot program Still Makes Sense

There is no such thing as perfect data. Every source has gaps, duplicates, or outdated entries—and this applies just as much to an in-house ERP system as it does to a public database. For a pilot project, the most important factor is whether a single, well-maintained source is sufficient for an initial use case.

The Limitations of the Available Data
A pilot project with a single, well-maintained data source is like a data project that never gets completed because it tries to clean up all the data sources in the company first. If you wait for perfect data, you’ll never get started.

In our own projects, we often see the opposite problem. Teams spend months cleaning up a database before the first prototype is even up and running, and in the process lose sight of the actual use case. A narrow use case with a single data source reveals more quickly whether the effort is worth it than a company-wide data project with no clear endpoint.

From the Database to the Pilot Project

A verified, well-defined data source paves the way for the next step. Our „Methodology for an AI Pilot Project in 90 Days“ describes how this leads to the creation of a prototype with real users. Anyone who would like to have their own data situation assessed in advance can do so in a Initial Consultation on AI Consulting An initial, independent assessment.

Classify the data available prior to the prototype
A quick sample check before prototyping can determine whether an existing data source is sufficient for an initial use case. This reveals whether specific fields need to be cleaned up before modeling begins.

Get in touch

Frequently Asked Questions

How do you check data quality?

The quickest way is to perform a spot check across the six dimensions: completeness, accuracy, consistency, timeliness, unambiguity, and accessibility. For a customer database, for example, this means manually checking 100 to 200 records and recording the error rate for each dimension, rather than cleaning the entire source in advance. The results fit on a single page and are sufficient to make the next decision.

Who is responsible?

Data quality is rarely the responsibility of a single role. The business department understands the meaning of the fields, IT understands the technical aspects, and the project team for the AI initiative must bring these two perspectives together. Without a designated person to coordinate this alignment, it remains an issue that, in the end, no one feels responsible for. In smaller teams, a single point of contact with access to both sides is sufficient; in larger organizations, a set format is needed—such as a brief weekly coordination meeting during the pilot phase.

Is a pilot worth it if the data is mediocre?

In most cases, yes—as long as the use case is narrowly defined enough. A pilot with known gaps reveals which errors are truly problematic in practice and which have little impact on the result. This insight is difficult to gain from a table with missing values, but can be obtained from a running prototype after just a few weeks. The key is to keep the scope of the pilot narrow enough so that a single missing data point does not immediately render the results unusable.

Why torck

torck builds AI applications with its own development teams in Maxhütte-Haidhof, Vienna, and Rabat, and begins every project by assessing the actual data situation—not an idealized version of it. This assessment often determines whether a use case will yield a prototype in six weeks or only after a lengthy data project. Anyone who wants to realistically assess their own data set can do so in a Initial Consultation on AI Consulting .

Schedule a meeting

Legal note
This article refers to laws and regulations to put technical decisions in context. It is not legal advice. Whether and how a rule applies to your company is a question for your legal department or a law firm.

Questions about this post?

Just a couple of sentences about your situation will suffice. The person responding builds these kinds of systems himself.

We'll respond within one business day.torck · code with torque
Florian Blischke
Managing Director of torck GmbH · Over 20 years of software development experience
Florian Blischke is the managing director of torck GmbH and has been working in software development for over 20 years. He is responsible for custom software solutions for industry and retail, ranging from the integration of physical processes and IoT to cloud architecture and data- and AI-driven systems. At torck, he oversees, among other projects, the Jouvoli energy platform and the KVM Fleet fleet management product. torck develops software at its locations in Maxhütte-Haidhof, Vienna, and Rabat, and places a strong emphasis on software that actually works in real-world operations.

Are you facing the same question?

We’ve been building software for industry and retail since 2017, based in Maxhütte-Haidhof, with teams in Vienna and Rabat. An initial consultation lasts 30 minutes and is free of charge. Afterward, you’ll know whether the project is worth pursuing—even if the answer is no.

More Articles