Risks of AI-Generated Code: Quality, Liability, Security

Risks of AI-Generated Code: Quality, Liability, Security

AI-generated code often looks finished before it actually is. This is precisely where the risks of AI-generated code lie—risks that differ from traditional programming errors. A language model generates syntactically correct, executable code that may nonetheless contain a security vulnerability, a license violation, or a lack of input validation—without this being apparent at first glance.

In a nutshell

AI-generated code carries three risks that are not present in traditional manual development: security vulnerabilities in executable code, omitted input validation, and licensing obligations arising from the model’s training data. All three are difficult to detect in a superficial review because the code appears unremarkable at first glance. An effective approach combines automated scanners, a second review by experienced developers, and a licensing analysis prior to deployment.

AI-Generated Code by the Numbers

A widely cited study by the University of Illinois In 2022, the company reviewed approximately 1,689 code examples generated by GitHub Copilot in security-related scenarios. According to this analysis, about 40 percent of the examples were vulnerable to flaws listed in the MITRE CWE Top 25, a list of the most common and dangerous software flaws. This figure comes from a controlled test setup using a single, now-outdated model. It cannot be directly applied to current tools, but the basic pattern has been confirmed in more recent studies.

40 percent

So many of the Copilot code examples that were tested were vulnerable to weaknesses from the MITRE CWE Top 25 in security-related scenarios.

University of Illinois, 2022

One Analysis on arXiv A study from 2026 examined 86,726 code examples across seven language models and four programming languages for compilation and runtime errors. According to this analysis, generated code conspicuously often omits input validation—precisely the point at which a program checks whether incoming data is even plausible. If this check is missing, a simple invalid input can cause the program to crash or, in the worst-case scenario, lead to a security vulnerability.

A typical example illustrates how this works in practice. A language model suggests code for a login function that verifies credentials but does not include a limit on the number of failed attempts. The code compiles, and the function works flawlessly in testing; nevertheless, this leaves the system vulnerable to automated login attempts without anyone noticing. Such vulnerabilities are rarely detected during a superficial review because the code appears unremarkable at first glance.

Three Types of Errors in Generated Code and How to Address Them
Risk How You Can Tell Countermeasure
Security vulnerability Only through targeted testing for known vulnerability patterns Automated scanner with every commit
Lack of Input Validation If an incorrect entry is made during testing or operation Review by an experienced developer
License Violation Often not until an audit or a company sale License Analysis Before Deployment

Licensing Risk: Who Is Liable for Adopted Code?

In addition to security, there is a second, often-overlooked risk: licensing law. Language models are trained using code from publicly accessible repositories, including code under copyleft licenses such as the GPL, which can impose certain obligations upon reuse—such as the disclosure of one’s own source code. If a developer inadvertently triggers such an obligation, the consequences fall on the company using the tool. The provider of the AI tool is generally not liable for this.

According to a Black Duck survey, only about 54 percent of organizations actually check AI-generated code for licensing and IP risks before deployment. This means that just under half of the companies surveyed put code into production without knowing what licensing obligations it might entail. During an audit or the sale of a company, this can become a problem, because a buyer or investor typically has the codebase reviewed for such legacy issues. This article describes how these factors can shift contractual obligations for clients overall. AI-Powered Software Development: What's Changing for Clients.

The Limits of the Copilot Number

The 40 percent figure from the Illinois study is often cited out of context, but it deserves some context. It dates from 2022 and refers to a single model and a limited set of test scenarios specifically designed to simulate security-related situations—not the average of all generated lines of code. Current models perform better in certain areas because providers have revised their training data and security filters multiple times since 2022.

Nevertheless, the fundamental problem remains, as shown by the 2026 analysis, which included 86,726 code examples. A language model is not familiar with the context of the target system and therefore cannot know which input is considered dangerous in a specific company. This gap can be narrowed through better training data, but it cannot be completely closed.

How to Reduce Risk in Practice

What works is a combination of several, rather unspectacular measures. Automated security scanners that check every commit against known vulnerability patterns catch some of the problems before a human even looks at them. A second review by an experienced developer is still necessary, however, because scanners only look for known patterns and may overlook novel combinations of errors. In security-critical areas such as authentication, payment processing, or interfaces to external systems, it’s also worth establishing a firm rule that AI suggestions are never adopted without review, regardless of how tight the deadline may be.

What This Means for Refactoring and Code Quality

Unreviewed AI code has a second, longer-term effect. It quickly becomes part of the technical debt of a system, because although it works, it is neither documented nor aligned with the patterns in the rest of the system. Anyone who regularly incorporates code from a language model should therefore also measure the contribution of that code to their own error rate, rather than relying on the impression that development as a whole has become faster.

Where this approach reaches its limits
Scanners can only detect what they recognize from a pattern. They consistently overlook novel combinations of errors and technical errors in business logic. Likewise, license analysis only detects code that is close enough to a known source; rewritten code falls through the cracks. And the 40 percent figure comes from a 2022 test setup using a single model; it does not represent the average performance of today’s tools.

Have Existing AI Code Reviewed
Security vulnerabilities, missing input validation, and unclear licensing requirements are rarely noticed during a cursory review. A look at your own codebase will reveal where it’s worth starting a more thorough review.

Get in touch

Frequently Asked Questions

Is the service provider liable for AI code?

In principle, yes, if the service provider delivers and accepts the code, regardless of whether it was written by a human or a language model. Liability under a contract for work under German law is based on the result, not on the tool used to create it. Nevertheless, it should be clearly stipulated how to handle license violations arising from the training of the model. A clause covering violations discovered even after the fact has proven effective, as such cases are often only detected months after acceptance during an external audit.

How do you check for licensing issues?

License analysis software, often referred to as software composition analysis, compares sections of code with known open-source projects and flags matches along with their associated licenses. In addition, a fixed rule in the development process helps: Code segments that are copied almost verbatim from a known source are flagged separately and reviewed before they are included in a product. This review must take place before deployment, not during a subsequent cleanup phase, because once code with an unresolved license has been released, it is virtually impossible to retrieve it.

Should the use of AI be banned?

A ban would merely postpone the problem rather than solve it, because its use is virtually impossible to control completely in practice, and the productivity gains are real. A more sensible approach is a set of rules that precisely defines where AI-generated code may be used without additional review and where a security review is mandatory—for example, for anything related to authentication, payment data, or external interfaces.

At torck, AI-generated code undergoes the same review process as any other code, overseen by the development teams in Maxhütte-Haidhof, Vienna, and Rabat. Since the contracting party is the German company torck GmbH, issues regarding liability and license verification can be addressed contractually before a project for industry and retail begins. Anyone who wants to have an existing codebase reviewed for precisely these risks can do so in a Initial Consultation address.

Schedule a meeting

Legal note
This article refers to laws and regulations to put technical decisions in context. It is not legal advice. Whether and how a rule applies to your company is a question for your legal department or a law firm.

Questions about this post?

Just a couple of sentences about your situation will suffice. The person responding builds these kinds of systems himself.

We'll respond within one business day.torck · code with torque
Florian Blischke
Managing Director of torck GmbH · Over 20 years of software development experience
Florian Blischke is the managing director of torck GmbH and has been working in software development for over 20 years. He is responsible for custom software solutions for industry and retail, ranging from the integration of physical processes and IoT to cloud architecture and data- and AI-driven systems. At torck, he oversees, among other projects, the Jouvoli energy platform and the KVM Fleet fleet management product. torck develops software at its locations in Maxhütte-Haidhof, Vienna, and Rabat, and places a strong emphasis on software that actually works in real-world operations.

Are you facing the same question?

We’ve been building software for industry and retail since 2017, based in Maxhütte-Haidhof, with teams in Vienna and Rabat. An initial consultation lasts 30 minutes and is free of charge. Afterward, you’ll know whether the project is worth pursuing—even if the answer is no.

More Articles