StackAttestTechnical Trust
Security

AI-generated code vulnerabilities: what actually goes wrong

Updated September 23, 202613 min read
Short answer

Mostly ordinary ones, arriving in a particular pattern. The common classes are broken access control, tenant data that is not scoped at the database, secrets that reach the browser, unvalidated input on state-changing routes, unverified webhooks, outdated dependencies, missing rate limits, leaky errors and production configuration nobody set. What is distinctive is not the exotic flaw. It is that these are omissions, and an omission looks exactly like working code until a second person uses it.

Mapped toOWASP ASVS 5OWASP API Security Top 10CWENIST SSDF

There is a version of this subject that is mostly alarm, and it is not very useful. The interesting thing about vulnerabilities in AI-generated code is not that they are novel. They are almost entirely the classes that have been on the same lists for a decade. What is different is the shape they arrive in.

A model produces code that is plausible. Plausible and correct overlap heavily, and where they come apart is usually a missing condition rather than a wrong one. That matters, because a missing check does not look like a bug. It looks like working software, and it goes on looking like working software right up until somebody who is not you sends a request.

So this is a catalogue rather than a warning. For each class: what it is, what it looks like in a generated codebase, and what it takes to establish whether you have it.

1. Broken access control

The most common serious finding, and the most expensive. Authentication asks who you are. Authorisation asks whether you may do this. Generated applications get the first right far more reliably than the second, because signing in is a feature somebody specified and ownership is an assumption nobody wrote down.

In practice it looks like a route that reads an identifier out of the request and fetches or updates that record without checking who is asking. While there is one user the identifier is always theirs, so the code is correct in every test anybody ran.

Establishing it: reading the routes shows a check is absent, which is a real finding. Sending the request as a second account shows that it was refused, which is a stronger one. The control is object ownership enforced server-side, and it maps to the access control sections of ASVS and to the broken object level authorisation entry in the API Security Top 10.

2. Data that is not scoped at the database

The same failure one layer down, and worth separating because the fix is in a different place. Where the browser talks to the database directly, the rules on the table are the only thing between a key that is public by design and the rows behind it.

Two states get mistaken for safe. A table with no row level security at all, which is readable by anyone holding the publishable key. And a table with a policy that only asks whether the caller is signed in, which scopes nothing at all if anybody can sign up.

Establishing it: read the policies per table and per command, then test what the deployment actually returns when an account asks for rows it does not own.

3. Secrets that reach the browser

Generated integrations need credentials, and the shortest path to a working integration is sometimes to put the credential where the code that needs it is running. When that code is in the browser, the credential is published.

The framework usually says so out loud. A build-time prefix marking a variable as public is an instruction to ship its value to every visitor, so a privileged key carrying that prefix has already left. The other common case is a key in an early commit that was removed later, which does nothing for anyone who cloned the repository in between.

Establishing it: scanning the working tree and the history settles this one cleanly, which makes it among the few on this list a machine can close by itself. The remediation order matters more than the detection: rotate first, then remove.

4. Injection

Old, well understood, and still present, because the place it survives is the one query somebody assembled by hand when the query builder did not do what they wanted. Generated code uses parameterised access most of the time, which is exactly why the exception gets no attention.

Establishing it: tracing user input to the point it reaches a query finds the construction. A deterministic test against the running system confirms it.

5. Validation that only exists in the form

Client-side validation is a user experience feature. It tells somebody their email is malformed before they wait for a round trip. It has never been a security boundary, because the form is running on their machine and the request does not have to come from it.

Generated applications produce good client validation reliably, because it is visible, and server validation less reliably, because it is not. The result is an endpoint that accepts a quantity of minus one.

6. Webhooks that trust whatever posts to them

A webhook endpoint on the public internet will be found. The failure is an endpoint that acts on a request because it arrived at the right address rather than because it was signed, and it is expensive in proportion to what the webhook does. The one that marks an order as paid is the one to look at first.

Three separate things are worth confirming, and generated handlers commonly have the first and not the others: signature verification, replay handling, and looking the record up on the provider's side rather than trusting the amounts in the payload.

7. Dependencies nobody has looked at

AI tools add packages readily and remove them rarely. The risk is less that something obscure is vulnerable and more that nobody knows either way, which turns every other answer about the software into a provisional one.

Establishing it: this is the other class a machine settles cleanly, provided a lockfile records what actually deployed. Without one, the build resolved something and that something is not written down anywhere.

8. No rate limit on the endpoints that need one

Authentication endpoints are attacked by volume rather than by cleverness, and a limit on them is rarely part of a first working version. The same applies to anything expensive: a generated endpoint that calls a paid model on every request is a bill somebody else can run up.

9. Errors that explain the system

A stack trace naming the framework, the version and the failing query is a map. Debug output is on by default in development and turning it off is a production step nobody was told to take.

Establishing it: this one is visible from outside with no access at all, which is why it is usually the first thing a technical reviewer looks at.

10. Production configuration that was never configured

Transport security that is available but not enforced, missing response headers, permissive cross-origin settings that made development easier and survived into production. Individually minor, collectively the clearest signal available about how much attention the deployment received.

Which of these a scanner can settle

Being precise about this is more useful than claiming everything. Static analysis is pattern matching, and it is genuinely good at patterns: a hardcoded credential, a concatenated query, a known-vulnerable package version. Those are the classes where a scanner produces an answer you can act on directly.

What pattern matching cannot do is judge context, and most of the list above is context. Whether a record belongs to the caller is a fact about your data model. Whether a policy scopes anything depends on who can sign up. Whether an endpoint should be rate limited depends on what it costs to call. A scanner can tell you a check is missing. It cannot tell you the check was the wrong one.

  • Pattern matching settles: exposed secrets, known-vulnerable dependencies, injection-prone query construction, missing headers, transport configuration.
  • Reading the source establishes: that an authorisation check is absent, that validation is client-side only, that a webhook handler verifies nothing. Absent is a real finding and it is not the same as refused.
  • Testing the running system establishes: that a request from the wrong account was actually refused, that the rate limit engages, that a backup restores. This is the only category that proves a control works rather than that it exists.
  • A person is still required for: whether the data model is right, whether the architecture will hold, whether the team can maintain it, and every question that is about a contract rather than a codebase.

Why the list is short and keeps being the same list

Nothing here is specific to code a model wrote. These are the same classes that appear in code written by a tired person on a deadline, which is the actual comparison rather than some idealised reviewed codebase. What generation changes is throughput. More code arrives, less of it is read closely, and the defaults get accepted rather than chosen.

That is a reason to check a small number of things carefully and repeatedly, rather than to audit everything once and call it done.

The free check tests a deployed application for the classes that are visible from outside, and names the ones a URL cannot reach as uncovered rather than passing them.Run the free check

Frequently asked questions

Is AI-generated code less secure than code written by a person?

The honest answer is that the comparison is usually made against an idealised baseline nobody ships. The classes of flaw are the same either way. What changes with generation is volume and review: more code arrives, less of it is read closely, and defaults get accepted rather than chosen. That is a process difference rather than a property of the code.

What is the most common vulnerability in AI-generated code?

Broken access control, by some distance, in its two forms: a route that trusts an identifier from the request, and a database table whose rules do not scope rows to their owner. Both are invisible while there is a single user, which is why they survive testing.

Will a security scanner find these?

Some of them cleanly. Exposed secrets, known-vulnerable dependencies and injection-prone query construction are pattern problems and a scanner is the right instrument. Access control, tenant isolation and business logic are context problems, and a scanner can at best report that a check is missing rather than that the right check is missing.

Where should I start if I can only do one thing?

Create a second account and try to read the first account's data with it. It takes twenty minutes, it needs no tooling, and it tests the class that causes the most expensive findings.

Keep reading

How to check an AI-built app before you launch

Guide · 11 min

Is my Lovable app secure enough to launch?

Guide · 9 min

Why AI-generated code ships with security flaws, and what to do about it

Article · 6 min

The vibe coding security checklist: 12 checks before you ship

Article · 9 min