StackAttestTechnical Trust
AI code audit

An independent audit of the code your AI tools wrote

AI writes correct-looking code very quickly. What it does not do is decide which assumptions are safe to ship, and nobody reviewing their own output catches the decisions they never noticed making. This is the independent pass.

Validate a repositoryScan a live appSee an example Passport

How the code is examined

Source analysis is the centre of this one, with the other two layers establishing what the code does once it is running.

Repository validation

E2
Repository

Reads the source: where secrets live, whether authorisation is enforced on the server or only in the interface, whether data access is injection safe, and which dependencies carry known critical vulnerabilities. Findings cite the file they came from, so a result is checkable rather than asserted.

Runtime validation

E4
Runtime

Builds and runs the application in an isolated environment and probes the running system. This is the layer that answers questions reading code cannot: whether the deployed configuration actually refuses an unauthorised request, and whether error responses leak internals under real conditions.

Public URL scan

E4
URL scan

Tests the deployed application from outside, the way any visitor reaches it: transport security, security response headers, and what the running service discloses about itself. It needs nothing but the address, and what it observes it observes against the running system. Its limit is reach rather than strength: a URL can never establish authorisation, data-layer or backup posture, and those controls stay uncovered.

What tends to be true of generated code

Generated code is not worse code. It is code whose assumptions were never argued about, and these are the assumptions that most often turn out to matter.

Architectural decisions nobody made on purpose

A generator picks a pattern because it fits the prompt. Whether that pattern suits a multi-tenant product with real customers is a question that was never asked, and it surfaces later as a boundary in the wrong place.

Insecure defaults that were never revisited

Permissive settings that make development frictionless are frequently the same settings that make production dangerous, and there is no moment in a generated workflow where someone is prompted to reconsider them.

Tests that cover behaviour, not boundaries

Generated test suites tend to assert that the happy path works. They rarely assert that a user cannot reach another user's record, which is the assertion that matters most and the one whose absence is invisible in a green test run.

Frontend and backend responsibilities blurred

Validation and permission logic drift into client code because that is where the generator was working. Anything enforced only there is advisory, since the client is under the caller's control.

Integrations wired for the demo

Payment and webhook handlers that work end to end in testing but skip signature verification, replay protection or idempotency, because none of those are needed to make the demo pass.

What the result looks like

An illustrative extract, not a real customer result. Every row carries the evidence grade behind it, and a control nobody could establish says so.

ControlStatusEvidenceNote
Server-side input validation on state-changing endpointsPartialE2Validation present on writes, absent on two updates
Injection-safe data accessVerifiedE2Parameterised throughout
Authentication endpoints rate limitedFailedE4No ceiling observed against the running service
Error responses do not leak internalsVerifiedE4Stack traces suppressed under induced failures
Deployment rollback path existsUncoveredNot gradedNot determinable from the layers that ran

Mapped to published standards

Each control is tested against a named clause, so a result means something outside our own vocabulary. Mapping is not certification, and none of these bodies endorse StackAttest.

OWASP ASVS 5CWENIST SSDFOWASP API Security Top 10SLSA

What this audit is not

  • It is not a penetration test. A deterministic check and a runtime probe are not an adversarial human engagement, and there are classes of weakness only a person hunting for them will find.
  • It is not SOC 2 and does not replace it. SOC 2 attests to organisational controls over time; this validates technical controls in the software. They answer different questions for different buyers.
  • It does not replace human due diligence. It gives a reviewer repeatable evidence to start from, which is a different thing from being the reviewer.
  • Controls the run could not reach are reported as uncovered rather than passed. A result you did not earn is worse than no result.

Questions

What is an AI code audit?

An independent examination of software largely written by AI tools, testing named controls against published standards rather than reading the code for style. Each result carries the strength of the evidence behind it, and controls that could not be reached are listed as uncovered.

Is AI-generated code less secure than hand-written code?

Not inherently, and we would not claim it. What is different is volume and review: far more software now reaches production without anyone having deliberated over its assumptions, so independent validation carries more weight than it used to.

Which standards do you map to?

Controls carry mappings to OWASP ASVS and ASVS 5, the OWASP API Security Top 10, the OWASP Testing Guide, CWE, NIST SSDF and SLSA. Mapping means a control is tested against a published definition. It is not certification, and none of those bodies endorse StackAttest.

Do you need the source code?

For source-level evidence, yes, through a read-only connection. Without it you can still run a URL scan and runtime validation, which produce evidence about the deployed system rather than the code, and the report is explicit about which was which.

Can you fix what you find?

Findings come with remediation guidance, and there is a reviewed path for applying proposed fixes. Nothing is changed in your codebase without review.

Related

Vibe code auditAI app security auditAI-generated code vulnerabilitiesThe control catalogueWhy AI-generated code ships with security flawsThe StackAttest Standard

Find out what is actually true of your application

Start with the layer you can run today. The report says which layers ran and what they could not reach, so the result is honest about its own limits.

Validate a repositoryScan a live app