StackAttestTechnical Trust
Claude Code security audit

Independent validation of what Claude Code built

The coding agent writes the software. StackAttest independently validates the software that was written. That is not a doubt about the tool. It is the ordinary reason a second reader exists: any review that shares an author's assumptions inherits the author's blind spots along with the strengths, and an agent reviewing its own output is the same reader twice.

Validate a repositoryScan a live appSee an example Passport

What gets read, run and probed

Repository validation leads here because the subject is a whole repository rather than one file. The other two layers establish what that repository does once it is running.

Repository validation

E2
Repository

Reads the source: where secrets live, whether authorisation is enforced on the server or only in the interface, whether data access is injection safe, and which dependencies carry known critical vulnerabilities. Findings cite the file they came from, so a result is checkable rather than asserted.

Runtime validation

E4
Runtime

Builds and runs the application in an isolated environment and probes the running system. This is the layer that answers questions reading code cannot: whether the deployed configuration actually refuses an unauthorised request, and whether error responses leak internals under real conditions.

Public URL scan

E4
URL scan

Tests the deployed application from outside, the way any visitor reaches it: transport security, security response headers, and what the running service discloses about itself. It needs nothing but the address, and what it observes it observes against the running system. Its limit is reach rather than strength: a URL can never establish authorisation, data-layer or backup posture, and those controls stay uncovered.

What drifts when an agent works across a repository

An agent works across the whole codebase over many sessions, so the findings worth having are rarely single mistakes. They are inconsistencies between sessions, each of them reasonable on its own. These are the ones that recur, and the audit establishes which are present in your repository. These are the ones that recur, not a claim about your application, and the audit establishes which of them are actually present.

Authorisation applied on some routes and not others

A session that adds a feature usually copies the pattern in front of it, and the pattern in front of it is whatever the previous session left. That is why the ownership check is present on the routes written early and missing on two written a month later. Nothing regressed, the convention simply never reached them, and a green test run does not notice.

One problem solved three different ways

Three sessions met the same question about validating input, or shaping an error, or deciding who may read a record, and each answered it well. The repository now holds three answers. Every one of them has to be right, changing the rule means finding all three, and the weakest of them is the one that gets exercised.

Tests that assert the feature works, and nothing else

Generated suites are good at proving the happy path. They rarely assert that a stranger cannot reach the record, because nobody asked for that behaviour in the prompt that produced the feature. The absence is invisible: the suite is green either way, and a passing test is exactly what a founder reads as reassurance.

Configuration and secrets moved for convenience

During a long refactor a value gets moved to where the code that needs it can see it, because that is the shortest path to the build passing. The move is fine in the session it happens in and is rarely revisited. A credential that ends up somewhere a client can read it is published from that moment, and taking it out later does not unpublish it.

Dependencies added to solve one problem, then left alone

A package arrives because it answered a question in one session, and no session afterwards has a reason to look at it again. Known vulnerabilities in those packages are public by definition, which makes them the cheapest class of finding to fix and the hardest to defend leaving in place.

The author and the reviewer being the same reader

Asking the agent to review its own work is genuinely useful and it is not the same as an independent check. It sees the files in its context rather than the repository, it cannot run the deployed system, and it carries forward the assumptions it made while writing. What it produces is a conversation. What a buyer, an investor or an enterprise customer asks for is evidence.

What the result looks like

An illustrative extract, not a real customer result. Every row carries the evidence grade behind it, and a control nobody could establish says so.

ControlStatusEvidenceNote
Object ownership enforced server-sidePartialE2Enforced on most routes, absent on two added later
Injection-safe data accessVerifiedE2Parameterised in every query the scan read
Authentication endpoints rate limitedFailedE4Repeated attempts accepted without a ceiling
Dependencies free of known-critical vulnerabilitiesPartialE3Two advisories, both with a fixed version published
Deployment rollback path existsUncoveredNot gradedNot determinable from the layers that ran

Mapped to published standards

Each control is tested against a named clause, so a result means something outside our own vocabulary. Mapping is not certification, and none of these bodies endorse StackAttest.

OWASP ASVS 5OWASP API Security Top 10CWENIST SSDFSLSA

What this audit is not

  • StackAttest is not affiliated with Anthropic and does not integrate with Claude Code. The agent writes the software; this validates the software that was written, which is a different job done by a different party.
  • It is not a penetration test. A deterministic check and a runtime probe are not an adversarial human engagement, and there are classes of weakness only a person hunting for them will find.
  • It is not SOC 2 and does not replace it. SOC 2 attests to organisational controls over time; this validates technical controls in the software. They answer different questions for different buyers.
  • It does not replace human due diligence. It gives a reviewer repeatable evidence to start from, which is a different thing from being the reviewer.
  • Controls the run could not reach are reported as uncovered rather than passed. A result you did not earn is worse than no result.
  • It is not a judgement about Claude Code. The result describes one repository at one commit. It is a claim about that codebase, not about the tool that helped write it, and a good tool used carefully still produces software somebody has to verify.

Questions

What is a Claude Code security audit?

An independent check of a repository that an AI coding agent largely wrote. It tests named controls against published standards, cites the file a source finding came from, and records how strong the evidence behind each result is, so the answer holds up with someone who was not in the sessions that produced the code.

Is code written by an agent less secure?

We would not claim that, and the audit is not built on it. What changed is throughput: far more software now reaches production without anyone having deliberated over the decisions inside it. That makes independent validation worth more than it used to be, whoever or whatever did the typing.

Why not just ask the agent to review the code?

Do that as well, it catches real things. It is not independent. The same model that made the assumptions is checking the assumptions, it sees what is in its context rather than the whole repository, and it cannot probe the running system. It ends in a transcript, and an audit ends in graded evidence you can hand to someone else.

Do you need access to my repository?

For source-level evidence, yes, through a read-only connection. Without it you can still run a URL scan and runtime validation, which establish facts about the deployed system rather than the code. How source is handled is described on the security page rather than summarised here.

What does the audit look at that a code review would miss?

Behaviour. Reading code tells you what should happen when an unauthorised request arrives; runtime validation builds and runs the application and finds out what does happen. Where a run cannot reach a control, it is reported as uncovered rather than assumed to pass.

What do I get at the end?

A result for each control with its evidence grade, an explicit list of the controls the run could not cover, remediation guidance on what failed, and a Passport you can publish if you want to. Publishing is your decision, and nothing becomes public because a scan ran.

Related

Vibe code auditAI app security auditAI code auditv0 security auditA security review prompt, and its limitsSecurity and data handling

Find out what is actually true of your application

Start with the layer you can run today. The report says which layers ran and what they could not reach, so the result is honest about its own limits.

Validate a repositoryScan a live app