StackAttestTechnical Trust
Cursor security audit

A Cursor security audit of what your sessions actually shipped

The coding agent writes the software. This is the pass that validates the software that was written, run against controls with published definitions rather than against the assumptions that produced the code in the first place.

Validate the repositoryScan the deployed appSee an example Passport

What the run looks at, whatever the stack is

Cursor is used by working developers across arbitrary stacks, so nothing here assumes a framework, a language or a backend. The controls are defined by behaviour rather than by file layout, and the report says which layers ran.

Repository validation

E2
Repository

Reads the source: where secrets live, whether authorisation is enforced on the server or only in the interface, whether data access is injection safe, and which dependencies carry known critical vulnerabilities. Findings cite the file they came from, so a result is checkable rather than asserted.

Public URL scan

E4
URL scan

Tests the deployed application from outside, the way any visitor reaches it: transport security, security response headers, and what the running service discloses about itself. It needs nothing but the address, and what it observes it observes against the running system. Its limit is reach rather than strength: a URL can never establish authorisation, data-layer or backup posture, and those controls stay uncovered.

Runtime validation

E4
Runtime

Builds and runs the application in an isolated environment and probes the running system. This is the layer that answers questions reading code cannot: whether the deployed configuration actually refuses an unauthorised request, and whether error responses leak internals under real conditions.

What generated code tends to leave unresolved

None of this is a criticism of the tool, which is very good at what it does, and none of it is a claim about your repository. These are the places independent validation most often finds something, so the run goes looking for them specifically and reports which of them are actually present in your application. These are patterns that recur in AI-assisted work, not a claim about your application, and the audit establishes which of them are actually present.

Authorisation that was implemented where it was visible

The component checks the role before rendering the button and the page redirects a user who should not be there, so the feature demonstrably works. The handler behind it frequently does not repeat the check. Requests do not arrive through your interface, and the interface is the one part of the system the caller can edit.

Route protection that covers most of the paths

A guard is written once, over the routes that existed when it was written. Routes added in later sessions land outside the matcher. Nothing about the protected code looks wrong, because the problem is not in the protected code: it is in the list of what the guard never mentions, which is not somewhere anyone thinks to look.

Secrets that moved into client-visible configuration

A value works locally, the browser build complains that it cannot see it, and the quickest way through is a public environment prefix. That prefix is an instruction to inline the value into the bundle. Not every key in a bundle is a problem and treating them all as one is how real exposures get buried, so what matters is whether the credential can bypass your access rules.

Tests that assert the feature and not the boundary

Generated suites reliably prove the happy path: the record is created, the response is correct. They rarely assert that a signed-in user cannot reach a record belonging to someone else. A green run is what a reviewer uses to decide a change is safe to merge, so a suite that never wrote that assertion hands out confidence it did not earn.

Dependencies added inside a task, evaluated by nobody

A package gets pulled in to solve a subproblem in the middle of a session and arrives with a transitive tree behind it. Known vulnerabilities in published packages are public by definition, which makes this the cheapest class of finding to fix and the hardest to justify shipping.

Architectural drift across sessions

Several sessions each solved a problem their own way: two ways of reading the session, three ways of talking to the database, error handling that differs by route. Each is defensible on its own. Together they mean there is no single place where a control is enforced, which is what makes review expensive and makes a gap easy to keep.

What the result looks like

An illustrative extract, not a real customer result. Every row carries the evidence grade behind it, and a control nobody could establish says so.

ControlStatusEvidenceNote
Object ownership enforced server-sidePartialE2Enforced on three handlers, absent on the two added most recently
Authentication endpoints are rate limitedFailedE4Repeated attempts accepted without a ceiling
Dependencies free of known-critical vulnerabilitiesFailedE3Two advisories with fixes already published
Error responses do not leak internalsVerifiedE4Generic bodies returned under induced failures
Verified backups and restore procedureUncoveredNot gradedNo attestation and nothing in the layers that ran establishes it

Mapped to published standards

Each control is tested against a named clause, so a result means something outside our own vocabulary. Mapping is not certification, and none of these bodies endorse StackAttest.

OWASP ASVS 5OWASP API Security Top 10CWENIST SSDFSLSA

What this audit is not

  • It is not a penetration test. A deterministic check and a runtime probe are not an adversarial human engagement, and there are classes of weakness only a person hunting for them will find.
  • It is not SOC 2 and does not replace it. SOC 2 attests to organisational controls over time; this validates technical controls in the software. They answer different questions for different buyers.
  • It does not replace human due diligence. It gives a reviewer repeatable evidence to start from, which is a different thing from being the reviewer.
  • Controls the run could not reach are reported as uncovered rather than passed. A result you did not earn is worse than no result.
  • It covers thirteen defined controls, which is a floor rather than a complete security review. It will not tell you whether the architecture was the right one, only whether the controls it tests hold in the code you have.
  • It is not affiliated with Cursor and does not connect to your editor. There is nothing to install and no session history to hand over. What it examines is the committed code and the running application, which is what your users get regardless of how it was written.

Questions

What is a Cursor security audit?

An independent examination of software written with an AI coding assistant, testing named controls against published standards instead of reviewing the code for style. Every result carries the strength of the evidence behind it, and controls the run could not establish are listed as uncovered.

Does this mean Cursor writes insecure code?

No, and we would not claim it. What changed is the ratio: code is now produced far faster than anyone reviews it, so more software reaches production without a person having deliberated over the assumptions inside it. The agent writes the software. This validates the software that was written.

Cursor can review my code. Why is this different?

A model reviewing a project from inside that project inherits the project's assumptions, sees the files that fit in its context rather than all of them, and cannot exercise the deployed system. It also produces a transcript. This produces a result per control with a grade attached, which is the part you can give to a customer or an investor.

Do you plug into Cursor?

No. There is no extension, no integration and nothing to authorise in the editor. We look at the repository you connect read-only and the application you deployed, which is the same thing a stranger interacts with.

My stack is unusual. Does that matter?

The controls are defined by behaviour rather than by framework: whether ownership is enforced on the server, whether secrets sit outside source code, whether errors disclose internals, whether dependencies carry known critical advisories. Where a check genuinely cannot be established for your stack, the control is reported as uncovered rather than guessed at.

Several of us used Cursor on the same repository. Is that a problem?

It is the ordinary case, and it is the reason a partial result is a real outcome here rather than a rounding error. Independent sessions solve the same problem independently, so a control can be enforced properly in one place and missing in another. The run tests the control everywhere it applies and reports partial when that is the honest answer.

Related

AI code auditLovable security auditVibe code auditWhy AI-generated code ships with security flawsA security review prompt, and its limits

Find out what is actually true of your application

Start with the layer you can run today. The report says which layers ran and what they could not reach, so the result is honest about its own limits.

Validate the repositoryScan the deployed app