Security audit for an AI-built application, with evidence attached
Most security answers a founder can give are assertions. This produces the other kind: named controls, tested at a stated strength, with the ones nobody could establish listed as uncovered rather than quietly passed.
The seven things a security audit has to separate
Buyers routinely receive one number covering all of these. They are different questions with different evidence behind them, and the report keeps them apart.
Public URL scan
E4Tests the deployed application from outside, the way any visitor reaches it: transport security, security response headers, and what the running service discloses about itself. It needs nothing but the address, and what it observes it observes against the running system. Its limit is reach rather than strength: a URL can never establish authorisation, data-layer or backup posture, and those controls stay uncovered.
Repository validation
E2Reads the source: where secrets live, whether authorisation is enforced on the server or only in the interface, whether data access is injection safe, and which dependencies carry known critical vulnerabilities. Findings cite the file they came from, so a result is checkable rather than asserted.
Runtime validation
E4Builds and runs the application in an isolated environment and probes the running system. This is the layer that answers questions reading code cannot: whether the deployed configuration actually refuses an unauthorised request, and whether error responses leak internals under real conditions.
Where AI-built SaaS most often fails a security review
Ordered roughly by how often the finding turns out to be real and how much it costs when it is.
Cross-tenant access
The single most consequential failure in multi-tenant software. If one customer can reach another customer's data by changing an identifier, nothing else about the security posture matters very much. It is also invisible to a scanner that only looks from outside.
Access control that was never tested as access control
Authentication and authorisation get conflated. Being signed in is not permission to act on a given record, and a system that only checks the first has no answer to the second.
Exposed privileged credentials
Not every key in a bundle is a problem, and treating them all as one is how real exposures get lost in noise. What matters is whether a credential that bypasses your access rules is reachable by a visitor.
Unverified webhook handling
Endpoints that accept and act on an unsigned payload, or that replay cleanly. Payment webhooks are the common case and the expensive one.
Missing ceilings on sensitive endpoints
Authentication and other costly endpoints without rate limits, which turns a slow guessing attack into a fast one.
Operational gaps that only appear under stress
Backups nobody has restored, no rollback path, and error handling that discloses internals when something genuinely breaks. Each is fine until the day it is not.
What the result looks like
An illustrative extract, not a real customer result. Every row carries the evidence grade behind it, and a control nobody could establish says so.
| Control | Status | Evidence | Note |
|---|---|---|---|
| Object ownership enforced server-side | Verified | E4 | Cross-account request refused by the running system |
| Security response headers present | Partial | E4 | Two of the expected headers absent on the live deployment |
| Authentication endpoints rate limited | Verified | E4 | Ceiling observed under repeated attempts |
| Secrets managed outside source code | Failed | E2 | A privileged key is reachable from client code |
| Health checks and basic observability | Uncovered | Not graded | No evidence available from the layers that ran |
Mapped to published standards
Each control is tested against a named clause, so a result means something outside our own vocabulary. Mapping is not certification, and none of these bodies endorse StackAttest.
What this audit is not
- It is not a penetration test. A deterministic check and a runtime probe are not an adversarial human engagement, and there are classes of weakness only a person hunting for them will find.
- It is not SOC 2 and does not replace it. SOC 2 attests to organisational controls over time; this validates technical controls in the software. They answer different questions for different buyers.
- It does not replace human due diligence. It gives a reviewer repeatable evidence to start from, which is a different thing from being the reviewer.
- Controls the run could not reach are reported as uncovered rather than passed. A result you did not earn is worse than no result.
Questions
What does an AI app security audit cover?
Public attack surface, source-level controls, dependencies, access control and the behaviour of the running system, tested against thirteen controls mapped to published standards. Each result records the strength of the evidence behind it.
Can you test whether one customer can reach another customer's data?
That is what runtime validation exists for. Reading code tells you what should happen; probing the running system establishes what does. Where a run cannot reach that behaviour, the control is reported as uncovered rather than assumed.
Is this a penetration test?
No, and it should not be sold as one. A penetration test is an adversarial human engagement with a scope and a report to match. This is repeatable automated validation of defined controls, which is a different instrument for a different question.
What is an evidence grade?
A record of how a result was established, from inferred at the weakest through source verified, deterministically tested, and runtime verified. Two products can both pass a control with very different amounts of evidence behind them, and hiding that difference would be the dishonest part.
Can I share the result with a customer or investor?
Yes, as a Passport, and only if you choose to publish. The founder controls what is public, and nothing is published because a scan ran.
Related
Find out what is actually true of your application
Start with the layer you can run today. The report says which layers ran and what they could not reach, so the result is honest about its own limits.