StackAttestTechnical Trust
Pre-launch

How to check an AI-built app before you launch

Updated September 23, 202611 min read
Short answer

Check it from outside first, then from the inside out. In order: what the deployment leaks to a stranger, whether your database rules scope data at all, whether one logged-in user can reach another's records, what your JavaScript bundle contains, what your dependencies carry, what your webhooks accept, and whether you can restore a backup. Most of it takes an afternoon. The order matters, because each step narrows what the next one has to consider.

Mapped toOWASP ASVS 5OWASP API Security Top 10CWEWSTG

A working demo proves one thing: that the application can complete its happy path while you are the only person using it. Everything that tends to go wrong afterwards lives outside that path. A second user arrives, or a stranger sends a request nobody designed for, and the assumptions that were never stated turn out to be the product.

This is not a property of AI-generated code specifically. It is a property of code that was written faster than it was reviewed, and AI tools are very good at producing that. The checks below are the ones that find the most in the least time, in the order that wastes the least of it.

Check in this order

Each step narrows the next. There is no point testing whether one user can read another user's records if the table is readable without logging in at all, and no point auditing dependencies if your service key is sitting in the browser. Start at the outside, where an attacker starts, and work inward.

1. What a stranger sees before logging in

The cheapest check, and the one that needs no access to anything. You are looking for what the deployment discloses to somebody who has done nothing but type the address.

How to check

  1. Load the site over plain http and confirm it redirects to https rather than answering.
  2. Look at the response headers. A missing content security policy is common and worth fixing; a missing strict transport security header after the redirect works is the next one.
  3. Break something on purpose. Request a page that does not exist, post malformed data to a form, and read what comes back. Stack traces, framework versions and database errors are all telling a stranger where to push.
  4. Check your cookies for the secure and httponly flags.

What this establishes: that the front door is configured. It says nothing at all about your data, which is the next three checks.

2. Whether the database scopes data at all

If the application uses Supabase, Firebase or anything else where the browser talks to the database with a public key, this is where the serious findings usually are. A table without row level security enabled is readable by anyone holding that key, and the key is in your bundle by design.

The subtle version is worse than the obvious one. A table with row level security enabled and a policy that only checks that the requester is authenticated scopes nothing: every logged-in user passes it, which in a product with open signup means everyone.

How to check

  1. List your tables and note which have row level security switched on. Anything without it is public.
  2. For each policy, ask what it compares. A policy that checks only for the presence of a session is not scoping to an owner.
  3. Test from outside your session: a private browser window, the public key, and a request for a table you believe is protected. Rows that are not yours coming back is the answer.
  4. Do the same for storage buckets. They have their own rules and are forgotten more often than tables are.

What this establishes: that the data layer refuses an anonymous request. It does not establish that it refuses the wrong logged-in user, which is a different test.

3. Whether one user can reach another user's records

This is the check that needs two accounts, and it is the one most self-assessments skip because of that. It is also where the findings that matter commercially tend to be, because an application that leaks between tenants is an application that cannot be sold to a second customer.

The pattern to look for is an identifier taken from the request and used without checking that the caller owns the thing it names. Generated code does this constantly, because at the time it was written there was one user and the identifier was always theirs.

How to check

  1. Create two accounts in different organisations, with different data.
  2. As the first account, find a request that names a record by its identifier. The network tab in developer tools is the fastest way to see one.
  3. Replay that request as the second account, with the first account's identifier. A record coming back is a finding. So is a successful update.
  4. Repeat for anything that changes state rather than reads it, and for any endpoint an administrator uses.

What this establishes: that a specific request was refused for a specific user. That is a narrow claim and a strong one, which is the opposite of most of what a checklist produces.

4. What your bundle is handing out

Everything your browser code can read, every visitor can read. The question is whether anything privileged ended up there, usually because an integration needed to work and the quickest way to make it work was to move the key.

How to check

  1. Open the deployed site, open developer tools, and search the loaded sources for service_role, secret, and the names of your providers.
  2. Read your environment variable names. A prefix such as NEXT_PUBLIC or VITE means the value is deliberately shipped to the browser, so a secret with that prefix is not a secret.
  3. Search the repository history, not just the working tree. A key deleted in a later commit is still in the clone somebody made.

If anything privileged was ever exposed, rotate it before you fix the code. Removing a key from the bundle does not retrieve it from everybody who already downloaded the bundle.

5. What your dependencies carry

AI tools add packages quickly and remove them rarely. The risk is less that something exotic is vulnerable and more that a common package is three major versions behind and one of those versions fixed something.

How to check

  1. Run your package manager's audit command and read it rather than counting it. A high-severity advisory in a build-time tool is a different problem from one in something that handles requests.
  2. Note anything that has not been published in over a year, which is a maintenance risk even with no advisory against it.
  3. Check what your lockfile actually resolved to, since that is what deployed, not what your manifest asked for.

6. What your webhooks accept

A webhook endpoint reachable on the public internet will be found. The failure is an endpoint that trusts a request because it arrived, rather than because it was signed, and it is expensive when the webhook in question is the one that marks an order as paid.

How to check

  1. Post a plausible payload to the endpoint with no signature and see what happens. Anything other than a rejection is the finding.
  2. Send the same signed request twice and check whether the effect happened twice.
  3. Confirm the handler looks the record up on the provider's side rather than trusting the amounts and identifiers in the payload.

7. Whether you can get back

The last check is the one that is never urgent until it is the only thing that matters. Backups that have never been restored are a belief rather than a control, and the same is true of a rollback path nobody has walked.

How to check

  1. Restore a backup into a scratch environment and look at the data. This is the whole test.
  2. Deploy a deliberately broken build to a staging environment and roll it back, timing how long it takes.
  3. Confirm you would find out about an outage from monitoring rather than from a customer.

What a check you ran yourself is worth

All seven checks are worth doing and most people should do them before they do anything else. It is still worth being precise about what you have at the end, because the answers are not equally strong and treating them as though they were is how a confident team gets surprised.

  • An answer you gave on a checklist is a belief about the system. It is useful for finding your own gaps and it is not evidence of anything.
  • Reading the code and finding no check is evidence that a check is absent. That is a real finding, and it is not the same as watching a request be refused.
  • Watching the running system refuse a request you actually sent is the strongest thing on this list, and it is the only one of these that proves the control works rather than that it exists.
  • None of it is independent. You checked your own work, which is the part a customer's security reviewer or an investor's technical diligence will discount, and reasonably so.

That last point is not an argument for buying something. It is the reason to write down what you tested and how, while you still remember: the same seven answers are worth considerably more to somebody else when they arrive with the method attached.

The free check runs the first of these seven against your deployed application and names the controls a URL cannot reach as uncovered rather than passing them.Run the free check

Frequently asked questions

How long does all of this take?

An afternoon if nothing is wrong. The two that expand are row level security, because writing correct policies per table and per command is real work, and the two-account testing, because setting up the second account properly usually surfaces the first finding.

I am not technical. Which of these can I still do?

The first, the fourth and the fifth: what a stranger sees, what is in the bundle, and what the dependencies carry. All three are a browser and a terminal command. The access control testing needs somebody who can read the code, and that is the point where the checking genuinely gets harder rather than just longer.

My AI builder has its own security scan. Is that enough?

It is worth running and it is not the same test. A scan inside the tool reads what the tool generated. It does not send a request to your deployment as a second user, and it does not produce anything you can hand to somebody who has no reason to take your word for it.

Does this replace a penetration test?

No. A penetration test is a person trying to defeat your application with intent and imagination, and it will find things a sequence of checks does not. This is the work that makes a penetration test worth paying for, by clearing out the findings that do not need one.

Keep reading

Is my Lovable app secure enough to launch?

Guide · 9 min

The startup technical audit: what it covers and when you need one

Guide · 10 min

Is your AI-built app production ready? A practical way to check

Article · 7 min

Why AI-generated code ships with security flaws, and what to do about it

Article · 6 min