StackAttestTechnical Trust
API rate limiting

How to rate limit an API, and which endpoints to do first

A rate limit is not a switch you turn on. It is three decisions: who gets counted, what they are counted against, and what happens when they go over. Each one can be wrong on its own. Get the first wrong and you lock out an office that shares an address. Get the second wrong and you protect your servers beautifully while somebody works through one account's password at a comfortable pace.

Run a runtime validationRun the free checkSee an example Passport

What a run can establish about a ceiling

One asymmetry decides this whole page. A request that comes back refused proves a ceiling exists. A request that comes back normally proves that this request was not refused, and nothing more. The report keeps those two apart instead of averaging them into a number.

Runtime validation

E4
Runtime

Builds and runs the application, sends a bounded burst of sign-in requests at the authentication path it found, and records what came back, including how many requests it sent. A 429 or a Retry-After settles the control wherever the ceiling is enforced, which is the thing reading code cannot do. A burst answered normally is recorded as a burst that stayed under the threshold, and the control is not passed on it.

Repository validation

E2
Repository

Reads the source for the machinery a limit needs: whether a rate limiting library is present, and where it is applied. It never fails this control on its own, because throttling is very often enforced at an edge or an API gateway that no reader of your repository can see. A source-only result here is a gap to go and confirm rather than a verdict.

Public URL scan

E4
URL scan

Sends a short burst of ordinary requests to the address and watches for the service to start refusing them. It is bounded to about a dozen, because it is a probe and not a load test, and on the free check it is bounded far harder than that. What it observes it observes against the running deployment, so the grade records that; what it fails to observe is not a finding.

The decisions a limit is actually made of

None of these is about which library you pick. They are the questions the library makes you answer, and its defaults answer most of them in whatever way suits a demonstration.

Only two kinds of endpoint earn a ceiling of their own

Limiting everything to the same number is the version of this that gets torn out again, because a threshold tight enough to protect your sign-in page makes your own product feel broken. Two kinds of endpoint deserve their own limit. The expensive ones, where a single call costs you money or a table scan: a model call you are billed for, a report, an export, an email or a message you pay to send. And the guessable ones, where somebody wins by repetition rather than cleverness: sign-in, sign-up, password reset, one-time codes, invitation and share tokens, discount codes. Everything else can sit behind one generous ceiling whose job is stopping accidents.

Counting the right caller is the hard part

A limit is a counter, and a counter needs a key. The network address is the obvious key and it is wrong in both directions at once. An office, a school, a mobile network or a corporate VPN is one address shared by hundreds of people, so a threshold low enough to matter starts refusing people who did nothing, and your support inbox hears about it before you do. Meanwhile anybody working through your login page has more addresses than your threshold and pays almost nothing for them, so the limit that punished the office never touched the attempt. Count against the account or the identifier being tried where you know it, against the address where you do not, and apply both rather than picking one.

Protecting the service and protecting an account are two controls

They have different keys, different thresholds and different purposes, and having the first is routinely mistaken for having both. A global or per-address ceiling protects your infrastructure: it stops one caller consuming what everybody else needs. It says nothing at all about a slow, patient run against a single account, which never approaches a service-level threshold. What makes guessing uneconomic is a per-account counter with a delay that grows on each failure. Work out which of the two you have, then go and write the other one.

What you tell a refused caller decides what they do next

Refuse with a 429 and put a Retry-After on it. That is not manners, it is control: a client told to wait a minute waits a minute, and a client told nothing retries immediately and then retries harder, so your limit has generated the load it existed to prevent. The same header is what stops your own mobile client or integration turning a small ceiling into an outage. Two things not to do. Do not let the refusal message confirm which accounts exist, and do not make producing the refusal more expensive than serving the request would have been.

A counter that lives in one process is not a limit

An in-memory counter is per instance, so a service running four copies has quietly multiplied every threshold you wrote by four, and a platform that starts an instance on demand has removed them. The count has to live somewhere every instance can see, or in front of them all. This is the same reason the question of where a limit is enforced matters so much: a ceiling at your edge covers everything behind it and is invisible to anyone reading your repository, and a ceiling in your application is visible in the code and absent for any request that never arrives there.

A quiet probe is not evidence that no limit exists

A burst that gets refused proves a ceiling exists. A burst that gets answered proves the burst was smaller than the ceiling. That cuts in both directions, and it is why a quiet result is never written up here as a failure, and why your own testing has to be honest about which threshold it actually crossed. The only way to know a limit works is to cross it deliberately against your own system and watch the refusal arrive.

How to add one, in the order that keeps it useful

The library is the last decision, not the first. These five are the ones that decide whether what you ship protects anything.

  1. Write down which endpoints are expensive and which are guessable

    Two columns. Expensive is anywhere one request costs you money or a table scan. Guessable is anywhere somebody wins by repeating themselves: credentials, reset links, one-time codes, tokens, discount codes. Those get their own ceilings. Everything else gets a single generous one so a runaway client cannot take the service down, and you have stopped arguing about the rest.

  2. Decide what each counter is keyed on

    On sign-in, count per account attempted and per address separately, and refuse when either is over. On an expensive endpoint, count per authenticated account, because that is the party you can bill, throttle or suspend. Never key a counter on something the caller can change for free, a self-issued key, a header, an identifier the client generates, because that is a limit with an opt-out built into it.

  3. Pick two windows, not one

    A short window stops a burst and a long one stops a grind, and no single threshold does both without being wrong for one of them. On credentials, add a delay that grows with consecutive failures. It is invisible to somebody who mistyped their password twice, it makes an automated run uneconomic, and unlike a hard lockout it does not let a stranger disable a real person's account by failing at it.

  4. Refuse with a 429 and a Retry-After, and log it

    Say no in the way every client library already understands, so callers back off instead of hammering. Keep the body free of anything that says whether the account exists. Then log each refusal with the key it counted against, because a graph of those is the first honest signal that somebody is working through your login page, and without it you find out from a customer.

  5. Cross the limit on purpose, then do it again after every deploy

    Send more than the threshold at your own system and confirm the refusal arrives, from more than one instance if you run more than one. A limit nobody has ever triggered is a configuration file rather than a control. This is also the step that catches the day the counter moved back into process memory and nothing else changed.

What the result looks like

An illustrative extract, not a real customer result. Every row carries the evidence grade behind it, and a control nobody could establish says so.

ControlStatusEvidenceNote
Authentication endpoints are rate limitedPartialE2A rate limiting library is present in the source, and the burst that ran was answered normally
Transport security (TLS) is enforcedVerifiedE4Valid certificate, and plain HTTP redirected to HTTPS
Error responses do not leak internalsVerifiedE4A refused request returned no framework detail
Health checks and basic observabilityPartialE2A health endpoint answers, and refusals are not recorded anywhere a person would see them
Deployment rollback path existsUncoveredNot gradedNeither a burst of requests nor a source read can settle it

Mapped to published standards

Each control is tested against a named clause, so a result means something outside our own vocabulary. Mapping is not certification, and none of these bodies endorse StackAttest.

OWASP ASVS 5OWASP API Security Top 10CWENIST SSDF

What this does not tell you

  • A probe can confirm a ceiling and can never disprove one. A burst that is refused settles the control at runtime. A burst that is answered may simply have stayed under the threshold, so a quiet result is reported as uncovered or as a gap to confirm, and never as a failure.
  • The free check cannot establish this at all. It sends only a couple of requests to any one address, on purpose, so that a tool anyone can point at any site cannot be used to generate traffic somebody did not ask for. There the control comes back uncovered, with that reason written next to it.
  • It is not a load test. Nothing described here measures what your service does under sustained pressure, and none of these bursts is large enough to try. Capacity is a different question answered with different tools.
  • It is not a penetration test. A deterministic check and a runtime probe are not an adversarial human engagement, and there are classes of weakness only a person hunting for them will find.
  • It is not SOC 2 and does not replace it. SOC 2 attests to organisational controls over time; this validates technical controls in the software. They answer different questions for different buyers.
  • It does not replace human due diligence. It gives a reviewer repeatable evidence to start from, which is a different thing from being the reviewer.
  • Controls the run could not reach are reported as uncovered rather than passed. A result you did not earn is worse than no result.

Questions

What are the best practices for API rate limiting?

Protect the expensive and the guessable endpoints rather than all of them equally. Key the counter on the account as well as the address. Use a short window for bursts and a longer one for a grind. Refuse with a 429 and a Retry-After so clients back off. Keep the counter somewhere every instance can see. Then cross your own threshold deliberately and watch the refusal arrive.

How do I rate limit authentication endpoints specifically?

Count against the account being attempted as well as the caller's address, and refuse when either is over. Add a delay that grows with consecutive failures, which costs somebody who mistyped their password almost nothing and makes an automated run uneconomic. Apply the same rule to password reset, one-time codes and invitation links, because each of those is a guess repeated.

Is limiting by IP address enough?

No, and it fails in both directions at once. One address can be a whole office or a mobile network, so a threshold low enough to matter refuses people who did nothing. And anybody serious rotates through addresses cheaply and stays under a per-address ceiling indefinitely. Use it as one signal next to a per-account counter rather than as the control itself.

Should the limit live at the edge or in the application?

Both work, and they answer different questions. A ceiling at your CDN or gateway is cheap, covers everything behind it, and stops traffic before it costs you anything, but it does not know who is signed in. A ceiling in your application knows the account, which is exactly what an authentication limit needs, and it only applies to requests that reach your application. Most teams end up with a broad one outside and a specific one inside.

Does StackAttest test whether I have one?

It sends a bounded burst and records the answer. A 429 or a Retry-After settles the control at runtime, wherever the ceiling is enforced, and reading your repository can show whether a rate limiting library is present and where it is applied. Source alone never fails this control, because a ceiling is very often enforced at an edge that no repository read can see, so silence is reported as a gap to confirm rather than as a broken control.

Does the free check tell me whether I am protected?

No. It sends only a couple of requests to any one address, far too few to cross a real threshold, so the control comes back uncovered with that reason attached. The bound is deliberate: an unauthenticated tool that anyone can point at anyone's address must not be usable to send a burst at somebody else's site.

Related

Security headers checkBroken access control testFree AI app security checkAI app security auditProduction readiness audit

Find out what is actually true of your application

Start with the layer you can run today. The report says which layers ran and what they could not reach, so the result is honest about its own limits.

Run a runtime validationRun the free check