StackAttestTechnical Trust
Reference

A score should have a standard behind it

Most trust products give you a grade and ask you to believe it. The StackAttest Standard is the other half: what was evaluated, how each result was established, how strong the evidence was, and what was not reached at all.

It is published because a result nobody can inspect is a result nobody should rely on, and that applies to ours too.

Controls

A control is a question with a defined way of being answered. The current baseline pack, StackAttest Core 1.0, carries 13 of them across security, reliability, operations and performance, and each one names what evidence can satisfy it and which external clauses it maps to. Controls are versioned data: a change in meaning requires a new control version rather than a quiet edit.

Read the control catalogue

Evidence grades

Every result carries a grade recording how it was established, from inferred at the weakest through source verified, deterministically tested and runtime verified. Two products can pass the same control with very different amounts behind them, and the grade is what makes that visible instead of hidden.

What each grade means
The validation ladderHow the score worksAI consensusChangelog

Four signals, not one number

A single grade hides the two things that decide what it is worth: how much evidence sat behind it, and how confident the review was. So the Standard reports four separate numbers and each answers a different question.

Score

The headline result, produced deterministically. The same inputs give the same score every time, with no model opinion mixed in. A failing production gate cannot be averaged away by strong results elsewhere.

Evidence Coverage

How much of the Standard could actually be checked for this product. A control with nothing behind it is reported uncovered rather than passed, and this number is what stops an absence of evidence reading as a clean result.

Validation Confidence

How strong the evidence behind the covered controls is. It separates a result that was watched happening from one that was read in a file, which is the distinction the evidence grades exist to make.

AI Consensus

How much a council of up to 6 model providers agreed when it examined the evidence. It is reported next to the score and it never changes it.

What the council does, and what it cannot do

The intended council spans OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral. Its job is to challenge the evidence: to find contradictions, unsupported claims and places where a verdict rests on less than it appears to.

What it does not do is move the score. The deterministic engine is authoritative, and AI Consensus is reported beside the result rather than folded into it. A panel that could vote a score upward would make the score a negotiation, and the disagreement itself is more useful published than averaged away.

Standards mapping

Controls map to OWASP ASVS 5, NIST SSDF, CWE, OWASP API Security Top 10, WSTG, SLSA. A mapping means a StackAttest control references a clause in that vocabulary. It is not a claim that StackAttest, or any product it validates, is certified, accredited or endorsed by the body that publishes it, and none of them reviews this work.

OWASP ASVSNIST SSDFCWEAPI Security Top 10WSTGAll mappings

What this is not

It is not a certification. StackAttest is not a certification body, issues no accreditation, and a passing result is a record of what was tested on a date rather than a seal that carries forward.

It is also not a replacement for a penetration test or for SOC 2. A penetration test is a person working against your application with intent. SOC 2 is an audit of organisational controls over a period. This tests technical claims about software and grades the evidence behind them, which is a third thing.

Keep reading