A score should have a standard behind it
Most trust products give you a grade and ask you to believe it. The StackAttest Standard is the other half: what was evaluated, how each result was established, how strong the evidence was, and what was not reached at all.
It is published because a result nobody can inspect is a result nobody should rely on, and that applies to ours too.
Controls
A control is a question with a defined way of being answered. The current baseline pack, StackAttest Core 1.0, carries 13 of them across security, reliability, operations and performance, and each one names what evidence can satisfy it and which external clauses it maps to. Controls are versioned data: a change in meaning requires a new control version rather than a quiet edit.
Read the control catalogueEvidence grades
Every result carries a grade recording how it was established, from inferred at the weakest through source verified, deterministically tested and runtime verified. Two products can pass the same control with very different amounts behind them, and the grade is what makes that visible instead of hidden.
What each grade meansFour signals, not one number
A single grade hides the two things that decide what it is worth: how much evidence sat behind it, and how confident the review was. So the Standard reports four separate numbers and each answers a different question.
Score
The headline result, produced deterministically. The same inputs give the same score every time, with no model opinion mixed in. A failing production gate cannot be averaged away by strong results elsewhere.
Evidence Coverage
How much of the Standard could actually be checked for this product. A control with nothing behind it is reported uncovered rather than passed, and this number is what stops an absence of evidence reading as a clean result.
Validation Confidence
How strong the evidence behind the covered controls is. It separates a result that was watched happening from one that was read in a file, which is the distinction the evidence grades exist to make.
AI Consensus
How much a council of up to 6 model providers agreed when it examined the evidence. It is reported next to the score and it never changes it.
What the council does, and what it cannot do
The intended council spans OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral. Its job is to challenge the evidence: to find contradictions, unsupported claims and places where a verdict rests on less than it appears to.
What it does not do is move the score. The deterministic engine is authoritative, and AI Consensus is reported beside the result rather than folded into it. A panel that could vote a score upward would make the score a negotiation, and the disagreement itself is more useful published than averaged away.
Standards mapping
Controls map to OWASP ASVS 5, NIST SSDF, CWE, OWASP API Security Top 10, WSTG, SLSA. A mapping means a StackAttest control references a clause in that vocabulary. It is not a claim that StackAttest, or any product it validates, is certified, accredited or endorsed by the body that publishes it, and none of them reviews this work.
What this is not
It is not a certification. StackAttest is not a certification body, issues no accreditation, and a passing result is a record of what was tested on a date rather than a seal that carries forward.
It is also not a replacement for a penetration test or for SOC 2. A penetration test is a person working against your application with intent. SOC 2 is an audit of organisational controls over a period. This tests technical claims about software and grades the evidence behind them, which is a third thing.