A score you cannot interrogate is a number, not a result
The scoring engine is pure arithmetic over recorded control evaluations. No database, no network, no model. Run it twice on the same evaluations under the same policy version and it returns the same answer, which is the property that makes a score worth arguing with.
Models may recommend a control outcome earlier in the pipeline. They never touch this calculation.
Why four numbers
A single grade hides the two things that decide what it is worth. So the result carries four separate numbers, and collapsing them would throw away the part that makes the first one readable.
Score is the quality of the controls that were verified. Evidence Coverage is how much of the applicable surface has source verified evidence or better behind it. Validation Confidence is how strong that evidence is. AI Consensus is computed separately and never feeds into any of the others.
A high score on thin coverage and a high score on deep coverage are different results. Reporting one number would make them look identical, which is precisely the confusion the product exists to remove.
What a critical failure does
A production-blocking critical failure cannot be compensated by unrelated categories. While it is unresolved it holds the production status at not production ready and caps the displayed score at 59.
That cap is the most useful single fact for reading one of these scores, which is why it is published. Above it, there is no unresolved critical blocker. It also means the score is not an average: averages let a strong operations category quietly pay for a broken authorisation boundary, and no arrangement of good results should be able to do that.
Why the policy is versioned
Every run records the policy version it was scored under, and changing any threshold requires a new version rather than an edit. Historical scores are never rewritten.
This matters more than it sounds. A trust product that silently restates old results has made its own history unciteable: a score somebody screenshotted last year would no longer mean what it meant. The cost is that two runs under different policies are not directly comparable, and saying so is better than hiding it.
What is not published, and why
The coefficients, the category weights, the outcome values and the exact thresholds stay private. Everything a reader needs in order to interpret a score is on this page and in the control catalogue: what is evaluated, how each result is established, what evidence is required, and what a critical failure does.
Publishing the formula would not make anybody better informed. It would make the number easier to aim at than to satisfy, and a score optimised against rather than earned is worth nothing to the person relying on it. If you want to check a specific result, the evidence behind every control is the thing to interrogate, and it is attached to the finding rather than summarised away.