StackAttestTechnical Trust
Consensus 1.0.0

A panel of models that cannot change the answer

After the evidence is gathered and before the score is computed, a council of models is asked to challenge what was found: to look for contradictions, for claims resting on less than they appear to, and for conclusions the evidence does not support.

The intended council spans OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral. What it produces is reported beside the score as a fourth number. It is never folded into it.

Why it cannot move the score

If a panel could vote a score upward, the score would be a negotiation. The whole argument for a deterministic engine is that the same evidence produces the same result every time, and a model with an opinion is the one thing that breaks that property.

So evidence-backed deterministic outcomes cannot be voted away. A control that failed against observed evidence has failed, whatever the panel thinks of it, and the disagreement is published rather than resolved in the score's favour. Disagreement is information. Averaging it into a number destroys it.

The stop rule

The aggregation refuses to report consensus at all while an evidence-backed critical objection is unresolved. This is the rule worth understanding, because the alternative is the failure mode every review panel has: a strong objection from one participant gets outvoted, the average looks healthy, and the objection disappears into it. Here it blocks the result instead, and the objection stays visible until it is answered.

A percentage without a denominator is not a number

Three models agreeing out of three configured and three agreeing out of 6 configured are both one hundred per cent, and only one of them is a full council. So the result carries how many models actually answered alongside how many seats the run was configured for, and a run where fewer answered than were expected is marked as a partial council.

It also reports the shape of the vote rather than only its average: how many agreed, how many disagreed, and how many said there was not enough here to have an opinion on. That last category is the one most products would quietly drop, and it is often the most informative: several reviewers declining to conclude is a finding about the evidence.

What this is not

It is not a second opinion on your software. It is a review of the evidence a validation gathered, which is a narrower and more useful job. A model that has not seen your system cannot tell you whether your architecture is sound; it can notice that a conclusion drawn three steps ago was not supported by what was actually observed.

It is also not a measure of how good the software is. High agreement on thin evidence means the panel agreed about very little. Read it next to Evidence Coverage or not at all.

Keep reading