FinMirror ยท deterministic meta-evaluation

Did the judge reward the right requirement?

A zero-model-call audit that separates checklist decomposition defects from permissive verifier judgments, then tests reward behavior under atomic omission, irrelevant context, and requirement reordering.

PASS

atomic-calibrated-verifier

Atomic checklist
100%
Item accuracy
100%
False-pass rate
0%
Item Brier
0.001
Paired relations
100%
Max reward inflation
-0.015
BLOCKED

atomic-permissive-verifier

Atomic checklist
100%
Item accuracy
94%
False-pass rate
100%
Item Brier
0.059
Paired relations
67%
Max reward inflation
+0.210
BLOCKED

collapsed-permissive-verifier

Atomic checklist
0%
Item accuracy
75%
False-pass rate
100%
Item Brier
0.232
Paired relations
0%
Max reward inflation
+0.210

Positive and negative controls: all three systems see the same oracle requirement states. Only the atomic calibrated verifier earns the release gate; permissive scores and collapsed checklists fail for different, inspectable reasons.

Reproduce with finmirror judge-demo. Supply external learned-judge outputs with finmirror judge-audit --input audit.json.