FinMirror · paired counterfactual evaluation
Does the agent change
for the right reason?
A zero-key offline comparison. The gold-reading oracle validates the harness; the non-gold evidence program validates the public contract; and the evidence-blind baseline demonstrates why ordinary accuracy misses groundedness failures.
| System | Audit score | Case accuracy | Pair reliability | Sensitivity | Invariance | Ablation | Gate |
|---|---|---|---|---|---|---|---|
| harness-oracle | 100.0 | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | PASS |
| evidence-program | 100.0 | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | PASS |
| memorized-evidence-blind | 49.5 | 71.4% | 0.0% | 0.0% | 0.0% | 0.0% | BLOCKED |
Oracle results are harness checks, not model results. Synthetic v0.1 is a developer preview and not a regulatory certification or investment recommendation.