Answers + canonical units
memorized-evidence-blind
Change one financial fact. Did the system change its answer for the right reason—or stay confident in a memorized story? FinMirror scores answers, evidence, abstention, calibration, and cross-language behavior together.
One score cannot hide a failure.
Every axis is deterministic in v0.1. The release gate blocks systems that are fluent but financially wrong, weakly cited, or unable to abstain.
Answer + evidence + replay
Changed for the right reason
Exact paired answer contract
Evidence followed each world
Allow-listed program executed
Inputs linked to exact evidence
Uncertainty followed evidence
Reacted to material facts
Ignored irrelevant changes
Abstained when proof vanished
Named the exact absent operand
Minimum sufficient evidence
English · French · Chinese
Brier 0.152
What broke under intervention?
Material changes demand sensitivity. Distractors, peer entities, stale periods, and document-borne instructions demand specificity.
Pass rate by transform
Failure taxonomy
- failed to abstain18
- incorrect operand provenance108
- insufficient evidence18
- invalid formula replay108
- missing requirement not identified18
- wrong answer18
Inspect every decision.
Filter by language, transform, or verdict. Evidence identifiers are stable, explicit, and machine-verifiable.
| Verification | Case | Question | Expected | Predicted | Confidence | Citation F1 | Formula |
|---|
Correct—and correct for the right reason.
A controlled change to a required operand must produce the exact new answer.
Peer entities, old periods, distractors, and injected instructions must not move the answer.
Answer, citations, confidence, abstention, and reported retrieval must move together.