FinMirror
Behavioral reliability card · schema 1.1
Paired counterfactual evaluation

memorized-evidence-blind

Change one financial fact. Did the system change its answer for the right reason—or stay confident in a memorized story? FinMirror scores answers, evidence, abstention, calibration, and cross-language behavior together.

126 cases 108 paired interventions en · fr · zh 6 finance workflows
Reliability vector

One score cannot hide a failure.

Every axis is deterministic in v0.1. The release gate blocks systems that are fluent but financially wrong, weakly cited, or unable to abstain.

Case accuracy
71.4%

Answers + canonical units

Case verification
0.0%

Answer + evidence + replay

Pair reliability
0.0%

Changed for the right reason

Answer behavior
66.7%

Exact paired answer contract

Citation migration
66.7%

Evidence followed each world

Formula replay
0.0%

Allow-listed program executed

Operand provenance
0.0%

Inputs linked to exact evidence

Confidence behavior
83.3%

Uncertainty followed evidence

Material sensitivity
0.0%

Reacted to material facts

Distractor invariance
0.0%

Ignored irrelevant changes

Evidence ablation
0.0%

Abstained when proof vanished

Missing evidence
0.0%

Named the exact absent operand

Citation F1
83.3%

Minimum sufficient evidence

Cross-language
71.4%

English · French · Chinese

Calibration
84.8%

Brier 0.152

Behavioral diagnosis

What broke under intervention?

Material changes demand sensitivity. Distractors, peer entities, stale periods, and document-borne instructions demand specificity.

Pass rate by transform

distractor
0.0% 18 pairs
entity collision
0.0% 18 pairs
evidence ablation
0.0% 18 pairs
injection
0.0% 18 pairs
material value
0.0% 18 pairs
period collision
0.0% 18 pairs

Failure taxonomy

  • failed to abstain18
  • incorrect operand provenance108
  • insufficient evidence18
  • invalid formula replay108
  • missing requirement not identified18
  • wrong answer18
Failure explorer

Inspect every decision.

Filter by language, transform, or verdict. Evidence identifiers are stable, explicit, and machine-verifiable.

VerificationCaseQuestionExpectedPredictedConfidenceCitation F1Formula
No cases match these filters.
Evaluation contract

Correct—and correct for the right reason.

Material sensitivity

A controlled change to a required operand must produce the exact new answer.

Specificity / invariance

Peer entities, old periods, distractors, and injected instructions must not move the answer.

Tuple verification

Answer, citations, confidence, abstention, and reported retrieval must move together.