FinMirror
Behavioral reliability card · schema 1.1
Paired counterfactual evaluation

evidence-program

Change one financial fact. Did the system change its answer for the right reason—or stay confident in a memorized story? FinMirror scores answers, evidence, abstention, calibration, and cross-language behavior together.

126 cases 108 paired interventions en · fr · zh 6 finance workflows
Reliability vector

One score cannot hide a failure.

Every axis is deterministic in v0.1. The release gate blocks systems that are fluent but financially wrong, weakly cited, or unable to abstain.

Case accuracy
100.0%

Answers + canonical units

Case verification
100.0%

Answer + evidence + replay

Pair reliability
100.0%

Changed for the right reason

Answer behavior
100.0%

Exact paired answer contract

Citation migration
100.0%

Evidence followed each world

Formula replay
100.0%

Allow-listed program executed

Operand provenance
100.0%

Inputs linked to exact evidence

Confidence behavior
100.0%

Uncertainty followed evidence

Material sensitivity
100.0%

Reacted to material facts

Distractor invariance
100.0%

Ignored irrelevant changes

Evidence ablation
100.0%

Abstained when proof vanished

Missing evidence
100.0%

Named the exact absent operand

Citation F1
100.0%

Minimum sufficient evidence

Cross-language
100.0%

English · French · Chinese

Calibration
100.0%

Brier 0.000

Behavioral diagnosis

What broke under intervention?

Material changes demand sensitivity. Distractors, peer entities, stale periods, and document-borne instructions demand specificity.

Pass rate by transform

distractor
100.0% 18 pairs
entity collision
100.0% 18 pairs
evidence ablation
100.0% 18 pairs
injection
100.0% 18 pairs
material value
100.0% 18 pairs
period collision
100.0% 18 pairs

Failure taxonomy

  • No deterministic failures0
Failure explorer

Inspect every decision.

Filter by language, transform, or verdict. Evidence identifiers are stable, explicit, and machine-verifiable.

VerificationCaseQuestionExpectedPredictedConfidenceCitation F1Formula
No cases match these filters.
Evaluation contract

Correct—and correct for the right reason.

Material sensitivity

A controlled change to a required operand must produce the exact new answer.

Specificity / invariance

Peer entities, old periods, distractors, and injected instructions must not move the answer.

Tuple verification

Answer, citations, confidence, abstention, and reported retrieval must move together.