Confidence must fit the evidence.
Confidence is not a decorative percentage and not a synonym for truth. It is a bounded statement about how strongly the current evidence architecture warrants one exact claim. Calibration begins when expressed confidence is forced to match observed reliability.
Quality, directness, independence and coverage.
A claim-level judgment bounded by current evidence.
Known limits, variance, conflict and missing observations.
The required level changes with consequence and reversibility.
Strong evidence can justify confidence. It cannot justify certainty.
Confidence calibration aligns the strength of a stated belief with the reliability of similarly supported claims. The unit of calibration is the claim—not the source, page, author or research project as a whole.
A high-confidence claim can still fail.
Calibration does not promise that every 80% claim is correct. Across comparable decisions, roughly eight of ten well-calibrated 80% claims should resolve as correct.
Evidence is an input, not the output.
Directness, independence and coverage describe the support structure. Confidence is the bounded judgment formed after conflict, time and uncertainty are included.
Action depends on consequence.
The same confidence may justify an inexpensive reversible test but fail a high-impact irreversible decision. Decision readiness is not identical to belief strength.
Change the evidence architecture. Watch the ceiling move.
This explanatory model separates five evidence dimensions from decision consequence. Its numeric output is illustrative and auditable—not a universal probability or hidden truth score.
Remove one assumption. Does confidence survive?
A robust claim should not collapse when one plausible weakness is introduced. Stress testing measures resilience by changing the evidence graph rather than polishing the final percentage.
The words must carry the same uncertainty as the evidence.
Calibrated language prevents a bounded evidence state from becoming an absolute claim during writing, summarization or decision handoff.
Use only when direct, convergent and temporally valid support covers every material claim component with no material unresolved conflict.
Strong mapped support remains after dependency, conflict and temporal controls. Residual uncertainty is still declared.
The claim is warranted but qualification, incomplete coverage or a meaningful alternative explanation remains.
Support exists but cannot carry an unqualified conclusion. Treat the statement as provisional and actively seek discriminating evidence.
A claim can remain plausible while evidentially unsupported. Plausibility must not be rewritten as a finding.
A number is calibrated only when outcomes can answer back.
Repeated probabilistic judgments should be grouped by stated confidence and compared with later resolutions. Without outcome feedback, a score may be structured and transparent—but it is not empirically calibrated.
Group comparable predictions.
Keep claim type, horizon and resolution rule consistent. A 70% group should not mix short-lived SERP states with durable historical events.
Record observable outcomes.
Predeclare what counts as correct, incorrect, unresolved or unresolvable. Do not change the resolution test after seeing the result.
Measure the calibration gap.
If only 55% of claims stated at 80% resolve as correct, the system is overconfident by 25 points for that defined class.
False precision has a recognizable shape.
The interface is designed to block four common failures before a numeric confidence state reaches a conclusion.
A strong source becomes a strong claim.
Authority is treated as universal even when the source is indirect, temporally stale or silent on a material subclaim.
Dependent repetition inflates confidence.
Ten downstream reports derived from one origin are counted as ten independent observations.
Precision exceeds the model.
A display such as 87.43% implies empirical resolution data and stable parameters that the research process does not possess.
The rule moves until the claim passes.
Decision thresholds are lowered after the evidence is seen, turning the desired answer into the calibration rule.
Seven controls. One defensible confidence state.
The protocol receives the mapped structure from EVD/09 and produces a bounded input for gap analysis and eventual synthesis.
Fix the target
Use one atomic, scoped and time-bound proposition.
Read support
Inspect every typed edge and evidence family.
Apply controls
Account for quality, dependence, coverage and time.
Set ceilings
Let unresolved conflict and gaps limit confidence.
Test resilience
Remove key assumptions and recalculate the state.
Match language
Use wording that preserves the uncertainty band.
Resolve outcomes
Compare stated confidence with later results.
Confidence exposes limits. Gap analysis names what is missing.
EVD/10 bounds what the mapped evidence can justify. EVD/11 identifies the missing observations and unknowns preventing a stronger or decision-ready conclusion.
Evidence objects, claims, context and support relations.
EVD / 02CLASSIFYEvidence Types & ClassesObserved, documentary, computational and derived evidence.
EVD / 03LINEAGESource ProvenanceOrigin, custody, transformation and derivation paths.
EVD / 04QUALITYSource QualityAuthority, directness, independence and claim fit.
EVD / 05CAPTUREEvidence Collection & PreservationAcquisition context, integrity and reproducible records.
EVD / 06CORROBORATECorroboration & TriangulationIndependent convergence across sources and methods.
EVD / 07CONFLICTConflicting Evidence ResolutionScope, time, meaning and method divergence.
EVD / 08TIMETemporal Validity & Evidence DecayFreshness windows, volatility and supersession.
EVD / 09MAPPINGClaim–Evidence MappingExplicit support paths from records to propositions.
EVD / 10CONFIDENCEConfidence CalibrationBounded confidence aligned with evidence strength.
EVD / 11GAPSEvidence Gaps & UnknownsMissing observations and unresolved states.
EVD / 12SYNTHESISEvidence Synthesis & Decision ReadinessIntegrated evidence states for bounded decisions.