Decision criteria.
Decision criteria translate objectives, protected conditions and material trade-offs into explicit properties used to distinguish feasible options. Good criteria clarify judgment. Bad criteria duplicate value, reward what is easiest to count and turn hidden preferences into authoritative-looking scores.
VALUE
MODEL
Criteria are comparison properties—not goals, data points or decoration.
The objective states what should improve. Criteria define the dimensions on which options are judged. Measures provide observations. Constraints remove prohibited states. Weights express value trade-offs only after these distinctions are stable.
Decision criteria are explicitly defined properties used to evaluate feasible options against the decision’s objectives, constraints and affected interests. Each criterion requires a meaning, direction, admissible evidence, scale, threshold or preference rule, time horizon and treatment of uncertainty.
Every criterion needs an operational specification.
A label such as “quality,” “risk” or “strategic fit” is not yet a criterion. It becomes usable only when its meaning and evaluation rule are explicit.
Meaning
The exact property being judged and why it matters to the objective.
Preference
Whether higher, lower, inside a range or threshold attainment is preferred.
Measure
Evidence, unit, source, population and observation window.
Value function
How observed differences translate into decision-relevant differences.
Threshold
Minimum, maximum or unacceptable region that changes option status.
Uncertainty
Range, quality, conflict and sensitivity of the underlying estimate.
Importance
Trade-off significance derived from the decision—not generic importance.
Owner
Who defines, validates and approves the criterion and its evidence.
Prevent five different concepts from collapsing into one score.
This separation is the foundation of a defensible evaluation model.
| Element | Function | Representation | Example | Do not use it as |
|---|---|---|---|---|
| Objective | Defines the valued state the decision should advance. | Outcome + beneficiary + horizon. | Maintain service continuity during growth. | A directly scored number. |
| Criterion | Defines a property on which options differ. | Meaning + direction + scale. | Recovery capability under peak load. | A vague heading such as “quality.” |
| Measure | Observes or estimates the criterion. | Unit + method + population + time. | Verified recovery time in load test. | The whole concept when it is only a proxy. |
| Hard constraint | Removes unauthorized or impossible states. | Pass/fail threshold with basis. | EU data residency required. | A low-weight criterion that can be compensated. |
| Guardrail | Protects a condition while optimizing another. | Acceptable range or no-worsening rule. | No increase in severe incident exposure. | A benefit to maximize. |
| Weight | Expresses relative value of criterion swings. | Declared trade-off basis. | Value of moving from worst to best plausible reliability. | A scientific fact about universal importance. |
Watch the preferred option change when value assumptions change.
The model is illustrative—not a recommendation. It demonstrates why weights must be visible and why a sensitivity check matters more than one polished total.
WEIGHT PROFILE
Move the sliders or choose a viewpoint. Values are normalized internally to show relative emphasis.
STAGED PILOT LEADS
Eight tests for a decision-ready criterion.
A criterion that fails these tests should be revised, decomposed, merged or removed before options are scored.
| Test | Question | Pass condition | Failure signature | Repair |
|---|---|---|---|---|
| Relevance | Does it represent a material objective or protected condition? | A clear causal or value link exists. | Included because it is conventional or easy to count. | Trace it to an objective or remove it. |
| Discrimination | Can feasible options differ meaningfully? | Plausible variation changes judgment. | Every option receives the same score. | Remove or use as a shared assumption. |
| Operational clarity | Could two reviewers apply it consistently? | Meaning, unit, evidence and rule are explicit. | “Quality” means different things to different reviewers. | Define subproperty and scale anchors. |
| Non-redundancy | Is the same value counted elsewhere? | Overlap is removed or modeled deliberately. | Reliability, uptime and availability all reward one effect. | Merge or define non-overlapping scopes. |
| Preferential independence | Can trade-offs be interpreted without hidden interaction? | Dependencies are absent or explicit. | Value of speed changes completely with safety state. | Use interaction rule, scenario or combined criterion. |
| Proportionality | Is analysis effort justified by decision exposure? | Evidence burden matches consequence and reversibility. | False precision for a low-stakes reversible choice. | Simplify scale or use qualitative bands. |
| Auditability | Can the score be traced? | Source, assessor, date and rationale are preserved. | Number appears without provenance. | Require a criterion record. |
| Sensitivity | Would plausible changes alter ranking? | Fragility is tested and reported. | One total shown as inevitable. | Vary weights, scores and uncertain inputs. |
Related criteria can silently count the same value more than once.
Correlation is not automatically duplication, but shared causes and overlapping meanings must be inspected before weighting.
observed service state
failure behavior
restoration capability
CONTINUITY
Not every criterion should be traded continuously.
Different value structures require different rules. A minimum compliance threshold, target range and monotonic preference are not interchangeable.
Hard threshold
Crossing the boundary makes the option infeasible or unauthorized.
EXAMPLE: mandatory residency unavailableMinimum standard
Improvement matters until an adequate level is reached; excess may add little value.
EXAMPLE: required support coverageMore or less is better
Preference moves consistently over the plausible range, subject to diminishing returns.
EXAMPLE: lower verified total costTarget interval
Both too little and too much can reduce value.
EXAMPLE: inventory or response intensityCriteria must be generated from the decision—not copied from a universal template.
Select a scenario. Each set separates value, hard limits, observation and review conditions.
A number is meaningful only when its anchors are defined.
A 1–5 or 1–10 scale does not create precision by itself. Each level must represent an observable state or defensible judgment boundary.
| Scale type | Best use | Required anchors | Main risk | Control |
|---|---|---|---|---|
| Natural unit | Directly measurable consequence. | Unit, period, population and uncertainty. | False comparability across different units. | Keep original units visible. |
| Threshold | Compliance, capacity or safety gate. | Pass boundary and authoritative basis. | Compensating for failure elsewhere. | Screen before aggregation. |
| Ordinal bands | Structured expert judgment. | Behavioral description for every band. | Treating intervals as equal. | Do not infer arithmetic distance. |
| Value function | Translate performance into decision value. | Worst/best plausible states and shape. | Hidden assumptions about marginal value. | Show curve and test alternatives. |
| Probability / range | Uncertain outcomes. | Reference class, time and confidence basis. | Single-point certainty. | Preserve distribution or interval. |
| Qualitative narrative | Complex or weakly observable effects. | Claim, evidence, mechanism and caveat. | Unstructured persuasion. | Use a fixed evidence template. |
Weights do not measure abstract importance.
A defensible weight reflects the decision value of moving across a defined performance range. Without common ranges, “reliability is twice as important as cost” has no stable operational meaning.
Range neglect
Weighting criterion names without specifying worst and best plausible performance makes trade-offs uninterpretable.
Compensatory error
A weighted total can allow exceptional strength on one criterion to mask an unacceptable weakness on another.
Precision theatre
Decimals and rankings may conceal fragile evidence, subjective anchors and unresolved disagreement.
Stakeholder averaging
Averaging incompatible values can erase genuine conflict instead of making it governable.
Proxy capture
The easiest observable metric can replace the condition the decision actually values.
Rank reversal
Small changes in weights, scales or candidate set can change the apparent winner.
Store the criterion with its meaning, scale and provenance.
Retrieving “reliability: 8” without its population, measure, anchors, evidence date and uncertainty is not decision intelligence. It is an orphaned number.
Carry the evaluation rule with every score.
This record makes definitions and assumptions retrievable. It does not turn judgment into fact; it makes judgment inspectable.
{
"decision_id": "DEC-05-001",
"criterion_id": "CRIT-REL-01",
"name": "recovery capability",
"objective_link": "service continuity",
"definition": "restore priority service after failure",
"direction": "lower is better",
"measure": {"unit": "minutes", "window": "peak load"},
"threshold": {"max": "60", "type": "hard"},
"scale_anchors": {"worst": "240", "best": "15"},
"weight_basis": "value of worst-to-best swing",
"overlap_with": ["availability"],
"uncertainty": {"range": "35–70"},
"provenance": [{"source": "test record URI", "date": "YYYY-MM-DD"}]
}Decision criteria, clarified.
Operational answers to the most common evaluation-model failures.
What makes a good decision criterion?
It represents a material objective or protected condition, distinguishes feasible options, has a clear direction and operational definition, uses appropriate evidence, avoids redundancy and can be applied consistently enough for the decision’s stakes.
How many criteria should be used?
Use the smallest set that adequately represents the material value and risk structure. Too few criteria omit important effects; too many increase overlap, dilute meaning and create false analytical burden. Merge duplicates and remove dimensions that do not discriminate.
Should cost always be a criterion?
Resource consequences should be represented, but the correct form depends on the decision. Cost may be a hard budget constraint, total lifecycle consequence, opportunity cost or one element of value for money. Avoid mixing price, total cost and affordability as if they were independent.
Can a hard constraint receive a weight?
Not if failure is genuinely non-compensable. Screen hard constraints before scoring. If a condition permits degrees of preference above a mandatory minimum, separate the threshold from the additional performance criterion.
Are equal weights neutral?
No. Equal weights are a value judgment, and their meaning still depends on performance ranges and scale design. They can be useful as one sensitivity scenario, but they should not be presented as assumption-free.
How should qualitative criteria be scored?
Define observable anchors or a fixed evidence narrative for each level. Preserve source, reasoning and uncertainty. Do not assume ordinal labels have equal numeric distance merely because they are coded 1, 2, 3 and 4.
What is sensitivity analysis?
It tests whether plausible changes to weights, scores, assumptions or uncertain inputs alter the result. A stable ranking under relevant variations provides different information from a ranking that reverses after a small change.
Does the highest weighted score identify the correct decision?
No. It identifies the leading option under the declared model. Decision-makers must still inspect constraints, uncertainty, distributional effects, omitted factors, model sensitivity and the accountable rationale for commitment.
Continue through the complete Decision Intelligence system.
Criteria follow a feasible option set and prepare it for evidence, trade-offs, uncertainty, commitment and review.
Decision Intelligence
Structure choices through objectives, alternatives, evidence, uncertainty, trade-offs, commitment and review.
Criteria should expose judgment—not hide it inside a score.
Trace each property to an objective. Separate gates from preferences. Define scales and ranges. Remove overlap. Preserve uncertainty. Test sensitivity. Then use the model to structure accountable judgment rather than impersonate certainty.