Sampling & Data Collection.
Sampling decides which eligible units become evidence; data collection decides how their states are captured. A defensible protocol preserves the sampling frame, selection rule, acquisition context, failed requests, missing values and transformation boundary—not merely the successful rows returned by a tool.
Returned data is not automatically a valid sample.
A provider response, crawl export or model answer becomes research evidence only after its relationship to the eligible population, selection procedure and acquisition conditions is known.
Sampling and data collection is the controlled selection of units from a defined population and the traceable acquisition of their observable states under declared, repeatable conditions.
The sampling frame is the operational list or mechanism through which those units can actually be reached.
Removing it silently changes the achieved sample and may change the conclusion.
Normalization, classification and scoring belong to a versioned transformation layer.
Five counts must remain separate.
If a dashboard reports only the final count, coverage errors, provider failures and validation exclusions disappear from view.
Eligible population
All units permitted by the MTH/03 definitions, context, time and inclusion rules.
Sampling frame
The list, index, endpoint or generation mechanism from which units can be selected.
Selected sample
Units chosen by the predeclared census, random, stratified, systematic or purposive rule.
Returned sample
Successful and null-bearing responses after bounded retries; failures remain registered.
Analytical sample
Validated observations eligible for computation after exclusions and duplicate resolution.
The selection rule determines what the sample can represent.
No strategy is universally best. The method must fit the question, population structure, operational access and intended claim.
Complete enumeration
Attempt every eligible unit. Useful for bounded query sets or canonical corpora, but failures still make the achieved sample incomplete.
Equal-probability selection
Select units using a preserved randomization procedure when a complete frame exists and population-level estimation is intended.
Selection within subgroups
Partition the frame by declared properties such as intent, market, page type or entity class, then sample within every stratum.
Fixed interval selection
Select every kth unit after a defined start. Efficient, but dangerous when frame ordering contains periodic structure.
Criterion-led cases
Select cases because they satisfy a declared analytical role: leading competitors, failure cases, high-value entities or citation events.
Available observations
Use what is accessible only when explicitly labeled. Convenience data supports exploration, not silent population generalization.
Sampling changes with the evidence system.
Select the environment. The console changes the frame, selection rule, acquisition state, missingness treatment and claim denominator. All initial content remains visible without JavaScript.
Attempt every query in the versioned research corpus
Preserve acquisition state row by row.
The ledger separates the intended unit from the request attempt, provider response, raw observation and analytical eligibility.
| FIELD | PURPOSE | EXAMPLE STATE | NEVER SILENTLY REPLACE WITH | CONTROL |
|---|---|---|---|---|
| unit_id | Stable identity of selected unit | query_0074 | Row order or displayed keyword | IMMUTABLE |
| selection_state | Why the unit entered | stratum:intent/informational | “Relevant” after inspection | PREDECLARED |
| request_context | Observation-changing parameters | location / language / device / depth | Interface defaults | EXPLICIT |
| attempt_state | Request and retry history | attempt 2 / transient failure resolved | Successful response only | TRACEABLE |
| raw_state | Provider or source observation | response preserved before scoring | TAO classification | SEPARATE |
| missing_state | Reason observation is absent | null / unavailable / failed / excluded | Zero | TYPED |
| validation_state | Analytical admission decision | valid / duplicate / unresolved | Deleted record | AUDITABLE |
| collected_at | Temporal provenance | ISO timestamp + snapshot id | Publication date | PRESERVED |
Most collection bias enters before analysis.
These failure modes can survive perfect formulas because the observed sample was already distorted when the data entered the system.
Coverage error
The frame cannot reach part of the eligible population, so some units have no chance of selection.
CHECK / eligible minus frameSelection bias
Entry probability depends on convenience, visibility or an outcome related to the intended conclusion.
CHECK / why each unit enteredAcquisition drift
Context, source parameters, prompt wording, crawl rules or collection time changes during the run.
CHECK / batch configuration diffSurvivorship
Only successful responses remain visible; failed, null, removed or inaccessible units disappear.
CHECK / selected vs returnedFour collection designs. Four distinct error surfaces.
Illustrative counts explain the mechanics. They are not live provider data or claims about an actual domain.
Null is not zero. Failure is not absence.
Missing states have different causes and different analytical consequences. Preserve the state before deciding whether it belongs in a calculation.
Freeze the acquisition logic with the dataset.
The manifest makes the achieved sample reconstructable and exposes the distance between planned collection and usable evidence.
Every method node. One controlled research route.
MTH/04 creates the evidence base. MTH/05 applies these controls specifically to search-result research.