Understanding Deeply · 07 · Fictional experiments

When the data tell the wrong story

Change who enters a comparison, when their clock starts, or how an outcome is recorded. Watch the apparent result move—even when the underlying effect stays the same.

What would address this?

The mathematical bridgeCurrent inputs → counts or clock times → apparent result → appropriate comparison.

Different mechanisms need different solutions

The experiments isolate one mechanism at a time. Actual studies can contain several at once.
BiasWhat changes?A useful design response
ConfoundingBaseline causes of outcome differ between groups.Design comparable groups; adjust appropriate pre-exposure variables.
Lead timeThe diagnosis clock starts earlier.Evaluate meaningful outcomes from comparable starting points.
Length timeLonger detectable states are overrepresented.Account for case mix; assess outcomes among people assigned to screening.
Immortal timeFuture exposure status requires surviving an earlier period.Align eligibility, strategies and follow-up; handle delayed starts explicitly.
SelectionInclusion depends on group and outcome-related factors.Preserve follow-up and examine the selection process and target population.
MeasurementDetection or classification differs between groups.Use comparable ascertainment, validation data and suitable bias analyses.

A bigger database does not automatically repair a distorted comparison. These are expected-value demonstrations without sampling noise or confidence intervals. The next lesson explores random error, power and p-values.

Model boundaries and sourcesWhat is invented, what is assumed, and what these examples cannot establish.

Six separate teaching worlds

All programs, conditions, people, risks and timelines are fictional. The risk experiments show expected counts, which may be fractional. The lead-time experiment follows two otherwise identical imagined trajectories. These are illustrations of mechanisms, not estimates of how common any bias is.

“Adjusted” does not mean automatically correct

The confounding example assumes baseline risk is measured without error, sufficient to control confounding, and observed in both groups. Its standardized result targets a 50:50 low/high-risk population. The selection and measurement panels show the simulator’s full-data benchmark; that benchmark is not recoverable from the distorted records alone without more information or assumptions.

Related problems

Confounding by indication and healthy-user bias concern different reasons treatment groups may differ. Attrition can introduce selection bias. Differential surveillance is one form of measurement bias. Overdiagnosis is distinct from simply starting the clock earlier: it identifies a condition that would not have caused symptoms or death during that person’s lifetime. No single adjustment resolves this entire list.