← Back to the series

Real-world data biases.

When the data tell the wrong story.

A comparison can change because the people, the clocks or the measurements change. Six fictional experiments make those mechanisms visible, with live results and calculations you can follow.

Open full view ↗

Ask what created the comparison

A large dataset can produce a precise answer to a poorly constructed question. Bias concerns systematic distortion, while random error concerns variation across samples. This playground isolates six ways a comparison can become misleading. It uses expected values so that sampling noise does not obscure the mechanism.

Every scenario is invented. You can change the selection process, risk mix, timing or measurement and compare the apparent result with an appropriate standardized comparison or a clearly labeled simulator benchmark. Those benchmarks are teaching privileges; a researcher would need a defensible design, additional data or assumptions to identify the corresponding effects.

Confounding can reverse the overall comparison

In the first example, a fictional prevention program reduces flare risk within both baseline categories. Yet the program group contains more high-risk people. Its overall event risk can therefore look worse. The raw average reflects both the program effect and the different composition of the groups.

Standardizing both groups to a 50:50 risk mix separates those contributions under this model’s assumptions. The controls let you balance the groups or change the true risk ratio. Scaling the expected cohort from 1,000 to 100,000 changes the counts, but preserves the distortion. For a closer look at creating comparable groups, visit the IPTW and matching lesson.

Lead time changes the clock; length time changes the cases

The lead-time experiment moves diagnosis earlier while leaving the death date unchanged. Measured survival after diagnosis increases, even though life is not extended. A separate slider adds actual life extension so that the two contributions remain visible.

The length-time experiment instead changes which cases a snapshot screen encounters. A slower-progressing form stays detectable longer and has a better prognosis. It can become overrepresented among detected cases, making that group look healthier without changing either type’s outcome risk. The example holds lead-time effects off to isolate this composition effect. The National Cancer Institute explains why these distinctions matter when interpreting screening outcomes.

Immortal time uses the future to classify the past

Suppose a person can start a program only after a waiting period without an event. A naive “ever used the program” group is guaranteed to have passed that period. Assigning people with early events to the non-user group can then manufacture an apparent advantage, even when the program has no effect.

The table follows those early events into the naive groups. It also shows an aligned comparison among people still eligible when the program becomes available. That later-start comparison has a different target population and follow-up window. It is not a universal way to estimate a day-0 strategy effect. The target trial emulation lesson develops the design question further.

Selection and measurement can alter the result without altering the outcome

In the selection experiment, both complete groups have the same true risk. Removing more event records from one group makes its retained subset appear safer. The data you analyze are no longer a neutral sample of the people you intended to compare. The missing-data lesson explores the assumptions needed to handle such losses.

In the measurement experiment, the true risks again match, but the probability of recording a real event differs. The recorded risk combines detected true events and false positives. Equalizing ascertainment removes the contrast in this particular example; it does not guarantee unbiased absolute risks or establish a general direction for measurement bias.

Use the mechanism to choose the response

Matching cannot repair every timing problem, and a smaller p-value cannot validate an outcome definition. The practical task is to identify the target comparison and the process that distorted it. The Cochrane Handbook’s discussion of non-randomized studies provides a broader framework for this assessment.

Begin with the guided tour, then use “Remove this distortion” to see what must change in each example. Expand the mathematical bridge to follow the counts, weighted risks or clock times into the final result.

Start with “Guide me through it,” then try your own settings. Expand the mathematical bridge to trace the results back to their denominators and assumptions.

Next: errors, power and p-values →