Understanding Deeply · 06
Missing data.
What did the unseen outcomes change?
A blank outcome is not a zero. Hide outcomes in a fictional study, compare the answers produced by different methods, and explore what you must assume about the people you cannot see.
Start with the question, then hide the outcomes
Imagine a prevention program A and a usual routine B, with 1,000 fictional people in each group. The question is the difference in their 12-month risk of a symptom flare among everyone enrolled. In the complete invented cohort, those risks are 25% and 40%: a difference of −15 percentage points. Baseline risk is recorded for everyone, but some final outcomes disappear.
The playground keeps that complete cohort fixed. You control why the outcomes become missing, how many disappear on average, and whether the analysis uses the recorded baseline information. The full-data result provides a teaching benchmark. A researcher with actual missing data would not have that privileged view.
MCAR, MAR and MNAR describe what missingness depends on
Missing completely at random (MCAR) means missingness is unrelated to the study variables, whether observed or unseen. The MCAR control uses one probability for every person. Complete-case analysis can preserve the comparison on average, while losing information. A particular random subset can still produce a different answer.
Missing at random (MAR) allows missingness to depend on recorded information. Here, higher baseline risk makes outcomes more likely to be missing, and the two groups can have different missingness rates. Within each group and baseline category, missingness no longer depends on whether a flare occurred. The word “random” does not mean that the observed subset represents the original cohort without adjustment.
Missing not at random (MNAR) means that dependence on the unseen outcome remains after accounting for the observed information. The MNAR scenario makes people with a flare more likely to have an unrecorded outcome even within the same baseline category. The labels describe the simulator’s known rules; observed data alone generally cannot distinguish MAR from MNAR.
Watch the numerator, denominator and assumptions change
Complete-case analysis removes missing outcomes from the denominator. Treating every missing outcome as “no flare” keeps those people in the denominator but adds no events. Filling with the group’s observed mean leaves the complete-case point estimate unchanged. These choices answer the calculation differently because they make different assumptions about the unknown values.
Conditional filling predicts the missing outcomes within group and baseline cells. Response weighting gives more influence to observed people from poorly observed cells. In this deliberately simple model, both produce the same point estimate when they use the same cells. The baseline switch shows why recording a useful predictor is not enough if the analysis ignores it. Response weights concern the probability of observing an outcome; compare this with treatment weights in the IPTW lesson.
Multiple imputation carries uncertainty forward
A single predicted value is not a new measurement. Multiple imputation creates several plausible completed datasets, analyzes each, and combines their estimates and uncertainty. The app lets you inspect those completions and see the within- and between-imputation components of Rubin’s rules. It uses a transparent binary-outcome model and displays an approximate large-sample interval.
More imputations reduce Monte Carlo noise, but do not make missing information reappear. Nor does a narrow interval rescue an inappropriate missing-data assumption. The expandable calculations show exactly how the completed outcomes enter the final risk difference.
Change the assumption about what you cannot see
The sensitivity panel holds the observed records fixed. You can assume that missing people have a flare risk a specified number of percentage points above or below the observed risk within their fitted cell. Moving those assumptions separately for A and B can weaken, strengthen or reverse the comparison.
The tipping-point button finds where the point estimate changes sign. That is not a test of statistical significance, and the assumptions still need a substantive justification. The panel also shows the extreme possible completion bounds without a missing-outcome model. The bounds and sensitivity scenarios are not confidence intervals.
A mechanism label is the beginning of the analysis
Real studies need a clear target question, thoughtful data collection, suitable models, diagnostics and sensitivity analyses. Complete-case analysis can be valid beyond MCAR in some settings; MAR methods can fail when their models omit important predictors. No method is guaranteed to be closest in every random draw. An outcome that cannot meaningfully exist is also different from an unrecorded outcome; explore that distinction in the intercurrent-events lesson.
Method references include Sterne and colleagues on multiple imputation, Seaman and colleagues on imputation and weighting, and Leurent and colleagues on MNAR sensitivity analysis. The scenario, counts and interactive calculations here are entirely fictional.
Start with “Guide me through it,” then try your own settings. Expand the mathematical bridge to trace the results back to their denominators and assumptions.
Explore all Understanding Deeply lessons →