Understanding Deeply · The trial design laboratory

Design the trial you wish you had.

Does starting a fictional follow-up support program reduce an unplanned visit? Begin with an imaginary randomized trial, then work out how an observational study could emulate it. Watch the result change when you align time zero and compare the same risk mix.

Entirely invented example · Expected counts, not real patient records · No sampling noise

Try a story
First, define the target

The question you are emulating

This example targets starting at a later decision point. Moving that point changes the eligible population. It does not estimate the effect of a strategy assigned at the initial record date.

1. Put everyone on the right clock

Each line represents a group of expected records. “Early event” means the first unplanned visit happened before the program was available.

2. Compare the resulting estimates

Risk difference = program risk − no-program risk. Negative values mean fewer unplanned visits with the program. pp = percentage points.

Adjustment after design

3. Give both groups the same baseline risk mix

At the aligned start, selection can still make the two groups different. Standardization averages each group’s risk-specific outcomes over the same eligible population.

Bars show the share of higher risk people. The target mix is defined by eligibility at the chosen start day.

Known only inside this simulation

The causal answer we built into the example

The adjusted result matches this reference because this model has only two measured risk groups, complete follow-up and correctly specified adjustment. Real observational data do not come with this answer.

From a question to a protocol

4. Build the seven parts of a target trial

Choose a component to see the intended trial beside the observational emulation. Your population, start day and follow-up controls update the protocol.

Open the mathematical bridgeSee every denominator, the risk strata and the standardized result

Step 1 · Account for the whole starting population

Expected records and outcomes. Decimals are intentional; calculations use unrounded values.
Risk groupInitial recordsEarly eventsEligible at startChance of starting

Step 2 · Follow eligible people forward from the same day

Risk groupProgram NProgram eventsProgram riskNo program NNo program eventsNo program risk

Step 3 · See how the misleading denominator creates a difference

This comparison labels people by future initiation. Everyone with an early event ends up in “never started.” Future starters must first remain event-free long enough to start. That guaranteed event-free time belongs to the group definition, not to a protective effect of the program.

Step 4 · Align the start, then calculate the unadjusted comparison

Both denominators now contain people eligible at the same decision point. This answers the later-start question. It still mixes higher and lower risk people in different proportions when selection is present.

Step 5 · Standardize both risks to the target mix

General form: Risk(a) = Σ over risk groups g [P(g among eligible people) × P(outcome | strategy a, g)]. This is the g-formula in a two-stratum setting. We estimate the second term as events divided by people within each group and strategy.

Explore the weighting mechanics in the IPTW lesson →

How would I do this in an actual study?A practical workflow, the assumptions, and the limits of this example
  1. Write the protocol before comparing outcomes.State who is eligible, the actions to compare, the decision date, the outcome, the horizon and the effect of interest. Separate assignment effects from effects under sustained adherence.
  2. Map each part to data you can observe.Check dates, prior treatment, baseline covariates, outcome validity and observation windows. Record where available data cannot represent the intended trial.
  3. Construct cohorts at the decision point.Assess eligibility and confounders before treatment starts. Give both strategies the same time zero and outcome window. Do not use future survival or future treatment to invent baseline groups.
  4. Choose adjustment for the causal structure.Consider standardization, weighting or other justified estimators. Assess overlap and covariate balance. A method cannot compensate for a group in which one strategy never occurs.
  5. Address what happens after baseline.Specify switching, adherence, censoring, missing outcomes and competing events according to the question. Sustained strategies may require time-varying confounding and censoring methods.
  6. Estimate risks and uncertainty.Report absolute risks as well as contrasts. Use an uncertainty procedure appropriate to the design, including repeated records or clones when present.
  7. Challenge the conclusion.Explore unmeasured confounding, measurement error, model choices and deviations from the intended protocol. Describe assumptions and remaining limitations.

What must be true for a causal interpretation?

Conditional exchangeability: measured pre-treatment variables are sufficient to control confounding. Positivity: both strategies are possible in each relevant covariate pattern. Consistency: each strategy is well defined and its observed implementation corresponds to the intervention of interest. Measurement, follow-up and statistical models must also support the analysis. Target trial emulation does not create randomization.

Why the aligned clock is not a universal repair

This example starts all program users on one common day and has no earlier users. The aligned analysis is a trial among people still eligible on that day, with both groups observed for the same duration. A generic landmark analysis of people who started at many earlier times can answer a different question and retain selection problems.

If your question instead compares “start within 30 days of initial eligibility” with “do not start,” specify that grace-period strategy from the original time zero. Depending on the protocol, clone–censor–weight methods may be appropriate: represent initially compatible strategies, censor a copy when its observed history becomes incompatible, then address the selection introduced by censoring. Repeated eligibility may call for sequential trial emulation. Those are different designs; this playground does not implement them.

How the fictional data are generated

There are 4,000 higher risk and 6,000 lower risk initial records (or just 4,000 with the restricted eligibility option). Their pre-start 30-day event risks are 20% and 4%. If the delay is d, pre-start risk is 1 − (1 − risk₃₀)^(d/30). Only people without an early event can start.

For selection setting s, initiation probability is 0.5 + s/250 in the higher risk group and 0.5 − s/250 in the lower risk group. This stays between 0.1 and 0.9. Among eligible people, no-program 90-day risks are 40% and 12%; for horizon h they become 1 − (1 − risk₉₀)^(h/90). Program risk equals this risk multiplied by your effect setting. These two periods use deliberately distinct invented risks.

All results are expectations, not random samples. We omit uncertainty intervals because there is no sampling experiment here. No deaths or competing events occur; the outcome is the first unplanned visit. Early events are counted once, and those people cannot enter the later trial.

Primary method references