A study with missing outcomes needs an explicit assumption about the people whose outcomes were not observed. Comparing respondents answers a different question from comparing everyone assigned to a program. Covariate adjustment can make a missing-at-random analysis more plausible, but it cannot prove the assumption. A tipping-point analysis shows how much unobserved participants would need to differ for the conclusion to change. The worked study below separates point estimates, uncertainty, and the decision threshold instead of treating imputation as recovered data.
A new learning tool earns an average score of 72 among respondents; the comparison tool earns 68. The four-point advantage looks straightforward until the analyst notices that one group lost twice as many participants before the final assessment. The missing people remain part of the question the organization thought it was studying.
The researcher can omit them, fill their scores with averages, or fit a sophisticated model. Each choice carries assumptions. The important work is to expose those assumptions in terms that the reader can challenge and to measure how the conclusion responds to plausible departures.
This article develops one original hypothetical study. Every count and score below is invented for the worked example; no results are attributed to a real trial or product. The methodological sources are primary research checked for the October 10, 2026 cutoff. Their clinical examples establish statistical ideas, while the article applies those ideas to a nonclinical learning decision.
The population is defined before the missing values
Suppose 400 volunteers are randomly assigned to learning tool A or B, with 200 in each group. The prespecified outcome is a day-28 assessment scored from zero to 100. The target quantity is the difference in mean day-28 scores among all assigned participants, regardless of whether they continued using the tool.
That final clause matters. A researcher might instead ask about persistent users or about participants who would comply under either program. Those are different questions and require different designs or assumptions. Switching to “people who finished” after observing dropout changes the target without repairing the original comparison.
At day 28, tool A has 160 measured outcomes with a mean of 72. Tool B has 180 with a mean of 68. The respondent difference is four points. The response rates are 80% and 90%. Neither percentage, by itself, reveals whether the absent outcomes are higher or lower than those observed.
A useful data table records assignment, baseline characteristics, whether the outcome was obtained, and reasons for absence where available. The outcome status should distinguish an unattempted assessment, a technical failure, and a participant who completed the test but whose record could not be linked. These mechanisms can imply different follow-up actions and different assumptions.
The researcher should also verify that “missing” is a genuine absence of the target measurement. A score of zero is an observed outcome if someone attempted the assessment and earned zero. An empty assessment is not automatically equivalent. Converting absence into failure can be a defensible policy for a different composite outcome, but it should be named and defined before analysis.
The target quantity gives the rest of the workflow a stable reference. Every sensitivity scenario below estimates the same all-assigned difference while changing assumptions about absent scores. That makes the scenarios comparable; changing both the population and the missing-data assumption would make the resulting differences hard to interpret.
Respondent means leave two unknown quantities
The observed score totals are 160 multiplied by 72, or 11,520, for A, and 180 multiplied by 68, or 12,240, for B. Let the mean among A’s 40 missing participants be uA and the mean among B’s 20 missing participants be uB.
The all-assigned mean for A is 57.6 plus 0.2 multiplied by uA. For B it is 61.2 plus 0.1 multiplied by uB. Subtracting gives the target difference: negative 3.6, plus 0.2 multiplied by uA, minus 0.1 multiplied by uB.
This identity does not require an imputation algorithm. It shows exactly what the data do and do not determine. If absent people have the same mean as their group’s respondents, setting uA to 72 and uB to 68 reproduces the four-point difference. That equality is an assumption, not an observation.
If both missing groups average 50, the all-assigned means become 67.6 and 66.2. The difference shrinks to 1.4. If A’s missing group averages 40 and B’s averages 60, the means become 65.6 and 67.2. The direction reverses to negative 1.6.
The bounded score scale permits a deliberately extreme range. Letting either unknown mean vary from zero to 100 yields a possible difference from negative 13.6 to positive 16.4. Those limits are identification bounds for this arithmetic, not a confidence interval and not a claim that the extremes are plausible.
The bounds have a practical role. They demonstrate that observed respondent scores alone cannot determine the direction of the target effect. Narrower conclusions need additional assumptions, covariate information, follow-up evidence, or some combination. Presenting a narrow standard error around the respondent mean does not supply the absent information.
Explain the missingness assumption in ordinary terms
Missing completely at random means that missingness does not systematically depend on the relevant values. Missing at random, or MAR, permits missingness to depend on observed information while assuming no remaining dependence on the missing value after conditioning on that information. Missing not at random allows such remaining dependence. Sterne and colleagues explain these distinctions and why observed data cannot establish MAR versus its alternative. Sterne et al., BMJ 2009.
For the learning study, an assessment server outage affecting a randomly selected subset would support a different story from participants avoiding the test because they expected to perform badly. A workload difference recorded at baseline might help explain absence. An unrecorded feeling of failure may not.
Do not attach the labels mechanically to reasons. “Too busy” can conceal several processes. People working longer hours may have less time for both learning and testing. People struggling with the tool may describe the same experience as being busy. Observed reason categories are evidence to discuss, not a diagnostic certificate for MAR.
A comparison of baseline characteristics among respondents and nonrespondents can show that a simple complete-case analysis is questionable. It cannot establish that adjustment removes every relevant difference. If baseline skill predicts both response and final score, including it is useful. If unmeasured frustration predicts both, the remaining problem survives that adjustment.
The most informative assumption statement is conditional and specific: among participants with the same assignment and observed baseline skill, absent day-28 scores are assumed to have the same mean as observed day-28 scores. A reader can then ask whether the collected skill measure is sufficient and whether the learning tool creates new reasons to avoid testing.
This formulation also suggests better data collection. An earlier assessment, baseline time availability, and a direct follow-up reason can make the assumed relationship more credible. Collecting those variables after seeing the result is possible in a follow-up, but their timing and availability should be clear rather than folded invisibly into the original design.
Baseline strata change the arithmetic
Now reveal the study’s baseline skill categories. For A, there are 80 novices and 120 experienced participants. All 120 experienced participants respond, with mean score 76. Only 40 novices respond, with mean score 60. Their combined score total is 11,520, matching the observed mean of 72.
For B, there are 100 novices and 100 experienced participants. All experienced participants respond, with mean score 74.4. Eighty novices respond, with mean score 60. Their combined total is 12,240, matching the observed mean of 68.
| Assignment and skill | Assigned | Observed | Observed mean | Missing |
|---|---|---|---|---|
| A, novice | 80 | 40 | 60 | 40 |
| A, experienced | 120 | 120 | 76 | 0 |
| B, novice | 100 | 80 | 60 | 20 |
| B, experienced | 100 | 100 | 74.4 | 0 |
Under the stated within-stratum MAR mean assumption, A’s missing novices receive an expected mean of 60 rather than the overall respondent mean of 72. Its all-assigned mean becomes 69.6. B’s mean becomes 67.2. The estimated all-assigned difference is 2.4 points.
This is a transparent mean-standardization calculation, not a complete multiple-imputation analysis. It ignores estimation uncertainty for the moment so that the consequence of the assumption is visible. It also retains the actual assigned populations rather than forcing both groups to have the same skill composition.
The baseline composition differs because random assignment does not guarantee identical realized samples. An adjusted causal analysis could account for baseline skill in another prespecified way. The current exercise asks a narrower question: how missing outcomes change the all-assigned sample mean difference under explicit assumptions.
Notice what the stratum information accomplishes. It explains why the respondent mean in A is especially optimistic: experienced participants make up a larger portion of its measured outcomes than of its assigned population. It does not establish that missing novices are equivalent to observed novices. That remaining assumption is the next quantity to stress.
Turn the unobserved difference into a sensitivity parameter
Define delta A as the difference between A’s missing novice mean and the predicted mean of 60. Define delta B similarly for B. A negative delta says that absent novices performed worse than the observed novices used for prediction.
The full-population difference becomes 2.4 plus 0.2 multiplied by delta A minus 0.1 multiplied by delta B. The coefficients are the missing fractions in each group. They show why identical departures do not necessarily cancel: one group has more missing outcomes.
| Delta A | Delta B | Estimated difference | Meaning under this scenario |
|---|---|---|---|
| 0 | 0 | 2.4 | Within-stratum MAR reference |
| -10 | 0 | 0.4 | A’s missing novices ten points worse |
| -12 | 0 | 0.0 | Point-estimate direction tipping point |
| -10 | -10 | 1.4 | Equal deterioration, unequal missing fractions |
| 0 | -10 | 3.4 | B’s missing novices worse |
The point-estimate direction changes when 0.2 multiplied by delta A minus 0.1 multiplied by delta B equals negative 2.4. If delta B is zero, A needs a twelve-point deterioration. If delta B is negative ten, the corresponding delta A is negative seventeen.
Delta-based sensitivity analysis is a documented approach in the controlled-imputation literature. Here it is expressed as a simple mean shift, while a full model would incorporate its distributional and estimation assumptions. Cro et al., Statistics in Medicine 2020.
The value of the table is interpretability. A researcher can ask instructors whether a twelve-point difference among otherwise similar novices is plausible. That conversation is more concrete than asking whether the data are “probably MAR.” It still requires evidence and judgment; the sensitivity parameter does not become known because it has a Greek-letter name.
Respect the score bounds. A delta implying a negative mean or a mean above 100 is incompatible with this outcome scale. Individual-level imputation also needs a suitable distribution. A convenient normal model can generate impossible scores unless its implementation handles the bounded outcome appropriately.
The organizational decision can tip before the sign does
Suppose adopting A costs more, and the organization requires an expected improvement of at least two points to justify the extra expenditure. That is a hypothetical decision threshold chosen for this example, not a universal educational benchmark.
Under the MAR reference calculation, the estimated difference of 2.4 just exceeds the threshold. With delta B held at zero, a delta A of negative two reduces the estimate to two. A deterioration slightly larger than two points among A’s missing novices makes the estimated benefit insufficient for this decision even though the effect remains positive.
The sign tipping point was negative twelve. The adoption tipping point is negative two. Conflating them would make the conclusion sound much more robust than the actual decision. A result can consistently favor A in direction while fail to support its additional cost.
Likewise, statistical significance is a different threshold. An uncertainty interval crossing zero depends on variance and sample information, not just the mean identity. The aggregated example does not contain the individual variances needed to compute a valid interval for the actual hypothetical study. It would be misleading to invent a p-value from the group means alone.
A clear report can therefore distinguish three claims: A has a positive estimated effect, the uncertainty evidence supports a positive effect under the model, and the expected benefit exceeds a practical threshold. They require different calculations. The sensitivity analysis should identify which claim changes first.
For a funding committee, the practical threshold often matters most. The analyst can present the region of delta values in which adoption is justified and ask whether the evidence supports that region. The committee then understands which belief about nonrespondents carries the decision, rather than treating an imputed estimate as an unconditional finding.
A weighting cross-check exposes the same assumption
The baseline-stratified calculation can be reconstructed by weighting respondents. In A, half the novices responded, so each observed novice represents two assigned novices. All experienced participants responded, so their weight is one. The weighted score total becomes 40 multiplied by two multiplied by 60, plus 120 multiplied by 76, or 13,920. Dividing by the represented population of 200 gives 69.6.
In B, 80% of novices responded, so each observed novice receives a weight of 1.25. The weighted total is 80 multiplied by 1.25 multiplied by 60, plus 100 multiplied by 74.4, or 13,440. Dividing by 200 gives 67.2. The resulting difference matches the reference calculation exactly because both use the same stratum information and the same assumption about missing novice means.
Agreement here is an algebraic cross-check, not independent evidence for MAR. Two methods can agree because they encode the same belief. A reader should not mistake that agreement for replication by a second data source. Conversely, disagreement under an intended equivalent specification can reveal an implementation or denominator error worth investigating.
The example also makes sparse response concrete. If only two of A’s 80 novices responded, each would represent forty people under this simple weighting scheme. A small amount of observed information would then carry much of the population estimate. Adding precision to the reported decimal places would not create information about the other 78 participants.
If no novice responded at all, this within-stratum calculation could not supply a novice mean. The researcher would need information from another stratum, another group, earlier measurements, or an explicit additional assumption. Calling a software procedure cannot remove that lack of support. The report should identify where it borrows information and why the borrowing is credible.
This cross-check is useful before fitting a larger model. It gives the analyst a small case whose denominator and predicted means are transparent. A more elaborate imputation or weighting result can then be compared with it to understand which additional predictors or assumptions changed the answer.
Multiple imputation represents uncertainty, not recovered truth
A single mean-filled dataset makes every missing novice look exactly like an average respondent. It also creates the illusion that the filled values were observed. Multiple imputation instead generates several plausible completed datasets, analyzes each, and combines estimates while accounting for variation between them. The mice paper describes chained-equation modeling for partially observed data. Van Buuren and Groothuis-Oudshoorn, JSS 2011.
A small separate arithmetic illustration explains pooling. Suppose five completed datasets yield effect estimates of 2, 3, 2.5, 3.5, and 1.5, each with estimated variance four. These are toy outputs demonstrating the formula, not fitted results for the learning study and not a recommendation to use only five imputations.
The pooled estimate is 2.5. The average within-imputation variance is four. The sample variance of the five estimates is 0.625. Rubin-style total variance adds the between-imputation contribution multiplied by one plus the reciprocal of the imputation count: four plus 1.2 multiplied by 0.625, or 4.75. The corresponding standard error is approximately 2.18.
Ignoring the between-imputation component would use a standard error of two. Stacking completed datasets and pretending they are five independent sets of observed people would be worse: it would artificially multiply the apparent sample size. Each dataset is a different plausible completion of the same incomplete sample.
The imputation model must serve the analysis model. If the analysis includes interactions or nonlinear relationships, an incompatible imputation specification can distort them. Primary research on substantive-model-compatible imputation addresses this issue. Bartlett et al., accommodating the substantive model.
For the learning study, the model would need assignment, baseline skill, relevant observed predictors, and an outcome distribution appropriate to the score. Additional measurements can help predict missing scores. They cannot eliminate the need for sensitivity analysis when absence may depend on information still unobserved.
Inspect the model where missingness is concentrated
A model can fit the dominant observed population well and still extrapolate poorly to the missing population. In the worked study, all missing outcomes belong to novices. Overall fit can therefore be dominated by experienced learners who require no imputation.
Inspect predictions and residual behavior among observed novices, by assignment. Check whether the model preserves the score scale and whether predicted distributions look plausible relative to earlier measurements. A mean prediction of 60 can arise from many distributions; one narrowly concentrated around 60 implies different uncertainty from one spanning a broad range.
Sparse strata create another problem. If nearly all low-skill participants in A were absent, the observed data provide weak support for their predicted outcome. Adding more predictors can increase this sparsity. A complex model may look sophisticated while rely on extrapolation that the report does not reveal.
A diagnostic masking exercise can hide a subset of observed outcomes and assess their predictions. It tests performance on values that were actually observed under the original response mechanism. It cannot prove good prediction for nonrespondents if they differ systematically. Use the exercise to find model defects, not to certify the missingness assumption.
Track computational variation separately from substantive sensitivity. If repeated runs under the same model yield meaningfully different pooled results, the imputation procedure or its run settings need attention. If stable runs change under different delta assumptions, that is the intended sensitivity result. Combining these sources of variation into one vague statement about uncertainty obscures what the researcher can improve.
The investigation should end with a reasoned model choice. Document why the included predictors matter, where prediction support is weak, and how the sensitivity range relates to those weaknesses. A list of software settings is useful for reproduction, but it does not explain whether the missing participants were modeled credibly.
A follow-up sample can reduce the important uncertainty
The tipping analysis points to a productive collection strategy. A’s missing novices carry twice the coefficient of B’s missing novices in the full-population difference. An additional assessment effort directed toward that group may have more decision value than collecting another large sample of experienced respondents.
Suppose the organization can attempt a short follow-up with all 40 absent novices in A. A randomly selected subsample can make follow-up estimates easier to interpret than choosing only people who answer immediately. Record which people were sampled, contacted, and measured. Nonresponse to follow-up introduces another missingness problem rather than erasing the first one.
The assessment should still measure the target as closely as possible. A brief self-report of learning satisfaction cannot simply replace the day-28 score. A later test changes the measurement time. Such information may help model or bound absent outcomes, but its relationship to the original target needs explanation.
Cost can determine the design. If a full assessment is expensive, earlier scores and structured reasons might help refine the sensitivity range. If the practical decision tips with only a two-point departure, even modest evidence about that departure can be consequential. The analysis identifies where information is valuable instead of automatically demanding a larger overall sample.
Importantly, follow-up should not be selected to rescue the preferred conclusion. The protocol can specify the sampling and analysis before new outcomes are obtained. That protects the evidential role of the additional measurements and makes it easier for another researcher to assess their contribution.
A study with unavoidable missingness can still produce an informative decision. The aim is to make the unresolved quantity smaller or better bounded, not to pretend the final dataset contains every answer. Sometimes the appropriate result is that the existing evidence supports continued experimentation but not adoption at the proposed cost.
Present the result as a map of assumptions
A defensible report begins with the all-assigned target, the observed response rates, and the respondent difference. It then shows the baseline-stratified reference estimate of 2.4 and explains its within-stratum MAR assumption. The sensitivity table follows as part of the result, not as a hidden technical appendix.
State the decision threshold alongside the sign threshold. In this example, the adoption conclusion is vulnerable to a modest two-point deterioration among A’s missing novices when B’s assumption remains fixed. The effect direction requires a much larger twelve-point deterioration to reach zero. Both statements are useful because they answer different questions.
If individual-level data and fitted models are available, report uncertainty intervals for each relevant scenario with their assumptions. In this aggregate demonstration, no study-specific interval is claimed. That boundary preserves the distinction between exact arithmetic on invented summaries and inference requiring additional information.
The final sentence should describe what the evidence permits. A researcher might recommend targeted follow-up, a lower-cost pilot, or adoption only under a stated range of assumptions. The choice comes from the organization’s threshold and the plausibility of the missing-outcome scenarios. Imputation supplies a way to analyze incomplete evidence; it does not transform an assumption into a measured fact.
Sources
- Multiple imputation for missing data in epidemiological and clinical research: potential and pitfalls, Sterne and colleagues, 2009.
- mice: Multivariate Imputation by Chained Equations in R, van Buuren and Groothuis-Oudshoorn, 2011.
- Sensitivity analysis for clinical trials with missing continuous outcome data using controlled multiple imputation: A practical guide, Cro and colleagues, 2020.
- Multiple imputation of covariates by fully conditional specification: accommodating the substantive model, Bartlett and colleagues; original methodological research.