Lumina
Applied Health Statistics · Module HEALTH.1
1/13
EN PT
HEALTH.1 · planning

From question to analysis plan

You will learn to choose the health question a table can answer, then write a short plan that makes the answer checkable. The outcome is not a test name: it is a stated comparison, an estimated effect, and an honest account of uncertainty.

Start with the clinic decision

A clinic adds six weeks of discharge support. Its follow-up table has one row per person: group, symptom score at discharge, and symptom score six weeks later. The team must decide what to report before reading a result.

Choose Question A when the clinic’s decision is which group had lower symptoms at the scheduled follow-up. Question B describes change within people; it does not compare the groups. Question C compares change, so it is appropriate only when that is the decision the clinic wants to make. The same table can support all three, but it cannot make them the same question.

what you’ll be able to doTurn a concrete health decision into one analysis plan: who is compared, what outcome is measured, when it is measured, what one row represents, and what result will answer the question.

How the same question may appear in a paper

You have already stated the population, comparison, outcome, and time in ordinary language. Some papers compress those parts into PICO: Population, Intervention, Comparator, Outcome. Here that means eligible adults leaving the clinic; the support program; usual care; and symptom score at six weeks. For a non-randomized exposure question, PECO changes Intervention to Exposure. Use the acronym as a reading aid, not as a substitute for the complete question.

The estimand is the exact effect or question the data must estimate. For Question A: the mean symptom-score difference at six weeks, support minus usual care, among eligible adults. An estimate is the number calculated from this sample for that target. The unit of analysis is what one row contributes: here, one participant, not one clinic visit. A confidence interval (CI) is the range of effect sizes compatible with the data and stated assumptions; it shows uncertainty, not whether the program is clinically proven.

clinic decisionsymptoms at 6 weekschosen estimandsupport − usual careone participant · 6 weeksreportestimate + CI

Write a plain-language analysis plan

A plan records the choices before a result can tempt you to change the question. Begin with ordinary language:

resolved planPopulation: eligible adults leaving this clinic.
Comparison: discharge support versus usual care.
Outcome and time: symptom score six weeks after discharge.
Unit: one participant.
Effect: mean score in support minus mean score in usual care.
Missing scores: record them and state the rule before analysis.

The plan chooses Question A. “Before and after” would answer Question B instead; “change in each group” would answer Question C unless it explicitly compares those changes. A negative support-minus-usual-care estimate would mean lower average symptoms in the support group; the CI would show how uncertain that estimated difference remains. This is a teaching decision, not evidence that the program works.

the clinic now asks whether the two groups changed differently from discharge to six weeks. What must the plan rename?
Rename the estimand as the between-group difference in change. Question A compares groups at six weeks; it cannot silently become a change-score question without changing the quantity, data structure and interpretation.

The canonical 12-item method audit

This numbered checklist is the stable audit used throughout the track. Answer each item in plain language before treating a method or result as ready; a missing or conflicting answer requires review.

  1. Target effect (estimand). State the exact contrast or quantity the analysis is meant to estimate.
  2. Population and unit. Name who can contribute and what one independent observation represents.
  3. Outcome scale. Define what is measured, its unit or range, and what higher and lower values mean.
  4. Groups and times. Name the groups, conditions, and time points that form the comparison.
  5. Dependence. Identify repeated measurements, shared participants, sites, families, or clusters that make observations related.
  6. Sampling and bias. Explain how records enter the analysis and how selection, exclusions, or missingness could distort the target effect.
  7. Data defects. Check impossible values, duplicates, units, dates, and cross-field inconsistencies while retaining an audit trail.
  8. Assumptions. State the model and missing-data assumptions and how their consequences will be challenged.
  9. Pre-specification and multiplicity. Name the primary claim before results and how the full family of comparisons will be handled.
  10. Direct-effect method. Choose a method that directly estimates item 1 rather than a convenient proxy.
  11. Effect, interval, and diagnostics. Report magnitude, uncertainty, and relevant model or data checks together.
  12. Causal and general limits. State what the design cannot establish about cause and about populations beyond the observed sample.

The item number is part of the contract: later modules link back here instead of silently renumbering the audit.

Try a new decision without code

A community clinic must choose between phone and text-message check-ins for people starting blood-pressure medicine. Its table has one row per person: check-in format and whether the person is taking the medicine as prescribed 14 days later. Write five lines: population, comparison, outcome and time, unit, and effect.

one defensible plan
Population: people starting the medicine at this clinic. Comparison: phone minus text-message check-ins. Outcome and time: taking medicine as prescribed at day 14. Unit: one person. Effect: difference in the proportion taking medicine as prescribed. This answers the between-format question at day 14. It is not a before-and-after comparison: the stated decision compares two formats at one scheduled follow-up, not each person’s change from a baseline measure.

Keep later complications in view, not at the entrance

Later modules add data integrity, missing-data assumptions, diagnosis, and reporting. In trials, an intercurrent event—for example, stopping treatment or starting rescue therapy—can change which question is being answered. ICH E9(R1) is a public extension for that planning problem. For now, record such events and do not silently switch the question after results appear.

Optional technical practice: AURORA-30

AURORA-30 is a fictional practice dossier made from small synthetic fixtures. It contains no real participants, institutions, interventions, or clinical outcomes, and results from it are not clinical evidence. You do not need it, JSON, or code to understand this module. A count or value carries forward only when the next module names the same file and field; the clinic example at six weeks above is a separate no-code example.

timeline contract · the dates are not aliases
  1. Baseline → day 30: the planning, data-audit, descriptive, and uncertainty fixtures in HEALTH.1–4 and the worked contract in HEALTH.5 use symptom_score_day_30 or the day-30 mean difference. HEALTH.5's bundled linter is an independent mean-change contract fixture.
  2. Baseline → week 4: HEALTH.6–8 use a separate comparison-and-model snapshot with w4_symptom and response_w4. Week 4 means 28 days here; it is not a renamed day-30 measurement.
  3. Day 0 → day 90: HEALTH.9 uses a distinct time-to-relapse outcome. A participant can have an event or last confirmed contact before day 90; that does not create a missing day-30 symptom score.
  4. Other views: HEALTH.10–11 use diagnostic and agreement measurements, while HEALTH.12–13 audit or assemble their artifacts. Those fields do not become extra symptom follow-ups.

The optional lab below checks the day-30 planning branch: whether a synthetic plan names its target, unit, time, and evidence.

{
  "population": "eligible synthetic participants",
  "comparison": "support minus usual care",
  "outcome": "symptom score at day 30",
  "unit": "participant",
  "estimand": "mean difference at day 30"
}
▶ optional technical practice · runnable lab

analysis-plan-contract

Validate a synthetic analysis plan. It is practice only; no clinical data, account, or download is required.

labs/analysis-plan-contract/ · python analysis_plan_contract.py --input samples/aurora30_plan.json --out analysis_plan.json --format json
what should make you pause before choosing a method?
Any unresolved choice about population, comparison, outcome, time, unit, or missing-score rule. A method cannot repair a question the plan has not stated.
key takeaways