← All notes

Statistical note

Difference-in-differences

Estimate differential before-after change under a defensible parallel-trends assumption.

6 min read

Which intervention, comparison, and assumptions make a causal effect identifiable from the observed data? This note turns that question into a practical analysis workflow.

Conceptual guide to Difference-in-differences

Core idea

Estimate differential before-after change under a defensible parallel-trends assumption. The method is useful only when its target, data-generating assumptions, and unit of analysis match the scientific question.

τ=(YˉT,postYˉT,pre)(YˉC,postYˉC,pre)\tau=(\bar Y_{T,post}-\bar Y_{T,pre})-(\bar Y_{C,post}-\bar Y_{C,pre})

The equation is a compact statement of the target; it is not a substitute for checking design and data quality. State what each observation represents, how it entered the sample, and which sources of dependence remain.

Practical workflow

  1. Write the scientific question and estimand in one sentence.
  2. Identify biological units, technical replicates, nesting, repeated measurements, and exclusions.
  3. Visualise raw observations and group structure before fitting the method.
  4. Check assumptions using design knowledge, diagnostic plots, and sensitivity analyses.
  5. Report the estimate, uncertainty, effect magnitude, sample sizes at every level, and limitations.
import numpy as np

# Standardised mean difference before adjustment.
treated = np.array([1.2, 1.4, 1.8, 2.0])
control = np.array([0.8, 1.0, 1.1, 1.3])
pooled_sd = np.sqrt((treated.var(ddof=1) + control.var(ddof=1)) / 2)
print((treated.mean() - control.mean()) / pooled_sd)

Interpretation and failure modes

Do not read a software output as an automatic scientific conclusion. Ask whether dependence was modeled, whether preprocessing used information from validation data, whether missingness or selection is informative, and whether the reported uncertainty covers every level of sampling. Prefer estimates and intervals to threshold-only language.

Microscopy case study

Estimate a software-upgrade effect using treated and unaffected microscopes only if their pre-upgrade trends are comparable.

Connections

Use the tags above to continue to notes on design, assumptions, uncertainty, diagnostics, and complementary methods. A robust analysis normally combines several concepts rather than selecting one test in isolation.