← Notes

Glossary

Use statistical language precisely.

Concise definitions for terms that are often blurred together. Search here, then open the connected notes for derivations and applications.

Alpha (α)

The prespecified long-run probability of a Type I error for a testing procedure.

Alternative hypothesis

A defined set of parameter values contrasted with the null hypothesis.

Bias

Systematic difference between an estimator’s expected value and its target.

Bootstrap

Resampling observations with replacement to approximate an estimator’s sampling distribution.

Calibration

Agreement between predicted probabilities or measured values and observed frequencies or references.

Censoring

Partial observation of a time-to-event outcome; the event time is known only to lie beyond or within limits.

Collider

A variable caused by two other variables; conditioning on it can create a spurious association.

Confidence interval

An interval produced by a procedure with stated repeated-sampling coverage under its assumptions.

Confounder

A common cause of an exposure and outcome that can bias an unadjusted causal comparison.

Consistency

Convergence of an estimator to its target as information grows.

Contrast

A prespecified comparison, often a weighted combination of group means or model parameters.

Covariance

Joint variation of two random variables around their expectations.

Credible interval

A region containing a stated posterior probability under a Bayesian model.

Cross-validation

Repeated separation of training and validation observations to estimate out-of-sample performance.

Effect size

A quantitative magnitude of difference, association, or model effect in interpretable or standardized units.

Estimand

The precisely defined population quantity a study intends to learn.

Estimate

The numerical value produced by applying an estimator to observed data.

Estimator

A rule mapping sample data to an estimate of an unknown quantity.

Exchangeability

A symmetry assumption under which joint probability is invariant to permissible reordering of units.

Experimental unit

The smallest unit independently assigned to a treatment or sampled for the target comparison.

False discovery rate

Expected proportion of false rejections among rejected hypotheses under the control procedure.

Fisher information

Expected curvature of the log-likelihood; a measure of information about a parameter.

Hazard

Instantaneous event rate among units that remain event-free immediately before a time.

Identifiability

Whether distinct parameter or causal values imply distinguishable observed-data distributions.

Independence

A factorization property stating that knowing one random quantity does not alter another’s distribution.

Interaction

A situation in which one factor’s effect differs across levels of another factor.

Likelihood

The observed-data probability or density considered as a function of unknown parameters.

Link function

A transformation connecting a modeled conditional mean to a linear predictor in a generalized model.

Loss function

A numerical penalty used to compare predictions, decisions, or parameter estimates with targets.

MAR

Missing at random: missingness may depend on observed data but not missing values after conditioning.

MCAR

Missing completely at random: missingness is independent of observed and unobserved values.

Mediator

A variable lying on a causal pathway from exposure to outcome.

MNAR

Missing not at random: missingness still depends on unseen values after conditioning on observed information.

Model

A set of probability distributions or structural relations proposed for how data arise.

Multiple testing

Simultaneous inference on several hypotheses requiring explicit control of an error criterion.

Null hypothesis

A precisely specified set of parameter values assessed by a hypothesis test.

Odds ratio

Ratio of two odds; not generally equal to a risk ratio, especially for common outcomes.

Overdispersion

Outcome variance exceeding that implied by a chosen probability model, often the Poisson model.

p-value

Under a specified null model, the probability of a test statistic at least as incompatible as observed.

Parameter

A fixed or modeled quantity indexing a population distribution or data-generating process.

Permutation test

A reference distribution obtained by rearrangements justified by exchangeability under the null.

Posterior

A Bayesian probability distribution for unknown quantities after combining prior and likelihood.

Power

Probability that a testing procedure rejects the null under a specified alternative.

Precision

Either inverse uncertainty of an estimate or, in classification, TP divided by predicted positives; context matters.

Prediction interval

A range intended to contain a future observation under the fitted model and sampling process.

Prior

A Bayesian probability model for unknown quantities before conditioning on current data.

Pseudoreplication

Treating dependent technical observations as independent evidence for a biological comparison.

Random effect

A group-specific modeled deviation drawn from a population distribution.

Randomization

Chance-based treatment allocation used to break systematic links with potential outcomes.

Recall

True-positive rate: the proportion of reference-positive cases correctly detected.

Regularization

A penalty or prior that constrains model complexity to improve stability or generalisation.

Residual

Observed outcome minus its fitted value, or a model-specific analogue used for diagnostics.

Robustness

Limited sensitivity of a procedure or conclusion to outliers or plausible assumption deviations.

Sampling distribution

The probability distribution of a statistic across hypothetical repeated samples.

Sensitivity analysis

Reanalysis under alternative plausible assumptions to assess conclusion stability.

Specificity

True-negative rate: the proportion of reference-negative cases correctly classified.

Standard deviation

Spread of observations around their mean in the measurement’s units.

Standard error

Estimated sampling variability of an estimator, not variability among raw observations.

Sufficient statistic

A statistic retaining all sample information about a parameter within a specified model.

Type I error

Rejecting a true null hypothesis.

Type II error

Failing to reject a false null hypothesis under a specified alternative.

Variance

Expected squared deviation of a random variable from its mean.

Zero inflation

More observed zeros than a baseline count model predicts, potentially reflecting a mixture process.