Covariate shift assumption — single invariant contrasts with → Training on multiple distributions
These sit at opposite ends of how much theoretical commitment a distributional-shift remedy requires. Covariate shift stakes everything on one explicit statistical claim — that p(y|x) does not change — and inherits the paper's own warning that this can fail silently (concrete-problems, §"Well-specified models: covariate shift and marginal likelihood.", p. 17). Training on multiple distributions makes no such claim at all: it is described as an engineering approach, training on several distinct distributions in the hope that a model good on all of them generalizes to a new one, with no assumption being asserted that could be true or false (concrete-problems, §"Training on multiple distributions.", p. 18). The cost of dropping the assumption is that there is no longer any theoretical account of why or when this should work; the paper's evidence for it is a practitioner's empirical report from speech recognition, not a proof.