z-of-a Zone of Avoidance

probability distributions experimental design

Accept First, Reveal Second

A distributional estimate can be mathematically tidy while the design that generated its observations left whole groups outside it. Distribution parameters and experimental allocation are two halves of the same claim.


The diagnostic model
You are seeing
  • A cross-sectional average is called a market estimate despite thin coverage
  • A model reports a tight confidence band after excluding hard-to-measure names
  • A loss distribution is fit to survivors and used for new issuance
  • A pooled analysis set is treated as the intended universe without an inclusion record
The mechanism
A distribution records the support, parameterization, and moment conditions of observed values, while a design fixes who can enter, what comparator applies, and how allocation is executed; a pooled statistic therefore describes the sampled assignment mechanism before it describes an intended population.
The older apparatus
Experimental design imposes pre-specified sampling and allocation, while distribution theory states support and validity ranges, preventing an average from being detached from who could enter it.
The false friend
A genuinely homogeneous population can make a pooled estimate appropriate, but balance and inclusion checks show that omitted strata do not differ materially.
The discriminating test
Compare the fitted distribution's support and group weights with the intended universe and inspect the execution flow for exclusions; mismatch identifies a wrong-population estimate.
On your own data
Publish sample flow, inclusion probabilities, support checks, and group-weighted estimates alongside every pooled statistic, then rerun results with explicit missing-strata bounds.

In 1948 the Medical Research Council compared streptomycin plus bed rest against bed rest alone. There was no placebo injection, because no other tuberculosis treatment existed to serve as an active comparator and no sham drug was made for the control arm.

The allocation was not a decoration added to an average afterwards. Random sampling numbers assigned accepted patients to treatment or control, the numbers were held in sealed envelopes, and an envelope was opened only after a patient had already been accepted.

Accept first. Reveal the assignment second.

That order is the whole point. An average calculated at the end is an average of whoever got through the door.

A distribution has a boundary before it has a mean #

A distribution record does not begin with a mean. It begins with support: the interval of values the model is permitted to produce at all.

Then the parameters, and their roles, and the fact that the roles are conventions. A gamma can be written with a shape and a scale, or with a shape and a rate, and the variance expression changes accordingly. The quantity has not moved. The notation has, and only one of the two forms matches the number sitting in the field.

Then the moments, each with its condition attached. Student’s t has a defined mean only above one degree of freedom and a defined variance only above two. Below that the variance is not large.

It is undefined.

At a single degree of freedom the t is the standard Cauchy, whose mean and variance are undefined at every parameter value, and no amount of additional sampling will produce them. Those conditions exist to stop a numerical summary from being detached from the family that made it mean anything.

The lognormal gives the quieter version of the same discipline. Its mode, median and mean are distinct and always in that order. Reporting the mean as the typical value is not a small loss of detail. It changes the thing being described, and it changes it in a known direction.

Allocation has a boundary before it has an outcome #

Experimental design puts its boundary at entry, before anything can be observed.

The first decision is a primary outcome and a direction of comparison — superiority, non-inferiority, equivalence — and these are three different questions that a single reported difference cannot distinguish between afterwards.

The vocabulary then separates things routinely compressed into the one word randomised. Simple randomization generates independent assignments. Block randomization restores the allocation ratio at the end of each block. Stratified randomization runs the procedure within subgroups such as trial site.

None of them determines who was eligible to be enrolled in the first place.

Concealment answers a different question again: whether the person enrolling the next unit can predict or influence where it will land. The MRC envelopes did that work and only that work. They did not turn every tuberculosis patient into an enrolled patient, and they did not make an eligible patient nobody reached appear in either arm.

A valid comparison is narrower than a pooled claim about everyone, and stronger for it.

Registration draws a second line, this one in time. Since 2005, a trial seeking publication in the major medical journals must be registered before enrollment, with its primary outcome, comparator and planned sample size stated in advance. A registry cross-check then compares the commitment against the published result.

That is not an argument about whether a result is attractive. It is a record of whether the question and the population stayed the ones first specified.

Finance inherits its selection #

Here the transfer breaks, and it breaks in the direction that matters.

A trial can decide allocation before any outcome occurs and preserve the decision in an envelope, a registry and a flow diagram. Finance receives an observational population in which inclusion has already depended on disclosure, distress, survival and whether anybody was measuring. There is no enrollment desk, no arm, and no moment at which a concealed sequence could have been imposed.

That makes the distributional record more important rather than less.

A loss distribution fitted to survivors has a support defined by survival. A cross-sectional average with patchy coverage has group weights defined by the coverage. A confidence interval can be narrow because the observed set is large and internally regular, while the population it is supposed to describe contains a stratum the procedure was never able to see.

The remedy is not to throw the pooled number away. It is to make it carry its full name: support, parameterisation, moment condition, sample flow, inclusion process, and the strata that are missing.

The normal distribution has maximum entropy among distributions with fixed variance. That is an excellent reason to choose a form when mean and variance are the constraints you actually have. It is not a reason to call the constraints you happen to have the population.

An envelope decides which accepted patient goes into which arm, and it is the part of the apparatus everybody can picture.

Nothing in finance opens an envelope. The population arrives pre-selected, and every one of those selections happened before the first number was written down.

So the interval can be narrow, the arithmetic can be exact, and the only honest label on the result names the population it was able to see.

Questions

What population does a pooled average actually describe?

A pooled average describes the observations that could enter the calculation under the actual inclusion process. It does not automatically describe a market, an issuer universe, or a future cohort. The fitted distribution must state its support and group weights, while the sample record must show who was screened, excluded, observed, and analyzed. A loss distribution fitted to survivors and carried to new issuance remains survival-selected unless the missing issuers are bounded explicitly.

Does randomization make an average representative of everyone?

No. Randomization balances assigned arms among units admitted to an experiment; it does not add units excluded before allocation. The 1948 Medical Research Council streptomycin trial assigned accepted patients by random sampling numbers held in sealed envelopes, but that procedure answers the treatment comparison for its enrolled population. A separate inclusion design is needed before that result can represent a wider population.

What should be checked before trusting a narrow confidence band?

The first check is whether the analysis set matches the intended universe. A narrow confidence band can be calculated from a selected set of observable survivors while hard-to-measure or distressed names never enter it. The distribution's support, its parameterization, and the weights of omitted groups must be compared with the target population before the interval is read as a market estimate.

Why are mean, median, and mode not interchangeable?

They coincide for some symmetric distributions but separate in skewed ones. For a lognormal distribution, the mode is below the median and the median is below the mean, because the right tail pulls the mean upward. A reported mean can therefore be a correct property of the sampled observations while overstating the typical observed value and saying nothing about groups excluded from the sample.

What record shows whether exclusions changed the question?

A CONSORT flow diagram follows participants from eligibility assessment through exclusion, randomization, allocated treatment, loss to follow-up, and final analysis. It makes visible whether the analyzed population differs from the enrolled population and whether the difference is concentrated in one arm. A registry cross-check then compares the pre-specified primary outcome and sample size with the final report, exposing later changes to the design.

Can a valid formula still answer the wrong question?

Yes. A moment formula is valid only over the parameter range that defines it: Student's t has variance ν/(ν−2) only when ν is greater than 2, while a Cauchy distribution has neither a defined mean nor variance. Even a formula used within its valid range answers only for the distribution fitted to the observed population; selection determines whose distribution that is.

Sources

  1. James Lind Library, "Medical Research Council (1948b)" (secondary, 2026-08-13)
  2. MathWorld, "Normal Distribution" (tertiary, 2026-08-13)
  3. NIST/SEMATECH e-Handbook of Statistical Methods, gallery of distributions (primary, 2026-08-13)
  4. R Project `medicaldata` package documentation, "RCT of Streptomycin Therapy for Tuberculosis" (secondary, 2026-08-13)
  5. Wikipedia, "Benford's law" · Wikimedia Foundation (tertiary, 2026-08-13)
  6. Wikipedia, "Blinded experiment" (2018 chronic pain blinding-assessment review; antidepressant and acupuncture unblinding rates) · Wikimedia Foundation (tertiary, 2026-08-13)
  7. Wikipedia, "Cauchy distribution" · Wikimedia Foundation (tertiary, 2026-08-13)
  8. Wikipedia, "Chi-squared distribution" · Wikimedia Foundation (tertiary, 2026-08-13)
  9. Wikipedia, "Clinical trial registration" (ICMJE 2005 policy, registry sizes as of August 2013) · Wikimedia Foundation (tertiary, 2026-08-13)
  10. Wikipedia, "CONSORT Statement" (history, flow diagram, 2010/2025 checklist counts, endorsing journals) · Wikimedia Foundation (tertiary, 2026-08-13)
  11. Wikipedia, "Differential entropy" · Wikimedia Foundation (tertiary, 2026-08-13)
  12. Wikipedia, "Gamma distribution" · Wikimedia Foundation (tertiary, 2026-08-13)
  13. Wikipedia, "Log-normal distribution" · Wikimedia Foundation (tertiary, 2026-08-13)
  14. Wikipedia, "Multivariate normal distribution" · Wikimedia Foundation (tertiary, 2026-08-13)
  15. Wikipedia, "Pareto distribution" · Wikimedia Foundation (tertiary, 2026-08-13)
  16. Wikipedia, "Publication bias" (significant-vs-null publication-rate figure, pre-registration as remedy) · Wikimedia Foundation (tertiary, 2026-08-13)
  17. Wikipedia, "Randomized controlled trial" (2011 review of RCT funding sources) · Wikimedia Foundation (tertiary, 2026-08-13)
  18. Wikipedia, "Student's t-distribution" · Wikimedia Foundation (tertiary, 2026-08-13)
  19. Wikipedia, "Zipf's law" · Wikimedia Foundation (tertiary, 2026-08-13)