# Would the measured difference matter outside the statistical test?

Draft • Research checked 4 October 2026 • AI-assisted DomDNA editorial content. No independent clinical review or publication is claimed. General education, not individual medical advice.

“Statistically significant” tells you something about an analysis. It does not tell you whether the change would be noticeable, useful or worth the effort involved.

That distinction becomes especially important when a study uses a score. A small difference on a large scale can look impressive in a headline that omits the units. A substantial-looking difference can also remain uncertain when the study is small or the measurement is poorly suited to the question.

## Begin with the scale

Before judging a score change, find what the score represents, its direction and the time period involved. A higher value might mean better function on one instrument and worse symptoms on another. Different questionnaires with similar names may measure different concepts.

[Cochrane's interpretation guidance](https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-15) separates effect size, uncertainty and importance. Those belong together. A significance label cannot stand in for the size of a benefit, and a favourable average cannot stand in for how participants experienced it.

The measurement needs its own evidence. [FDA's final guidance on fit-for-purpose clinical outcome assessments](https://www.fda.gov/regulatory-information/search-fda-guidance-documents/patient-focused-drug-development-selecting-developing-or-modifying-fit-purpose-clinical-outcome) addresses suitability for a defined research purpose. A score does not become validated simply because it uses a familiar zero-to-one-hundred format.

## An invented result in everyday terms

Suppose a hypothetical work-scheduling study uses a self-rated convenience scale from zero to one hundred. One group averages two points higher than the other. The report calls the difference statistically significant.

You still do not know whether two points means anything participants would value. Perhaps it corresponds to a barely detectable wording preference. Perhaps the score is sensitive to a practical improvement that is difficult to describe briefly. The scale's interpretation matters.

Now imagine the same report says participants saved time completing a defined administrative task. Minutes have a more familiar unit, but importance still depends on frequency, effort, costs and uncertainty. Saving two minutes once is a different practical proposition from saving two minutes repeatedly with no additional burden.

These examples are hypothetical. The numbers are not meaningful-change thresholds, and the convenience score is not a validated clinical instrument.

## Meaningful change is a measurement question

The [1989 paper by Jaeschke, Singer and Guyatt](https://pubmed.ncbi.nlm.nih.gov/2691207/) is an early methodological source on identifying a minimal clinically important difference. Only its bibliographic record and title were accessible in this research pass; no numerical rule or detailed finding from it is asserted here.

A meaningful-change estimate from one instrument and population should not be transferred automatically to another. The context, concept being measured and method used to establish importance can differ.

There is also a distinction between a meaningful change for an individual and a difference between group averages. A threshold discussed for one purpose cannot simply be pasted onto the other. If a summary uses a threshold, look for what it was developed to interpret.

## Costs and unwanted effects stay in the picture

A useful benefit is not a benefit measured in isolation from everything else. Time, inconvenience, unwanted experiences and expense may change how people value the same outcome. A trial's statistical result does not decide those preferences for everyone.

This does not mean that a researcher must produce one universal answer to “worth it”. It means the report should describe the outcome clearly enough for practical interpretation, with uncertainty and relevant trade-offs still visible.

Avoid the opposite shortcut too: a non-significant result does not prove that the true effect is trivial. A wide interval may include effects that would matter as well as little or no difference. Imprecision and practical unimportance are not synonyms.

## What you can do with this

Rewrite the headline in units. Instead of “significant improvement”, write “a reported average difference of X on this named scale, at this follow-up point”. Then look for evidence explaining what that difference means in the studied population.

If the interpretation is unavailable, retain the measured result without inventing a meaningfulness claim. “Importance not established in the material I read” is a useful note.

That next step can change what you seek from further research. Another significance label adds little if the unresolved question is whether the outcome captures something people value. Evidence about the instrument, direct functional outcomes and participant priorities may be more informative than a smaller P value.

For an educational starting point, the [DomDNA lifestyle quiz](https://domdna.com/quiz) records lifestyle answers, not DNA analysis or a clinical diagnosis.
