# An imputed genotype is a prediction, not another measurement

Draft • Research checked 4 October 2026 • AI-assisted DomDNA editorial content. No independent clinical review or publication is claimed. General education, not individual medical advice.

A DNA file can gain millions of rows without the laboratory taking another sample. That is possible when an analysis uses imputation: a statistical process that estimates unobserved genotypes from measured data and reference patterns. The extra rows can be useful in research. They need a different label from directly observed results.

Before comparing two files, look for a description of how each was produced. A larger file might contain more measurements, more predictions, a different export format or some combination. File size alone cannot tell you which.

## Where the extra information comes from

The primary methods paper by [Howie and colleagues](https://pubmed.ncbi.nlm.nih.gov/22384356/) describes using reference haplotypes to predict genotypes not observed in a study dataset. A reference panel supplies examples of genetic patterns; it is not an additional sample from the person whose file is being analysed. The paper concerns research methods and performance, not permission to treat any predicted rare variant as a clinical result.

The distinction survives improvements in software. The [Minimac4 project](https://github.com/statgen/minimac4) describes an implementation of genotype-imputation algorithms. Faster processing or a newer version can change what researchers can analyse, but does not turn inference into a direct laboratory observation.

For an everyday reader, the useful question is: does this line represent a measured call, an inferred call or an expected allele count? Those are different kinds of information even when an export presents them in similar-looking columns.

## A hypothetical file comparison

Imagine that a laboratory export contains a measured result at marker A and leaves marker B unmeasured. A later analysis estimates the result at B using nearby measured markers and a reference dataset. In this editorial example, a summary export prints B beside A without an obvious distinction.

Nothing fraudulent has necessarily happened. The problem is that the display removed information you need. You cannot compare the two rows as if both had been observed under the same laboratory conditions.

A more useful comparison would preserve a provenance column. It would say which result was observed, which was imputed, the reference panel and software version, and the available quality measure. Do not invent a certainty percentage where the export supplies none. A blank provenance field should remain a question, not become “directly measured” by default.

This also explains why a decimal dosage should not automatically be rounded into a clinical genotype. An expected allele count summarises uncertainty in a model. Rounding can hide that uncertainty while producing a deceptively tidy answer.

## Quality checks belong before interpretation

The [EMBL-EBI eQTL Catalogue methods](https://www.ebi.ac.uk/eqtl/Methods/) describe genotype quality control and imputation as steps in a research workflow. That ordering matters: researchers do not simply add predicted rows and assume every row is equally useful. Their downstream question and filtering rules shape which data can be used.

These sources do not establish one universal quality cutoff for a consumer export. Different methods can supply different metrics, and suitability depends on what the analysis is trying to do. A quality label useful for a population association analysis is not automatically a clinical confirmation standard.

If someone advertises a report as analysing “more variants”, ask what proportion was directly observed and what proportion was inferred. Ask whether the report makes that distinction visible at the individual result level. An answer about total variant count is not an answer about provenance.

## Keep the original and the derived file separate

A practical next step is to make a short file inventory before interpreting anything: original provider export, derived analysis, creation date and method description. Store the original unchanged. Give any derived file a name that makes its status obvious, rather than replacing the original with the apparently more complete version.

If a derived report raises a health concern, preserve the exact row and its provenance for a qualified genetics service. Do not make a medicine, screening or family-planning decision from an inferred row alone. The purpose of this inventory is to retain the uncertainty that a simplified display may have hidden.

You do not need to learn an imputation algorithm to ask a good question about it. You need to know which information was measured, which was estimated and which decision the estimate was built to support.

For general lifestyle learning, [DomDNA’s educational quiz](https://domdna.com/quiz) uses answers, not DNA-file analysis.
