Skip to content
DomDNA LAB

Research library · Genomics

An imputed genotype is a prediction, not another measurement

DomDNA editorial resources · Released · Research 2026-10-04 · 4 min read

Draft • Research checked 4 October 2026 • AI-assisted DomDNA editorial content. No independent clinical review or publication is claimed. General education, not individual medical advice.

A DNA file can gain millions of rows without the laboratory taking another sample. That is possible when an analysis uses imputation: a statistical process that estimates unobserved genotypes from measured data and reference patterns. The extra rows can be useful in research. They need a different label from directly observed results.

Before comparing two files, look for a description of how each was produced. A larger file might contain more measurements, more predictions, a different export format or some combination. File size alone cannot tell you which.

Where the extra information comes from

The primary methods paper by Howie and colleagues describes using reference haplotypes to predict genotypes not observed in a study dataset. A reference panel supplies examples of genetic patterns; it is not an additional sample from the person whose file is being analysed. The paper concerns research methods and performance, not permission to treat any predicted rare variant as a clinical result.

The distinction survives improvements in software. The Minimac4 project describes an implementation of genotype-imputation algorithms. Faster processing or a newer version can change what researchers can analyse, but does not turn inference into a direct laboratory observation.

For an everyday reader, the useful question is: does this line represent a measured call, an inferred call or an expected allele count? Those are different kinds of information even when an export presents them in similar-looking columns.

A hypothetical file comparison

Imagine that a laboratory export contains a measured result at marker A and leaves marker B unmeasured. A later analysis estimates the result at B using nearby measured markers and a reference dataset. In this editorial example, a summary export prints B beside A without an obvious distinction.

Nothing fraudulent has necessarily happened. The problem is that the display removed information you need. You cannot compare the two rows as if both had been observed under the same laboratory conditions.

A more useful comparison would preserve a provenance column. It would say which result was observed, which was imputed, the reference panel and software version, and the available quality measure. Do not invent a certainty percentage where the export supplies none. A blank provenance field should remain a question, not become “directly measured” by default.

This also explains why a decimal dosage should not automatically be rounded into a clinical genotype. An expected allele count summarises uncertainty in a model. Rounding can hide that uncertainty while producing a deceptively tidy answer.

Quality checks belong before interpretation

The EMBL-EBI eQTL Catalogue methods describe genotype quality control and imputation as steps in a research workflow. That ordering matters: researchers do not simply add predicted rows and assume every row is equally useful. Their downstream question and filtering rules shape which data can be used.

These sources do not establish one universal quality cutoff for a consumer export. Different methods can supply different metrics, and suitability depends on what the analysis is trying to do. A quality label useful for a population association analysis is not automatically a clinical confirmation standard.

If someone advertises a report as analysing “more variants”, ask what proportion was directly observed and what proportion was inferred. Ask whether the report makes that distinction visible at the individual result level. An answer about total variant count is not an answer about provenance.

Keep the original and the derived file separate

A practical next step is to make a short file inventory before interpreting anything: original provider export, derived analysis, creation date and method description. Store the original unchanged. Give any derived file a name that makes its status obvious, rather than replacing the original with the apparently more complete version.

If a derived report raises a health concern, preserve the exact row and its provenance for a qualified genetics service. Do not make a medicine, screening or family-planning decision from an inferred row alone. The purpose of this inventory is to retain the uncertainty that a simplified display may have hidden.

You do not need to learn an imputation algorithm to ask a good question about it. You need to know which information was measured, which was estimated and which decision the estimate was built to support.

For general lifestyle learning, DomDNA’s educational quiz uses answers, not DNA-file analysis.

Original source and access ledger

  1. Howie et al.: genotype imputation with thousands of genomes

    sourceDate: 2011

    type: Primary methods study

    population: Genetic research datasets and reference haplotypes

    endpoint: Imputation performance

    supportedClaimAndLimit: Reference haplotypes predict unobserved genotypes; study performance is not clinical confirmation of a personal inferred row.

    fundingAndConflicts: Study-specific funding and conflicts not fully extracted from accessed abstract or indexed text; no independence or absence-of-conflict claim.

    accessEvidence: Official source page, documentation or primary abstract text accessed via web search/open on 2026-10-04. Indexed text was used where direct opening was incomplete; no complete study-methods or supplement appraisal claimed.

    researchDate: 2026-10-04

    correctionStatus: Access-date source check only; no comprehensive correction, retraction, policy-version or guideline surveillance claimed.

    sourceWordLimit: 200

    quoteWords: 0

    sourceUseBudget: 200-word aggregate source-derived limit across article and adaptations; no verbatim quotations. Short central claims retained; hypothetical examples and administrative suggestions are original editorial material, not study findings.

  2. Minimac4 official project

    sourceDate: Accessed 2026-10-04; release not pinned

    type: Official software project documentation

    population: Users of genotype-imputation software

    endpoint: Algorithm implementation

    supportedClaimAndLimit: Implementation of imputation algorithms; no claim that inference becomes laboratory observation or that a particular release was run.

    fundingAndConflicts: Official resource or project documentation; not a personal-benefit trial or endorsement. Institutional authorship does not establish clinical review of this draft.

    accessEvidence: Official source page, documentation or primary abstract text accessed via web search/open on 2026-10-04. Indexed text was used where direct opening was incomplete; no complete study-methods or supplement appraisal claimed.

    researchDate: 2026-10-04

    correctionStatus: Access-date source check only; no comprehensive correction, retraction, policy-version or guideline surveillance claimed.

    sourceWordLimit: 200

    quoteWords: 0

    sourceUseBudget: 200-word aggregate source-derived limit across article and adaptations; no verbatim quotations. Short central claims retained; hypothetical examples and administrative suggestions are original editorial material, not study findings.

  3. EMBL-EBI eQTL Catalogue methods

    sourceDate: Accessed 2026-10-04; page date not shown

    type: Official research-resource methods

    population: Catalogue research datasets

    endpoint: Genotype quality-control and imputation workflow

    supportedClaimAndLimit: Quality control and imputation are research-processing steps; does not define a universal consumer or clinical cutoff.

    fundingAndConflicts: Official resource or project documentation; not a personal-benefit trial or endorsement. Institutional authorship does not establish clinical review of this draft.

    accessEvidence: Official source page, documentation or primary abstract text accessed via web search/open on 2026-10-04. Indexed text was used where direct opening was incomplete; no complete study-methods or supplement appraisal claimed.

    researchDate: 2026-10-04

    correctionStatus: Access-date source check only; no comprehensive correction, retraction, policy-version or guideline surveillance claimed.

    sourceWordLimit: 200

    quoteWords: 0

    sourceUseBudget: 200-word aggregate source-derived limit across article and adaptations; no verbatim quotations. Short central claims retained; hypothetical examples and administrative suggestions are original editorial material, not study findings.

Download exact original article Markdown · Download exact source and unsent social pack

Finished companion media

Silent films, slides and source captions on the separate dashboard. Original preview and historical product limits remain in their captions. No social campaign has been sent.