# An allele frequency counts copies, not people

Draft • Research checked 4 October 2026 • AI-assisted DomDNA editorial content. No independent clinical review or publication is claimed. General education, not individual medical advice.

Suppose a database reports a variant frequency of one per cent. It is tempting to read that as “one person in a hundred has this variant”. That is not necessarily what the number says. An allele-frequency figure usually counts copies of a particular allele against the copies examined at that location.

The denominator matters before the rarity label does. A percentage without its counted object can make an accurate database field sound like a different finding.

## Work through the count

Here is a hypothetical teaching example, not a database result. Consider an autosomal location with complete calls in 100 people, each contributing two copies. There are 200 examined copies. Ten people each carry one copy of the allele being counted, so the allele count is ten and its frequency is ten divided by 200, or five per cent.

Now imagine five people each carrying two copies instead. The allele count is still ten, and the allele frequency is still five per cent. The proportion of people carrying it has changed. The same allele frequency can therefore describe different carrier counts.

That arithmetic is deliberately limited to a simple diploid autosomal example. Missing calls, sex chromosomes and other details can change the denominator. It would be a mistake to apply “double the allele frequency” to every database field or to infer an individual's inheritance pattern from this example.

The [gnomAD production team's explanation](https://gnomad.broadinstitute.org/news/2023-11-genetic-ancestry) defines frequency using allele copies and examined chromosomes. Its grouped frequency estimates describe an aggregated dataset, not an exact count of everyone worldwide.

## Check which allele the column names

A second ambiguity is the identity of the counted allele. “Alternate”, “reference” and “minor” are not interchangeable names. The reference allele is a comparison label; it is not guaranteed to be the most common one in every dataset.

The [NCBI Variation Viewer FAQ](https://www.ncbi.nlm.nih.gov/variation/view/faq/) explains a case where a minor-allele display differs from an alternate-allele frequency. This is a documentation example about field definitions, not a reason to choose whichever number looks more alarming.

When two websites disagree, first write down the allele named next to each number. Then record the dataset and release. If one page counts A and another counts G at a two-allele location, a difference may reflect complementary counts rather than conflicting experiments. With more than two alleles, the comparison needs additional care.

## A missing denominator is not a worldwide zero

Some sites are well called in many samples; others are not. The [gnomAD v4.1 release notes](https://gnomad.broadinstitute.org/news/2024-04-gnomad-v4-1) describe reporting the examined allele number even at sites without non-reference calls in that release. That helps distinguish a searched denominator from an assumption based only on average coverage.

A variant absent from a particular dataset has not been proven absent from humanity. The dataset may not represent every ancestry, age group or clinical setting. “Not observed among the examined copies” is a more faithful description than “unique to me”.

Keep the release beside the frequency when saving a result. Later database growth can change the count without changing anyone's original DNA. Comparisons across releases should not be presented as a biological change in the person.

## Rarity is one piece of interpretation

A low frequency can be relevant to investigating a rare genetic condition, but rarity does not explain what a variant does. Many uncommon differences are not harmful. Conversely, a common allele can be associated with a complex trait without supplying a personal diagnosis.

For a practical next step, make a four-field note from one database entry: counted allele, allele count, examined allele number and dataset release. Add the named group if the figure is group-specific. If a field is unavailable, write “not supplied” rather than calculating a substitute from a different page.

This note is useful when asking why two resources show different frequencies. It is not a diagnostic worksheet. Leave clinical classification, inheritance and the fit with a person's history to the appropriate assessment.

The important improvement is modest: translate “rare” back into a count, a denominator and a dataset. That gives the percentage a meaning you can actually compare.

For general lifestyle learning, [DomDNA’s educational quiz](https://domdna.com/quiz) does not calculate allele frequencies or interpret DNA.
