Read the genetic history written in your populations
Microsatellites, AFLP bands, chloroplast sequences or morphological traits — the same guided workflow takes you from the raw genotype table to diversity indices, Hardy–Weinberg tests, AMOVA, genetic distances, Bayesian clustering and haplotype networks, and ends in editable, journal-ready figures. No programming, no installation.
How it works
A guided route, in the order a population genetics paper is actually written. Each step explains the idea in plain language, checks whether your data can support it, and recommends what to do next.
Try it: watch drift, migration and selection at work
Six populations start with the same allele at p = 0.5. Every generation the model applies selection, migration, mutation and then draws 2Ne gene copies at random — the Wright–Fisher recursion, exactly as written in the textbooks. The right-hand panels are computed from the simulated frequencies: gene diversity within demes (HS), in the whole metapopulation (HT) and the FST that grows between them. Turn migration down and watch the populations drift apart.
What's inside
Nine analysis blocks, built in order and released one at a time. Click a card to jump to it.
Every kind of data a plant study produces
The marker decides which statistics are even defined. PopGeneticsPro types your loci on import and then offers only the analyses that make sense for them — no silently wrong heterozygosities from an AFLP gel, no χ² test on haplotypes.
Methods covered
The classical toolkit of population genetics — diversity, equilibrium, differentiation and Bayesian clustering — plus the plant-specific analyses rarely found together: mating system, fine-scale spatial genetic structure and PST (or QST) versus FST.
What it brings together
One guided path from the data sheet to the figure. The formulas come from the published methods, and every one is cited where it is used.
The theory, in plain language
Enough to run the analyses and defend them in a seminar. Every section names the papers the formulas come from; the full list is at the bottom of this page.
What population genetics actually asks start here
A species is rarely one big interbreeding unit. It is a set of populations that exchange genes at some rate, lose variation at some rate, and accumulate differences at some rate. Population genetics measures those three rates from a sample of individuals, and asks four questions in this order:
- How much genetic variation is there, and is it the same in every population? (Block 3)
- Is that variation organised the way random mating predicts, or is there inbreeding, selfing, clonality or a scoring problem? (Block 4)
- How is the variation distributed — mostly within populations, or between them? (Blocks 5–7)
- What history produced that pattern — gene flow, isolation, a bottleneck, an expansion? (Block 9)
Hardy–Weinberg: the null model everything is measured against Block 4
With two alleles at frequencies p and q, random mating in a large population with no selection, mutation or migration gives genotype frequencies p², 2pq and q² — and they stay there, generation after generation. That is the Hardy–Weinberg principle, and it is useful precisely because real populations depart from it.
A positive F (fewer heterozygotes than expected) is the everyday result in plants, and it has several possible causes that the analysis must separate:
- Selfing or biparental inbreeding — the biological explanation, consistent across all loci.
- The Wahlund effect — you pooled two genetically distinct groups into one "population".
- Null alleles — a primer site mutation hides one allele, so true heterozygotes are scored as homozygotes. Affects one locus, not all of them.
The four forces: drift, migration, mutation, selection simulator
Drift is the random sampling of gametes between generations. It is not a small correction: in a population of effective size Ne, heterozygosity decays as Ht = H0(1 − 1/2Ne)t, and alleles are eventually lost or fixed by chance alone. Small and fragmented populations lose variation fast.
Migration pushes the other way. In Wright's island model the equilibrium differentiation is FST ≈ 1/(1 + 4Nem), which is why one successful migrant per generation (Nem = 1) is enough to keep populations from diverging by drift.
Mutation creates new variation, slowly (10⁻⁵–10⁻³ per locus per generation for microsatellites, much lower for sequences). Selection changes frequencies systematically, but only at the loci it acts on — which is what makes neutral markers useful for reading demography.
F-statistics and AMOVA: where the variation sits Block 5
Wright's three F-statistics describe the same sample at three levels: FIS compares an individual with its own population (inbreeding), FST compares populations with the total (differentiation), and FIT compares an individual with the total.
With highly polymorphic markers such as microsatellites, HS can be so high that GST is mathematically capped far below 1 — two populations sharing no alleles at all may still give GST = 0.1. That is the reason to report G′ST or Jost's D alongside it, not instead of it.
AMOVA (Excoffier et al. 1992) does the same job as an analysis of variance on a distance matrix, and it is the right tool when your design is hierarchical: individuals within populations, within regions. It gives a percentage of variation at each level and tests each one by permutation — and because it works from distances, it handles dominant markers and haplotypes as easily as microsatellites.
Dominant markers: what AFLP and ISSR can and cannot tell you binary data
A dominant band is present or absent. The heterozygote looks exactly like the dominant homozygote, so Ho cannot be observed and FIS cannot be estimated from the data alone. Allele frequencies must be inferred from the frequency of the null phenotype:
- Analyses that still work: ΦPT and AMOVA, Nei's gene diversity, Jaccard/Dice distances, PCoA, Bayesian clustering, Mantel tests, spatial autocorrelation.
- Analyses that do not: observed heterozygosity, HWE tests, FIS, assignment based on genotype likelihoods, most bottleneck tests.
Plants are different: selfing, clones, pollen and seed plant-specific
Most population genetics theory was written for animals that mate at random and move as adults. Plants do neither, and four consequences show up in every dataset:
- Mixed mating. Outcrossing rates (tm) range from near 0 to near 1, often within the same species. Selfing reduces Ne, inflates FIS and speeds up differentiation between populations.
- Clonality. Repeated multilocus genotypes must be detected before anything else — a clone counted as 40 individuals will distort every index. Block 2 flags them.
- Two dispersal vehicles. Pollen carries nuclear genes, seed carries everything. Comparing nuclear with maternally inherited chloroplast markers separates the two, and their differentiation ratio estimates the pollen-to-seed flow ratio.
- Fine-scale spatial structure. Limited seed dispersal makes neighbours relatives. That is what spatial autocorrelation and the Sp statistic quantify (Block 9) — and it is why a "population" sampled along a transect may not be one population at all.
Sampling design: how many individuals, how many loci before the lab
The honest answer is that it depends on the question, but the literature converges on useful minima:
| Goal | Individuals per population | Loci |
|---|---|---|
| Diversity indices (He, Na) | 20–30 | 8–10 SSR |
| Allelic richness, rare alleles | 30–50 | 10–15 SSR |
| FST / AMOVA | 20–25, ≥5 populations | 10+ SSR or 100+ AFLP bands |
| Bayesian clustering | 25–30 | 15+ SSR |
| Bottleneck tests | 30+ | 15–20 SSR |
| Fine-scale spatial structure | 100+ mapped individuals | 8+ SSR |
Which analysis answers which question decision guide
| Your question | The analysis | Block |
|---|---|---|
| Which population is most diverse? | He, uHe, allelic richness by rarefaction | 3 |
| Is this population inbred? | FIS with bootstrap CI, HWE tests | 4 |
| Are my populations differentiated? | FST/θ, G′ST, Jost's D, AMOVA | 5 |
| Does my hierarchical design matter? | Three-level AMOVA with permutations | 5 |
| How do populations relate to each other? | Nei's D, NJ/UPGMA tree, PCoA | 6 |
| Is differentiation just distance? | Mantel test, isolation by distance | 6 |
| How many genetic groups are there really? | Bayesian clustering, ΔK, DAPC | 7 |
| Where does this individual come from? | Assignment test, migrant detection | 7 |
| What is the history of my haplotypes? | Network, π, Tajima's D, mismatch | 8 |
| Did the population crash recently? | Heterozygosity excess, M-ratio, mode shift | 9 |
| Are neighbours relatives? | Spatial autocorrelation, kinship, Sp | 9 |
How to cite
If PopGeneticsPro helped with an analysis in a thesis or a paper, please cite the software and the original method papers it implements — every result and the methods paragraph of Block 10 name them by author and year.
References behind the calculations
Every statistic in this app comes from a published method, and each one is named again at the point where it is used.
Data & quality control
Load the file, agree on what it contains, and screen it before a single statistic is computed.
1 · Load your data
Everything stays on your computer: the file is read by the browser and never uploaded.
Drop your file here or click to choose one
File loaded: none yet
Allele frequencies & diversity
How much genetic variation there is, and how it is spread among your populations.
1 · Settings
The defaults follow the most common conventions of the literature. Change them and the whole block recomputes.
Hardy–Weinberg, inbreeding & linkage
Are the genotypes what random mating predicts — and if not, is it inbreeding, a null allele, or a mixture of populations?
1 · Settings
Differentiation & AMOVA
How much of the variation separates your populations, which pairs differ, and whether the regional grouping is real.
1 · Settings
Genetic distances, ordination & trees
How far apart the populations and individuals are, drawn as a map (PCoA), as a tree, and against geography.
1 · Settings
Bayesian clustering, DAPC & assignment
How many genetic groups the data themselves support, who belongs where, and which individuals are not from where they were collected.
1 · Bayesian clustering (admixture model)
The model of Pritchard, Stephens & Donnelly (2000): K clusters with their own allele frequencies, and for every individual the fraction of its genome drawn from each cluster (Q). It runs a Gibbs sampler in a background thread, so the page stays usable.
4 · Discriminant analysis of principal components (DAPC)
Jombart, Devillard & Balloux (2010). A model-free alternative: a PCA summarises the genotypes, a discriminant analysis then finds the axes that best separate the groups — either your populations, or groups found in the data by k-means clustering with the number chosen by BIC. No Hardy–Weinberg or linkage assumptions, and fast.
6 · Assignment tests and migrant detection
Each genotype is assigned to the population where it is most likely (Paetkau et al. 1995; Rannala & Mountain 1997), leaving the individual out of its own population. First-generation migrants are flagged by comparing Lhome/Lmax with the distribution obtained from genotypes simulated in the home population (Paetkau et al. 2004).
DNA sequences & haplotypes
What an alignment says about diversity, demographic history and the geography of lineages.
1 · Settings
Demography & spatial structure
Has the population crashed, how large is it genetically, are neighbours relatives, how much does it self, and is trait divergence more than drift?
1 · Settings
Figures & report
Everything you ran, as you left it: tables, interpretations, every figure as edited, the methods paragraph with its citations, and a ZIP for the paper.
1 · What goes into the report
Each figure is exported exactly as you edited it in its block (title, palette, fonts, size). Go back, change anything, and build the report again.