Unsupervised learning for biology, ecology and agronomy

Find the natural groups hidden in your data

From the raw table to a publication-ready dendrogram: check whether your data are clusterable, pick the right similarity coefficient for quantitative, binary, categorical, mixed or ecological data, compare hierarchical and partitioning methods, decide how many clusters to keep, validate them and export beautiful, fully editable figures. No programming required.

Runs entirely in your browser Your data never leave your computer Figures up to 900 dpi · PNG, TIFF, SVG Free for teaching and research

How it works

A guided workflow. Every step explains the idea in plain language, checks your data and recommends what to do next; a dendrogram is the last step of the process, not the first.

STEP 1
Load & explore
types, scaling, outliers, clustering tendency
STEP 2
Measure similarity
the right coefficient for your data
STEP 3
Cluster
hierarchical, partitioning, density, model-based
STEP 4
Choose k
elbow, silhouette, gap, consensus
STEP 5
Validate
internal, stability, external
STEP 6
Interpret & publish
profiles, prediction, editable figures, report

Try it: k-means and a dendrogram on the same points

Pick a dataset (or click inside the left panel to add points), choose k and a linkage rule, and watch k-means move its centres step by step while the tree on the right is recomputed from the very same points. Notice how the two moons and the ring defeat k-means but not single linkage.

k-means · centroids and hulls
Hierarchical · real dendrogram, cut at k

What's inside

Eight blocks, built in order. Click a card to jump to it.

Every kind of data

Clustering starts with a dissimilarity, and the dissimilarity depends on what your variables are. ClusteringPro types each column and proposes the coefficients that make sense for it.

Methods covered

The classical families and the modern ones, with the tools to compare them on your data rather than trusting a single tree.

Clustering, explained without the fog

Thirteen short lessons with figures. Open the ones you need; every block of the app repeats the part that matters at that step.

1 · What is a cluster, and what is clustering? basics
Small distances inside a group, large distances between groups.

Clustering is the search for groups of objects that are more similar to each other than to the rest. Nobody tells the algorithm which groups exist or how many there are: that is why it is called unsupervised. Compare with classification (supervised), where the groups are known in advance and the goal is to predict them.

In biology and agronomy the "objects" are typically genotypes, accessions, sites, plots, soil profiles, samples, patients or genes, and the "variables" are the traits, environmental measurements, species counts, markers or expression values recorded on them.

A good clustering has two properties at once: compactness (members of a cluster are close) and separation (clusters are far apart). Almost every method, index and figure in this app is a way of measuring one or both.

The four families of methods, plus hybrids and fuzzy variants. ClusteringPro implements all of them.
Typical questions Do my 120 maize landraces fall into distinct morphological groups? Which sampling sites share a similar community of insects? Are there soil-fertility types in the region? Do these gene-expression profiles define patient subtypes? Which cultivars respond alike across environments?
2 · Before clustering: are my data clusterable at all? exploratory analysis
Uniform noise (left) versus real structure (right). The Hopkins statistic tells them apart.

Every clustering algorithm returns clusters, even when there are none. Cutting a dendrogram of uniformly scattered points still yields groups; k-means will happily split a single cloud into three. So the first step is to ask whether the data contain structure worth finding.

  • Hopkins statistic compares the distances from random locations to their nearest data point with the distances from real points to their nearest neighbour. Around 0.5 the data look uniform; values approaching 1 indicate strong clustering. A common rule: below 0.75 be cautious, above 0.9 be confident.
  • VAT image (visual assessment of tendency): the distance matrix is reordered so that similar objects sit together; dark diagonal blocks reveal clusters, a smooth gradient reveals none.
  • PCA or MDS preview: a scatter of the first two dimensions often shows groups, gradients or a single blob at a glance.
  • Correlations and redundancy: highly correlated variables count several times in a Euclidean distance; consider dropping one, using Mahalanobis distance or clustering on principal components.
  • Outliers and missing values: a single extreme object becomes its own cluster in hierarchical methods and drags k-means centres. Decide how to treat them before, not after.
Rows and columns The number of objects should comfortably exceed the number of clusters you expect (at least 5–10 objects per cluster). Many variables are fine, but they should all be relevant: irrelevant "noise" variables hide the structure that the useful ones carry.
3 · Similarity, dissimilarity and distance the heart of the matter
Three ways to measure how far two objects are in variable space.

Algorithms never see your variables; they see a matrix of dissimilarities between pairs of objects. Change the coefficient and you change the clusters, so this choice deserves more attention than the choice of algorithm. A similarity s (1 = identical) becomes a dissimilarity through d = 1 − s or d = √(1 − s).

Quantitative variables

  • Euclidean: the straight line. The default for k-means and Ward, sensitive to scale and to outliers.
  • Manhattan (city block): sum of absolute differences; more robust to extreme values.
  • Minkowski generalises both; Chebyshev keeps only the largest difference.
  • Canberra: differences relative to magnitude, useful for skewed positive data.
  • Correlation-based (Pearson, Spearman, Kendall, cosine): two objects are similar when their profiles rise and fall together, regardless of level. The standard choice for gene expression and for shape rather than size.
  • Mahalanobis: Euclidean after removing correlations among variables.
Binary coefficients are built from the 2 × 2 table of matches and mismatches.

Binary variables (presence/absence, 0/1)

  • Jaccard and Dice–Sørensen ignore double zeros: two sites that both lack a species are not made similar by that absence. Preferred in ecology and for markers.
  • Simple matching, Rogers–Tanimoto, Sokal–Sneath count double zeros as agreement: right when 0 and 1 are symmetric states (male/female, allele A/B).
  • Russell–Rao, Ochiai, Kulczynski, Hamming: further variants offered in Block 3 with their formulas.

Nominal, ordinal and mixed variables

  • Nominal: simple matching across categories (equivalently Hamming on dummy columns).
  • Ordinal: replace levels by ranks and use Manhattan, or treat them inside Gower.
  • Gower distance handles any mixture: each variable contributes a 0–1 dissimilarity of the appropriate kind and the average is taken. It is the natural partner of PAM.

Ecological counts and abundances

  • Bray–Curtis: the workhorse for species × site tables; ignores double zeros, bounded 0–1.
  • Hellinger and chord transformations followed by Euclidean distance make abundance data suitable for k-means and Ward.
  • Chi-square distance, Morisita–Horn, Canberra, Kulczynski (quantitative) and Jaccard on abundances (Ružička) complete the set.
The double-zero problem In community data most cells are zero. A coefficient that treats "both absent" as evidence of similarity will group sites by the species they lack. Asymmetric coefficients (Jaccard, Bray–Curtis, Hellinger) are the safe choice.
4 · Standardisation: putting variables on the same footing preprocessing
In raw units the yield axis dominates; after standardisation both variables count.

Euclidean and Manhattan distances add differences across variables. A variable measured in kilograms per hectare (thousands) swamps one measured in pH units (5–8). Unless the units are already comparable, standardise:

  • z-score: subtract the mean, divide by the standard deviation. Each variable gets mean 0 and variance 1. The default.
  • Range (0–1): divides by max − min; keeps the shape of the distribution, sensitive to outliers.
  • Robust: median and interquartile range, when outliers are present.
  • Log or square-root: for skewed counts before any distance.
  • Hellinger, chord, proportions: transformations specific to abundance data (see lesson 3).
When not to standardise If all variables share a unit and the differences in variance are meaningful (for example, 12 morphological traits all in millimetres), scaling may erase real information. Binary and Gower distances do not need it. ClusteringPro shows the variances so you can decide.
5 · Hierarchical clustering and linkage rules agglomerative · divisive
Five ways to define the distance between two groups. The red line or lines show what each rule measures.

Agglomerative methods start with every object in its own cluster and repeatedly merge the two closest clusters until one remains. Divisive methods (DIANA) start from one cluster and split. Both produce a hierarchy: a tree with a merge height at every node, drawn as a dendrogram.

"Closest clusters" must be defined, and that definition is the linkage rule:

RuleDistance between clustersCharacter
Singlenearest pair of membersfinds elongated shapes; prone to chaining
Completefarthest paircompact, similar-sized clusters; sensitive to outliers
Average (UPGMA)mean of all pairwise distancesbalanced; the usual choice in taxonomy and ecology
Weighted (WPGMA)mean of the two sub-cluster distanceslike UPGMA but ignores cluster sizes
Centroid (UPGMC)distance between centroidscan produce reversals (a merge lower than a previous one)
Median (WPGMC)centroid of the two centroidssame caveat
Ward (Ward.D2)increase in within-cluster sum of squarescompact spherical clusters; needs Euclidean distances
Flexible betatunable compromiseβ = −0.25 behaves like a space-conserving average

Two diagnostics tell you how well a tree represents the original distances:

  • Cophenetic correlation: correlation between the original distances and the heights at which pairs are joined. Above 0.75 the tree is a faithful summary; UPGMA usually scores highest here.
  • Agglomerative coefficient: how strongly clustered the objects are (close to 1 = clear structure). It grows with sample size, so compare it only among trees of the same data.
Ward.D versus Ward.D2 Ward's criterion is defined on squared Euclidean distances. Ward.D2 squares them for you and reports heights in the original units; Ward.D applies the same formula to whatever you feed it. Prefer Ward.D2 unless you are reproducing an old analysis.
6 · Reading (and drawing) a dendrogram figures
Anatomy of a dendrogram. Cutting at a height gives a partition.

The height of a node is the dissimilarity at which its two branches were merged. Long vertical branches mean well-separated groups; a cluster is convincing when the branch below the cut is short and the branch above it is long.

  • The left–right order of the leaves is largely arbitrary: any node can be flipped without changing the tree (2n−1 equivalent drawings). Do not read "neighbours" horizontally; read the height at which two objects join.
  • Cutting the tree at a height, or asking for k groups, produces a partition. The cut can be at a single height or at different heights in different branches (dynamic cut).
  • For large trees, zoom into a branch, collapse sub-clusters into triangles, or show only the top part.

What ClusteringPro draws. Rectangular, triangular, horizontal and radial (circular) layouts; branches coloured by cluster or by a height gradient; leaf labels coloured, rotated, resized or replaced by symbols; a movable cut line; cluster boxes; a legend you can position; editable titles, axes, fonts, palettes (categorical, viridis-type, colour-blind safe) and line widths; export to PNG, TIFF, SVG, JPG and WEBP at up to 900 dpi. Every figure remembers your edits for the report.

7 · Comparing dendrograms tanglegram · cophenetic · Baker
A tanglegram links the same leaves in two trees; fewer crossings means better agreement.

Different linkage rules, distances or variable sets give different trees for the same objects. Instead of trusting one, compare them:

  • Tanglegram: two trees face to face with the leaves connected. Leaves are re-ordered to minimise crossings; the entanglement (0 = perfect match, 1 = total disagreement) summarises the result.
  • Cophenetic correlation between trees: correlation of the two sets of merge heights.
  • Baker's gamma: rank correlation of the level at which each pair of objects first joins in each tree.
  • Fowlkes–Mallows index Bk: agreement of the partitions obtained by cutting each tree into k groups, for every k.
  • Correlation matrix of methods: run all eight linkages and see which ones agree; a stable structure appears in most of them.
8 · Partitioning: k-means, PAM, CLARA and fuzzy clustering non-hierarchical
One k-means run: assign each point to the nearest centre, move the centres, repeat.

k-means chooses k centres and alternates two steps until nothing changes: assign every object to its nearest centre, then move each centre to the mean of its objects. It minimises the total within-cluster sum of squares, is fast for large data and has no tree. Variants: Hartigan–Wong (the usual default, usually best), Lloyd/Forgy, MacQueen; smarter starts with k-means++; many random starts (nstart) to avoid poor local optima.

  • k-medoids (PAM): centres are real observations (medoids), so any distance works, including Gower and Bray–Curtis, and outliers weigh less.
  • CLARA: PAM on repeated samples, for thousands of objects.
  • Fuzzy c-means (FANNY): every object receives a membership degree in each cluster instead of a hard label. Useful for gradients and for objects sitting between groups; the Dunn partition coefficient measures how crisp the result is.
  • Hierarchical k-means: a Ward tree supplies the initial centres, k-means refines them; removes the dependence on random starts.
Two datasets where the k-means answer is wrong by construction.

Limits. k-means assumes round clusters of similar size and variance, in Euclidean space, and needs k in advance. Curved, elongated or nested shapes (moons, rings) are split incorrectly; single linkage, DBSCAN or spectral clustering recover them. The live demo above lets you see this happen.

9 · Density-based and model-based clustering DBSCAN · Gaussian mixtures
DBSCAN classifies points as core, border or noise using a radius ε and a minimum count.

DBSCAN grows clusters from core points, those with at least minPts neighbours within radius ε. Clusters can have any shape, the number of clusters is not fixed in advance, and isolated points are labelled noise rather than forced into a group. The k-nearest-neighbour distance plot suggests ε: sort the distance to the minPts-th neighbour and look for the knee. HDBSCAN* removes ε altogether and extracts clusters of varying density from a hierarchy.

Model-based clustering assumes the data come from a mixture of Gaussian distributions, one per cluster, with its own mean and covariance. The EM algorithm estimates the parameters and gives each object a probability of belonging to each cluster; the BIC chooses both the number of clusters and the covariance shape (spherical, diagonal, ellipsoidal; equal or varying volume, shape and orientation). It is the principled way to get uncertainty with your labels, and it handles elliptical clusters that k-means cannot.

10 · How many clusters? elbow · silhouette · gap · consensus
Three classical criteria evaluated for k = 1…7. Here all three point to k = 3.

There is no single correct answer; there are criteria, and the domain has the last word. ClusteringPro computes them for every k in a range and for any method, then reports a consensus:

  • Elbow: total within-cluster sum of squares drops fast until the true k, then flattens. Subjective when the bend is gentle.
  • Average silhouette: how well each object sits in its cluster compared with the nearest alternative (lesson 11); pick the k that maximises the mean.
  • Gap statistic: compares the within-cluster dispersion with what is expected under no structure (bootstrap of uniform data); choose the smallest k whose gap is at least the gap of k + 1 minus one standard error. Unlike the other criteria it can also answer k = 1: no structure.
  • Calinski–Harabasz (ratio of between to within dispersion), Davies–Bouldin (lower is better), Dunn (higher is better), plus the C-index, McClain–Rao, PBM, Ratkowsky–Lance, Hartigan, Krzanowski–Lai and Ball–Hall indices: thirteen criteria in all, and the app tallies their votes.
  • For hierarchies: the largest jumps in merge heights, and the support of each node by multiscale bootstrap.
  • For model-based clustering: BIC. For DBSCAN: the knee of the kNN-distance plot.
Then look A clustering that scores well but cannot be described in terms of the variables is not useful. Block 7 profiles every cluster so you can name it.
11 · Validating clusters internal · stability · external
The silhouette of one object compares its own cluster with the nearest other cluster.

Internal validation uses the data alone:

  • Silhouette width s(i) for every object: near 1 = well placed, near 0 = between two clusters, negative = probably misassigned. The silhouette plot shows them sorted within each cluster; the mean is the overall quality (0.7+ strong, 0.5 reasonable, below 0.25 no substantial structure).
  • Dunn index (smallest between-cluster distance / largest cluster diameter) and connectivity (whether neighbours share a cluster).

Stability asks whether the clusters survive perturbation:

  • Bootstrap Jaccard stability: resample the objects, recluster, and measure how often each cluster reappears (above 0.75 = stable, below 0.5 = doubtful).
  • Multiscale bootstrap p-values for hierarchical clusters (AU / BP): which branches of the tree are supported by the data.
  • Variable removal: clusters that vanish when one variable is dropped depend on that variable alone.

External validation compares the clusters with a known grouping (species, origin, treatment): Rand and adjusted Rand index, normalised mutual information, purity, variation of information, plus a cross-table with a chi-square test. The same indices compare two clusterings with each other.

Comparing algorithms. Run hierarchical, k-means, PAM and model-based side by side over a range of k and rank them by connectivity, Dunn and silhouette before you commit.

12 · From clusters to meaning: profiles, indicators and supervised follow-up interpretation

A cluster is only useful once you can say what characterises it. Block 7 provides:

  • Cluster profiles: means, medians and dispersion of every variable per cluster, as a table, a heat map of standardised means, a radar chart and parallel coordinates.
  • Which variables define each cluster: v-tests comparing the cluster mean with the overall mean, one-way ANOVA or Kruskal–Wallis per variable, and effect sizes. For categorical variables, the categories that are over- or under-represented.
  • Indicator species / indicator variables for ecological tables: the IndVal statistic with permutation test.
  • Supervised follow-up: a linear discriminant analysis trained on the clusters gives the axes that separate them, a confusion matrix showing how sharp the boundaries are, and a rule to assign new observations to the existing clusters (also by nearest centroid or k-nearest neighbours). A classification tree yields plain-language rules such as "cluster 2 = plant height > 180 cm and early flowering".
  • Cluster map on PCA / MDS axes with convex hulls, confidence ellipses, centroid labels and supplementary variables.
13 · Pitfalls checklist before you publish
Do
  • Check clustering tendency first (Hopkins, VAT).
  • Match the coefficient to the data type; use asymmetric coefficients for presence/absence and abundances.
  • Standardise when units differ; say so in the methods.
  • Try several linkages and algorithms and compare them (tanglegram, cophenetic, ARI).
  • Choose k with several criteria and report them all.
  • Report the silhouette and a stability measure with every partition.
  • Describe each cluster in terms of the variables.
Don't
  • Read the left–right order of a dendrogram as proximity.
  • Use Ward or k-means on non-Euclidean distances (Bray–Curtis, Gower) without a transformation.
  • Let one variable in big units decide the clusters.
  • Take a single k-means run with one random start as the answer.
  • Use double-zero coefficients (simple matching) on species tables.
  • Test differences between clusters with ANOVA on the very variables used to build them and call the result "significant".
  • Present a tree without its cophenetic correlation or a cut without its silhouette.

Which method should I use?

A starting point, not a rule. Block 2 makes this recommendation automatically after typing your columns.

Your dataDissimilarityFirst choiceAlso tryValidate with
Quantitative traits, similar unitsEuclidean (scaled if needed)Ward.D2 tree, then k-means from the treeUPGMA, model-based (GMM)Silhouette, gap, bootstrap stability
Quantitative profiles (expression, spectra)1 − Pearson correlationUPGMA or complete linkagePAM, k-means on standardised rowsCophenetic, silhouette, AU p-values
Presence/absence (species, markers)Jaccard or DiceUPGMAPAM, Ward on √(1 − s)Cophenetic, bootstrap, indicator species
Species abundancesBray–Curtis or Hellinger + EuclideanUPGMA (Bray–Curtis) or Ward (Hellinger)PAM, k-means on HellingerSilhouette, IndVal, ANOSIM/PERMANOVA
Mixed quantitative + categoricalGowerPAMUPGMA, hierarchical on GowerSilhouette, stability
Large tables (thousands of rows)Euclideank-means (k-means++, many starts)CLARA, DBSCAN, mini-batchSilhouette on a sample, CH index
Irregular shapes, noise, outliersEuclideanDBSCAN / HDBSCAN*Single linkage, spectralSilhouette (with caution), stability
Overlapping groups, gradientsEuclideanFuzzy c-means or GMMPAMPartition coefficient, BIC

Why ClusteringPro

🌿

Built for life-science data

Binary markers, species tables, morphometrics, soil and yield data, gene expression: each with the right coefficient and the right warnings.

🧭

Guided, not just computed

Every block explains the theory in plain language and tells you whether your data are clusterable, which distance fits and how many groups to keep.

🔬

Validated numerics

Distances, linkages, indices and p-values computed in the browser and checked against independent reference implementations.

🌳

Dendrograms worth publishing

Rectangular, radial and phylogenic layouts, palettes, gradients, coloured labels, cut lines, legends, boxes: all editable, all exportable at up to 900 dpi.

⚖️

Compare, don't trust

Tanglegrams, cophenetic and Baker correlations, consensus of indices, bootstrap stability and side-by-side algorithms.

🔒

Private and offline

Nothing is uploaded. Works without internet, from a USB stick or a shared folder.

Glossary at a glance

agglomerativedivisivelinkageUPGMAWarddendrogramcophenetic correlationtanglegramentanglement centroidmedoidk-means++membershipε · minPtsBIC HopkinsVATelbowsilhouettegap statisticDunnCalinski–Harabaszadjusted RandNMIbootstrap stabilityIndValGowerBray–CurtisJaccard

How to cite ClusteringPro

If the platform contributes to a publication, a thesis or a report, please cite it. The DOI is a concept DOI registered in Zenodo: it always resolves to the latest archived version.

Open the Zenodo record ↗ Source code on GitHub ↗
BibTeX entry

Version-specific DOIs and the full metadata are in the CITATION.cff file of the repository (GitHub shows them under "Cite this repository"). The same reference is printed in every report generated in Block 8.

2

Data: import, variable types, preprocessing and clustering tendency

Load your table, confirm what each column is, choose how to treat missing values and scale, and find out whether the data are worth clustering.

Before you upload: what this block checks and why

Short and practical. The long version is in the lessons on the home page.

1 · How to arrange the table
  • One row per object to be clustered (accession, site, sample, plot, gene) and one column per variable. The first row holds the column names.
  • An identifier column (name or code) is optional but recommended: it becomes the label of the dendrogram leaves.
  • A known grouping column (region, species, treatment) is optional. It is never used to build the clusters; it colours the plots and is used for external validation later.
  • Missing values may be blank or coded NA, ND, ?, –.
  • A distance matrix is accepted as a square table whose first column repeats the column names (zeros on the diagonal). Scaling is skipped and the matrix is used directly.
AccessionRegionPlantHeight_cmEarLength_cmKernelRows…
ACC-001Highland182.711.112…
ACC-002Highland206.114.614…
ACC-016Valley241.018.214…
2 · Variable types and how they are recognised

The type decides which dissimilarity makes sense in Block 3, so check the automatic guess:

  • Quantitative: numbers with many distinct values. Integers with very few distinct values are flagged, because they may be codes.
  • Count / abundance: non-negative integers with many zeros (or a name that suggests species or counts). Recommended route: Bray–Curtis or Hellinger.
  • Binary: exactly two states such as 0/1, yes/no, present/absent, +/−. The second state is coded 1.
  • Nominal: text with up to 40 categories. Text where every value is different is taken as the object label.
  • Ordinal: recognised sets such as low < medium < high or poor < moderate < good; otherwise set it by hand and type the order of the levels.
3 · Missing values, transformation and scaling

Missing values. Distances need complete rows. Either drop the objects with gaps (safe when they are few) or impute them (mean for quantitative, median for counts and ordinal levels, mode for categories). Gower distance in Block 3 can also skip missing pairs variable by variable.

Transformation acts before scaling: log(x + 1) or √x tame skewed variables; the row-wise Hellinger, chord and relative-abundance transformations turn a species × site table into something Euclidean distance, Ward and k-means can handle.

Scaling puts variables on comparable footing: z-scores (mean 0, SD 1) are the default, range 0–1 keeps the shape, robust scaling resists outliers. Binary variables are never scaled; nominal variables are dummy-coded only for the exploratory preview (the real categorical distance is chosen in Block 3).

4 · What the tendency diagnostics mean
  • Hopkins statistic H compares nearest-neighbour distances of random points with those of real objects. Uniform data give H ≈ 0.5; the grey band in the figure is where H falls 95 % of the time when there is no structure. Values above 0.75 indicate a clear tendency to cluster.
  • VAT image: the distance matrix reordered so that neighbours sit together. Dark square blocks on the diagonal are candidate clusters; their number is a first guess of k. A smooth gradient means no clusters. If you provided a known grouping, the coloured strip shows whether the blocks match it.
  • PCA / MDS map: the objects in the two directions of largest variance. Look for separate clouds, gradients or a single blob.
  • Mahalanobis D²: how far each object is from the multivariate centre, accounting for correlations. Objects beyond the 99.9 % chi-square quantile will form their own branch in a tree or drag k-means centres.
  • Correlation heat map: variables with |r| above 0.95 are redundant and count twice in a Euclidean distance.

1 · Load your data

Drag a file or click to browse. The first row must contain the column names; one row per object.

Drop your file here or click to choose

.xlsx.xlsm.ods.csv.tsv.txt.json.dat
Only for .csv / .txt files.
Spreadsheets set to Spanish or Portuguese often export with a comma.
Paste from the clipboard instead

Or try an example dataset

3

Similarity and distance

Choose how "far apart" two objects are. Everything that follows (trees, partitions, validation) is computed from this matrix.

Choosing a coefficient: the short version

Lesson 3 on the home page has the full story; these are the rules that matter here.

1 · Match the coefficient to the data type
DataUseAvoid
Quantitative, comparable units (scaled)Euclidean, Manhattan; Mahalanobis if variables are correlated; Euclidean on PCs to drop noiseCorrelation distances with fewer than ~5 variables
Profiles: expression, spectra, time series1 − Pearson, 1 − Spearman, cosineEuclidean when only the shape matters
Presence/absence where 0 = absentJaccard, Dice, Ochiai, KulczynskiSimple matching (rewards shared absences)
Binary with two meaningful statesSimple matching, Rogers–TanimotoJaccard (ignores one of the states)
Abundances, countsBray–Curtis, Ružička, Morisita–Horn; Hellinger or chord if you need Euclidean geometryRaw Euclidean (dominated by abundant species and double zeros)
Nominal / ordinalSimple matching on categories, ranks for ordinal, GowerTreating category codes as numbers
Mixed typesGower (weights optional)Dummy coding + Euclidean unless categories are few
2 · Metric, Euclidean, and why it matters

A dissimilarity is a metric when d(x, z) ≤ d(x, y) + d(y, z) for every triple. It is Euclidean-embeddable when the objects can be placed as points in some Euclidean space reproducing every d exactly. Euclidean ⇒ metric, not the reverse.

  • Ward's method, k-means on coordinates and principal coordinates (PCoA) assume Euclidean geometry. On a non-Euclidean matrix (Bray–Curtis, Jaccard, most similarities converted as 1 − s) they still run but distort; the fix is usually √d, which is Euclidean for Bray–Curtis, Jaccard, Sørensen and simple matching.
  • Single, complete and average linkage and PAM only need the matrix: any dissimilarity, metric or not, is fine.
  • This block checks both properties: the share of triples violating the triangle inequality and the negative eigenvalues of the double-centred matrix.
3 · Reading the figures
  • Heat map: objects reordered so that close ones sit together; dark diagonal blocks are candidate clusters. The strips show your known grouping, if any.
  • PCoA map: the best 2-D picture of the matrix. Read it with the Shepard diagram and the stress: below 0.1 the map is faithful, above 0.2 it is only a sketch.
  • Distribution of dissimilarities: two humps (near and far pairs) suggest groups; a single narrow hump means every pair is about equally far, which is bad news for clustering.
  • Nearest-neighbour network: which objects would join first; edges crossing between known groups reveal where the grouping and the data disagree.
  • Comparison of coefficients: correlation between the matrices produced by several coefficients. If they agree, the choice is harmless; if not, your clusters depend on it, and you should say so.

1 · Choose the coefficient

4

Hierarchical clustering and the dendrogram studio

Build the tree from the dissimilarity matrix of Block 3, judge it, cut it, draw it the way you want, and compare linkage rules.

What to look at before trusting a tree

Lessons 5–7 on the home page explain the linkage rules and how to read a dendrogram.

1 · Three numbers per tree
  • Cophenetic correlation between the original dissimilarities and the heights at which each pair joins. Above 0.75 the tree is a faithful summary; UPGMA typically scores highest, Ward and complete linkage lower because they favour compact clusters over fidelity.
  • Agglomerative coefficient (or divisive coefficient for DIANA): mean over objects of 1 − (height of the object's first merge / height of the last merge). Near 1 = clear structure; it grows with n, so compare only trees of the same data.
  • Reversals: merges that happen at a lower height than an earlier one (centroid and median linkage). They make cutting by height ambiguous.
2 · Cutting the tree

Cut by number of clusters k or by height. The bar chart of merge heights shows where the biggest jumps are: a long vertical stretch without merges means that the clusters below it are well separated. Block 6 adds silhouette, gap statistic, stability and many other criteria; use the jumps here as a first idea, then let Block 6 and your knowledge of the objects decide.

If you supplied a known grouping, the cross-table and the adjusted Rand index tell you how well the cut recovers it (1 = identical, 0 = what random labels give).

3 · Comparing linkage rules

Run several rules on the same matrix and compare them with Baker's gamma (rank correlation of the level at which each pair joins), the correlation between cophenetic matrices, the Fowlkes–Mallows index at the chosen k, and a tanglegram that draws two trees face to face with the leaves connected; the entanglement (0 = identical orders after rotating branches, 1 = reversed) and the number of crossings summarise the disagreement. A structure that survives every linkage is real; one that appears with a single rule deserves suspicion.

1 · Build the tree

5

Partitioning, fuzzy, model-based, density-based and spectral clustering

Non-hierarchical methods on the same data: choose one, judge it with the silhouette, and compare it with the tree of Block 4 and with the other algorithms.

Which algorithm, and what it assumes

Lessons 8 and 9 on the home page explain each method; here is what matters when you press Run.

1 · The methods at a glance
MethodWorks onAssumesNeeds k?Strength
k-meanscoordinatesround clusters, similar size, Euclideanyesfast, well understood; many starts avoid bad optima
Hierarchical k-meanscoordinatesas k-meansyesstarts from the tree centroids: reproducible
PAMany dissimilaritymedoid-shaped clustersyesrobust to outliers; medoids are real objects; Gower, Bray–Curtis, Jaccard all fine
CLARAany dissimilarityas PAMyesthousands of objects
Fuzzy c-meanscoordinatesas k-means, soft boundariesyesmembership degrees reveal intermediate objects and gradients
Gaussian mixturecoordinatesellipsoidal Gaussian componentsBIC choosesprobabilities and uncertainty per object; k and shape selected by BIC
DBSCANany dissimilaritydense regions separated by sparse onesno (ε, minPts)arbitrary shapes, labels outliers as noise
Spectralany dissimilarityconnected neighbourhood graphyescurved, nested or elongated shapes (moons, rings)
2 · Coordinates versus the matrix

k-means, fuzzy c-means and Gaussian mixtures need coordinates. Two sources are offered: the working matrix of Block 2 (the natural choice when your Block 3 dissimilarity is Euclidean on those variables), or the principal coordinates (PCoA axes) of the chosen dissimilarity, which turns any matrix, even Bray–Curtis or Gower, into Euclidean coordinates that reproduce it as closely as possible. PAM, CLARA, DBSCAN and spectral clustering use the matrix directly.

The silhouette is always computed on the Block 3 dissimilarity, so every method is judged on the same footing.

3 · Reading the results
  • Average silhouette: above 0.7 strong, 0.5 reasonable, 0.25 weak, below that no substantial structure. Negative values flag objects placed in the wrong cluster.
  • Between-SS / total-SS: share of the variance explained by the partition (coordinate methods). It always grows with k, so it cannot choose k on its own.
  • Adjusted Rand index against the tree cut and against your known grouping: 1 = identical, 0 = chance agreement.
  • Membership / posterior probabilities (fuzzy, mixtures): objects with a maximum below about 0.6 sit between clusters; the partition coefficient summarises how crisp the solution is.
  • k-NN distance plot (DBSCAN): sort the distance to the (minPts − 1)-th neighbour; the knee is a good ε. Objects beyond it are noise.

1 · Choose the method

Same seed → same result.
6

Optimal number of clusters and validation

How many clusters, how stable, how well supported, and how much they agree with what you already know.

Four questions, four tools

Lessons 10 and 11 on the home page explain each criterion in detail.

1 · How many clusters? Thirteen criteria and a vote

For every k in a range, a partition is produced (the tree cut, k-means or PAM) and each criterion selects its own k: silhouette, Calinski–Harabasz, PBM, Ratkowsky–Lance, Dunn and Krzanowski–Lai at their maximum; Davies–Bouldin, C-index and McClain–Rao at their minimum; Hartigan at the first k with H ≤ 10; Ball–Hall at the largest drop; the elbow at the knee of the WSS curve; and the gap statistic by Tibshirani's first-SE rule against uniform reference data. The tally of votes is the consensus.

2 · Are the clusters stable?

Bootstrap stability: the objects are resampled, the clustering is recomputed and each original cluster is matched to its most similar bootstrap cluster by the Jaccard index. Mean Jaccard above 0.85 = highly stable, 0.75 = stable, 0.6 = some pattern, below 0.5 = the cluster dissolves and should not be interpreted.

Multiscale bootstrap for trees: the variables are resampled at several sizes, the tree is rebuilt and the frequency of each cluster is tracked. The AU value corrects the bias of the plain bootstrap probability BP; clusters with AU ≥ 95 % are strongly supported by the data.

3 · External validation and comparison of algorithms

Against a known grouping (or between two partitions): Rand and adjusted Rand, NMI, variation of information, Jaccard, Fowlkes–Mallows, purity and a χ² test with Cramér's V.

Comparison of algorithms: hierarchical, k-means and PAM over the range of k, scored by connectivity (lower is better), Dunn and silhouette (higher is better).

1 · Optimal number of clusters

2 · Stability of the clusters (bootstrap)

3 · Support of the tree clusters (multiscale bootstrap, AU / BP)

10 scales × B trees; 30 is quick, use 200 or more to report the values.

4 · External validation

5 · Comparison of algorithms over k

7

Cluster profiles, interpretation and prediction

What makes each cluster different, which variables and categories define it, which species indicate it, and how to assign new objects to the clusters.

From labels to meaning

Lesson 12 on the home page introduces these tools.

1 · Describing clusters with test values

For every quantitative variable the test value v compares the cluster mean with the overall mean, accounting for the cluster size: |v| ≥ 1.96 marks a variable that characterises the cluster (higher when v > 0, lower when v < 0). A one-way ANOVA (and its rank-based counterpart, Kruskal–Wallis) tells which variables separate the clusters at all, with η² as effect size. Categorical variables get the same treatment on proportions, plus a χ² test with Cramér's V.

A caveat Clusters were built to differ on the active variables, so their p-values are descriptive, not inferential. Supplementary variables (not used for clustering) are the ones that support genuine claims.
2 · Indicator species

The indicator value (Dufrêne & Legendre) of a species for a cluster is the product of its specificity A (share of the species' abundance found in that cluster) and its fidelity B (share of the cluster's sites where the species occurs). Species with IndVal above about 50 % and a significant permutation test are good indicators. Available whenever count or presence/absence variables are active.

3 · Supervised follow-up: discriminant analysis, rules and new objects

Linear discriminant analysis finds the axes that best separate the clusters, gives each variable a weight (standardised coefficient) and a correlation with the axis (structure), and yields a rule to classify objects. The leave-one-out accuracy is the honest estimate of how well new objects will be assigned; Wilks' Λ tests the separation. A classification tree translates the same task into plain rules such as "Cluster 2 if plant height ≤ 200 cm". New observations are assigned by three rules at once (LDA posterior, nearest centroid in standardised space, tree); disagreement flags objects between clusters.

1 · Profile the clusters

Permutations:

Quantitative variables

Categorical variables

8

Report and export

A self-contained HTML report with every table, diagnostic and figure exactly as you edited them; print it to PDF; or download a ZIP with data, CSV tables and figures at publication resolution.

1 · What goes into the report

Sections