Find the natural groups hidden in your data
From the raw table to a publication-ready dendrogram: check whether your data are clusterable, pick the right similarity coefficient for quantitative, binary, categorical, mixed or ecological data, compare hierarchical and partitioning methods, decide how many clusters to keep, validate them and export beautiful, fully editable figures. No programming required.
How it works
A guided workflow. Every step explains the idea in plain language, checks your data and recommends what to do next; a dendrogram is the last step of the process, not the first.
Try it: k-means and a dendrogram on the same points
Pick a dataset (or click inside the left panel to add points), choose k and a linkage rule, and watch k-means move its centres step by step while the tree on the right is recomputed from the very same points. Notice how the two moons and the ring defeat k-means but not single linkage.
What's inside
Eight blocks, built in order. Click a card to jump to it.
Every kind of data
Clustering starts with a dissimilarity, and the dissimilarity depends on what your variables are. ClusteringPro types each column and proposes the coefficients that make sense for it.
Methods covered
The classical families and the modern ones, with the tools to compare them on your data rather than trusting a single tree.
Clustering, explained without the fog
Thirteen short lessons with figures. Open the ones you need; every block of the app repeats the part that matters at that step.
1 · What is a cluster, and what is clustering? basics
Clustering is the search for groups of objects that are more similar to each other than to the rest. Nobody tells the algorithm which groups exist or how many there are: that is why it is called unsupervised. Compare with classification (supervised), where the groups are known in advance and the goal is to predict them.
In biology and agronomy the "objects" are typically genotypes, accessions, sites, plots, soil profiles, samples, patients or genes, and the "variables" are the traits, environmental measurements, species counts, markers or expression values recorded on them.
A good clustering has two properties at once: compactness (members of a cluster are close) and separation (clusters are far apart). Almost every method, index and figure in this app is a way of measuring one or both.
2 · Before clustering: are my data clusterable at all? exploratory analysis
Every clustering algorithm returns clusters, even when there are none. Cutting a dendrogram of uniformly scattered points still yields groups; k-means will happily split a single cloud into three. So the first step is to ask whether the data contain structure worth finding.
- Hopkins statistic compares the distances from random locations to their nearest data point with the distances from real points to their nearest neighbour. Around 0.5 the data look uniform; values approaching 1 indicate strong clustering. A common rule: below 0.75 be cautious, above 0.9 be confident.
- VAT image (visual assessment of tendency): the distance matrix is reordered so that similar objects sit together; dark diagonal blocks reveal clusters, a smooth gradient reveals none.
- PCA or MDS preview: a scatter of the first two dimensions often shows groups, gradients or a single blob at a glance.
- Correlations and redundancy: highly correlated variables count several times in a Euclidean distance; consider dropping one, using Mahalanobis distance or clustering on principal components.
- Outliers and missing values: a single extreme object becomes its own cluster in hierarchical methods and drags k-means centres. Decide how to treat them before, not after.
3 · Similarity, dissimilarity and distance the heart of the matter
Algorithms never see your variables; they see a matrix of dissimilarities between pairs of objects. Change the coefficient and you change the clusters, so this choice deserves more attention than the choice of algorithm. A similarity s (1 = identical) becomes a dissimilarity through d = 1 − s or d = √(1 − s).
Quantitative variables
- Euclidean: the straight line. The default for k-means and Ward, sensitive to scale and to outliers.
- Manhattan (city block): sum of absolute differences; more robust to extreme values.
- Minkowski generalises both; Chebyshev keeps only the largest difference.
- Canberra: differences relative to magnitude, useful for skewed positive data.
- Correlation-based (Pearson, Spearman, Kendall, cosine): two objects are similar when their profiles rise and fall together, regardless of level. The standard choice for gene expression and for shape rather than size.
- Mahalanobis: Euclidean after removing correlations among variables.
Binary variables (presence/absence, 0/1)
- Jaccard and Dice–Sørensen ignore double zeros: two sites that both lack a species are not made similar by that absence. Preferred in ecology and for markers.
- Simple matching, Rogers–Tanimoto, Sokal–Sneath count double zeros as agreement: right when 0 and 1 are symmetric states (male/female, allele A/B).
- Russell–Rao, Ochiai, Kulczynski, Hamming: further variants offered in Block 3 with their formulas.
Nominal, ordinal and mixed variables
- Nominal: simple matching across categories (equivalently Hamming on dummy columns).
- Ordinal: replace levels by ranks and use Manhattan, or treat them inside Gower.
- Gower distance handles any mixture: each variable contributes a 0–1 dissimilarity of the appropriate kind and the average is taken. It is the natural partner of PAM.
Ecological counts and abundances
- Bray–Curtis: the workhorse for species × site tables; ignores double zeros, bounded 0–1.
- Hellinger and chord transformations followed by Euclidean distance make abundance data suitable for k-means and Ward.
- Chi-square distance, Morisita–Horn, Canberra, Kulczynski (quantitative) and Jaccard on abundances (Ružička) complete the set.
4 · Standardisation: putting variables on the same footing preprocessing
Euclidean and Manhattan distances add differences across variables. A variable measured in kilograms per hectare (thousands) swamps one measured in pH units (5–8). Unless the units are already comparable, standardise:
- z-score: subtract the mean, divide by the standard deviation. Each variable gets mean 0 and variance 1. The default.
- Range (0–1): divides by max − min; keeps the shape of the distribution, sensitive to outliers.
- Robust: median and interquartile range, when outliers are present.
- Log or square-root: for skewed counts before any distance.
- Hellinger, chord, proportions: transformations specific to abundance data (see lesson 3).
5 · Hierarchical clustering and linkage rules agglomerative · divisive
Agglomerative methods start with every object in its own cluster and repeatedly merge the two closest clusters until one remains. Divisive methods (DIANA) start from one cluster and split. Both produce a hierarchy: a tree with a merge height at every node, drawn as a dendrogram.
"Closest clusters" must be defined, and that definition is the linkage rule:
| Rule | Distance between clusters | Character |
|---|---|---|
| Single | nearest pair of members | finds elongated shapes; prone to chaining |
| Complete | farthest pair | compact, similar-sized clusters; sensitive to outliers |
| Average (UPGMA) | mean of all pairwise distances | balanced; the usual choice in taxonomy and ecology |
| Weighted (WPGMA) | mean of the two sub-cluster distances | like UPGMA but ignores cluster sizes |
| Centroid (UPGMC) | distance between centroids | can produce reversals (a merge lower than a previous one) |
| Median (WPGMC) | centroid of the two centroids | same caveat |
| Ward (Ward.D2) | increase in within-cluster sum of squares | compact spherical clusters; needs Euclidean distances |
| Flexible beta | tunable compromise | β = −0.25 behaves like a space-conserving average |
Two diagnostics tell you how well a tree represents the original distances:
- Cophenetic correlation: correlation between the original distances and the heights at which pairs are joined. Above 0.75 the tree is a faithful summary; UPGMA usually scores highest here.
- Agglomerative coefficient: how strongly clustered the objects are (close to 1 = clear structure). It grows with sample size, so compare it only among trees of the same data.
6 · Reading (and drawing) a dendrogram figures
The height of a node is the dissimilarity at which its two branches were merged. Long vertical branches mean well-separated groups; a cluster is convincing when the branch below the cut is short and the branch above it is long.
- The left–right order of the leaves is largely arbitrary: any node can be flipped without changing the tree (2n−1 equivalent drawings). Do not read "neighbours" horizontally; read the height at which two objects join.
- Cutting the tree at a height, or asking for k groups, produces a partition. The cut can be at a single height or at different heights in different branches (dynamic cut).
- For large trees, zoom into a branch, collapse sub-clusters into triangles, or show only the top part.
What ClusteringPro draws. Rectangular, triangular, horizontal and radial (circular) layouts; branches coloured by cluster or by a height gradient; leaf labels coloured, rotated, resized or replaced by symbols; a movable cut line; cluster boxes; a legend you can position; editable titles, axes, fonts, palettes (categorical, viridis-type, colour-blind safe) and line widths; export to PNG, TIFF, SVG, JPG and WEBP at up to 900 dpi. Every figure remembers your edits for the report.
7 · Comparing dendrograms tanglegram · cophenetic · Baker
Different linkage rules, distances or variable sets give different trees for the same objects. Instead of trusting one, compare them:
- Tanglegram: two trees face to face with the leaves connected. Leaves are re-ordered to minimise crossings; the entanglement (0 = perfect match, 1 = total disagreement) summarises the result.
- Cophenetic correlation between trees: correlation of the two sets of merge heights.
- Baker's gamma: rank correlation of the level at which each pair of objects first joins in each tree.
- Fowlkes–Mallows index Bk: agreement of the partitions obtained by cutting each tree into k groups, for every k.
- Correlation matrix of methods: run all eight linkages and see which ones agree; a stable structure appears in most of them.
8 · Partitioning: k-means, PAM, CLARA and fuzzy clustering non-hierarchical
k-means chooses k centres and alternates two steps until nothing changes: assign every object to its nearest centre, then move each centre to the mean of its objects. It minimises the total within-cluster sum of squares, is fast for large data and has no tree. Variants: Hartigan–Wong (the usual default, usually best), Lloyd/Forgy, MacQueen; smarter starts with k-means++; many random starts (nstart) to avoid poor local optima.
- k-medoids (PAM): centres are real observations (medoids), so any distance works, including Gower and Bray–Curtis, and outliers weigh less.
- CLARA: PAM on repeated samples, for thousands of objects.
- Fuzzy c-means (FANNY): every object receives a membership degree in each cluster instead of a hard label. Useful for gradients and for objects sitting between groups; the Dunn partition coefficient measures how crisp the result is.
- Hierarchical k-means: a Ward tree supplies the initial centres, k-means refines them; removes the dependence on random starts.
Limits. k-means assumes round clusters of similar size and variance, in Euclidean space, and needs k in advance. Curved, elongated or nested shapes (moons, rings) are split incorrectly; single linkage, DBSCAN or spectral clustering recover them. The live demo above lets you see this happen.
9 · Density-based and model-based clustering DBSCAN · Gaussian mixtures
DBSCAN grows clusters from core points, those with at least minPts neighbours within radius ε. Clusters can have any shape, the number of clusters is not fixed in advance, and isolated points are labelled noise rather than forced into a group. The k-nearest-neighbour distance plot suggests ε: sort the distance to the minPts-th neighbour and look for the knee. HDBSCAN* removes ε altogether and extracts clusters of varying density from a hierarchy.
Model-based clustering assumes the data come from a mixture of Gaussian distributions, one per cluster, with its own mean and covariance. The EM algorithm estimates the parameters and gives each object a probability of belonging to each cluster; the BIC chooses both the number of clusters and the covariance shape (spherical, diagonal, ellipsoidal; equal or varying volume, shape and orientation). It is the principled way to get uncertainty with your labels, and it handles elliptical clusters that k-means cannot.
10 · How many clusters? elbow · silhouette · gap · consensus
There is no single correct answer; there are criteria, and the domain has the last word. ClusteringPro computes them for every k in a range and for any method, then reports a consensus:
- Elbow: total within-cluster sum of squares drops fast until the true k, then flattens. Subjective when the bend is gentle.
- Average silhouette: how well each object sits in its cluster compared with the nearest alternative (lesson 11); pick the k that maximises the mean.
- Gap statistic: compares the within-cluster dispersion with what is expected under no structure (bootstrap of uniform data); choose the smallest k whose gap is at least the gap of k + 1 minus one standard error. Unlike the other criteria it can also answer k = 1: no structure.
- Calinski–Harabasz (ratio of between to within dispersion), Davies–Bouldin (lower is better), Dunn (higher is better), plus the C-index, McClain–Rao, PBM, Ratkowsky–Lance, Hartigan, Krzanowski–Lai and Ball–Hall indices: thirteen criteria in all, and the app tallies their votes.
- For hierarchies: the largest jumps in merge heights, and the support of each node by multiscale bootstrap.
- For model-based clustering: BIC. For DBSCAN: the knee of the kNN-distance plot.
11 · Validating clusters internal · stability · external
Internal validation uses the data alone:
- Silhouette width s(i) for every object: near 1 = well placed, near 0 = between two clusters, negative = probably misassigned. The silhouette plot shows them sorted within each cluster; the mean is the overall quality (0.7+ strong, 0.5 reasonable, below 0.25 no substantial structure).
- Dunn index (smallest between-cluster distance / largest cluster diameter) and connectivity (whether neighbours share a cluster).
Stability asks whether the clusters survive perturbation:
- Bootstrap Jaccard stability: resample the objects, recluster, and measure how often each cluster reappears (above 0.75 = stable, below 0.5 = doubtful).
- Multiscale bootstrap p-values for hierarchical clusters (AU / BP): which branches of the tree are supported by the data.
- Variable removal: clusters that vanish when one variable is dropped depend on that variable alone.
External validation compares the clusters with a known grouping (species, origin, treatment): Rand and adjusted Rand index, normalised mutual information, purity, variation of information, plus a cross-table with a chi-square test. The same indices compare two clusterings with each other.
Comparing algorithms. Run hierarchical, k-means, PAM and model-based side by side over a range of k and rank them by connectivity, Dunn and silhouette before you commit.
12 · From clusters to meaning: profiles, indicators and supervised follow-up interpretation
A cluster is only useful once you can say what characterises it. Block 7 provides:
- Cluster profiles: means, medians and dispersion of every variable per cluster, as a table, a heat map of standardised means, a radar chart and parallel coordinates.
- Which variables define each cluster: v-tests comparing the cluster mean with the overall mean, one-way ANOVA or Kruskal–Wallis per variable, and effect sizes. For categorical variables, the categories that are over- or under-represented.
- Indicator species / indicator variables for ecological tables: the IndVal statistic with permutation test.
- Supervised follow-up: a linear discriminant analysis trained on the clusters gives the axes that separate them, a confusion matrix showing how sharp the boundaries are, and a rule to assign new observations to the existing clusters (also by nearest centroid or k-nearest neighbours). A classification tree yields plain-language rules such as "cluster 2 = plant height > 180 cm and early flowering".
- Cluster map on PCA / MDS axes with convex hulls, confidence ellipses, centroid labels and supplementary variables.
13 · Pitfalls checklist before you publish
Do
- Check clustering tendency first (Hopkins, VAT).
- Match the coefficient to the data type; use asymmetric coefficients for presence/absence and abundances.
- Standardise when units differ; say so in the methods.
- Try several linkages and algorithms and compare them (tanglegram, cophenetic, ARI).
- Choose k with several criteria and report them all.
- Report the silhouette and a stability measure with every partition.
- Describe each cluster in terms of the variables.
Don't
- Read the left–right order of a dendrogram as proximity.
- Use Ward or k-means on non-Euclidean distances (Bray–Curtis, Gower) without a transformation.
- Let one variable in big units decide the clusters.
- Take a single k-means run with one random start as the answer.
- Use double-zero coefficients (simple matching) on species tables.
- Test differences between clusters with ANOVA on the very variables used to build them and call the result "significant".
- Present a tree without its cophenetic correlation or a cut without its silhouette.
Which method should I use?
A starting point, not a rule. Block 2 makes this recommendation automatically after typing your columns.
| Your data | Dissimilarity | First choice | Also try | Validate with |
|---|---|---|---|---|
| Quantitative traits, similar units | Euclidean (scaled if needed) | Ward.D2 tree, then k-means from the tree | UPGMA, model-based (GMM) | Silhouette, gap, bootstrap stability |
| Quantitative profiles (expression, spectra) | 1 − Pearson correlation | UPGMA or complete linkage | PAM, k-means on standardised rows | Cophenetic, silhouette, AU p-values |
| Presence/absence (species, markers) | Jaccard or Dice | UPGMA | PAM, Ward on √(1 − s) | Cophenetic, bootstrap, indicator species |
| Species abundances | Bray–Curtis or Hellinger + Euclidean | UPGMA (Bray–Curtis) or Ward (Hellinger) | PAM, k-means on Hellinger | Silhouette, IndVal, ANOSIM/PERMANOVA |
| Mixed quantitative + categorical | Gower | PAM | UPGMA, hierarchical on Gower | Silhouette, stability |
| Large tables (thousands of rows) | Euclidean | k-means (k-means++, many starts) | CLARA, DBSCAN, mini-batch | Silhouette on a sample, CH index |
| Irregular shapes, noise, outliers | Euclidean | DBSCAN / HDBSCAN* | Single linkage, spectral | Silhouette (with caution), stability |
| Overlapping groups, gradients | Euclidean | Fuzzy c-means or GMM | PAM | Partition coefficient, BIC |
Why ClusteringPro
Built for life-science data
Binary markers, species tables, morphometrics, soil and yield data, gene expression: each with the right coefficient and the right warnings.
Guided, not just computed
Every block explains the theory in plain language and tells you whether your data are clusterable, which distance fits and how many groups to keep.
Validated numerics
Distances, linkages, indices and p-values computed in the browser and checked against independent reference implementations.
Dendrograms worth publishing
Rectangular, radial and phylogenic layouts, palettes, gradients, coloured labels, cut lines, legends, boxes: all editable, all exportable at up to 900 dpi.
Compare, don't trust
Tanglegrams, cophenetic and Baker correlations, consensus of indices, bootstrap stability and side-by-side algorithms.
Private and offline
Nothing is uploaded. Works without internet, from a USB stick or a shared folder.
Glossary at a glance
How to cite ClusteringPro
If the platform contributes to a publication, a thesis or a report, please cite it. The DOI is a concept DOI registered in Zenodo: it always resolves to the latest archived version.
BibTeX entry
Version-specific DOIs and the full metadata are in the CITATION.cff file of the repository (GitHub shows them under "Cite this repository"). The same reference is printed in every report generated in Block 8.
Data: import, variable types, preprocessing and clustering tendency
Load your table, confirm what each column is, choose how to treat missing values and scale, and find out whether the data are worth clustering.
Before you upload: what this block checks and why
Short and practical. The long version is in the lessons on the home page.
1 · How to arrange the table
- One row per object to be clustered (accession, site, sample, plot, gene) and one column per variable. The first row holds the column names.
- An identifier column (name or code) is optional but recommended: it becomes the label of the dendrogram leaves.
- A known grouping column (region, species, treatment) is optional. It is never used to build the clusters; it colours the plots and is used for external validation later.
- Missing values may be blank or coded NA, ND, ?, –.
- A distance matrix is accepted as a square table whose first column repeats the column names (zeros on the diagonal). Scaling is skipped and the matrix is used directly.
| Accession | Region | PlantHeight_cm | EarLength_cm | KernelRows | … |
|---|---|---|---|---|---|
| ACC-001 | Highland | 182.7 | 11.1 | 12 | … |
| ACC-002 | Highland | 206.1 | 14.6 | 14 | … |
| ACC-016 | Valley | 241.0 | 18.2 | 14 | … |
2 · Variable types and how they are recognised
The type decides which dissimilarity makes sense in Block 3, so check the automatic guess:
- Quantitative: numbers with many distinct values. Integers with very few distinct values are flagged, because they may be codes.
- Count / abundance: non-negative integers with many zeros (or a name that suggests species or counts). Recommended route: Bray–Curtis or Hellinger.
- Binary: exactly two states such as 0/1, yes/no, present/absent, +/−. The second state is coded 1.
- Nominal: text with up to 40 categories. Text where every value is different is taken as the object label.
- Ordinal: recognised sets such as low < medium < high or poor < moderate < good; otherwise set it by hand and type the order of the levels.
3 · Missing values, transformation and scaling
Missing values. Distances need complete rows. Either drop the objects with gaps (safe when they are few) or impute them (mean for quantitative, median for counts and ordinal levels, mode for categories). Gower distance in Block 3 can also skip missing pairs variable by variable.
Transformation acts before scaling: log(x + 1) or √x tame skewed variables; the row-wise Hellinger, chord and relative-abundance transformations turn a species × site table into something Euclidean distance, Ward and k-means can handle.
Scaling puts variables on comparable footing: z-scores (mean 0, SD 1) are the default, range 0–1 keeps the shape, robust scaling resists outliers. Binary variables are never scaled; nominal variables are dummy-coded only for the exploratory preview (the real categorical distance is chosen in Block 3).
4 · What the tendency diagnostics mean
- Hopkins statistic H compares nearest-neighbour distances of random points with those of real objects. Uniform data give H ≈ 0.5; the grey band in the figure is where H falls 95 % of the time when there is no structure. Values above 0.75 indicate a clear tendency to cluster.
- VAT image: the distance matrix reordered so that neighbours sit together. Dark square blocks on the diagonal are candidate clusters; their number is a first guess of k. A smooth gradient means no clusters. If you provided a known grouping, the coloured strip shows whether the blocks match it.
- PCA / MDS map: the objects in the two directions of largest variance. Look for separate clouds, gradients or a single blob.
- Mahalanobis D²: how far each object is from the multivariate centre, accounting for correlations. Objects beyond the 99.9 % chi-square quantile will form their own branch in a tree or drag k-means centres.
- Correlation heat map: variables with |r| above 0.95 are redundant and count twice in a Euclidean distance.
1 · Load your data
Drag a file or click to browse. The first row must contain the column names; one row per object.
Drop your file here or click to choose
Paste from the clipboard instead
Or try an example dataset
Similarity and distance
Choose how "far apart" two objects are. Everything that follows (trees, partitions, validation) is computed from this matrix.
Choosing a coefficient: the short version
Lesson 3 on the home page has the full story; these are the rules that matter here.
1 · Match the coefficient to the data type
| Data | Use | Avoid |
|---|---|---|
| Quantitative, comparable units (scaled) | Euclidean, Manhattan; Mahalanobis if variables are correlated; Euclidean on PCs to drop noise | Correlation distances with fewer than ~5 variables |
| Profiles: expression, spectra, time series | 1 − Pearson, 1 − Spearman, cosine | Euclidean when only the shape matters |
| Presence/absence where 0 = absent | Jaccard, Dice, Ochiai, Kulczynski | Simple matching (rewards shared absences) |
| Binary with two meaningful states | Simple matching, Rogers–Tanimoto | Jaccard (ignores one of the states) |
| Abundances, counts | Bray–Curtis, Ružička, Morisita–Horn; Hellinger or chord if you need Euclidean geometry | Raw Euclidean (dominated by abundant species and double zeros) |
| Nominal / ordinal | Simple matching on categories, ranks for ordinal, Gower | Treating category codes as numbers |
| Mixed types | Gower (weights optional) | Dummy coding + Euclidean unless categories are few |
2 · Metric, Euclidean, and why it matters
A dissimilarity is a metric when d(x, z) ≤ d(x, y) + d(y, z) for every triple. It is Euclidean-embeddable when the objects can be placed as points in some Euclidean space reproducing every d exactly. Euclidean ⇒ metric, not the reverse.
- Ward's method, k-means on coordinates and principal coordinates (PCoA) assume Euclidean geometry. On a non-Euclidean matrix (Bray–Curtis, Jaccard, most similarities converted as 1 − s) they still run but distort; the fix is usually √d, which is Euclidean for Bray–Curtis, Jaccard, Sørensen and simple matching.
- Single, complete and average linkage and PAM only need the matrix: any dissimilarity, metric or not, is fine.
- This block checks both properties: the share of triples violating the triangle inequality and the negative eigenvalues of the double-centred matrix.
3 · Reading the figures
- Heat map: objects reordered so that close ones sit together; dark diagonal blocks are candidate clusters. The strips show your known grouping, if any.
- PCoA map: the best 2-D picture of the matrix. Read it with the Shepard diagram and the stress: below 0.1 the map is faithful, above 0.2 it is only a sketch.
- Distribution of dissimilarities: two humps (near and far pairs) suggest groups; a single narrow hump means every pair is about equally far, which is bad news for clustering.
- Nearest-neighbour network: which objects would join first; edges crossing between known groups reveal where the grouping and the data disagree.
- Comparison of coefficients: correlation between the matrices produced by several coefficients. If they agree, the choice is harmless; if not, your clusters depend on it, and you should say so.
1 · Choose the coefficient
Hierarchical clustering and the dendrogram studio
Build the tree from the dissimilarity matrix of Block 3, judge it, cut it, draw it the way you want, and compare linkage rules.
What to look at before trusting a tree
Lessons 5–7 on the home page explain the linkage rules and how to read a dendrogram.
1 · Three numbers per tree
- Cophenetic correlation between the original dissimilarities and the heights at which each pair joins. Above 0.75 the tree is a faithful summary; UPGMA typically scores highest, Ward and complete linkage lower because they favour compact clusters over fidelity.
- Agglomerative coefficient (or divisive coefficient for DIANA): mean over objects of 1 − (height of the object's first merge / height of the last merge). Near 1 = clear structure; it grows with n, so compare only trees of the same data.
- Reversals: merges that happen at a lower height than an earlier one (centroid and median linkage). They make cutting by height ambiguous.
2 · Cutting the tree
Cut by number of clusters k or by height. The bar chart of merge heights shows where the biggest jumps are: a long vertical stretch without merges means that the clusters below it are well separated. Block 6 adds silhouette, gap statistic, stability and many other criteria; use the jumps here as a first idea, then let Block 6 and your knowledge of the objects decide.
If you supplied a known grouping, the cross-table and the adjusted Rand index tell you how well the cut recovers it (1 = identical, 0 = what random labels give).
3 · Comparing linkage rules
Run several rules on the same matrix and compare them with Baker's gamma (rank correlation of the level at which each pair joins), the correlation between cophenetic matrices, the Fowlkes–Mallows index at the chosen k, and a tanglegram that draws two trees face to face with the leaves connected; the entanglement (0 = identical orders after rotating branches, 1 = reversed) and the number of crossings summarise the disagreement. A structure that survives every linkage is real; one that appears with a single rule deserves suspicion.
1 · Build the tree
Partitioning, fuzzy, model-based, density-based and spectral clustering
Non-hierarchical methods on the same data: choose one, judge it with the silhouette, and compare it with the tree of Block 4 and with the other algorithms.
Which algorithm, and what it assumes
Lessons 8 and 9 on the home page explain each method; here is what matters when you press Run.
1 · The methods at a glance
| Method | Works on | Assumes | Needs k? | Strength |
|---|---|---|---|---|
| k-means | coordinates | round clusters, similar size, Euclidean | yes | fast, well understood; many starts avoid bad optima |
| Hierarchical k-means | coordinates | as k-means | yes | starts from the tree centroids: reproducible |
| PAM | any dissimilarity | medoid-shaped clusters | yes | robust to outliers; medoids are real objects; Gower, Bray–Curtis, Jaccard all fine |
| CLARA | any dissimilarity | as PAM | yes | thousands of objects |
| Fuzzy c-means | coordinates | as k-means, soft boundaries | yes | membership degrees reveal intermediate objects and gradients |
| Gaussian mixture | coordinates | ellipsoidal Gaussian components | BIC chooses | probabilities and uncertainty per object; k and shape selected by BIC |
| DBSCAN | any dissimilarity | dense regions separated by sparse ones | no (ε, minPts) | arbitrary shapes, labels outliers as noise |
| Spectral | any dissimilarity | connected neighbourhood graph | yes | curved, nested or elongated shapes (moons, rings) |
2 · Coordinates versus the matrix
k-means, fuzzy c-means and Gaussian mixtures need coordinates. Two sources are offered: the working matrix of Block 2 (the natural choice when your Block 3 dissimilarity is Euclidean on those variables), or the principal coordinates (PCoA axes) of the chosen dissimilarity, which turns any matrix, even Bray–Curtis or Gower, into Euclidean coordinates that reproduce it as closely as possible. PAM, CLARA, DBSCAN and spectral clustering use the matrix directly.
The silhouette is always computed on the Block 3 dissimilarity, so every method is judged on the same footing.
3 · Reading the results
- Average silhouette: above 0.7 strong, 0.5 reasonable, 0.25 weak, below that no substantial structure. Negative values flag objects placed in the wrong cluster.
- Between-SS / total-SS: share of the variance explained by the partition (coordinate methods). It always grows with k, so it cannot choose k on its own.
- Adjusted Rand index against the tree cut and against your known grouping: 1 = identical, 0 = chance agreement.
- Membership / posterior probabilities (fuzzy, mixtures): objects with a maximum below about 0.6 sit between clusters; the partition coefficient summarises how crisp the solution is.
- k-NN distance plot (DBSCAN): sort the distance to the (minPts − 1)-th neighbour; the knee is a good ε. Objects beyond it are noise.
1 · Choose the method
Optimal number of clusters and validation
How many clusters, how stable, how well supported, and how much they agree with what you already know.
Four questions, four tools
Lessons 10 and 11 on the home page explain each criterion in detail.
1 · How many clusters? Thirteen criteria and a vote
For every k in a range, a partition is produced (the tree cut, k-means or PAM) and each criterion selects its own k: silhouette, Calinski–Harabasz, PBM, Ratkowsky–Lance, Dunn and Krzanowski–Lai at their maximum; Davies–Bouldin, C-index and McClain–Rao at their minimum; Hartigan at the first k with H ≤ 10; Ball–Hall at the largest drop; the elbow at the knee of the WSS curve; and the gap statistic by Tibshirani's first-SE rule against uniform reference data. The tally of votes is the consensus.
2 · Are the clusters stable?
Bootstrap stability: the objects are resampled, the clustering is recomputed and each original cluster is matched to its most similar bootstrap cluster by the Jaccard index. Mean Jaccard above 0.85 = highly stable, 0.75 = stable, 0.6 = some pattern, below 0.5 = the cluster dissolves and should not be interpreted.
Multiscale bootstrap for trees: the variables are resampled at several sizes, the tree is rebuilt and the frequency of each cluster is tracked. The AU value corrects the bias of the plain bootstrap probability BP; clusters with AU ≥ 95 % are strongly supported by the data.
3 · External validation and comparison of algorithms
Against a known grouping (or between two partitions): Rand and adjusted Rand, NMI, variation of information, Jaccard, Fowlkes–Mallows, purity and a χ² test with Cramér's V.
Comparison of algorithms: hierarchical, k-means and PAM over the range of k, scored by connectivity (lower is better), Dunn and silhouette (higher is better).
1 · Optimal number of clusters
2 · Stability of the clusters (bootstrap)
3 · Support of the tree clusters (multiscale bootstrap, AU / BP)
4 · External validation
5 · Comparison of algorithms over k
Cluster profiles, interpretation and prediction
What makes each cluster different, which variables and categories define it, which species indicate it, and how to assign new objects to the clusters.
From labels to meaning
Lesson 12 on the home page introduces these tools.
1 · Describing clusters with test values
For every quantitative variable the test value v compares the cluster mean with the overall mean, accounting for the cluster size: |v| ≥ 1.96 marks a variable that characterises the cluster (higher when v > 0, lower when v < 0). A one-way ANOVA (and its rank-based counterpart, Kruskal–Wallis) tells which variables separate the clusters at all, with η² as effect size. Categorical variables get the same treatment on proportions, plus a χ² test with Cramér's V.
2 · Indicator species
The indicator value (Dufrêne & Legendre) of a species for a cluster is the product of its specificity A (share of the species' abundance found in that cluster) and its fidelity B (share of the cluster's sites where the species occurs). Species with IndVal above about 50 % and a significant permutation test are good indicators. Available whenever count or presence/absence variables are active.
3 · Supervised follow-up: discriminant analysis, rules and new objects
Linear discriminant analysis finds the axes that best separate the clusters, gives each variable a weight (standardised coefficient) and a correlation with the axis (structure), and yields a rule to classify objects. The leave-one-out accuracy is the honest estimate of how well new objects will be assigned; Wilks' Λ tests the separation. A classification tree translates the same task into plain rules such as "Cluster 2 if plant height ≤ 200 cm". New observations are assigned by three rules at once (LDA posterior, nearest centroid in standardised space, tree); disagreement flags objects between clusters.
1 · Profile the clusters
Quantitative variables
Categorical variables
Report and export
A self-contained HTML report with every table, diagnostic and figure exactly as you edited them; print it to PDF; or download a ZIP with data, CSV tables and figures at publication resolution.