Encuentra los ejes que ordenan tus datos
Del archivo crudo a una figura lista para publicar: comprueba si tus datos admiten un ACP, decide cuántos componentes retener con ocho criterios en lugar de uno, rota si hace falta, lee los mapas factoriales y exporta un informe completo. Sin programar y sin instalar nada.
Cómo funciona
Un recorrido guiado. Cada paso explica la idea en palabras llanas, revisa tus datos y recomienda qué hacer después. La decisión siempre la tomas tú: la plataforma pone delante el argumento, no la respuesta.
Qué incluye
Seis bloques, en orden. Pulsa una tarjeta para ir directo.
Por qué existe
El cálculo de un ACP nunca es el problema: cualquier programa devuelve los valores propios. Lo que decide si el análisis se sostiene son las decisiones de alrededor —cómo tratas los faltantes, si estandarizas, cuántos componentes retienes, qué cargas interpretas— y ahí el software habitual calcula la cifra y se detiene. PCAPro pone el argumento metodológico junto al número que lo justifica, en el momento en que hay que decidir.
Marco teórico del Análisis de Componentes Principales
Todo lo que conviene tener claro antes de tocar los datos. Despliega la sección que necesites; las secciones marcadas con ▸ Bloque N se aplican más adelante en la plataforma.
1 · ¿Qué es el ACP y qué problema resuelve?
El Análisis de Componentes Principales (ACP; PCA en inglés) es una técnica de estadística multivariada descriptiva que resume la información contenida en muchas variables cuantitativas correlacionadas en un número pequeño de variables nuevas —los componentes principales— construidas como combinaciones lineales de las originales, no correlacionadas entre sí y ordenadas por la cantidad de varianza que retienen.
Fue formulado por Karl Pearson (1901) como el problema de encontrar la recta (o el plano) que mejor se ajusta a una nube de puntos en el sentido de mínimos cuadrados ortogonales, y desarrollado independientemente por Harold Hotelling (1933) desde el punto de vista de la maximización de la varianza. Ambas formulaciones conducen al mismo resultado: los ejes principales son los vectores propios de la matriz de covarianzas (o de correlaciones) de los datos.
Para qué se usa realmente:
- Reducción de dimensiones. Pasar de 20 variables a 2 o 3 ejes que conserven el 70–90 % de la información.
- Visualización. Representar en un plano individuos y variables que viven en un espacio de p dimensiones.
- Detección de estructura. Descubrir grupos de variables que se comportan juntas (síndromes, gradientes, ejes ecológicos).
- Depuración de la multicolinealidad. Generar predictores ortogonales para regresión (Principal Component Regression).
- Detección de atípicos y control de calidad de una base de datos.
- Construcción de índices sintéticos (calidad, productividad, condición) a partir del primer componente.
Principal Component Analysis (PCA) is a technique of descriptive multivariate statistics that summarises the information contained in many correlated quantitative variables into a small number of new variables —the principal components— built as linear combinations of the original ones, uncorrelated with each other and ordered by the amount of variance they retain.
It was formulated by Karl Pearson (1901) as the problem of finding the line (or the plane) that best fits a cloud of points in the orthogonal least-squares sense, and developed independently by Harold Hotelling (1933) from the point of view of variance maximisation. Both formulations lead to the same result: the principal axes are the eigenvectors of the covariance (or correlation) matrix of the data.
What it is actually used for:
- Dimension reduction. Going from 20 variables to 2 or 3 axes that keep 70–90% of the information.
- Visualisation. Displaying on a plane the individuals and variables that live in a p-dimensional space.
- Detecting structure. Finding groups of variables that behave together (syndromes, gradients, ecological axes).
- Cleaning up multicollinearity. Generating orthogonal predictors for regression (principal component regression).
- Outlier detection and quality control of a data set.
- Building composite indices (quality, productivity, condition) from the first component.
2 · La matemática, en una página
Partimos de una matriz de datos X con n filas (individuos) y p columnas (variables cuantitativas). Se centra cada columna restando su media, y opcionalmente se divide entre su desviación estándar:
Con los datos centrados se calcula la matriz de covarianzas (o de correlaciones si además se dividió entre s):
El ACP resuelve el problema de valores propios:
- λ_k (valor propio o eigenvalue) = varianza que captura el componente k.
Se cumple
Σ λ_k = traza(S), es decir la varianza total. Con matriz de correlaciones,traza = py por eso el criterio de Kaiser es λ > 1. - v_k (vector propio o eigenvector) = dirección del componente en el espacio de las variables. Sus elementos son los coeficientes o pesos de la combinación lineal.
- Puntuaciones (scores): coordenadas de los individuos sobre el nuevo eje,
F_k = Xᶜ · v_k. - Cargas (loadings): correlación entre la variable original y el componente,
carga_jk = v_jk · √λ_k. Son lo que se interpreta y lo que se dibuja en el círculo de correlaciones.
Numéricamente es preferible obtener lo mismo por descomposición en valores singulares (SVD) de la matriz centrada, que evita construir S y es más estable:
Propiedades que conviene recordar:
- Los componentes son ortogonales (no correlacionados) por construcción.
- El primer componente es la dirección de máxima varianza; el segundo, la de máxima varianza entre las direcciones perpendiculares al primero, y así sucesivamente.
- El ACP es la mejor aproximación de rango k a la matriz de datos en el sentido de mínimos cuadrados (teorema de Eckart–Young).
- Los signos de los vectores propios son arbitrarios: un eje espejado no cambia la interpretación.
We start from a data matrix X with n rows (individuals) and p columns (quantitative variables). Each column is centred by subtracting its mean and, optionally, divided by its standard deviation:
With the centred data, the covariance matrix is computed (or the correlation matrix if the data were also divided by s):
PCA solves the eigenvalue problem:
- λ_k (the eigenvalue) = variance captured by component k.
It holds that
Σ λ_k = trace(S), that is, the total variance. With a correlation matrix,trace = p, and that is why Kaiser's criterion is λ > 1. - v_k (the eigenvector) = direction of the component in the space of the variables. Its elements are the coefficients or weights of the linear combination.
- Scores: coordinates of the individuals on the new axis,
F_k = Xᶜ · v_k. - Loadings: the correlation between the original variable and the component,
loading_jk = v_jk · √λ_k. They are what is interpreted and what is drawn in the correlation circle.
Numerically it is preferable to obtain the same result by singular value decomposition (SVD) of the centred matrix, which avoids building S and is more stable:
Properties worth remembering:
- The components are orthogonal (uncorrelated) by construction.
- The first component is the direction of maximum variance; the second is the direction of maximum variance among those perpendicular to the first, and so on.
- PCA is the best rank-k approximation to the data matrix in the least-squares sense (Eckart–Young theorem).
- The signs of the eigenvectors are arbitrary: a mirrored axis does not change the interpretation.
3 · Postulados y supuestos del ACP esencial
El ACP es una técnica descriptiva: no plantea un modelo probabilístico ni contrasta hipótesis, así que sus «supuestos» son menos rígidos que los de un ANOVA. Pero para que el resultado sea interpretable y reproducible deben cumplirse las siguientes condiciones. La app verifica automáticamente las que están marcadas con ✔.
- Variables cuantitativas continuas ✔. El ACP clásico opera sobre variables de intervalo o de razón. Con variables categóricas hay que usar otras técnicas: AC (análisis de correspondencias) para dos variables cualitativas, ACM para varias, AFDM para datos mixtos. Las escalas Likert con 5 o más niveles se aceptan en la práctica, pero conviene reportarlo.
- Linealidad de las relaciones ✔ (parcial). El ACP solo captura estructura lineal, porque parte de covarianzas o correlaciones de Pearson. Si dos variables se relacionan en forma de U, el ACP no lo verá. Revisa los diagramas de dispersión y, si hace falta, transforma (log, raíz) o usa ACP no lineal / kernel PCA.
- Correlaciones suficientes entre variables ✔. Es el requisito de fondo: si R ≈ I (matriz identidad) no hay redundancia que resumir. Se comprueba con la prueba de esfericidad de Bartlett, el índice KMO y la inspección de la matriz de correlaciones (debe haber varios |r| ≥ 0.30).
- Tamaño de muestra adecuado ✔. Reglas de uso común: n ≥ 100 (o al menos 50); razón de al menos 5 individuos por variable, preferiblemente 10:1; y siempre n > p. Muestras pequeñas producen cargas inestables que no se replican.
- Ausencia de multicolinealidad extrema y de singularidad ✔. Correlaciones |r| > 0.90 o determinante de R ≈ 0 indican que hay variables que son combinación lineal casi exacta de otras (por ejemplo incluir largo, ancho y área = largo × ancho). Infla el primer componente y hace inestable la rotación.
- Escalas comparables o estandarización previa ✔. El ACP sobre covarianzas está dominado por la variable de mayor varianza. Si las unidades difieren (mm, g, µm, %), hay que estandarizar y trabajar sobre la matriz de correlaciones.
- Ausencia de valores atípicos influyentes ✔. Como el método maximiza varianza, un solo caso extremo puede generar un componente propio. Se detectan con la distancia de Mahalanobis y se reporta el análisis con y sin ellos.
- Tratamiento explícito de los datos faltantes ✔. El ACP clásico necesita una matriz completa; hay que decidir entre eliminar filas o imputar, y declararlo.
- Normalidad multivariante: no es obligatoria. El ACP descriptivo funciona sin ella. Se vuelve necesaria si se van a usar pruebas inferenciales sobre los valores propios o si se interpreta la prueba de Bartlett de forma estricta, ya que esta última la supone. Sí conviene evitar asimetrías extremas (|g₁| > 2), porque distorsionan las correlaciones de Pearson.
- Homogeneidad de la muestra. El ACP asume que todos los individuos provienen de la misma población. Si mezclas poblaciones muy distintas, el primer eje solo reflejará esa separación; considera analizar por separado o usar las variables de grupo como suplementarias.
- Independencia de las observaciones. Cada fila debe ser un individuo distinto. Con medidas repetidas o datos espaciales/temporales la interpretación cambia y existen variantes específicas.
PCA is a descriptive technique: it poses no probability model and tests no hypotheses, so its "assumptions" are less rigid than those of an ANOVA. But for the result to be interpretable and reproducible, the following conditions must hold. The app automatically checks the ones marked with ✔.
- Continuous quantitative variables ✔. Classical PCA works on interval or ratio variables. With categorical variables other techniques are needed: CA (correspondence analysis) for two qualitative variables, MCA for several, FAMD for mixed data. Likert scales with 5 or more levels are accepted in practice, but this should be reported.
- Linear relationships ✔ (partly). PCA only captures linear structure, because it starts from Pearson covariances or correlations. If two variables are related in a U shape, PCA will not see it. Check the scatterplots and, if necessary, transform (log, root) or use non-linear PCA / kernel PCA.
- Enough correlation between variables ✔. This is the underlying requirement: if R ≈ I (identity matrix) there is no redundancy to summarise. It is checked with Bartlett's test of sphericity, the KMO index and inspection of the correlation matrix (several |r| ≥ 0.30 should be present).
- Adequate sample size ✔. Rules in common use: n ≥ 100 (or at least 50); a ratio of at least 5 individuals per variable, preferably 10:1; and always n > p. Small samples produce unstable loadings that do not replicate.
- No extreme multicollinearity and no singularity ✔. Correlations |r| > 0.90 or a determinant of R ≈ 0 mean that some variables are an almost exact linear combination of others (for instance including length, width and area = length × width). This inflates the first component and makes the rotation unstable.
- Comparable scales, or prior standardisation ✔. PCA on covariances is dominated by the variable with the largest variance. If the units differ (mm, g, µm, %), you must standardise and work on the correlation matrix.
- No influential outliers ✔. Because the method maximises variance, a single extreme case can generate a component of its own. They are detected with the Mahalanobis distance and the analysis is reported with and without them.
- Explicit handling of missing data ✔. Classical PCA needs a complete matrix; you must decide between deleting rows and imputing, and declare it.
- Multivariate normality is not required. Descriptive PCA works without it. It becomes necessary if inferential tests on the eigenvalues are to be used, or if Bartlett's test is interpreted strictly, since the latter assumes it. Extreme skewness (|g₁| > 2) should still be avoided, because it distorts Pearson correlations.
- Homogeneity of the sample. PCA assumes that all individuals come from the same population. If you mix very different populations, the first axis will merely reflect that separation; consider analysing them separately, or using the grouping variables as supplementary.
- Independence of the observations. Each row must be a distinct individual. With repeated measures or spatial/temporal data the interpretation changes, and specific variants exist.
4 · Pruebas previas obligatorias: Bartlett, KMO, determinante
Prueba de esfericidad de Bartlett (1950). Contrasta la hipótesis nula de que la matriz de correlaciones poblacional es la identidad, es decir, que las variables no están correlacionadas:
Se busca rechazar H₀: un valor p < 0.05 indica que existen correlaciones suficientes y que el ACP tiene sentido. Cuidado: con n grande casi siempre resulta significativa, así que es una condición necesaria pero no suficiente. Supone normalidad multivariante.
Índice de Kaiser–Meyer–Olkin (KMO) y MSA. Compara la magnitud de las correlaciones observadas con la de las correlaciones parciales. Si las variables comparten factores comunes, las correlaciones parciales serán pequeñas y el KMO alto:
| KMO | Calificación de Kaiser | Decisión |
|---|---|---|
| ≥ 0.90 | Excelente («maravilloso») | Adelante |
| 0.80 – 0.89 | Meritorio | Adelante |
| 0.70 – 0.79 | Aceptable («mediano») | Adelante |
| 0.60 – 0.69 | Mediocre | Con reservas |
| 0.50 – 0.59 | Bajo | Revisar variables |
| < 0.50 | Inaceptable | No hacer ACP |
El mismo cálculo aplicado a cada variable da su MSA (measure of sampling adequacy). Las variables con MSA < 0.50 se eliminan una a una, recalculando después de cada eliminación.
Determinante de R. Debe ser mayor que 0 pero no demasiado grande. Un valor cercano a 0 (< 10⁻⁵) señala multicolinealidad severa o singularidad; un valor cercano a 1 significa que R ≈ I y no hay nada que resumir.
Bartlett's test of sphericity (1950). It tests the null hypothesis that the population correlation matrix is the identity, that is, that the variables are uncorrelated:
The aim is to reject H₀: a value of p < 0.05 indicates that there are enough correlations and that PCA makes sense. Careful: with a large n it is almost always significant, so it is a necessary but not sufficient condition. It assumes multivariate normality.
Kaiser–Meyer–Olkin (KMO) index and MSA. It compares the magnitude of the observed correlations with that of the partial correlations. If the variables share common factors, the partial correlations will be small and the KMO high:
| KMO | Kaiser's label | Decision |
|---|---|---|
| ≥ 0.90 | Excellent ("marvellous") | Go ahead |
| 0.80 – 0.89 | Meritorious | Go ahead |
| 0.70 – 0.79 | Middling (acceptable) | Go ahead |
| 0.60 – 0.69 | Mediocre | With reservations |
| 0.50 – 0.59 | Low | Review the variables |
| < 0.50 | Unacceptable | Do not run a PCA |
The same computation applied to each variable gives its MSA (measure of sampling adequacy). Variables with MSA < 0.50 are removed one at a time, recomputing after each removal.
Determinant of R. It must be greater than 0 but not too large. A value close to 0 (< 10⁻⁵) signals severe multicollinearity or singularity; a value close to 1 means R ≈ I and there is nothing to summarise.
5 · Tamaño de muestra: cuánta gente / cuántas parcelas hacen falta
- Regla del número absoluto. Comrey y Lee: 50 = muy deficiente, 100 = deficiente, 200 = aceptable, 300 = bueno, 500 = muy bueno, 1000 = excelente.
- Regla de la razón n:p. Entre 5:1 (mínimo) y 20:1 (ideal); 10:1 es el consenso habitual.
- Regla de la comunalidad (más moderna, MacCallum et al. 1999). Lo que importa no es tanto n como la fuerza de la estructura: con comunalidades altas (> 0.60) y 3–4 variables marcadoras por componente, n = 60–100 puede bastar; con comunalidades bajas (< 0.40) pueden hacer falta n > 300.
- Regla práctica de campo. Si vas a interpretar cargas, exige al menos 3 variables con carga alta por cada componente que quieras retener.
Umbral de significación de una carga según el tamaño de muestra (Hair et al., α = 0.05, potencia 80 %):
| n | 50 | 60 | 70 | 85 | 100 | 120 | 150 | 200 | 250 | 350 |
|---|---|---|---|---|---|---|---|---|---|---|
| |carga| mínima | 0.75 | 0.70 | 0.65 | 0.60 | 0.55 | 0.50 | 0.45 | 0.40 | 0.35 | 0.30 |
- Absolute-number rule. Comrey and Lee: 50 = very poor, 100 = poor, 200 = fair, 300 = good, 500 = very good, 1000 = excellent.
- n:p ratio rule. Between 5:1 (minimum) and 20:1 (ideal); 10:1 is the usual consensus.
- Communality rule (more modern, MacCallum et al. 1999). What matters is not so much n as the strength of the structure: with high communalities (> 0.60) and 3–4 marker variables per component, n = 60–100 may be enough; with low communalities (< 0.40) n > 300 may be needed.
- Practical field rule. If you are going to interpret loadings, require at least 3 variables with a high loading for every component you want to retain.
Threshold for a loading to be significant given the sample size (Hair et al., α = 0.05, power 80%):
| n | 50 | 60 | 70 | 85 | 100 | 120 | 150 | 200 | 250 | 350 |
|---|---|---|---|---|---|---|---|---|---|---|
| minimum |loading| | 0.75 | 0.70 | 0.65 | 0.60 | 0.55 | 0.50 | 0.45 | 0.40 | 0.35 | 0.30 |
6 · Datos faltantes: qué hacer con los huecos
El ACP clásico exige una matriz completa. Las opciones, de menor a mayor sofisticación:
- Eliminación por lista (listwise, la opción por defecto aquí). Se descartan las filas con algún vacío. Es honesta y no distorsiona correlaciones, pero puede costar mucha muestra. Aceptable si se pierde menos del 5–10 % de los individuos y los vacíos son aleatorios (MCAR).
- Imputación por la media o la mediana. Rápida y disponible en la app. Preserva el tamaño de muestra pero reduce artificialmente la varianza y atenúa las correlaciones; la mediana es preferible si la variable es asimétrica. Declara siempre cuántas celdas imputaste.
- Imputación por regresión o por vecinos más cercanos (kNN). Conserva mejor la estructura de correlación.
- ACP iterativo regularizado (Josse y Husson, 2012). Es la opción recomendada en la literatura cuando hay más de un 10 % de vacíos: estima los valores faltantes usando la propia estructura de componentes, de forma iterativa.
Classical PCA requires a complete matrix. The options, from least to most sophisticated:
- Listwise deletion (the default option here). Rows with any gap are discarded. It is honest and does not distort correlations, but it can cost a lot of sample. Acceptable if less than 5–10% of the individuals are lost and the gaps are random (MCAR).
- Mean or median imputation. Fast and available in the app. It preserves the sample size but artificially reduces the variance and attenuates the correlations; the median is preferable if the variable is skewed. Always report how many cells you imputed.
- Regression or nearest-neighbour (kNN) imputation. It preserves the correlation structure better.
- Regularised iterative PCA (Josse & Husson, 2012). This is the option recommended in the literature when more than 10% of the values are missing: it estimates the missing values iteratively, using the component structure itself.
7 · Estandarización, transformaciones y la elección covarianza / correlación
Es la decisión más importante de la preparación, porque cambia por completo el resultado.
| Opción | Fórmula | Cuándo usarla |
|---|---|---|
| Sin escalar (ACP sobre covarianzas) | x | Todas las variables en la misma unidad y con varianzas comparables. Conserva las unidades originales. |
| Solo centrar | x − x̄ | Equivalente a la anterior; el centrado es obligatorio en cualquier caso. |
| Estandarización z (ACP sobre correlaciones) | (x − x̄) / s | Opción por defecto. Unidades distintas o varianzas muy dispares. Toda variable pesa lo mismo. |
| Pareto | (x − x̄) / √s | Metabolómica y espectros. Reduce la dominancia de las variables grandes sin igualarlas del todo; conserva parte de la estructura de magnitud. |
| VAST | (x − x̄) · x̄ / s² | Cuando interesa dar más peso a las variables estables (con bajo coeficiente de variación). |
| Rango [0, 1] | (x − mín) / (máx − mín) | Índices y variables acotadas. Muy sensible a atípicos. |
| Robusta | (x − mediana) / MAD | Presencia de atípicos que no se quieren eliminar. Reduce su influencia sobre el escalado. |
Transformaciones previas (se aplican antes del escalado):
- Logarítmica — variables con asimetría positiva fuerte, crecimiento multiplicativo, conteos grandes, concentraciones. Linealiza relaciones de potencia y estabiliza la varianza.
- Raíz cuadrada — conteos y abundancias con asimetría moderada (varianza proporcional a la media).
- Inversa (1/x) — tasas y tiempos con colas muy largas.
- Box–Cox / Yeo–Johnson — familia general que estima el exponente óptimo; útil cuando ninguna de las anteriores basta.
This is the most important decision of the preparation stage, because it completely changes the result.
| Option | Formula | When to use it |
|---|---|---|
| No scaling (PCA on covariances) | x | All variables in the same unit and with comparable variances. It keeps the original units. |
| Centring only | x − x̄ | Equivalent to the previous one; centring is mandatory in any case. |
| z standardisation (PCA on correlations) | (x − x̄) / s | Default option. Different units or very unequal variances. Every variable carries the same weight. |
| Pareto | (x − x̄) / √s | Metabolomics and spectra. It reduces the dominance of the large variables without fully equalising them; it keeps part of the magnitude structure. |
| VAST | (x − x̄) · x̄ / s² | When you want to give more weight to stable variables (with a low coefficient of variation). |
| Range [0, 1] | (x − min) / (max − min) | Indices and bounded variables. Very sensitive to outliers. |
| Robust | (x − median) / MAD | Outliers present that you do not want to remove. It reduces their influence on the scaling. |
Prior transformations (applied before scaling):
- Logarithmic — variables with strong positive skew, multiplicative growth, large counts, concentrations. It linearises power relationships and stabilises the variance.
- Square root — counts and abundances with moderate skew (variance proportional to the mean).
- Reciprocal (1/x) — rates and times with very long tails.
- Box–Cox / Yeo–Johnson — a general family that estimates the optimal exponent; useful when none of the above is enough.
8 · Variables e individuos activos frente a suplementarios
Distinción central en la escuela francesa de análisis de datos (Benzécri, Lebart, Escofier–Pagès), que es la que sigue esta plataforma:
- Variables activas. Construyen los ejes. Son las que entran en el cálculo de los valores propios.
- Variables cuantitativas suplementarias. No participan en el cálculo, pero se proyectan después sobre el círculo de correlaciones para ayudar a interpretar los ejes. Ideales para variables de resultado (rendimiento, biomasa final) que no quieres que definan los componentes.
- Variables cualitativas (grupos). Se usan para colorear los individuos, dibujar elipses de confianza por grupo y calcular los centroides de cada categoría sobre el plano factorial.
- Individuos suplementarios. Casos que se proyectan sin haber intervenido en la construcción de los ejes (por ejemplo, una nueva campaña de muestreo o un testigo).
En el Bloque 1 tú asignas ese papel a cada columna. La app propone una asignación automática, pero la decisión es tuya y es una decisión teórica, no estadística: las variables activas deben responder a una misma pregunta conceptual.
A central distinction in the French school of data analysis (Benzécri, Lebart, Escofier–Pagès), the one followed by this platform:
- Active variables. They build the axes. They are the ones entering the eigenvalue computation.
- Supplementary quantitative variables. They take no part in the computation, but are projected afterwards onto the correlation circle to help interpret the axes. Ideal for outcome variables (yield, final biomass) that you do not want to define the components.
- Qualitative variables (groups). Used to colour the individuals, draw confidence ellipses per group and compute the centroid of each category on the factor plane.
- Supplementary individuals. Cases projected without having taken part in building the axes (for example, a new sampling campaign or a control).
In Block 1 you assign that role to each column. The app proposes an automatic assignment, but the decision is yours, and it is a theoretical decision, not a statistical one: the active variables must answer one and the same conceptual question.
9 · ¿Cuántos componentes retener? ▸ Bloque 2
- Criterio de Kaiser (λ > 1). Solo válido sobre matriz de correlaciones: retiene los componentes que explican más que una variable original. Es el más usado y el que más tiende a sobreestimar el número de componentes, sobre todo con p > 30.
- Gráfico de sedimentación (scree plot, Cattell 1966). Se retienen los componentes que quedan antes del «codo». Es visual y algo subjetivo, pero muy informativo.
- Porcentaje de varianza acumulada. Retener hasta alcanzar un 70–80 % (ciencias naturales) o un 60 % (ciencias sociales). Nunca como criterio único.
- Análisis paralelo de Horn (1965). El más recomendado hoy. Compara los valores propios observados con los que se obtendrían de datos aleatorios del mismo tamaño; se retiene mientras λ_observado > λ_aleatorio.
- Bastón roto (broken stick). Compara con el reparto esperado de un segmento partido al azar. Muy usado en ecología; es conservador.
- MAP de Velicer y validación cruzada (predicción de celdas omitidas). Los más rigurosos.
- Criterio de interpretabilidad. Un componente que no se puede nombrar no sirve, aunque cumpla todos los criterios numéricos.
La buena práctica es combinar al menos tres criterios y reportarlos todos.
- Kaiser's criterion (λ > 1). Valid only on a correlation matrix: it retains the components that explain more than one original variable. It is the most widely used, and the one that most tends to overestimate the number of components, especially with p > 30.
- Scree plot (Cattell 1966). The components kept are those before the "elbow". It is visual and somewhat subjective, but very informative.
- Cumulative percentage of variance. Retain until 70–80% is reached (natural sciences) or 60% (social sciences). Never as the sole criterion.
- Horn's parallel analysis (1965). The most recommended today. It compares the observed eigenvalues with those that would be obtained from random data of the same size; components are retained while λ_observed > λ_random.
- Broken stick. It compares against the expected split of a stick broken at random. Widely used in ecology; it is conservative.
- Velicer's MAP and cross-validation (prediction of omitted cells). The most rigorous.
- Interpretability criterion. A component that cannot be named is useless, even if it meets every numerical criterion.
Good practice is to combine at least three criteria and report them all.
10 · Rotaciones: varimax, quartimax, equamax, promax, oblimin ▸ Bloque 3
Los ejes que salen del ACP maximizan varianza, pero no tienen por qué ser fáciles de interpretar: es habitual que todas las variables carguen sobre el primer componente. La rotación gira los ejes retenidos —sin cambiar la varianza total explicada por el conjunto— buscando una estructura simple en el sentido de Thurstone: que cada variable cargue fuerte en un solo componente y casi cero en los demás.
Rotaciones ortogonales (los componentes siguen sin correlacionarse):
- Varimax (Kaiser, 1958). La más usada. Maximiza la varianza de las cargas al cuadrado dentro de
cada componente, es decir, empuja las cargas hacia 0 o hacia ±1 columna por columna. Tiende a repartir las
variables entre varios componentes bien definidos.
maximizar Σ_k [ p·Σ_j (a²ⱼₖ)² − ( Σ_j a²ⱼₖ )² ] / p²
- Quartimax. Simplifica por filas: busca que cada variable cargue en el menor número posible de componentes. Suele producir un primer factor general muy dominante.
- Equamax. Compromiso entre varimax y quartimax (simplifica filas y columnas a la vez). Menos estable.
- Parsimax y Orthomax: familia general que engloba a las anteriores según un parámetro γ (γ = 0 quartimax, γ = 1 varimax, γ = p/2 equamax).
Rotaciones oblicuas (permiten que los componentes se correlacionen entre sí):
- Promax. Parte de una solución varimax y la eleva a una potencia κ (habitualmente 4) para exagerar el contraste entre cargas altas y bajas. Rápida y muy usada en muestras grandes.
- Oblimin directo. Controla el grado de oblicuidad con el parámetro δ (δ = 0 es el más usado).
- Geomin y Quartimin. Alternativas modernas, frecuentes en análisis factorial exploratorio.
The axes produced by PCA maximise variance, but they need not be easy to interpret: it is common for every variable to load on the first component. Rotation turns the retained axes —without changing the total variance explained by the set— seeking a simple structure in Thurstone's sense: each variable loading strongly on a single component and near zero on the rest.
Orthogonal rotations (the components remain uncorrelated):
- Varimax (Kaiser, 1958). The most widely used. It maximises the variance of the squared loadings within
each component, that is, it pushes the loadings towards 0 or towards ±1 column by column. It tends to spread the
variables over several well-defined components.
maximise Σ_k [ p·Σ_j (a²ⱼₖ)² − ( Σ_j a²ⱼₖ )² ] / p²
- Quartimax. It simplifies by rows: it tries to make each variable load on as few components as possible. It usually produces a very dominant first general factor.
- Equamax. A compromise between varimax and quartimax (simplifying rows and columns at once). Less stable.
- Parsimax and orthomax: a general family that encompasses the previous ones through a parameter γ (γ = 0 quartimax, γ = 1 varimax, γ = p/2 equamax).
Oblique rotations (they let the components correlate with each other):
- Promax. It starts from a varimax solution and raises it to a power κ (usually 4) to exaggerate the contrast between high and low loadings. Fast and widely used with large samples.
- Direct oblimin. It controls the degree of obliqueness through the parameter δ (δ = 0 is the most used).
- Geomin and quartimin. Modern alternatives, frequent in exploratory factor analysis.
11 · ACP frente a Análisis Factorial (y otras confusiones frecuentes)
| ACP | Análisis factorial exploratorio | |
|---|---|---|
| Objetivo | Resumir la varianza total | Explicar la varianza común (covarianza) |
| Modelo | Sin modelo: transformación de los datos | Modelo latente con error específico |
| Diagonal de R | Unos (varianza total) | Comunalidades estimadas |
| Dirección | Componente = función de las variables | Variable = función de los factores |
| Uso típico | Reducción, índices, visualización | Identificar constructos latentes |
Otras técnicas emparentadas, por si tu pregunta no es de ACP:
- Análisis de correspondencias (AC / ACM) — datos categóricos, tablas de contingencia.
- Análisis factorial de datos mixtos (AFDM) — cuantitativas y cualitativas juntas.
- Análisis factorial múltiple (AFM) — variables organizadas en grupos o bloques.
- Análisis discriminante — hay grupos conocidos y se quiere separarlos (el ACP no busca separar).
- Escalamiento multidimensional (NMDS) y PCoA — se parte de una matriz de distancias.
- Regresión por componentes principales (RCP) — los componentes se usan como predictores en regresión para esquivar la multicolinealidad.
| PCA | Exploratory factor analysis | |
|---|---|---|
| Goal | Summarise the total variance | Explain the common variance (covariance) |
| Model | No model: a transformation of the data | Latent model with a specific error term |
| Diagonal of R | Ones (total variance) | Estimated communalities |
| Direction | Component = function of the variables | Variable = function of the factors |
| Typical use | Reduction, indices, visualisation | Identifying latent constructs |
Other related techniques, in case your question is not a PCA question:
- Correspondence analysis (CA / MCA) — categorical data, contingency tables.
- Factor analysis of mixed data (FAMD) — quantitative and qualitative variables together.
- Multiple factor analysis (MFA) — variables organised in groups or blocks.
- Discriminant analysis — known groups exist and the goal is to separate them (PCA does not seek separation).
- Multidimensional scaling (NMDS) and PCoA — they start from a distance matrix.
- Principal component regression (PCR) — the components are used as predictors in a regression to sidestep multicollinearity.
12 · Errores frecuentes que arruinan un ACP
- Meter variables derivadas de otras (área = largo × ancho, porcentajes que suman 100). Genera dependencia lineal exacta y ejes artificiales.
- No estandarizar cuando las unidades difieren, y luego «descubrir» que el primer componente es simplemente la variable con números más grandes.
- Interpretar cargas pequeñas. Fija un umbral (0.30–0.40, o el de la tabla del punto 5) y respétalo.
- Retener componentes solo por el criterio de Kaiser.
- Interpretar el plano 1–2 cuando explica el 35 % de la varianza, sin advertirlo.
- Olvidar que el signo de un componente es arbitrario.
- Confundir la proximidad de dos individuos con la de dos variables: en el biplot se leen con reglas distintas (distancias entre puntos, ángulos entre vectores).
- Usar el ACP como si fuera una prueba de hipótesis o como sustituto de un análisis discriminante.
- No reportar el preprocesamiento (transformación, escalado, tratamiento de faltantes y de atípicos).
- Including variables derived from others (area = length × width, percentages that add to 100). This creates exact linear dependence and artificial axes.
- Not standardising when the units differ, and then "discovering" that the first component is simply the variable with the largest numbers.
- Interpreting small loadings. Set a threshold (0.30–0.40, or the one in the table of point 5) and stick to it.
- Retaining components on Kaiser's criterion alone.
- Interpreting the 1–2 plane when it explains 35% of the variance, without saying so.
- Forgetting that the sign of a component is arbitrary.
- Confusing the proximity of two individuals with that of two variables: in the biplot they are read with different rules (distances between points, angles between vectors).
- Using PCA as if it were a hypothesis test, or as a substitute for discriminant analysis.
- Not reporting the preprocessing (transformation, scaling, handling of missing data and of outliers).
13 · Glosario y lecturas
| Término | Significado |
|---|---|
| Valor propio (λ) | Varianza capturada por un componente. |
| Vector propio | Dirección del componente; sus elementos son los coeficientes. |
| Carga (loading) | Correlación entre una variable y un componente. |
| Puntuación (score) | Coordenada de un individuo sobre un componente. |
| Comunalidad | Proporción de la varianza de una variable explicada por los componentes retenidos. |
| cos² | Calidad de representación de un punto sobre un eje o plano. |
| Contribución | Aporte porcentual de un individuo o variable a la construcción de un eje. |
| Círculo de correlaciones | Representación de las variables en el plano factorial; radio 1. |
| Biplot | Individuos y variables superpuestos en el mismo plano. |
| Estructura simple | Situación ideal tras la rotación: cada variable carga en un solo componente. |
Referencias base de esta plataforma
- Kassambara, A. (2017). Practical Guide to Principal Component Methods in R. STHDA. — criterios de visualización, cos², contribuciones, elementos suplementarios.
- Gray, V. (ed.). Principal Component Analysis: Methods, Applications and Technology. Nova Science.
- Suryanarayana, T.M.V. y Mistry, P.B. Principal Component Regression for Crop Yield Estimation. Springer.
- Jolliffe, I.T. (2002). Principal Component Analysis, 2ª ed. Springer. — referencia canónica.
- Hair, J.F. et al. Multivariate Data Analysis. — reglas prácticas de tamaño de muestra y cargas.
- Lebart, Morineau y Piron. Statistique exploratoire multidimensionnelle. — escuela francesa, elementos activos y suplementarios.
| Term | Meaning |
|---|---|
| Eigenvalue (λ) | Variance captured by a component. |
| Eigenvector | Direction of the component; its elements are the coefficients. |
| Loading | Correlation between a variable and a component. |
| Score | Coordinate of an individual on a component. |
| Communality | Share of a variable's variance explained by the retained components. |
| cos² | Quality of representation of a point on an axis or plane. |
| Contribution | Percentage share of an individual or variable in building an axis. |
| Correlation circle | Representation of the variables on the factor plane; radius 1. |
| Biplot | Individuals and variables superimposed on the same plane. |
| Simple structure | The ideal situation after rotation: each variable loads on a single component. |
Reference works behind this platform
- Kassambara, A. (2017). Practical Guide to Principal Component Methods in R. STHDA. — visualisation criteria, cos², contributions, supplementary elements.
- Gray, V. (ed.). Principal Component Analysis: Methods, Applications and Technology. Nova Science.
- Suryanarayana, T.M.V. and Mistry, P.B. Principal Component Regression for Crop Yield Estimation. Springer.
- Jolliffe, I.T. (2002). Principal Component Analysis, 2nd ed. Springer. — the canonical reference.
- Hair, J.F. et al. Multivariate Data Analysis. — practical rules for sample size and loadings.
- Lebart, Morineau and Piron. Statistique exploratoire multidimensionnelle. — the French school, active and supplementary elements.
1.1 · Sube tu base de datos
Formato .xlsx, .xls o .csv. La primera fila debe contener los nombres de las variables y cada fila siguiente una observación (una planta, una parcela, un individuo, una muestra…). Una sola tabla por hoja, sin filas ni columnas en blanco intercaladas y sin celdas combinadas.
Arrastra tu archivo aquí o haz clic para elegirlo
.xlsx · .xls · .csv · .tsv · .txt
Extracción de componentes: teoría
Qué se calcula exactamente en este paso y cómo se decide cuántos ejes conservar.
1 · Qué produce la extracción
Sobre la matriz ya preparada en el Bloque 1 se calcula la matriz de dispersión —de correlaciones si estandarizaste, de covarianzas si no— y se resuelve su descomposición espectral. De ahí salen cuatro objetos que se usarán en todos los bloques siguientes:
- Valores propios λ₁ ≥ λ₂ ≥ … ≥ λ_p. Cuánta varianza captura cada eje. Su suma es la varianza total (igual a p si se trabaja con correlaciones, porque cada variable estandarizada aporta varianza 1).
- Vectores propios. Los coeficientes de la combinación lineal que define cada componente.
- Puntuaciones (scores). La coordenada de cada individuo sobre cada eje: es la nueva base de datos reducida, y se puede descargar y usar en regresión, ANOVA o clasificación.
- Cargas (loadings). La correlación entre cada variable original y cada componente,
carga_jk = v_jk · √λ_k / s_j. Es lo que se interpreta y lo que dibuja el círculo de correlaciones.
El porcentaje de varianza de un componente es λ_k / Σλ. El acumulado hasta k
es la fracción de información original que conservarías si te quedaras con esos k ejes.
On the matrix already prepared in Block 1, the dispersion matrix is computed —the correlation matrix if you standardised, the covariance matrix if you did not— and its spectral decomposition is solved. That yields four objects that will be used in every following block:
- Eigenvalues λ₁ ≥ λ₂ ≥ … ≥ λ_p. How much variance each axis captures. Their sum is the total variance (equal to p when working with correlations, because each standardised variable contributes a variance of 1).
- Eigenvectors. The coefficients of the linear combination that defines each component.
- Scores. The coordinate of each individual on each axis: this is the new reduced data set, and it can be downloaded and used in regression, ANOVA or classification.
- Loadings. The correlation between each original variable and each component,
loading_jk = v_jk · √λ_k / s_j. This is what is interpreted and what the correlation circle draws.
The percentage of variance of a component is λ_k / Σλ. The cumulative value up to k
is the fraction of the original information you would keep if you retained those k axes.
2 · Los criterios de retención, uno por uno
Kaiser–Guttman (λ > 1). Retiene los componentes que explican más varianza que una variable original estandarizada. Simple y universal, pero es el que más sobreestima: con p grande tiende a retener aproximadamente p/3 componentes aunque los datos sean puro ruido. Solo tiene sentido sobre la matriz de correlaciones; con covarianzas el equivalente es «λ mayor que el promedio de los λ».
Jolliffe (λ > 0.7). Versión relajada de Kaiser propuesta por Jolliffe al observar que el corte en 1 descartaba componentes útiles por variación muestral.
Gráfico de sedimentación / codo (Cattell, 1966). Se grafican los λ en orden y se busca el punto donde la caída se aplana: los componentes anteriores al codo son estructura, los posteriores son ruido. La app estima el codo con la máxima segunda diferencia (aceleración) de la curva, pero conviene mirarlo a ojo: es un criterio visual por naturaleza.
Varianza acumulada. Retener hasta alcanzar un porcentaje objetivo: 70–80 % en ciencias naturales, 60 % en ciencias sociales, más del 90 % si el ACP es un paso previo a una compresión o a una regresión. Nunca debe usarse como criterio único, porque siempre se puede alcanzar cualquier umbral añadiendo ejes.
Análisis paralelo de Horn (1965). Es hoy el criterio mejor evaluado en los estudios de simulación. La idea: generar muchas matrices de datos sin estructura del mismo tamaño (mismo n, mismo p), calcular sus valores propios y quedarse solo con los componentes cuyo λ observado supera el λ que produciría el azar. La app ofrece dos formas de generar esos datos:
- Permutación de tus propios datos (opción por defecto): baraja cada columna por separado. Destruye la correlación entre variables pero conserva las distribuciones marginales reales y las varianzas, así que no supone normalidad y funciona igual con covarianzas que con correlaciones.
- Datos normales aleatorios: la versión clásica de Horn, sobre matrices de correlación simuladas.
Se compara contra el percentil 95 de los λ aleatorios (criterio de Glorfeld), más estricto y estable que compararlos contra la media.
Bastón roto (broken stick, Frontier 1976). Si se parte un bastón de longitud 1 en p
trozos al azar, la longitud esperada del trozo k-ésimo es
b_k = (1/p) · Σ_{i=k..p} (1/i). Se retienen los componentes cuyo λ supere ese reparto aleatorio.
Es conservador —suele retener menos ejes que los demás criterios— y es muy usado en ecología numérica.
MAP de Velicer (1976, revisado en 2000). Para cada número posible de componentes extraídos se calcula el promedio de las correlaciones parciales al cuadrado entre las variables una vez removidos esos componentes. Mientras se extrae varianza común, ese promedio baja; cuando se empieza a extraer varianza específica, vuelve a subir. Se retiene el número que da el mínimo. La versión con cuarta potencia (MAP4) es algo más precisa con muestras pequeñas.
Criterio de interpretabilidad. El último y el más importante: un componente que no se puede nombrar ni conectar con la teoría del problema no sirve para nada, cumpla o no los criterios numéricos.
EE(λ) ≈ λ·√(2/(n−1)), la aproximación asintótica de Anderson. Solo es válida bajo
normalidad multivariante y con valores propios bien separados; tómala como una referencia del orden de
magnitud de la incertidumbre, no como un intervalo exacto. Si dos λ consecutivos difieren menos que su error
estándar, el plano que forman es inestable y su orientación no debe interpretarse.Kaiser–Guttman (λ > 1). It retains the components explaining more variance than one standardised original variable. Simple and universal, but it is the one that overestimates the most: with a large p it tends to retain roughly p/3 components even when the data are pure noise. It only makes sense on the correlation matrix; with covariances the equivalent is "λ greater than the average of the λ".
Jolliffe (λ > 0.7). A relaxed version of Kaiser's rule, proposed by Jolliffe after noting that a cut at 1 discarded useful components because of sampling variation.
Scree plot / elbow (Cattell, 1966). The λ are plotted in order and one looks for the point where the decline flattens out: the components before the elbow are structure, those after it are noise. The app estimates the elbow with the maximum second difference (acceleration) of the curve, but it is worth looking at it by eye: it is a visual criterion by nature.
Cumulative variance. Retain until a target percentage is reached: 70–80% in the natural sciences, 60% in the social sciences, more than 90% if the PCA is a step before compression or regression. It must never be used as the only criterion, because any threshold can always be reached by adding axes.
Horn's parallel analysis (1965). Today this is the best-evaluated criterion in simulation studies. The idea: generate many data matrices without structure of the same size (same n, same p), compute their eigenvalues, and keep only the components whose observed λ exceeds the λ that chance would produce. The app offers two ways of generating those data:
- Permutation of your own data (default): it shuffles each column separately. This destroys the correlation between variables but preserves the real marginal distributions and variances, so it assumes no normality and works equally well with covariances and with correlations.
- Random normal data: Horn's classical version, on simulated correlation matrices.
The comparison is against the 95th percentile of the random λ (Glorfeld's criterion), stricter and more stable than comparing against the mean.
Broken stick (Frontier 1976). If a stick of length 1 is broken into p
pieces at random, the expected length of the k-th piece is
b_k = (1/p) · Σ_{i=k..p} (1/i). The components whose λ exceeds that random share are retained.
It is conservative —it usually retains fewer axes than the other criteria— and it is widely used in numerical ecology.
Velicer's MAP (1976, revised in 2000). For each possible number of extracted components, the average of the squared partial correlations between the variables is computed once those components have been removed. While common variance is being extracted, that average falls; when specific variance starts to be extracted, it rises again. The number giving the minimum is retained. The fourth-power version (MAP4) is somewhat more accurate with small samples.
Interpretability criterion. The last and the most important: a component that cannot be named or connected to the theory of the problem is of no use, whether or not it meets the numerical criteria.
SE(λ) ≈ λ·√(2/(n−1)), Anderson's asymptotic approximation. It is valid only under
multivariate normality and with well-separated eigenvalues; take it as a guide to the order of
magnitude of the uncertainty, not as an exact interval. If two consecutive λ differ by less than their standard
error, the plane they form is unstable and its orientation should not be interpreted.3 · Errores frecuentes en este paso
- Usar solo el criterio de Kaiser y acabar con seis componentes de los cuales tres son ruido.
- Retener tantos ejes que la varianza acumulada llegue al 95 % — eso ya no es reducir dimensiones.
- Interpretar el plano 1–2 sin decir qué porcentaje explica.
- Aplicar el corte λ > 1 a un ACP hecho sobre covarianzas, donde la varianza total no es p y el número 1 no significa nada.
- Olvidar que el signo de cada componente es arbitrario: si tu CP1 sale invertido respecto a lo esperado, puedes multiplicarlo por −1 sin cambiar nada, solo hay que declararlo.
- Using Kaiser's criterion alone and ending up with six components, three of which are noise.
- Retaining so many axes that the cumulative variance reaches 95% — that is no longer dimension reduction.
- Interpreting the 1–2 plane without saying what percentage it explains.
- Applying the λ > 1 cut to a PCA run on covariances, where the total variance is not p and the number 1 means nothing.
- Forgetting that the sign of each component is arbitrary: if your PC1 comes out inverted with respect to what you expected, you can multiply it by −1 without changing anything, as long as you declare it.
2.1 · Extraer los componentes
Se usará la matriz preparada en el Bloque 1.
Rotación: teoría
Por qué rotar, qué método elegir y cómo se reporta.
1 · El problema que resuelve la rotación
Los ejes que salen del ACP maximizan varianza, no interpretabilidad. El resultado típico es que casi todas las variables cargan fuerte en el CP1 —porque ese eje recoge el «tamaño» o el nivel general del fenómeno— y los componentes siguientes quedan como contrastes difíciles de nombrar.
La rotación aprovecha un hecho geométrico: el subespacio de los k componentes retenidos representa la misma información sea cual sea la orientación de los ejes dentro de él. Se pueden girar esos ejes buscando una posición más interpretable sin perder nada.
Cambia: el reparto de la varianza entre los componentes, las cargas individuales, y el orden —tras rotar los ejes ya no van de mayor a menor λ, por eso aquí se renombran CPR1, CPR2…
The axes that come out of PCA maximise variance, not interpretability. The typical result is that almost every variable loads strongly on PC1 —because that axis picks up the "size" or the general level of the phenomenon— and the following components are left as contrasts that are hard to name.
Rotation exploits a geometric fact: the subspace of the k retained components represents the same information whatever the orientation of the axes inside it. Those axes can be turned in search of a more interpretable position without losing anything.
Changed: the split of the variance between the components, the individual loadings, and the order —after rotating, the axes no longer run from the largest λ to the smallest, which is why they are renamed RPC1, RPC2… here.
2 · Estructura simple: el objetivo de Thurstone
Thurstone (1947) definió los criterios de la estructura simple, que siguen siendo el objetivo:
- Cada variable debe tener al menos una carga próxima a cero.
- Cada componente debe tener varias variables con carga próxima a cero.
- Para cada par de componentes, debe haber variables con carga alta en uno y nula en el otro.
- Con cuatro o más componentes, la mayoría de las variables deben tener carga nula en la mayoría de ellos.
- Los solapamientos (cargas altas en varios componentes) deben ser pocos.
La app mide qué tan cerca estás de ese ideal con tres indicadores:
- Variables limpias: superan el umbral en un solo componente.
- Cargas cruzadas: superan el umbral en dos o más. Son las que estropean la interpretación.
- Complejidad de Hofmann:
c = (Σa²)² / Σa⁴. Vale 1 si la variable carga en un único componente y crece hacia k si se reparte por igual entre todos. Una media por debajo de 1.3 indica una solución muy limpia.
Thurstone (1947) defined the criteria for simple structure, which are still the goal:
- Each variable must have at least one loading close to zero.
- Each component must have several variables with a loading close to zero.
- For each pair of components, there must be variables loading high on one and zero on the other.
- With four or more components, most variables must have a zero loading on most of them.
- Overlaps (high loadings on several components) must be few.
The app measures how close you are to that ideal with three indicators:
- Clean variables: they exceed the threshold on a single component.
- Cross-loadings: they exceed the threshold on two or more. These are what spoil the interpretation.
- Hofmann complexity:
c = (Σa²)² / Σa⁴. It equals 1 if the variable loads on a single component and grows towards k if it is spread equally over all of them. A mean below 1.3 indicates a very clean solution.
3 · Rotaciones ortogonales: la familia ortomax
Mantienen los ejes perpendiculares, así que los componentes siguen sin correlacionarse. Todas maximizan el mismo criterio general con distinto parámetro γ:
| Método | γ | Qué simplifica | Cuándo usarla |
|---|---|---|---|
| Quartimax | 0 | Filas (variables) | Cuando esperas un factor general y quieres que cada variable cargue en pocos ejes. |
| Varimax | 1 | Columnas (componentes) | La opción por defecto. Da componentes bien diferenciados, cada uno con su grupo de variables. |
| Equamax | k/2 | Ambas | Compromiso. Puede ser inestable con pocas variables o pocos componentes. |
| Parsimax | p(k−1)/(p+k−2) | Ambas, con más peso | Busca la máxima parsimonia global; poco frecuente pero bien fundamentada. |
Varimax (Kaiser, 1958) maximiza la varianza de las cargas al cuadrado dentro de cada columna: empuja cada carga hacia 0 o hacia ±1. Es, con diferencia, la rotación más reportada en la literatura.
Normalización de Kaiser. Antes de rotar, cada fila de la matriz de cargas se divide entre su comunalidad (se lleva a longitud 1) y después se restaura. Así las variables mal representadas no pesan menos en la rotación. Es el comportamiento por defecto habitual en los programas de estadística, y aquí también.
They keep the axes perpendicular, so the components remain uncorrelated. They all maximise the same general criterion with a different parameter γ:
| Method | γ | What it simplifies | When to use it |
|---|---|---|---|
| Quartimax | 0 | Rows (variables) | When you expect a general factor and want each variable to load on few axes. |
| Varimax | 1 | Columns (components) | The default option. It gives well-differentiated components, each with its own group of variables. |
| Equamax | k/2 | Both | A compromise. It can be unstable with few variables or few components. |
| Parsimax | p(k−1)/(p+k−2) | Both, with more weight | It seeks maximum overall parsimony; uncommon but well founded. |
Varimax (Kaiser, 1958) maximises the variance of the squared loadings within each column: it pushes each loading towards 0 or towards ±1. It is by far the most reported rotation in the literature.
Kaiser normalisation. Before rotating, each row of the loading matrix is divided by its communality (brought to length 1) and afterwards restored. That way poorly represented variables do not count for less in the rotation. It is the usual default in statistical software, and here as well.
4 · Rotaciones oblicuas: cuando los ejes pueden correlacionarse
Renuncian a la perpendicularidad. En fenómenos biológicos, ecológicos y sociales las dimensiones reales suelen estar relacionadas, así que forzar la ortogonalidad puede ser un artificio.
- Quartimin — la oblicua más simple; oblimin con δ = 0.
- Oblimin directo — familia general con parámetro δ: valores negativos producen ejes más próximos a la ortogonalidad, positivos los hacen más oblicuos. δ = 0 es el uso habitual y casi siempre suficiente.
- Promax (Hendrickson y White, 1964) — parte de una solución varimax y eleva las cargas a una potencia κ (habitualmente 4) para exagerar el contraste entre altas y bajas; luego ajusta por mínimos cuadrados. Es muy rápida y la preferida con muestras grandes.
Estructura: la correlación simple variable–componente. Incluye el efecto indirecto de los otros componentes, así que sus valores son más altos y menos «limpios».
Φ (phi): la matriz de correlaciones entre los propios componentes. Si todas sus correlaciones fuera de la diagonal son menores que ≈0.15, la oblicuidad no aportó nada y conviene volver a varimax; si superan 0.32 (10 % de varianza compartida), la rotación oblicua está claramente justificada.
Con rotación ortogonal patrón y estructura coinciden, y Φ es la identidad.
They give up perpendicularity. In biological, ecological and social phenomena the real dimensions are usually related, so forcing orthogonality can be an artefact.
- Quartimin — the simplest oblique rotation; oblimin with δ = 0.
- Direct oblimin — a general family with a parameter δ: negative values produce axes closer to orthogonality, positive ones make them more oblique. δ = 0 is the usual choice and almost always enough.
- Promax (Hendrickson and White, 1964) — it starts from a varimax solution and raises the loadings to a power κ (usually 4) to exaggerate the contrast between high and low ones; it then fits by least squares. It is very fast and the preferred choice with large samples.
Structure: the simple variable–component correlation. It includes the indirect effect of the other components, so its values are higher and less "clean".
Φ (phi): the matrix of correlations between the components themselves. If all its off-diagonal correlations are smaller than ≈0.15, obliqueness added nothing and it is better to go back to varimax; if they exceed 0.32 (10% shared variance), the oblique rotation is clearly justified.
With an orthogonal rotation, pattern and structure coincide, and Φ is the identity.
5 · ¿Es legítimo rotar componentes principales?
Es una pregunta razonable y conviene tenerla clara. La rotación nació en el análisis factorial, donde la orientación de los ejes es indeterminada por definición del modelo. En ACP los ejes sí están determinados: son las direcciones de máxima varianza.
Al rotar componentes principales se pierden dos propiedades: los componentes dejan de maximizar varianza individualmente y dejan de estar ordenados por λ. Lo que sí se conserva es la varianza total del subconjunto retenido y las comunalidades.
La práctica está muy extendida y es aceptada —los principales programas de estadística la ofrecen por defecto— siempre que se declare con claridad. Fórmula habitual en la sección de métodos:
It is a reasonable question and worth being clear about. Rotation was born in factor analysis, where the orientation of the axes is undetermined by the very definition of the model. In PCA the axes are determined: they are the directions of maximum variance.
When principal components are rotated, two properties are lost: the components no longer maximise variance individually and they are no longer ordered by λ. What is preserved is the total variance of the retained subset and the communalities.
The practice is very widespread and accepted —the main statistical packages offer it by default— provided it is clearly declared. The usual wording in the methods section:
6 · Cómo elegir el umbral de las cargas
Una carga solo se interpreta si supera un umbral. Criterios de uso común:
- 0.30 — mínimo para considerarla; explica un 9 % de la varianza de la variable.
- 0.40 — el más habitual; el que trae la app por defecto.
- 0.50 — prácticamente significativa; la variable comparte un 25 % con el componente.
- 0.55–0.75 — necesario con muestras pequeñas (ver la tabla de Hair del Bloque 1, sección 5).
El umbral debe fijarse antes de mirar los resultados y declararse en el reporte. Cambiarlo hasta que la solución «se vea bien» es una forma de sobreajuste.
A loading is only interpreted if it exceeds a threshold. Criteria in common use:
- 0.30 — the minimum for considering it; it explains 9% of the variable's variance.
- 0.40 — the most usual; the app's default.
- 0.50 — practically significant; the variable shares 25% with the component.
- 0.55–0.75 — needed with small samples (see Hair's table in Block 1, section 5).
The threshold must be set before looking at the results and declared in the report. Changing it until the solution "looks good" is a form of overfitting.
3.1 · Elegir la rotación
Mapas factoriales: teoría
Cómo se construye cada gráfico y, sobre todo, cómo se lee sin equivocarse.
1 · Las dos nubes: individuos y variables
Un ACP produce dos representaciones distintas del mismo análisis, y se leen con reglas diferentes. Confundirlas es el error más común al interpretar un ACP.
| Nube de individuos | Nube de variables | |
|---|---|---|
| Qué se dibuja | Un punto por observación | Un vector por variable |
| Coordenada | Puntuación (score) sobre el eje | Carga = correlación con el eje |
| Qué se lee | Distancias entre puntos: dos individuos próximos tienen perfiles parecidos | Ángulos entre vectores respecto al origen |
| Escala | La de las puntuaciones (varía por eje) | Siempre entre −1 y 1 |
Reglas de lectura del círculo de correlaciones:
- Vectores que apuntan en la misma dirección (ángulo pequeño): variables correlacionadas positivamente.
- Vectores opuestos (ángulo de 180°): correlación negativa.
- Vectores perpendiculares (90°): variables no correlacionadas en ese plano.
- Vectores largos, cerca del círculo unitario: la variable está bien representada en el plano. Un vector corto no significa que la variable sea poco importante, sino que su información está en otros ejes; no interpretes su dirección.
A PCA produces two different representations of the same analysis, and they are read with different rules. Confusing them is the most common mistake when interpreting a PCA.
| Cloud of individuals | Cloud of variables | |
|---|---|---|
| What is drawn | One point per observation | One vector per variable |
| Coordinate | Score on the axis | Loading = correlation with the axis |
| What is read | Distances between points: two nearby individuals have similar profiles | Angles between vectors from the origin |
| Scale | That of the scores (it varies by axis) | Always between −1 and 1 |
Rules for reading the correlation circle:
- Vectors pointing in the same direction (small angle): positively correlated variables.
- Opposite vectors (180° angle): negative correlation.
- Perpendicular vectors (90°): variables uncorrelated in that plane.
- Long vectors, close to the unit circle: the variable is well represented in the plane. A short vector does not mean the variable is unimportant, but that its information lies on other axes; do not interpret its direction.
2 · cos² y contribuciones: las dos medidas que evitan malinterpretar
cos² (coseno al cuadrado) — calidad de representación. Es el cuadrado del coseno del ángulo entre el punto y el eje; equivale a la proporción de la información de ese elemento que el eje recoge:
Para las variables, con ACP sobre correlaciones, cos² = carga². Sumado sobre todos los ejes da 1.
Interpretación: cos² ≥ 0.7 muy buena, 0.5–0.7 aceptable, 0.3–0.5 pobre, < 0.3 no interpretable en
ese plano. Un punto con cos² bajo está mal proyectado: aparece en el mapa donde no le corresponde.
Contribución — cuánto pesa el elemento en la construcción del eje.
Suma 100 % por eje. La referencia es el valor esperado si todos contribuyeran por igual: 100/p para variables y 100/n para individuos (la línea roja de la figura de contribuciones). Los elementos por encima de esa línea son los que definen el eje.
cos² (squared cosine) — quality of representation. It is the squared cosine of the angle between the point and the axis; it equals the share of that element's information captured by the axis:
For variables, with PCA on correlations, cos² = loading². Summed over all axes it gives 1.
Interpretation: cos² ≥ 0.7 very good, 0.5–0.7 acceptable, 0.3–0.5 poor, < 0.3 not interpretable in
that plane. A point with a low cos² is badly projected: it appears on the map where it does not belong.
Contribution — how much the element weighs in building the axis.
It adds up to 100% per axis. The benchmark is the value expected if all contributed equally: 100/p for variables and 100/n for individuals (the red line in the contributions figure). The elements above that line are the ones that define the axis.
3 · El biplot y su escala
El biplot (Gabriel, 1971) superpone las dos nubes en un solo gráfico. Es el más informativo y el más fácil de sobreinterpretar, porque individuos y variables viven en escalas distintas y hay que multiplicar los vectores por un factor arbitrario para que se vean juntos. Ese factor no tiene significado estadístico: por eso aquí es un control deslizante, no un número fijo.
Lo que sí se puede leer en un biplot:
- La proyección de un individuo sobre la dirección de una variable estima su valor en ella: si un punto cae lejos, en el sentido de la flecha, ese individuo tiene valores altos en esa variable.
- Los grupos de individuos que se sitúan en la dirección de un grupo de variables se caracterizan por ellas.
Lo que no se debe leer: la longitud absoluta de las flechas comparada con la distancia entre puntos, ni distancias entre un punto y una punta de flecha.
The biplot (Gabriel, 1971) superimposes the two clouds on a single plot. It is the most informative and the easiest to over-read, because individuals and variables live on different scales and the vectors have to be multiplied by an arbitrary factor for them to be seen together. That factor has no statistical meaning: this is why it is a slider here, not a fixed number.
What can be read in a biplot:
- The projection of an individual onto the direction of a variable estimates its value on that variable: if a point falls far along the direction of the arrow, that individual has high values on that variable.
- Groups of individuals lying in the direction of a group of variables are characterised by them.
What must not be read: the absolute length of the arrows compared with the distance between points, nor distances between a point and an arrowhead.
4 · Elipses: cuál dibujar y qué significa cada una
Las elipses resumen dónde cae cada grupo, pero responden a preguntas distintas y confundirlas cambia por completo la conclusión.
| Tipo | Qué encierra | Cuándo usarla |
|---|---|---|
| Elipse de concentración (de los datos) |
Aproximadamente el 95 % de las observaciones del grupo, bajo normalidad bivariante. Su tamaño no depende de n. | Para describir la dispersión y el solapamiento real entre grupos. Es la opción por defecto aquí. |
| Elipse de confianza de la media | La región donde está el centroide del grupo con un 95 % de confianza. Es √n veces más pequeña. | Para argumentar que dos grupos difieren en promedio. Con n grande se vuelve diminuta. |
| Envolvente convexa | El polígono mínimo que contiene todos los puntos del grupo. No supone ninguna distribución. | Cuando los grupos no son elípticos o hay pocos individuos. |
Recuerda además que el ACP no busca separar grupos: es una técnica no supervisada que ignora la variable de agrupación. Si tu objetivo es discriminar, la herramienta es el análisis discriminante o el PLS-DA.
Ellipses summarise where each group falls, but they answer different questions, and confusing them completely changes the conclusion.
| Type | What it encloses | When to use it |
|---|---|---|
| Concentration ellipse (of the data) |
Roughly 95% of the group's observations, under bivariate normality. Its size does not depend on n. | To describe the spread and the real overlap between groups. It is the default option here. |
| Confidence ellipse of the mean | The region where the group's centroid lies with 95% confidence. It is √n times smaller. | To argue that two groups differ on average. With a large n it becomes tiny. |
| Convex hull | The smallest polygon containing all the group's points. It assumes no distribution. | When the groups are not elliptical or there are few individuals. |
Remember, too, that PCA does not seek to separate groups: it is an unsupervised technique that ignores the grouping variable. If your goal is to discriminate, the tool is discriminant analysis or PLS-DA.
5 · Elementos suplementarios
Las variables que marcaste como suplementarias en el Bloque 1 no intervinieron en el cálculo de los ejes, pero se proyectan ahora sobre ellos:
- Cuantitativas suplementarias: su coordenada sobre cada eje es simplemente su correlación con las puntuaciones de ese eje. Aparecen en el círculo con trazo discontinuo. Sirven para «etiquetar» los ejes con variables de resultado sin dejar que los definan.
- Categorías suplementarias: cada nivel se sitúa en el centroide (la media) de los individuos que lo comparten. Una categoría alejada del origen caracteriza esa zona del plano.
Es la forma correcta de responder a «¿se separan mis tratamientos?» sin contaminar el análisis: los ejes se construyen solo con las mediciones, y el factor se proyecta después.
The variables you marked as supplementary in Block 1 took no part in computing the axes, but they are now projected onto them:
- Supplementary quantitative variables: their coordinate on each axis is simply their correlation with the scores of that axis. They appear in the circle with a dashed stroke. They serve to "label" the axes with outcome variables without letting them define the axes.
- Supplementary categories: each level is placed at the centroid (the mean) of the individuals sharing it. A category far from the origin characterises that region of the plane.
This is the correct way to answer "do my treatments separate?" without contaminating the analysis: the axes are built from the measurements alone, and the factor is projected afterwards.
6 · Convenciones al publicar una figura de ACP
- Los rótulos de los ejes siempre con el porcentaje de varianza: «CP1 (41.2 %)».
- Declara si la solución está rotada y con qué método.
- Indica el tipo de elipse y su nivel de confianza en el pie de figura.
- El signo de un eje es arbitrario: puedes invertirlo para que la lectura sea más natural, pero decláralo.
- Usa una paleta apta para daltónicos si la figura va a impresión (Okabe–Ito está disponible en el editor).
- Exporta en SVG si la revista lo acepta; si pide mapa de bits, mínimo 300 ppp (opción 4×) y preferiblemente 600 ppp (8×).
- Axis labels always with the percentage of variance: "PC1 (41.2%)".
- State whether the solution is rotated and by which method.
- Give the type of ellipse and its confidence level in the figure caption.
- The sign of an axis is arbitrary: you may invert it to make the reading more natural, but declare it.
- Use a colour-blind-safe palette if the figure goes to print (Okabe–Ito is available in the editor).
- Export in SVG if the journal accepts it; if it asks for a raster, use at least 300 dpi (the 4× option) and preferably 600 dpi (8×).
4.1 · Construir los mapas
Interpretación: teoría
De los números a las frases: cómo se nombra un eje, cómo se caracteriza y hasta dónde se puede llegar sin salirse de lo que el ACP permite afirmar.
1 · Qué significa «interpretar» un componente
Un componente es una combinación lineal; por sí mismo no significa nada. Interpretarlo es encontrar el concepto de tu disciplina que explica por qué esas variables varían juntas, y darle un nombre. El procedimiento estándar:
- Mira las variables con |carga| por encima del umbral que fijaste (aquí, el mismo del Bloque 3).
- Separa las de carga positiva de las de carga negativa. Si hay de ambos signos, el eje es un contraste: opone dos conjuntos de características, no mide «más o menos» de una sola cosa.
- Busca qué tienen en común: ¿tamaño?, ¿velocidad de crecimiento?, ¿estrategia adquisitiva frente a conservativa?, ¿un gradiente ambiental?
- Ponle un nombre corto y sustantivo. Si no consigues nombrarlo, ese componente probablemente no deba retenerse, por muchos criterios numéricos que lo respalden.
- Confirma el nombre con los elementos suplementarios y con los individuos extremos: ¿los que están en el lado positivo son los que tu conocimiento del sistema esperaría?
A component is a linear combination; on its own it means nothing. Interpreting it means finding the concept from your discipline that explains why those variables vary together, and giving it a name. The standard procedure:
- Look at the variables with a |loading| above the threshold you set (here, the same as in Block 3).
- Separate those with a positive loading from those with a negative one. If there are both signs, the axis is a contrast: it opposes two sets of characteristics, it does not measure "more or less" of a single thing.
- Look for what they have in common: size? growth rate? an acquisitive versus a conservative strategy? an environmental gradient?
- Give it a short, substantive name. If you cannot name it, that component probably should not be retained, however many numerical criteria support it.
- Confirm the name with the supplementary elements and with the extreme individuals: are those on the positive side the ones your knowledge of the system would expect?
2 · El valor test (v.test): cómo se caracteriza un eje con categorías
Es el estadístico de la escuela francesa de análisis de datos, el que se usa en la descripción automática de dimensiones (Lebart et al., 2006). Responde a: ¿el centroide de esta categoría está más lejos del origen de lo que cabría esperar por azar?
Bajo la hipótesis nula de que los n_q individuos de la categoría son una muestra al azar del total, v se distribuye aproximadamente como una normal estándar. Por tanto:
- |v| ≥ 1.96 → la categoría se sitúa significativamente a un lado del eje (p < 0.05).
- |v| ≥ 2.58 → p < 0.01.
- El signo indica de qué lado del eje cae la categoría.
Es especialmente útil porque convierte «el Sitio C está a la derecha» en una afirmación cuantificada, y porque se puede ordenar: las categorías con |v| más alto son las que mejor caracterizan el eje.
This is the statistic of the French school of data analysis, the one used in the automatic description of dimensions (Lebart et al., 2006). It answers: is the centroid of this category further from the origin than one would expect by chance?
Under the null hypothesis that the n_q individuals of the category are a random sample of the whole, v is approximately standard normal. Therefore:
- |v| ≥ 1.96 → the category lies significantly on one side of the axis (p < 0.05).
- |v| ≥ 2.58 → p < 0.01.
- The sign indicates which side of the axis the category falls on.
It is especially useful because it turns "Site C is on the right" into a quantified statement, and because it can be ordered: the categories with the highest |v| are the ones that best characterise the axis.
3 · ¿Los grupos difieren en los componentes? — y la trampa de la circularidad
Una vez tienes las puntuaciones, es natural comparar los grupos sobre cada eje con un ANOVA de una vía. La app lo hace y añade el tamaño del efecto, que importa más que el valor p:
- η² (eta cuadrada) = proporción de la varianza del componente explicada por el factor. Referencias de Cohen: 0.01 pequeño, 0.06 medio, 0.14 grande.
- ω² (omega cuadrada) es una versión menos sesgada de lo mismo; con n pequeña conviene reportarla.
- Se incluye además el Kruskal–Wallis, que no supone normalidad, por si las puntuaciones son asimétricas o hay grupos muy desiguales.
• Estás haciendo k pruebas (una por componente) sin corrección por comparaciones múltiples.
• Los ejes fueron elegidos a posteriori por maximizar varianza, así que el p-valor es optimista.
Úsalos como descripción. Si tu pregunta de investigación es «¿difieren los grupos?», la respuesta correcta es un MANOVA, un PERMANOVA o un análisis discriminante sobre las variables originales.
Lo mismo vale para los valores p de las correlaciones variable–eje: los ejes se construyeron con esas variables, así que esas pruebas son circulares. Están ahí como ayuda de lectura, no como evidencia.
Once you have the scores, it is natural to compare the groups on each axis with a one-way ANOVA. The app does this and adds the effect size, which matters more than the p value:
- η² (eta squared) = share of the component's variance explained by the factor. Cohen's benchmarks: 0.01 small, 0.06 medium, 0.14 large.
- ω² (omega squared) is a less biased version of the same thing; with a small n it is worth reporting.
- The Kruskal–Wallis test is also included, which assumes no normality, in case the scores are skewed or the groups are very unequal.
• You are running k tests (one per component) without correction for multiple comparisons.
• The axes were chosen a posteriori by maximising variance, so the p value is optimistic.
Use them as description. If your research question is "do the groups differ?", the correct answer is a MANOVA, a PERMANOVA or a discriminant analysis on the original variables.
The same holds for the p values of the variable–axis correlations: the axes were built from those variables, so those tests are circular. They are there as a reading aid, not as evidence.
4 · Calidad del ajuste: matriz reproducida y RMSR
Los k componentes retenidos permiten reconstruir la matriz de correlaciones original:
La diferencia entre lo observado y lo reproducido son los residuos. Su resumen es el RMSR (residuo cuadrático medio, sobre los elementos fuera de la diagonal):
- RMSR < 0.05 — ajuste bueno; los componentes retenidos reproducen bien la estructura.
- 0.05 – 0.08 — aceptable.
- > 0.08 — se está perdiendo estructura: considera retener un componente más.
Regla complementaria muy usada: que menos del 10 % de los residuos superen 0.05 en valor absoluto. Un par concreto con residuo grande señala dos variables cuya relación no queda recogida por ningún eje.
The k retained components allow the original correlation matrix to be reconstructed:
The difference between what is observed and what is reproduced gives the residuals. Their summary is the RMSR (root mean square residual, over the off-diagonal elements):
- RMSR < 0.05 — good fit; the retained components reproduce the structure well.
- 0.05 – 0.08 — acceptable.
- > 0.08 — structure is being lost: consider retaining one more component.
A widely used complementary rule: fewer than 10% of the residuals should exceed 0.05 in absolute value. A particular pair with a large residual points to two variables whose relationship is captured by no axis.
5 · Cómo se escriben los resultados
La sección de resultados de un ACP suele tener esta estructura, y la app te genera un borrador con tus propios números al final del bloque:
- Adecuación: n, p, KMO, Bartlett, tratamiento de faltantes y escalado.
- Extracción: cuántos componentes, con qué criterio, y qué porcentaje explican.
- Rotación: método y normalización, si se usó.
- Interpretación: cada componente con su porcentaje, sus variables marcadoras y su nombre.
- Ajuste: RMSR o comunalidades.
- Figuras: círculo de correlaciones y mapa de individuos, con el porcentaje en los rótulos de eje.
Errores frecuentes al redactar: llamar «factores» a los componentes (son cosas distintas), decir que un componente «explica» una variable cuando la relación es descriptiva, afirmar que el ACP «demostró» que los grupos difieren, y omitir el preprocesamiento.
The results section of a PCA usually has this structure, and the app generates a draft with your own numbers at the end of the block:
- Adequacy: n, p, KMO, Bartlett, handling of missing data and scaling.
- Extraction: how many components, by which criterion, and what percentage they explain.
- Rotation: method and normalisation, if used.
- Interpretation: each component with its percentage, its marker variables and its name.
- Fit: RMSR or communalities.
- Figures: correlation circle and map of individuals, with the percentage in the axis labels.
Common mistakes in the write-up: calling the components "factors" (they are different things), saying that a component "explains" a variable when the relationship is descriptive, claiming that the PCA "proved" that the groups differ, and omitting the preprocessing.
5.1 · Configuración
Agrupamiento sobre los componentes: teoría
El análisis factorial ordena a los individuos en un espacio continuo. Agruparlos es una operación distinta y opcional: convierte esa nube en un pequeño número de clases con nombre.
1 · Por qué agrupar sobre los ejes y no sobre las variables
El HCPC —hierarchical clustering on principal components, de Husson, Josse y Pagès— agrupa sobre las coordenadas factoriales, no sobre la tabla original. Hay dos razones y las dos son prácticas:
- Quita el ruido. Los ejes que no retuviste son, en su mayor parte, variación sin estructura. Agrupar sobre los primeros ejes es agrupar sobre una versión limpia de la tabla.
- Unifica los tipos de dato. Sobre las coordenadas de un ACM o de un AFDM se puede usar la distancia euclídea aunque las variables originales sean cualitativas. Sin ese paso previo habría que elegir un coeficiente de similitud entre las decenas que existen, y esa elección cambia el resultado.
HCPC — hierarchical clustering on principal components, by Husson, Josse and Pagès — clusters on the factor coordinates, not on the original table. There are two reasons and both are practical:
- It removes noise. The axes you did not retain are, for the most part, variation without structure. Clustering on the first axes means clustering on a cleaned-up version of the table.
- It unifies data types. Euclidean distance can be used on the coordinates of an MCA or a FAMD even when the original variables are qualitative. Without that step you would have to pick a similarity coefficient among the dozens that exist, and that choice changes the result.
2 · El criterio de Ward
En cada paso se unen los dos grupos cuya fusión menos inercia intra añade. El coste de unir A con B es
Δ(A, B) = [wA · wB / (wA + wB)] · d²(gA, gB)
con g los centros y w las masas. Es la altura a la que aparece la fusión en el dendrograma: una rama alta significa que unir esos dos grupos costó caro, es decir, que eran distintos.
Ward es el criterio natural aquí porque habla el mismo idioma que el análisis factorial: los dos reparten la inercia total entre una parte intra y una parte entre. La inercia entre grupos dividida por la total es la proporción de la nube que la partición explica.
At each step the two clusters whose merger adds the least within-cluster inertia are joined. The cost of merging A with B is
Δ(A, B) = [wA · wB / (wA + wB)] · d²(gA, gB)
with g the centres and w the masses. This is the height at which the merger appears in the dendrogram: a tall branch means that joining those two clusters was expensive, that is, that they were different.
Ward is the natural criterion here because it speaks the same language as factor analysis: both split the total inertia into a within part and a between part. Between-cluster inertia divided by the total is the share of the cloud the partition accounts for.
3 · Cuántos grupos
No existe una respuesta única, igual que con el número de componentes. Se calculan tres reglas y se muestra si coinciden:
- Mayor salto de altura. El corte se pone donde la siguiente fusión sería desproporcionadamente cara.
- Mayor pérdida relativa de inercia. La misma idea, pero relativa a la altura anterior; es la regla que usan Husson y Josse y la que decide cuando las tres discrepan.
- Máxima silueta media. Para cada individuo compara su distancia media a los suyos con la distancia media al grupo vecino más próximo. Por encima de 0.5 la estructura es clara; por debajo de 0.25 no hay estructura que reportar, y conviene decirlo.
There is no single answer, just as with the number of components. Three rules are computed and their agreement is shown:
- Largest height jump. The cut goes where the next merger would be disproportionately expensive.
- Largest relative loss of inertia. The same idea, but relative to the previous height; this is the rule Husson and Josse use, and the one that decides when the three disagree.
- Maximum average silhouette. For each individual it compares the mean distance to its own cluster with the mean distance to the nearest neighbouring cluster. Above 0.5 the structure is clear; below 0.25 there is no structure to report, and it is worth saying so.
4 · La consolidación por k-medias
El árbol tiene una limitación: una fusión, una vez hecha, no se deshace. Un individuo que quedó del lado equivocado en un paso temprano se queda ahí. La consolidación arregla eso: partiendo de los centros del corte, reasigna cada individuo al centro más cercano y repite hasta que nadie se mueva.
Nunca empeora la inercia intra —si lo hiciera sería un error de programación, y la suite de pruebas lo comprueba—. A cambio, la partición final puede dejar de coincidir con el dendrograma: si algún individuo cambió de grupo, el árbol y el mapa discrepan en esos casos y aquí se te avisa de cuántos son.
The tree has one limitation: once a merger is made, it is never undone. An individual that ended up on the wrong side at an early step stays there. Consolidation fixes that: starting from the centres of the cut, it reassigns each individual to the nearest centre and repeats until nobody moves.
It can never worsen the within-cluster inertia — if it did, that would be a programming error, and the test suite checks for it. In exchange, the final partition may stop matching the dendrogram: if any individual changed cluster, the tree and the map disagree in those cases, and you are told how many.
5 · Cómo se describe un grupo
Un grupo se describe diciendo en qué se aparta del conjunto. El valor test mide ese apartamiento en desviaciones típicas, bajo el modelo de que los individuos del grupo se hubieran sacado al azar y sin reemplazo del total. Por encima de |1.96| la diferencia no se explica por ese reparto al azar.
Además se dan dos clases de individuos:
- Paragones: los más cercanos al centro de su grupo. Son los ejemplares típicos, los que enseñarías para explicar qué es ese grupo.
- Individuos específicos: los más alejados de todos los demás centros. Son los que menos se confunden con otro grupo, útiles cuando hay que elegir casos para un estudio de seguimiento.
A cluster is described by saying how it departs from the whole. The test value measures that departure in standard deviations, under the model that the individuals in the cluster had been drawn at random and without replacement from the total. Above |1.96| the difference is not explained by that random draw.
Two kinds of individuals are also given:
- Paragons: those closest to the centre of their cluster. They are the typical members, the ones you would show to explain what that cluster is.
- Specific individuals: those furthest from every other centre. They are the ones least confused with another cluster, useful when cases have to be picked for a follow-up study.
6 · Lo que un agrupamiento no demuestra
El algoritmo siempre devuelve grupos, también sobre datos sin ninguna estructura. Tres cautelas que conviene tener escritas antes de mirar el resultado:
- Los valores test describen la partición, no la ponen a prueba. Los grupos se construyeron precisamente para maximizar esas diferencias, así que presentarlos como contrastes de hipótesis es circular.
- Una silueta baja con una inercia entre grupos alta significa que la partición reparte bien la varianza pero que los grupos no están separados: la nube es continua y la has cortado en rebanadas.
- Si los grupos coinciden con una variable que ya conocías —el tratamiento, la especie, el sitio—, eso no es un descubrimiento del agrupamiento: compruébalo con esa variable como suplementaria y dilo así.
The algorithm always returns clusters, including on data with no structure at all. Three cautions worth writing down before looking at the result:
- Test values describe the partition, they do not test it. The clusters were built precisely to maximise those differences, so presenting them as hypothesis tests is circular.
- A low silhouette with a high between-cluster inertia means the partition splits the variance well but the clusters are not separated: the cloud is continuous and you have sliced it.
- If the clusters coincide with a variable you already knew — treatment, species, site — that is not a discovery made by the clustering: check it with that variable as supplementary and say so.
5b.1 · Configuración del agrupamiento
Informe y exportación
Reúne todo lo hecho en los cinco bloques anteriores en un documento único y en un paquete de archivos listo para archivar o compartir.
Qué produce este bloque
- Informe en HTML autocontenido: portada, métodos redactados automáticamente con tus propios parámetros, tablas con formato de publicación, las figuras tal como las editaste (vectoriales, incrustadas en el propio archivo), interpretación, limitaciones y referencias. Un solo archivo, sin dependencias: se abre en cualquier navegador y se puede enviar por correo.
- PDF: el informe trae hoja de estilo de impresión (evita cortar tablas y figuras a mitad de página). Usa el botón de imprimir y elige «Guardar como PDF» en el diálogo del navegador.
- Paquete ZIP completo: el informe, todas las figuras en el formato y la resolución que elijas, y hasta 14 tablas de resultados en CSV, organizados en carpetas.
- A self-contained HTML report: cover page, methods drafted automatically with your own parameters, publication-formatted tables, the figures exactly as you edited them (vector graphics, embedded in the file itself), interpretation, limitations and references. A single file, with no dependencies: it opens in any browser and can be sent by e-mail.
- PDF: the report carries a print stylesheet (it avoids splitting tables and figures across pages). Use the print button and choose "Save as PDF" in the browser dialog.
- A complete ZIP package: the report, every figure in the format and resolution you choose, and up to 14 result tables as CSV, organised in folders.
6.1 · Estado del análisis
Solo se incluirán en el informe los bloques que hayas ejecutado.
6.2 · Configuración del informe
Secciones a incluir
Figuras
6.4 · Cómo citar PCAPro
Si publicas resultados obtenidos con esta plataforma, cita el software. La misma referencia va en el informe generado y en el archivo CITATION.bib del paquete ZIP.