Appendix A. Canonical variation partitioning: statistical details.
This Appendix shows that canonical variation partitioning, as per Borcard et al. (1992), provides a correct partitioning of the variation of a response data table Y.
The variation in a single response variable y is measured by the sum of the squared differences of the values to the mean of that variable (SS). If y was centered before the calculations, SS(y) = ∑ yi2. If we analyze y by ordinary least-squares regression on an explanatory variable x, also centered, the fitted values can be computed as
= x[x'x]¯1x'y. The amount of variation in the fitted values, SS(
), can be computed in the same way as for SS(y). The portion of the variation of y explained by x is given by the coefficient of determination, R2 = SS(
)/SS(y).
Let us move to multivariate data. Canonical redundancy analysis (RDA) of a response table Y by an explanatory table X consists in two steps: (1) a series of multiple regressions of the individual variables of Y on X, which produces a table of fitted values
; this is followed by (2) a principal component analysis of
which produces the canonical eigenvalues and eigenvectors and the canonical ordination scores (Legendre and Legendre 1998, Section 11.1). The second step is not necessary for variation partitioning. Assuming that the individual variables in Y were centered on their means, the total variation in Y, SS(Y), is obtained by computing the sum of the squared values in table Y, just as in the univariate case. Because Y was centered, the individual columns of
are also centered on zero; SS(
) is obtained in the same way as SS(Y), by computing the sum of the squared values in
. The portion of the variation of Y explained by X, SS(
)/SS(Y), is the bimultivariate redundancy statistic R2Y|X (Miller and Farr 1971), which is the canonical equivalent of the coefficient of determination R2; it is called the RDA "trace" statistic in the program Canoco (ter Braak and Smilauer 2002). Canonical correspondence analysis (CCA) is similar to RDA, with two small differences: (1) the table subjected to analysis is not Y but a table, called
by Legendre and Legendre (1998, Section 11.2), which contains a transformation of the original species presence-absence or abundance data into contributions to chi-square; and (2) the regression involves weights, given by the row sums of table Y divided by the sum total of the values in Y, and written in a diagonal matrix of weights which intervenes in the regression equations. These weights are also taken into account when standardizing the explanatory variables X at the beginning of the analysis. The rest of the calculations are similar to RDA. This short exposé shows that canonical analysis produces estimates of the portion of the variation of Y explained by X that are similar to the familiar coefficients of determination of regression analysis.
All the fractions of variation that will be displayed in the simulation result tables (see the "Simulation study" section of the main paper) were obtained from 3 multiple regressions (for a single y response variable) or 3 canonical analyses (when analyzing a multivariate response table Y), followed by simple calculations: (1) RDA(Y|environmental matrix X1) produces the bimultivariate R2 (or "trace") statistic SS(
)/SS(Y) for [a+b], (2) RDA(Y|spatial matrix X2) produces the statistic SS(
)/SS(Y) for [b+c], and (3) RDA(Y|matrices X1 and X2) produces the statistic SS(
)/SS(Y) for [a+b+c]. From these results, one can calculate [a] = [a+b+c] – [b+c], [c] = [a+b+c] – [a+b], and [b] = [a+b] + [b+c] – [a+b+c] (Fig. 1 of the main paper). The residual variation is given by the bimultivariate coefficient of nondetermination, [d] = 1 – [a+b+c]. The fractions [a], [b], [c], and [d] are additive and sum to 1, as in partial regression analysis (see for instance Legendre and Legendre 1998, Subsection 10.3.5).
While calculating the fractions of variation only involves simple canonical analyses, partial canonical analyses are necessary to test the significance of fractions [a] and [c] of variation partitioning. Partial canonical analysis is the direct extension of partial regression to multivariate response data.
LITERATURE CITED
Borcard, D., P. Legendre, and P. Drapeau. 1992. Partialling out the spatial component of ecological variation. Ecology 73:10451055.
Legendre, P., and L. Legendre. 1998. Numerical ecology, 2nd English edition. Elsevier, Amsterdam, The Netherlands.
Miller, J. K., and S. D. Farr. 1971. Bimultivariate redundancy: a comprehensive measure of interbattery relationship. Multivariate Behavioral Research 6:313324.
ter Braak, C. J. F. and P. Smilauer. 2002. Canoco reference manual and CanoDraw for Windows user's guide: software for canonical community ordination (version 4.5). Microcomputer Power, Ithaca, New York, USA.