Multi-table methods

Classical approaches for integrating multiple data tables

Wednesday, September 30, 2026

Overview

Multi-omics studies measure several tables on the same samples (taxa, metabolites, host markers, …).

Before fitting latent factor models such as MOFA2 or Joint-RPCA, classical multivariate statistics offers simple tools to

  • test whether two tables are associated at all
  • select the features that drive the association
  • combine more than two tables into a common view
  • use prior structure, such as the phylogenetic tree

Adapted from Holmes (2019) and Holmes and Huber (2019, Ch. 9).

Example data

Rat cecum samples with microbiota, metabolites and biomarkers (HintikkaXOData in mia)

library(mia)
library(MultiAssayExperiment)

data("HintikkaXOData", package = "mia")
mae <- HintikkaXOData

# Microbiota: genus level, clr
mae[[1]] <- agglomerateByRank(mae[[1]], rank = "Genus")
mae[[1]] <- transformAssay(mae[[1]], method = "clr", pseudocount = TRUE)

# Metabolites: log10 NMR signals
mae[[2]] <- transformAssay(mae[[2]], assay.type = "nmr", method = "log10")

# Sample x feature matrices with matching samples
x <- t(assay(mae[[1]], "clr"))
y <- t(assay(mae[[2]], "log10"))
y <- y[, colSums(is.na(y)) == 0]   # drop metabolites with missing values

40 samples; 262 genera; 30 metabolites

Are two tables associated?

Global association between tables

Test the tables as a whole before looking at feature pairs.

  • Mantel test (Mantel 1967): correlation between the two sample distance matrices; significance by permutation
  • RV coefficient (Robert and Escoufier 1976): a “correlation between tables”, between 0 and 1

\[ RV(X, Y) = \frac{\mathrm{tr}(XX^\top YY^\top)} {\sqrt{\mathrm{tr}\big((XX^\top)^2\big)\,\mathrm{tr}\big((YY^\top)^2\big)}} \]

  • Co-inertia analysis (Dolédec and Chessel 1994): shared axes of co-variation between two tables measured on the same samples

Review: Josse and Holmes (2016)

Global association in R

# Mantel test: Euclidean distances on clr / log10 data
vegan::mantel(dist(x), dist(y), permutations = 999)

# RV coefficient with a permutation test
ade4::RV.rtest(as.data.frame(x), as.data.frame(y), nrepet = 999)

In this data:

  • Mantel r = 0.23 (p = 0.001)
  • RV = 0.48 (p = 0.001; expected under permutation 0.20)

➡️ Microbiota and metabolite profiles are associated, but which features drive this?

Sparse CCA

Many more features than samples

  • Canonical correlation analysis (CCA) finds linear combinations of the features in each table that are maximally correlated
  • Classical CCA needs more samples than features: it overfits omics data
  • Sparse CCA adds an L1 penalty and keeps only a few features from each table (Witten, Tibshirani, and Hastie 2009)
  • Use it as a screening step, then visualize the selected features (e.g. PCA biplot)

Example: microbiota and metabolites

Mouse gut, two diets, wild type vs. knockout (Kashyap et al. 2013):

  • 12 samples; 637 metabolite features; 20,609 taxa (filtered first)
  • Sparse CCA selected 5 taxa and 15 metabolites
  • Correlation between the two selected combinations: 0.97
  • PCA of the 20 selected features separated the two diets

Analysis from Holmes and Huber (2019, Ch. 9).

Sparse CCA in R

cc <- PMA::CCA(x, y,
               typex = "standard", typez = "standard",
               penaltyx = 0.15, penaltyz = 0.15)

# Selected features
colnames(x)[cc$u[, 1] != 0]
colnames(y)[cc$v[, 1] != 0]
cc$cors

# PCA of the selected features only
selected <- cbind(x[, cc$u[, 1] != 0], y[, cc$v[, 1] != 0])
pca <- prcomp(selected, scale. = TRUE)

In HintikkaXOData: 10 genera and butyrate selected; correlation 0.88

More than two tables

A “compromise” across tables

  • Compute the RV coefficient between every pair of tables
  • Weight each table by how well it agrees with the others
  • The weighted combination is the compromise: ordinate it and project each table onto it to see where tables agree or differ
  • Classical methods: STATIS (Lavit et al. 1994; Abdi et al. 2012), multiple factor analysis (Escofier and Pagès 1994), multiple co-inertia analysis (Chessel and Hanafi 1996)
  • Latent factor models (MOFA2, Joint-RPCA) pursue a related goal in a probabilistic or robust low-rank framework

Prior knowledge

The phylogenetic tree

  • Taxa are not independent features: related taxa often behave alike
  • DPCoA (double principal coordinate analysis) uses phylogenetic distances between taxa in the ordination (Pavoine, Dufour, and Chessel 2004; Purdom 2011)
  • Adaptive generalized PCA tunes how strongly the tree is used, from ignoring it (PCA) to relying on it fully (DPCoA) (Fukuyama 2019)
  • Example: repeated antibiotic courses in three adults (Dethlefsen and Relman 2011); tree-aware ordination separated disturbed from undisturbed samples in some subjects (Fukuyama et al. 2017)

DPCoA in mia

data("GlobalPatterns", package = "mia")
tse <- agglomerateByRank(GlobalPatterns, rank = "Genus", update.tree = TRUE)
tse <- tse[rowSums(assay(tse, "counts")) > 0, ]

# Uses the phylogenetic tree stored in rowTree(tse)
tse <- addDPCoA(tse, name = "DPCoA")

scater::plotReducedDim(tse, "DPCoA", colour_by = "SampleType")

Summary

  • Global tests (Mantel, RV, co-inertia): are two tables associated?
  • Sparse CCA: which features drive the association?
  • Compromise methods (STATIS, MFA, MCIA): more than two tables
  • Tree-aware ordination (DPCoA, adaptive gPCA): use prior knowledge

➡️ Multi-omics slides · Joint-RPCA slides

References

Abdi, Hervé, Lynne J. Williams, Dominique Valentin, and Mohammed Bennani-Dosse. 2012. “STATIS and DISTATIS: Optimum Multitable Principal Component Analysis and Three Way Metric Multidimensional Scaling.” WIREs Computational Statistics 4 (2): 124–67. https://doi.org/10.1002/wics.198.
Chessel, D., and M. Hanafi. 1996. “Analyses de La Co-Inertie de K Nuages de Points.” Revue de Statistique Appliquée 44 (2): 35–60.
Dethlefsen, Les, and David A. Relman. 2011. “Incomplete Recovery and Individualized Responses of the Human Distal Gut Microbiota to Repeated Antibiotic Perturbation.” Proceedings of the National Academy of Sciences 108 (Supplement 1): 4554–61. https://doi.org/10.1073/pnas.1000087107.
Dolédec, S., and D. Chessel. 1994. “Co-Inertia Analysis: An Alternative Method for Studying Species–Environment Relationships.” Freshwater Biology 31: 277–94. https://doi.org/10.1111/j.1365-2427.1994.tb01741.x.
Escofier, Brigitte, and Jérôme Pagès. 1994. “Multiple Factor Analysis (AFMULT Package).” Computational Statistics & Data Analysis 18 (1): 121–40. https://doi.org/10.1016/0167-9473(94)90135-X.
Fukuyama, Julia. 2019. “Adaptive gPCA: A Method for Structured Dimensionality Reduction with Applications to Microbiome Data.” The Annals of Applied Statistics 13 (2). https://doi.org/10.1214/18-AOAS1227.
Fukuyama, Julia, Laurie Rumker, Kris Sankaran, Pratheepa Jeganathan, Les Dethlefsen, David A. Relman, and Susan P. Holmes. 2017. “Multidomain Analyses of a Longitudinal Human Microbiome Intestinal Cleanout Perturbation Experiment.” PLOS Computational Biology 13 (8): e1005706. https://doi.org/10.1371/journal.pcbi.1005706.
Holmes, Susan. 2019. “Multidomain Methods: Using the Data, All the Data.” Lecture, Statistics for Microbiome Data workshop, Pune, India.
Holmes, Susan, and Wolfgang Huber. 2019. Modern Statistics for Modern Biology. Cambridge University Press. https://www.huber.embl.de/msmb/.
Josse, Julie, and Susan Holmes. 2016. “Measuring Multivariate Association and Beyond.” Statistics Surveys 10: 132–67. https://doi.org/10.1214/16-SS116.
Kashyap, Purna C., Angela Marcobal, Luke K. Ursell, Samuel A. Smits, et al. 2013. “Genetically Dictated Change in Host Mucus Carbohydrate Landscape Exerts a Diet-Dependent Effect on the Gut Microbiota.” Proceedings of the National Academy of Sciences 110 (42): 17059–64. https://doi.org/10.1073/pnas.1306070110.
Lavit, Christine, Yves Escoufier, Robert Sabatier, and Pierre Traissac. 1994. “The ACT (STATIS Method).” Computational Statistics & Data Analysis 18 (1): 97–119. https://doi.org/10.1016/0167-9473(94)90134-1.
Mantel, Nathan. 1967. “The Detection of Disease Clustering and a Generalized Regression Approach.” Cancer Research 27 (2): 209–20.
Pavoine, Sandrine, Anne-Béatrice Dufour, and Daniel Chessel. 2004. “From Dissimilarities Among Species to Dissimilarities Among Communities: A Double Principal Coordinate Analysis.” Journal of Theoretical Biology 228 (4): 523–37. https://doi.org/10.1016/j.jtbi.2004.02.014.
Purdom, Elizabeth. 2011. “Analysis of a Data Matrix and a Graph: Metagenomic Data and the Phylogenetic Tree.” The Annals of Applied Statistics 5 (4). https://doi.org/10.1214/10-AOAS402.
Robert, P., and Y. Escoufier. 1976. “A Unifying Tool for Linear Multivariate Statistical Methods: The RV-Coefficient.” Applied Statistics 25 (3): 257–65. https://doi.org/10.2307/2347233.
Witten, Daniela M., Robert Tibshirani, and Trevor Hastie. 2009. “A Penalized Matrix Decomposition, with Applications to Sparse Principal Components and Canonical Correlation Analysis.” Biostatistics 10 (3): 515–34. https://doi.org/10.1093/biostatistics/kxp008.