$ec2ko
ec ko
1 1.1.1.1 K00001
2 1.1.1.2 K00002
3 1.1.1.3 K00003
$ec2cpd
ec cpd
1 1.1.1.1 C00001
2 1.1.1.115 C00001
3 1.1.1.132 C00001
[1] 6754 7270
Knowledge-based integration of multi-omics data: annotation-based analysis of specific interactions
Wednesday, September 30, 2026
Annotation-based Analysis of Specific Interactions
Correlate every feature of data set \(Y\) with every feature of data set \(X\), then show a heatmap or a list of “significant” pairs.
Two problems:
Microbe–metabolite example: instead of correlating taxa with metabolites, correlate the microbial genes (functions) that make or use each metabolite. Then pathway knowledge tells us which pairs to test.
9 metabolites × 8 enzymes: of the 72 possible pairs, only 17 are known to interact. The other 55 tests cannot be interpreted, but still count in the FDR correction.
Two data sets on the same \(N\) samples: \(\mathbf{Y}\) (\(M_Y\) features) and \(\mathbf{X}\) (\(M_X\) features).
All-vs-all association matrix (\(M_Y \times M_X\)):
\[ \rho(\mathbf{Y}, \mathbf{X})_{ij} = \rho\big(f^{(Y)}_i, f^{(X)}_j\big) \]
Binary (bi)adjacency matrix from a knowledge database:
\[ A_{ij} = \begin{cases} 1 & \text{if features } i \text{ and } j \text{ are known to interact} \\ 0 & \text{otherwise} \end{cases} \]
Masked association matrix, the core output:
\[ \mathbf{R} = \rho(\mathbf{Y}, \mathbf{X}) \circ \mathbf{A} \]
Masked pairs are never tested, so they don’t enter the FDR correction.
Knowledge databases list which features interact:
kegg_link()$ec2ko
ec ko
1 1.1.1.1 K00001
2 1.1.1.2 K00002
3 1.1.1.3 K00003
$ec2cpd
ec cpd
1 1.1.1.1 C00001
2 1.1.1.115 C00001
3 1.1.1.132 C00001
[1] 6754 7270
Does the association between \(y\) and \(x\) depend on a covariate, such as treatment, disease status or a clinical score?
| Type | What changes | Tested as |
|---|---|---|
| Full | total association of \(y\) with \(x\), incl. interactions | \(y \sim x \times (\text{covariates})\) |
| Disjointed | the slope of the association | the \(x{:}\text{covariate}\) interaction term |
| Emergent | the strength of the association (how tightly points follow the line) | how the spread around the line depends on the covariate |
Error(subject), as in stats::aov()Simulated Krebs-cycle data with spiked-in effects (krebsDemoWeb()).
weaveWeb(): combine the two feature tables, the linking information and the sample metadata into an AnansiWeb object; the features are matched through the biadjacency matrixanansi(): fit the models for every linked feature pair; returns one row per pair with correlations, full, disjointed and emergent statistics, p-values and adjusted p-valuesplotAnansi(), or your own ggplot2 code: visualizeasMAE() and asTSE() convert backpairwiseApply(): apply any function to every linked pairstats::lm() / aov() conventionsBoehme et al. (2021): can faecal microbiota transplantation (FMT) from young donor mice reverse ageing-related changes in aged mice?
anansi::AnansiWeb S7_object with 36 observations:
tableY: cpd (3 features)
tableX: ko (2476 features)
Use @ to access: tableX, tableY, dictionary, metadata.
all_vs_all tested
7428 65
feature_Y feature_X All_r.values All_t.values All_p.values
1 C00186 K00016 -0.08985529 0.5260699 0.60225456
2 C00186 K00101 0.28661046 1.7443940 0.09012569
3 C00041 K00259 -0.21708011 1.2967052 0.20346437
4 C00064 K00265 -0.25811198 1.5578255 0.12853530
Aged oFMT Aged yFMT All disjointed_Legend
4 4 4 4
emergent_Legend full Young yFMT
4 4 4
| Association | Method | Statistics | Significance |
|---|---|---|---|
| Pooled (All) | correlation | \(\rho\), \(t\) | p, adjusted p |
| Per group | correlation | \(\rho\), \(t\) | p, adjusted p |
| Full | linear model | \(R^2\), \(F\) | p, adjusted p |
| Disjointed | linear model | \(R^2\), \(F\) | p, adjusted p |
| Emergent | linear model | \(R^2\), \(F\) | p, adjusted p |
pairs <- getFeaturePairs(web)
grp_cols <- c("Young yFMT" = "#2166ac", "Aged oFMT" = "#b2182b",
"Aged yFMT" = "#ef8a62")
plots <- lapply(pairs[seq_len(6)], function(p) {
ggplot(p, aes(y = .data[[colnames(p)[1]]], x = .data[[colnames(p)[2]]])) +
facet_grid(colnames(p)[1] ~ colnames(p)[2])
})
wrap_plots(plots, ncol = 3) &
geom_point(aes(fill = FMT_metadata$Legend), shape = 21,
show.legend = FALSE) &
geom_smooth(aes(colour = FMT_metadata$Legend), method = "lm",
se = FALSE, show.legend = FALSE) &
scale_fill_manual(values = grp_cols, aesthetics = c("colour", "fill")) &
labs(x = NULL, y = NULL) &
theme_bw(base_size = 12)# Continuous covariates and several terms
anansi(web, formula = ~ group + score, groups = "group")
# Repeated measures: random intercept per subject
anansi(web, formula = ~ group + Error(subject_id), groups = "group")
# Bioconductor containers in and out
mae <- asMAE(web)
web <- weaveWeb(mae, tableY = "cpd", tableX = "ko", link = kegg_link())
# Custom links: any two-column data.frame of feature IDs
web <- weaveWeb(y ~ x, link = my_links, tableY = tY, tableX = tX)
# Apply any function to each linked pair
pairwiseApply(web, FUN = function(x, y) cor(x, y, method = "spearman"))| Approach | Idea | Uses prior knowledge? |
|---|---|---|
| All-vs-all correlation | every pair, then FDR | no |
| SparCC, SPIEC-EASI, proportionality | compositional association measures | no |
| HAllA | cluster features first, then test clusters | no (data-driven clusters) |
| DIABLO, MINT (mixOmics) | (sparse) PLS to discriminate phenotypes | no |
| mmvec | neural network for microbe–metabolite co-occurrence | no |
| anansi | test only documented feature pairs | yes |
anansi constrains the hypothesis space instead of letting the data decide which pairs to test.
Package and documentation: https://thomazbastiaanssen.github.io/anansi/