anansi

Knowledge-based integration of multi-omics data: annotation-based analysis of specific interactions

Wednesday, September 30, 2026

anansi

Annotation-based Analysis of Specific Interactions

  • Integrates two omics data sets, e.g. microbial functions and metabolites
  • Tests only the feature pairs that a knowledge database (e.g. KEGG) says can interact
  • Tests whether associations differ between groups or along a covariate

Bastiaanssen, Quinn, and Cryan (2023) · R/Bioconductor package anansi

anansi logo

Motivation

The usual approach: all-vs-all correlation

Correlate every feature of data set \(Y\) with every feature of data set \(X\), then show a heatmap or a list of “significant” pairs.

Two problems:

  1. Hard to interpret: results come without biological context; many significant pairs cannot be explained mechanistically
  2. Wastes statistical power: every test adds a p-value to the multiple-testing correction (FDR), including tests of pairs that cannot biologically interact

Microbe–metabolite example: instead of correlating taxa with metabolites, correlate the microbial genes (functions) that make or use each metabolite. Then pathway knowledge tells us which pairs to test.

Example: the Krebs cycle

9 metabolites × 8 enzymes: of the 72 possible pairs, only 17 are known to interact. The other 55 tests cannot be interpreted, but still count in the FDR correction.

The method

Masking the association matrix

Two data sets on the same \(N\) samples: \(\mathbf{Y}\) (\(M_Y\) features) and \(\mathbf{X}\) (\(M_X\) features).

All-vs-all association matrix (\(M_Y \times M_X\)):

\[ \rho(\mathbf{Y}, \mathbf{X})_{ij} = \rho\big(f^{(Y)}_i, f^{(X)}_j\big) \]

Binary (bi)adjacency matrix from a knowledge database:

\[ A_{ij} = \begin{cases} 1 & \text{if features } i \text{ and } j \text{ are known to interact} \\ 0 & \text{otherwise} \end{cases} \]

Masked association matrix, the core output:

\[ \mathbf{R} = \rho(\mathbf{Y}, \mathbf{X}) \circ \mathbf{A} \]

Masked pairs are never tested, so they don’t enter the FDR correction.

Where does the adjacency matrix come from?

Knowledge databases list which features interact:

  • KEGG (Kanehisa and Goto 2000): compounds (cpd), enzymes (EC numbers), orthologues (KO); shipped with anansi as kegg_link()
  • Also MetaCyc, CAZy, HMDB, EggNOG, …; or any custom two-column table
library(anansi)

lapply(kegg_link(), head, 3)
$ec2ko
       ec     ko
1 1.1.1.1 K00001
2 1.1.1.2 K00002
3 1.1.1.3 K00003

$ec2cpd
         ec    cpd
1   1.1.1.1 C00001
2 1.1.1.115 C00001
3 1.1.1.132 C00001
# Biadjacency matrix linking KEGG compounds to KEGG orthologues
web_demo <- weaveWeb(cpd ~ ko, link = kegg_link())
dim(dictionary(web_demo))
[1] 6754 7270

Beyond correlation: differential associations

Does the association between \(y\) and \(x\) depend on a covariate, such as treatment, disease status or a clinical score?

Type What changes Tested as
Full total association of \(y\) with \(x\), incl. interactions \(y \sim x \times (\text{covariates})\)
Disjointed the slope of the association the \(x{:}\text{covariate}\) interaction term
Emergent the strength of the association (how tightly points follow the line) how the spread around the line depends on the covariate
  • The paradigm comes from differential proportionality (Erb et al. 2017; Quinn et al. 2017), here applied outside the simplex
  • Covariates can be categorical or continuous
  • Repeated measures: random intercepts via Error(subject), as in stats::aov()
  • Each model reports \(R^2\) or \(\rho\), the test statistic, p-values and adjusted p-values

Disjointed and emergent associations

Simulated Krebs-cycle data with spiked-in effects (krebsDemoWeb()).

The anansi workflow

  1. weaveWeb(): combine the two feature tables, the linking information and the sample metadata into an AnansiWeb object; the features are matched through the biadjacency matrix
  2. anansi(): fit the models for every linked feature pair; returns one row per pair with correlations, full, disjointed and emergent statistics, p-values and adjusted p-values
  3. plotAnansi(), or your own ggplot2 code: visualize
  • Input: plain tables, or MultiAssayExperiment / TreeSummarizedExperiment; asMAE() and asTSE() convert back
  • pairwiseApply(): apply any function to every linked pair
  • Linear models throughout, following stats::lm() / aov() conventions

Example: FMT and ageing

The study

Boehme et al. (2021): can faecal microbiota transplantation (FMT) from young donor mice reverse ageing-related changes in aged mice?

  • Three groups: young mice given young FMT, aged mice given aged FMT, aged mice given young FMT
  • Microbial functions (KEGG orthologues) inferred from 16S rRNA data
  • Hippocampal metabolites
  • An early version of anansi linked each metabolite to the microbial functions that produce or consume it

anansi ships curated subsets of these data as tutorial data:

data(FMT_data)
dim(FMT_KOs)     # KOs x samples
[1] 6468   36
dim(FMT_metab)   # metabolites x samples
[1]  3 36
table(FMT_metadata$Legend)

 Aged oFMT  Aged yFMT Young yFMT 
        12         12         12 

Weave the web

web <- weaveWeb(
  cpd ~ ko,                  # link KEGG compounds (Y) to orthologues (X)
  link     = kegg_link(),
  tableY   = t(FMT_metab),   # samples as rows
  tableX   = t(FMT_KOs),
  metadata = FMT_metadata
)
web
anansi::AnansiWeb S7_object with 36 observations:
    tableY: cpd (3 features)
    tableX: ko (2476 features)
Use @ to access: tableX, tableY, dictionary, metadata.

Run anansi

out <- anansi(
  web,
  formula = ~ Legend,        # does the association depend on the group?
  adjust.method = "BH",
  verbose = FALSE
)

# All-vs-all would test every metabolite with every KO
c(all_vs_all = ncol(tableY(web)) * ncol(tableX(web)),
  tested     = nrow(out))    # one row per linked feature pair
all_vs_all     tested 
      7428         65 
head(out[, 1:5], 4)
  feature_Y feature_X All_r.values All_t.values All_p.values
1    C00186    K00016  -0.08985529    0.5260699   0.60225456
2    C00186    K00101   0.28661046    1.7443940   0.09012569
3    C00041    K00259  -0.21708011    1.2967052   0.20346437
4    C00064    K00265  -0.25811198    1.5578255   0.12853530

The output

# Columns per model: statistic, test statistic, p-value, adjusted p-value
table(sub("_[^_]+$", "", colnames(out)[-(1:2)]))

        Aged oFMT         Aged yFMT               All disjointed_Legend 
                4                 4                 4                 4 
  emergent_Legend              full        Young yFMT 
                4                 4                 4 
Association Method Statistics Significance
Pooled (All) correlation \(\rho\), \(t\) p, adjusted p
Per group correlation \(\rho\), \(t\) p, adjusted p
Full linear model \(R^2\), \(F\) p, adjusted p
Disjointed linear model \(R^2\), \(F\) p, adjusted p
Emergent linear model \(R^2\), \(F\) p, adjusted p

Visualize differential associations

out_sig <- out[out$full_q.values < 0.2, ]   # pairs where the full model fits

plotAnansi(
  out_sig,
  association.type = "disjointed",
  model.var = "Legend",
  signif.threshold = 0.05,
  fill_by = "group"
)

Visualize differential associations

Look at individual pairs

Code
pairs <- getFeaturePairs(web)
grp_cols <- c("Young yFMT" = "#2166ac", "Aged oFMT" = "#b2182b",
              "Aged yFMT" = "#ef8a62")

plots <- lapply(pairs[seq_len(6)], function(p) {
  ggplot(p, aes(y = .data[[colnames(p)[1]]], x = .data[[colnames(p)[2]]])) +
    facet_grid(colnames(p)[1] ~ colnames(p)[2])
})

wrap_plots(plots, ncol = 3) &
  geom_point(aes(fill = FMT_metadata$Legend), shape = 21,
             show.legend = FALSE) &
  geom_smooth(aes(colour = FMT_metadata$Legend), method = "lm",
              se = FALSE, show.legend = FALSE) &
  scale_fill_manual(values = grp_cols, aesthetics = c("colour", "fill")) &
  labs(x = NULL, y = NULL) &
  theme_bw(base_size = 12)

More options

# Continuous covariates and several terms
anansi(web, formula = ~ group + score, groups = "group")

# Repeated measures: random intercept per subject
anansi(web, formula = ~ group + Error(subject_id), groups = "group")

# Bioconductor containers in and out
mae <- asMAE(web)
web <- weaveWeb(mae, tableY = "cpd", tableX = "ko", link = kegg_link())

# Custom links: any two-column data.frame of feature IDs
web <- weaveWeb(y ~ x, link = my_links, tableY = tY, tableX = tX)

# Apply any function to each linked pair
pairwiseApply(web, FUN = function(x, y) cor(x, y, method = "spearman"))

Limitations

  • Database coverage: interactions that are real but not catalogued are never tested; many microbial metabolites and xenobiotics are still unmapped
  • Inferred functions are nested: PICRUSt2 or HUMAnN function abundances depend on taxon abundances, so features are not independent; metatranscriptomics or metaproteomics would reduce this
  • Binary links: interaction strength is a spectrum (e.g. binding affinities), but the adjacency matrix is 0/1
  • Associations are linear models; compositional data should be transformed appropriately first

Take-home messages

  • anansi replaces all-vs-all testing with knowledge-based testing of feature pairs that are known to interact
  • Results are structured by pathways and easier to interpret
  • Fewer tests means more power after FDR correction
  • Differential associations: disjointed (slope changes) and emergent (strength changes), with covariates and repeated measures
  • Works with any pair of data sets that has a dictionary: metabolite–function, receptor–ligand, phage–bacterium, …
  • R/Bioconductor, with MultiAssayExperiment and TreeSummarizedExperiment support

References

Bastiaanssen, Thomaz F S, Thomas P Quinn, and John F Cryan. 2023. “Knowledge-Based Integration of Multi-Omic Datasets with Anansi: Annotation-Based Analysis of Specific Interactions.” arXiv. https://doi.org/10.48550/arXiv.2305.10832.
Boehme, Marcus, Katherine E. Guzzetta, Thomaz F. S. Bastiaanssen, Marcel van de Wouw, Gerard M. Moloney, others, and John F. Cryan. 2021. “Microbiota from Young Mice Counteracts Selective Age-Associated Behavioral Deficits.” Nature Aging 1 (8): 666–76. https://doi.org/10.1038/s43587-021-00093-9.
Erb, Ionas, Thomas Quinn, David Lovell, and Cedric Notredame. 2017. “Differential Proportionality – a Normalization-Free Approach to Differential Gene Expression.” bioRxiv. https://doi.org/10.1101/134536.
Kanehisa, Minoru, and Susumu Goto. 2000. “KEGG: Kyoto Encyclopedia of Genes and Genomes.” Nucleic Acids Research 28 (1): 27–30. https://doi.org/10.1093/nar/28.1.27.
Quinn, Thomas P., Mark F. Richardson, David Lovell, and Tamsyn M. Crowley. 2017. “propr: An R-Package for Identifying Proportionally Abundant Features Using Compositional Data Analysis.” Scientific Reports 7: 16252. https://doi.org/10.1038/s41598-017-16520-0.

Package and documentation: https://thomazbastiaanssen.github.io/anansi/