Orchestrating Hologenome Analysis with Bioconductor

ECCB2026

Tuomas Borman, Mahkameh Salehi, Sneha Das, Himmi Lindgren, Leo Lahti

Monday, August 31, 2026

Orchestrating Hologenome Analysis with Bioconductor

mia logo.

Orchestrating Hologenome Analysis with Bioconductor

  • Introduces microbiome analysis framework, focusing multiomics methods
  • For participants with sufficient programming skills
mia logo.

Time outline

Time Activity
9:00-9:30 Introduction to hologenomics
9:30-10:30 Bioconductor’s data science framework
10:30-10:45 Coffee break
10:45-11:45 Tabular data analysis for microbiomes
11:45-12:45 Incorporating hierarchical side information
12:45-13:45 Lunch
13:45-14:30 Programmatic access to hologenome data collections
14:30-15:15 Multi-table data structures
15:15-15:30 Coffee break
15:30-17:00 Methods for taxonomic, functional, and host omic integration
17:00-17:30 Q&A and closing session

Participation

NCI logo.

Questions

Objectives

  • Analyze and apply methods: Apply the TreeSummarizedExperiment and MultiAssayExperiment ecosystem to process, integrate, and analyze hologenome and multiomics data.

  • Create visualizations: Generate and interpret visualizations for multiomics data.

  • Explore documentation: Use the OMA to explore additional tools and methods for hologenome analysis.

(9:00-9:30) Introduction to hologenomics

Moreno-Indias et al. (2021) Statistical and Machine Learning Techniques in Human Microbiome Studies: Contemporary Challenges and Solutions. Frontiers in Microbiology.

Hologenomics

  • Holobiont = Host + microbiome (Nyholm et al. 2020)
  • Understanding the host often requires understanding its microbes.
  • Hologenomic approaches integrate information across the host–microbiome system

Multi-omics

Multi-omics = combining different types of molecular data

Multi-omics integration

There are different ways to combine omics data:

  • Analyse separately: Each omics layer independently

  • Find associations: Identify relationships between omics layers

  • Jointly model: Combine omics layers to explain outcomes or mechanisms

Why Bioconductor?

Bioconductor provides tools for:

  • Data handling
  • Analysis
  • Visualization
  • Integration
  • Reproducible workflows
Bioconductor logo.

(9:30-10:30) Bioconductor’s data science framework

Bioconductor

  • Community-driven, global open-source project
  • Started in 2001 from genomics
  • High impact across bioinformatics
Bioconductor logo.

Community is the key

  • Training programs & workshops
  • Community support
  • Conferences

Software

  • ~2,300 R packages
  • Review, testing, documentation

Data containers

  • The core of software
  • Structured, standardized way to manage complex data
  • Enables modular, efficient workflows

Optimal container for microbiome data?

  • Multiple assays: seamless interlinking
  • Hierarchical data: supporting samples & features
  • Side information: extended capabilities & data types
  • Optimized: for speed & memory
  • Integrated: with other applications & frameworks

TreeSummarizedExperiment

(Huang et al. 2021)

TreeSummarizedExperiment class

SummarizedExperiment

(Huber et al. 2015)

Data import

Microbiome data science workflow

Microbiome data science workflow

Importers and converters

  • mia includes importers and converters for standard data formats
    • Importer: Import file
    • Converter: Convert R object to other format
  • You can also create TreeSE object manually

(10:45-12:45) Tabular data analysis for microbiomes & incorporating hierarchical side information

Microbiome Analysis (mia)

  • Microbiome data science ecosystem
  • “Downstream”, statistical analysis
  • > 15,000 yearly downloads from distinct IPs (top 8.05% in Bioconductor)

Bioconductor sticker mia logo

Community-driven ecosystem of tools

mia logo. MGnifyR logo. HoloFoodR logo. iSEE logo. MAE logo. SE logo. SCE logo. scater logo. benchdamic logo. netcomi logo. radEmu logo. DESeq2 logo. Biobakery logo. anansi logo.

Orchestrating Microbiome Analysis with Bioconductor

  • Resources and tutorials for microbiome analysis
  • Community-built best practices
  • Open to contributions!

Go to the Orchestrating Microbiome Analysis (OMA) online book

(13:45-14:30) Programmatic access to hologenome data collections

Several resources provide microbiome multi-omics data

(14:30-15:15) Multi-table data structures

MultiAssayExperiment

(Ramos et al. 2017)

  • For more complex sample mapping
TreeSummarizedExperiment class

(15:30-17:00) Methods for taxonomic, functional, and host omic integration

A systematic benchmark of integrative strategies for microbiome-metabolome data (Mangnier et al. 2025)

Multi-omics statistical approaches
Aim Method
Global associations Mantel
Data summarization RDA
Individual associations MiRKAT
Feature selection CODA-LASSO

Mantel test

Question: Are two omics datasets related?

  • Compares distance matrices
  • Tests for a global association between omics layers

Microbiome Regression-based Kernel Association Test (MiRKAT)

Question: Is a pathway associated with the global microbiome profile?

  • Represents microbiome similarity between samples
  • Tests its association with a pathway

Pairwise association

Question: Which features are associated?

  • Tests feature–feature associations
  • Can reveal candidate biological interactions

Joint robust principal component analysis (joint RPCA) (Cordazzo Vargas et al. 2026)

Question: What are the main shared patterns?

  1. Denoise: learn a lower-dimensional representation (samples × dimensions)
  2. PCA: find components that capture maximum variance
  3. Visualize: samples in PCA space and identify the loadings driving the variation

Integrated machine learning for multi-omics prediction and classification (Mallick et al. 2024)

Multi-omics data fusion

  • Early fusion: Concatenate the different data tables into a single table and train one machine-learning model on the combined features.

  • Intermediate fusion: Train separate models for each omics layer and then use a meta-model to combine the information or predictions from these models. This approach can often capture complementary information more flexibly than simple early or late fusion.

  • Late fusion: Train a separate model for each omics layer and combine their predictions, for example by averaging or weighting them.

(17:00-17:30) Q&A and closing session

Other approaches

Thank you for your time!

Join us!

mia logo

References

Argelaguet, Ricard, Damien Arnol, Danila Bredikhin, Yonatan Deloro, Britta Velten, John C. Marioni, and Oliver Stegle. 2020. MOFA+: A Statistical Framework for Comprehensive Integration of Multi-Modal Single-Cell Data.” Genome Biology 21 (1): 111. https://doi.org/10.1186/s13059-020-02015-1.
Bastiaanssen, Thomaz F S, Thomas P Quinn, and John F Cryan. 2023. “Knowledge-Based Integration of Multi-Omic Datasets with Anansi: Annotation-Based Analysis of Specific Interactions.” arXiv. https://doi.org/10.48550/arXiv.2305.10832.
Cordazzo Vargas, Bianca, Cameron Martino, Amanda Hazel Dilmore, Jessica Metcalf, Zachary Burcham, Leo Lahti, Aituar Bektanov, et al. 2026. Joint-RPCA: Domain-Aware Multi-Omics Integration for Systems Microbiology.” Molecular Systems Biology.
Huang, Ruizhu, Charlotte Soneson, Felix G. M. Ernst, et al. 2021. “TreeSummarizedExperiment: A S4 Class for Data with Hierarchical Structure.” F1000Research 9: 1246. https://doi.org/10.12688/f1000research.26669.2.
Huber, W., V. J. Carey, R. Gentleman, S. Anders, M. Carlson, B. S. Carvalho, H. C. Bravo, et al. 2015. Orchestrating High-Throughput Genomic Analysis with Bioconductor.” Nature Methods 12 (2): 115–21. http://www.nature.com/nmeth/journal/v12/n2/full/nmeth.3252.html.
Mallick, Himel, Anupreet Porwal, Satabdi Saha, Piyali Basak, Vladimir Svetnik, and Erina Paul. 2024. “An Integrated Bayesian Framework for Multi-Omics Prediction and Classification.” Statistics in Medicine 43 (5): 983–1002. https://doi.org/10.1002/sim.9953.
Mangnier, Loïc, Antoine Bodein, Margaux Mariaz, Alban Mathieu, Marie-Pier Scott-Boyer, Neerja Vashist, Matthew S. Bramble, and Arnaud Droit. 2025. “A Systematic Benchmark of Integrative Strategies for Microbiome-Metabolome Data.” Communications Biology 8 (1). https://doi.org/10.1038/s42003-025-08515-9.
Nyholm, Lasse, Adam Koziol, Sofia Marcos, Amanda Bolt Botnen, Ostaizka Aizpurua, Shyam Gopalakrishnan, Morten T. Limborg, M.Thomas P. Gilbert, and Antton Alberdi. 2020. “Holo-Omics: Integrated Host-Microbiota Multi-Omics for Basic and Applied Biological Research.” iScience 23 (8): 101414. https://doi.org/10.1016/j.isci.2020.101414.
Ramos, Marcel, Lucas Schiffer, Angela Re, Rimsha Azhar, Azfar Basunia, Carmen Rodriguez, Tiffany Chan, et al. 2017. “Software for the Integration of Multiomics Experiments in Bioconductor.” Cancer Research 77 (21): e39–42. https://doi.org/10.1158/0008-5472.CAN-17-0344.
Rohart, Florian, Benoı̂t Gautier, Amrit Singh, and Kim-Anh Lê Cao. 2017. mixOmics: An R Package for ‘Omics Feature Selection and Multiple Data Integration.” PLoS Computational Biology 13 (11): e1005752. https://doi.org/10.1371/journal.pcbi.1005752.
Verma, Shivangi, Nalin Arora, Chandra Prakash Ajay, Pankhuri Singh, Himel Mallick, and Tarini Shankar Ghosh. 2026. “HuMMANet: A Harmonized Cross-Study Resource for Integrative Analysis of Human Gut Microbiome Metabolome Associations.” bioRxiv. https://doi.org/10.64898/2026.08.24.746727.