Human gut metagenome and metatranscriptome in the inflammatory bowel disease (iHMP/HMP2)
Download from source ↗Dataset overview
Clustering co-abundant genes identifies components of the gut microbiome that are reproducibly associated with colorectal cancer and inflammatory bowel disease.
Abstract
<h4>Background</h4>Whole-genome "shotgun" (WGS) metagenomic sequencing is an increasingly widely used tool for analyzing the metagenomic content of microbiome samples. While WGS data contains gene-level information, it can be challenging to analyze the millions of microbial genes which are typically found in microbiome experiments. To mitigate the ultrahigh dimensionality challenge of gene-level metagenomics, it has been proposed to cluster genes by co-abundance to form Co-Abundant Gene groups (CAGs). However, exhaustive co-abundance clustering of millions of microbial genes across thousands of biological samples has previously been intractable purely due to the computational challenge of performing trillions of pairwise comparisons.<h4>Results</h4>Here we present a novel computational approach to the analysis of WGS datasets in which microbial gene groups are the fundamental unit of analysis. We use the Approximate Nearest Neighbor heuristic for near-exhaustive average linkage clustering to group millions of genes by co-abundance. This results in thousands of high-quality CAGs representing complete and partial microbial genomes. We applied this method to publicly available WGS microbiome surveys and found that the resulting microbial CAGs associated with inflammatory bowel disease (IBD) and colorectal cancer (CRC) were highly reproducible and could be validated independently using multiple independent cohorts.<h4>Conclusions</h4>This powerful approach to gene-level metagenomics provides a powerful path forward for identifying the biological links between the microbiome and human health. By proposing a new computational approach for handling high dimensional metagenomics data, we identified specific microbial gene groups that are associated with disease that can be used to identify strains of interest for further preclinical and mechanistic experimentation.
doi:10.1186/s40168-019-0722-6 ↗ PMID 31370880 ↗ PMC6670193 ↗
Study facts
- Organism
- —
- Platform
- —
- Age group
- —
- Disease groups
- —
- Anatomical sites
- —
Data availability
- Analysis code
Strengths & limitations for reuse
Strengths
- Raw reads are advertised
- Feature/OTU tables are advertised
- Taxonomic tables are advertised
- Analysis code is available
- Participant counts are documented
Extraction evidence & provenance
Each extracted field is shown with the source excerpt and location used to resolve it.
Assay
| Field | Value | Evidence |
|---|---|---|
assay.sequencing_type |
shotgun_metagenomics |
All microbiome WGS data were analyzed using a Docker-based workflow Section |
Cohort
| Field | Value | Evidence |
|---|---|---|
cohort.study_design |
longitudinal |
100 individuals sampled over a one year period Section |
cohort.total_participants |
100 |
the human gut microbiome among 100 individuals sampled over a one year period Section |
Data_Assets
| Field | Value | Evidence |
|---|---|---|
data_assets.analysis_code |
True |
as well as the Jupyter notebooks used to analyze those datasets and produce the figures and tables presented here Section |
data_assets.environment_or_container_info |
True |
All microbiome WGS data were analyzed using a Docker-based workflow Section |
data_assets.feature_or_otu_table |
True inferred |
includes all of the outputs from the bioinformatic pipeline used for gene-level metagenomic analysis Section |
data_assets.pipeline_or_tool_versions |
True |
Software version(s): SPAdes-3.11.1-Linux Section |
data_assets.raw_reads |
True |
Each sample was individually downloaded from NCBI SRA Section |
data_assets.representative_sequences |
True |
create a set of non-redundant protein sequences Section |
data_assets.taxonomic_table |
True |
The non-redundant protein sequences were analyzed via the taxonomic assignment functionality of DIAMOND Section |
Specimens
| Field | Value | Evidence |
|---|---|---|
specimens.body_site |
human gut |
profiling metagenomic and metatranscriptomic sequencing of the human gut microbiome Section |
specimens.longitudinal_sampling |
True |
100 individuals sampled over a one year period Section |
specimens.sample_type |
stool |
each has been studied by multiple groups who have collected stool samples and performed metagenomic whole-genome “shotgun” (WGS) sequencing Section |