Foundry120 atlas

Reproducibility of associations between the human gut microbiome and colorectal cancer assessed in a patient population from Washington, DC, USA

Download from source ↗

Dataset overview

Participants 104
Samples 110
Reuse readiness 6.6/10 evidence-backed score

Clustering co-abundant genes identifies components of the gut microbiome that are reproducibly associated with colorectal cancer and inflammatory bowel disease.

Abstract

<h4>Background</h4>Whole-genome "shotgun" (WGS) metagenomic sequencing is an increasingly widely used tool for analyzing the metagenomic content of microbiome samples. While WGS data contains gene-level information, it can be challenging to analyze the millions of microbial genes which are typically found in microbiome experiments. To mitigate the ultrahigh dimensionality challenge of gene-level metagenomics, it has been proposed to cluster genes by co-abundance to form Co-Abundant Gene groups (CAGs). However, exhaustive co-abundance clustering of millions of microbial genes across thousands of biological samples has previously been intractable purely due to the computational challenge of performing trillions of pairwise comparisons.<h4>Results</h4>Here we present a novel computational approach to the analysis of WGS datasets in which microbial gene groups are the fundamental unit of analysis. We use the Approximate Nearest Neighbor heuristic for near-exhaustive average linkage clustering to group millions of genes by co-abundance. This results in thousands of high-quality CAGs representing complete and partial microbial genomes. We applied this method to publicly available WGS microbiome surveys and found that the resulting microbial CAGs associated with inflammatory bowel disease (IBD) and colorectal cancer (CRC) were highly reproducible and could be validated independently using multiple independent cohorts.<h4>Conclusions</h4>This powerful approach to gene-level metagenomics provides a powerful path forward for identifying the biological links between the microbiome and human health. By proposing a new computational approach for handling high dimensional metagenomics data, we identified specific microbial gene groups that are associated with disease that can be used to identify strains of interest for further preclinical and mechanistic experimentation.

Study facts

Organism
Platform
Illumina HiSeq 2000
Age group
Disease groups
Anatomical sites

Data availability

  • Analysis code

Strengths & limitations for reuse

Strengths

  • Raw reads are advertised
  • Feature/OTU tables are advertised
  • Analysis code is available
  • Participant counts are documented
  • Sample counts are documented

Limitations

  • Not documented: taxonomic tables are advertised
Extraction evidence & provenance

Each extracted field is shown with the source excerpt and location used to resolve it.

Assay

FieldValueEvidence
assay.platform Illumina HiSeq 2000 from source
ENA instrument_model=Illumina HiSeq 2000

Section ENA study report, offset —

assay.sequencing_type shotgun_metagenomics
we performed whole-genome shotgun metagenomics sequencing

Section dataset-authority study description, offset —

Cohort

FieldValueEvidence
cohort.total_participants 104 inferred
whole-genome shotgun metagenomics sequencing on fecal samples from 52 pre-treatment colorectal cancer cases and 52 matched controls

Section dataset-authority study description, offset —

Data_Assets

FieldValueEvidence
data_assets.analysis_code True
The repository includes ... the Jupyter notebooks used to analyze those datasets and produce the figures and tables presented here.

Section Data availability, offset 19000

data_assets.environment_or_container_info True
All microbiome WGS data were analyzed using a Docker-based workflow

Section Methods — Gene-level metagenomic analysis pipeline, offset 6900

data_assets.feature_or_otu_table True inferred
The final HDF file creation step (9) includes the results of that quantification step for the validation datasets

Section Methods — Gene-level metagenomic analysis pipeline, offset 8600

data_assets.open_access True
"isOpenAccess": "Y"

Section publication metadata, offset —

data_assets.pipeline_or_tool_versions True
Software version(s): SPAdes-3.11.1-Linux ... Prokka v1.12; barrnap v0.9

Section Methods — Gene-level metagenomic analysis pipeline, offset 7000

data_assets.raw_reads True
Each sample was individually downloaded from NCBI SRA

Section Methods — Gene-level metagenomic analysis pipeline, offset 7900

Specimens

FieldValueEvidence
specimens.number_of_samples 110 from source
ENA sample_count=110

Section ENA study report, offset —

specimens.sample_type stool
whole-genome shotgun metagenomics sequencing on fecal samples

Section dataset-authority study description, offset —