Foundry120 atlas

Differences in microbiome populations following hematopoietic stem cell transplantation in Crohn?s disease patients

Download from source ↗

Dataset overview

Participants 37
Samples 158
Reuse readiness 9.4/10 evidence-backed score

Multi-omic modelling of inflammatory bowel disease with regularized canonical correlation analysis.

Abstract

<h4>Background</h4>Personalized medicine requires finding relationships between variables that influence a patient's phenotype and predicting an outcome. Sparse generalized canonical correlation analysis identifies relationships between different groups of variables. This method requires establishing a model of the expected interaction between those variables. Describing these interactions is challenging when the relationship is unknown or when there is no pre-established hypothesis. Thus, our aim was to develop a method to find the relationships between microbiome and host transcriptome data and the relevant clinical variables in a complex disease, such as Crohn's disease.<h4>Results</h4>We present here a method to identify interactions based on canonical correlation analysis. We show that the model is the most important factor to identify relationships between blocks using a dataset of Crohn's disease patients with longitudinal sampling. First the analysis was tested in two previously published datasets: a glioma and a Crohn's disease and ulcerative colitis dataset where we describe how to select the optimum parameters. Using such parameters, we analyzed our Crohn's disease data set. We selected the model with the highest inner average variance explained to identify relationships between transcriptome, gut microbiome and clinically relevant variables. Adding the clinically relevant variables improved the average variance explained by the model compared to multiple co-inertia analysis.<h4>Conclusions</h4>The methodology described herein provides a general framework for identifying interactions between sets of omic data and clinically relevant variables. Following this method, we found genes and microorganisms that were related to each other independently of the model, while others were specific to the model used. Thus, model selection proved crucial to finding the existing relationships in multi-omics datasets.

Study facts

Organism
Homo sapiens
Platform
MiSeq (Illumina)
Age group
Disease groups
Crohn's disease, Non-IBD controls
Anatomical sites

Data availability

  • Analysis code

Strengths & limitations for reuse

Strengths

  • Raw reads are advertised
  • Feature/OTU tables are advertised
  • Taxonomic tables are advertised
  • Analysis code is available
  • Participant-to-sample mapping is available
  • Participant counts are documented
  • Sample counts are documented
Extraction evidence & provenance

Each extracted field is shown with the source excerpt and location used to resolve it.

Assay

FieldValueEvidence
assay.paired_end True
sequencing was performed with pooled samples in paired-end modus (PE275)

Section Methods > High throughput 16S ribosomal RNA (rRNA) gene sequencing, offset 9500

assay.platform MiSeq (Illumina)
sequencing was performed... using a MiSeq system (Illumina, Inc.)

Section High throughput 16S ribosomal RNA (rRNA) gene sequencing, offset —

assay.primers_reported True
using forward and reverse primers 341F-785R

Section Methods > High throughput 16S ribosomal RNA (rRNA) gene sequencing, offset 9200

assay.read_length 275
paired-end modus (PE275)

Section Methods > High throughput 16S ribosomal RNA (rRNA) gene sequencing, offset 9500

assay.sequencing_type amplicon_16s
High throughput 16S ribosomal RNA (rRNA) gene sequencing

Section Methods > High throughput 16S ribosomal RNA (rRNA) gene sequencing, offset 8700

assay.target_region V3-V4
the V3-V4 regions of 16S rRNA gene were amplified

Section Methods > High throughput 16S ribosomal RNA (rRNA) gene sequencing, offset 9200

Cohort

FieldValueEvidence
cohort.crohns_disease_participants 18
The HSCT CD cohort involved 158 samples ... from 18 CD patients undergoing HSCT

Section Methods—Datasets, offset 16500

cohort.disease_activity_metadata_available True
segmental simple endoscopic score for Crohn’s disease (SES-CD) ... were collected

Section Methods—Datasets, offset 16850

cohort.non_ibd_controls 19
biopsies were taken from the ileum and colon regions of 19 non-IBD controls

Section Methods—Patients and biopsies processing, offset 4300

cohort.study_design longitudinal
Patients were followed-up for 4 years and biopsies were collected every six or twelve months after HSCT.

Section Methods—Patients and biopsies processing, offset 3300

cohort.total_participants 37 computed
The HSCT CD cohort involved 158 samples ... from 18 CD patients undergoing HSCT ... and 19 non-IBD controls

Section Methods—Datasets, offset 16500

cohort.treatment_exposure_documented True
clinical information such as age, sex, treatment, years since disease diagnosis, prior surgery

Section Methods—Datasets, offset 16850

cohort.treatment_response_metadata_available True
time of the HSCT and response to treatment were collected

Section Methods—Datasets, offset 16850

Data_Assets

FieldValueEvidence
data_assets.analysis_code True
Analysis code of our HSCT CD dataset available at: https://github.com/llrs/TRIM

Section Data Availability, offset 47200

data_assets.environment_or_container_info True
Analysis was performed using R (version 3.6.1) and Bioconductor (Version 3.10) on Ubuntu 18.04.

Section Methods > Mucosal transcriptome, offset 7600

data_assets.feature_or_otu_table True
The resulting OTUs table was normalized using edgeR (Version 3.28).

Section Methods > Microbial profiling, offset 10700

data_assets.open_access True from source
isOpenAccess: Y

Section publication metadata, offset —

data_assets.pipeline_or_tool_versions True
using the IMNGS (version 1.0 Build 2007) pipeline based on the UPARSE approach

Section Methods > Microbial profiling, offset 9900

data_assets.qc_or_negative_controls_reported True
Sequences with less than 300 and more than 600 nucleotides and paired reads with an expected error >3 were excluded from the analysis.

Section Microbial profiling, offset 9900

data_assets.raw_reads True
Processing of raw-reads was performed by using the IMNGS

Section Microbial profiling, offset 8400

data_assets.sample_metadata True
clinical information such as age, sex, treatment ... location of the biopsies ... and response to treatment were collected.

Section Methods > Datasets, offset 13300

data_assets.taxonomic_table True
Taxonomy assignment was performed at 80% confidence level using the RDP classifier

Section Methods—Microbial profiling, offset 12400

Specimens

FieldValueEvidence
specimens.body_site colon and ileum
Colonic and ileal biopsies were obtained at several time points

Section Patients and biopsies processing, offset —

specimens.inflamed_status_available True
Samples were obtained when possible from both uninvolved and involved areas.

Section Patients and biopsies processing, offset 2400

specimens.longitudinal_sampling True
Patients were followed-up for 4 years and biopsies were collected every six or twelve months after HSCT.

Section Methods > Patients and biopsies processing, offset 5000

specimens.number_of_samples 158
The HSCT CD cohort involved 158 samples (both host RNA and microbial DNA)

Section Methods—Datasets, offset 16500

specimens.participant_to_sample_mapping_available True
the HSCT CD dataset included the following variables: patient ID, sex, age

Section Models used, offset 23700

specimens.sample_type mucosal_biopsy
Colonic and ileal biopsies were obtained at several time points during ileocolonoscopy.

Section Methods—Patients and biopsies processing, offset 2700