Differences in microbiome populations following hematopoietic stem cell transplantation in Crohn?s disease patients
Download from source ↗Dataset overview
Multi-omic modelling of inflammatory bowel disease with regularized canonical correlation analysis.
Abstract
<h4>Background</h4>Personalized medicine requires finding relationships between variables that influence a patient's phenotype and predicting an outcome. Sparse generalized canonical correlation analysis identifies relationships between different groups of variables. This method requires establishing a model of the expected interaction between those variables. Describing these interactions is challenging when the relationship is unknown or when there is no pre-established hypothesis. Thus, our aim was to develop a method to find the relationships between microbiome and host transcriptome data and the relevant clinical variables in a complex disease, such as Crohn's disease.<h4>Results</h4>We present here a method to identify interactions based on canonical correlation analysis. We show that the model is the most important factor to identify relationships between blocks using a dataset of Crohn's disease patients with longitudinal sampling. First the analysis was tested in two previously published datasets: a glioma and a Crohn's disease and ulcerative colitis dataset where we describe how to select the optimum parameters. Using such parameters, we analyzed our Crohn's disease data set. We selected the model with the highest inner average variance explained to identify relationships between transcriptome, gut microbiome and clinically relevant variables. Adding the clinically relevant variables improved the average variance explained by the model compared to multiple co-inertia analysis.<h4>Conclusions</h4>The methodology described herein provides a general framework for identifying interactions between sets of omic data and clinically relevant variables. Following this method, we found genes and microorganisms that were related to each other independently of the model, while others were specific to the model used. Thus, model selection proved crucial to finding the existing relationships in multi-omics datasets.
doi:10.1371/journal.pone.0246367 ↗ PMID 33556098 ↗ PMC7870068 ↗
Study facts
- Organism
- Homo sapiens
- Platform
- MiSeq (Illumina)
- Age group
- —
- Disease groups
- Crohn's disease, Non-IBD controls
- Anatomical sites
- —
Data availability
- Analysis code
Strengths & limitations for reuse
Strengths
- Raw reads are advertised
- Feature/OTU tables are advertised
- Taxonomic tables are advertised
- Analysis code is available
- Participant-to-sample mapping is available
- Participant counts are documented
- Sample counts are documented
Extraction evidence & provenance
Each extracted field is shown with the source excerpt and location used to resolve it.
Assay
| Field | Value | Evidence |
|---|---|---|
assay.paired_end |
True |
sequencing was performed with pooled samples in paired-end modus (PE275) Section |
assay.platform |
MiSeq (Illumina) |
sequencing was performed... using a MiSeq system (Illumina, Inc.) Section |
assay.primers_reported |
True |
using forward and reverse primers 341F-785R Section |
assay.read_length |
275 |
paired-end modus (PE275) Section |
assay.sequencing_type |
amplicon_16s |
High throughput 16S ribosomal RNA (rRNA) gene sequencing Section |
assay.target_region |
V3-V4 |
the V3-V4 regions of 16S rRNA gene were amplified Section |
Cohort
| Field | Value | Evidence |
|---|---|---|
cohort.crohns_disease_participants |
18 |
The HSCT CD cohort involved 158 samples ... from 18 CD patients undergoing HSCT Section |
cohort.disease_activity_metadata_available |
True |
segmental simple endoscopic score for Crohn’s disease (SES-CD) ... were collected Section |
cohort.non_ibd_controls |
19 |
biopsies were taken from the ileum and colon regions of 19 non-IBD controls Section |
cohort.study_design |
longitudinal |
Patients were followed-up for 4 years and biopsies were collected every six or twelve months after HSCT. Section |
cohort.total_participants |
37 computed |
The HSCT CD cohort involved 158 samples ... from 18 CD patients undergoing HSCT ... and 19 non-IBD controls Section |
cohort.treatment_exposure_documented |
True |
clinical information such as age, sex, treatment, years since disease diagnosis, prior surgery Section |
cohort.treatment_response_metadata_available |
True |
time of the HSCT and response to treatment were collected Section |
Data_Assets
| Field | Value | Evidence |
|---|---|---|
data_assets.analysis_code |
True |
Analysis code of our HSCT CD dataset available at: https://github.com/llrs/TRIM Section |
data_assets.environment_or_container_info |
True |
Analysis was performed using R (version 3.6.1) and Bioconductor (Version 3.10) on Ubuntu 18.04. Section |
data_assets.feature_or_otu_table |
True |
The resulting OTUs table was normalized using edgeR (Version 3.28). Section |
data_assets.open_access |
True from source |
isOpenAccess: Y Section |
data_assets.pipeline_or_tool_versions |
True |
using the IMNGS (version 1.0 Build 2007) pipeline based on the UPARSE approach Section |
data_assets.qc_or_negative_controls_reported |
True |
Sequences with less than 300 and more than 600 nucleotides and paired reads with an expected error >3 were excluded from the analysis. Section |
data_assets.raw_reads |
True |
Processing of raw-reads was performed by using the IMNGS Section |
data_assets.sample_metadata |
True |
clinical information such as age, sex, treatment ... location of the biopsies ... and response to treatment were collected. Section |
data_assets.taxonomic_table |
True |
Taxonomy assignment was performed at 80% confidence level using the RDP classifier Section |
Specimens
| Field | Value | Evidence |
|---|---|---|
specimens.body_site |
colon and ileum |
Colonic and ileal biopsies were obtained at several time points Section |
specimens.inflamed_status_available |
True |
Samples were obtained when possible from both uninvolved and involved areas. Section |
specimens.longitudinal_sampling |
True |
Patients were followed-up for 4 years and biopsies were collected every six or twelve months after HSCT. Section |
specimens.number_of_samples |
158 |
The HSCT CD cohort involved 158 samples (both host RNA and microbial DNA) Section |
specimens.participant_to_sample_mapping_available |
True |
the HSCT CD dataset included the following variables: patient ID, sex, age Section |
specimens.sample_type |
mucosal_biopsy |
Colonic and ileal biopsies were obtained at several time points during ileocolonoscopy. Section |