• No se han encontrado resultados

Motif enrichment was performed on selected regions by inputting them to the

developed by Ian Sudbery and contains some fixes done by myself. Briefly, the analysis

pipeline runs MEME (Bailey and Elkan, 1994) and DREME (Bailey, 2011) version 4.12.0 from the MEME suite. MEME was used to detect 40 maximum motifs, using DNA alphabet, allowing sites on both strands, finding distribution of motifs with any number of repetitions. DREME was used with minimum width of core motif 5 and maximum 30. This was performed to find novel motifs enriched in the provided regions. The motifs found were inspected for similarities between them using Tomtom (Gupta et al., 2007), which is also included in the MEME suite. First clusters of motifs were created by linking de novo motifs that were significantly similar (q- value less than 0.05). The de novo motif having the most significant E-value and number of found binding sites was selected as the representative motif of the cluster. Then clusters were merged if their representative de novo motifs were similar (q-value less than 0.1).

The reference motifs for each cluster were then checked for similarity to consensus binding sites for TFs from multiple databases using Tomtom. The databases used are:

 Jaspar databases: JASPAR_CORE_2016, JASPAR_CORE_REDUNDANT_2016, JASPAR_CORE_2016_vertebrates, JASPAR_CORE_REDUNDANT_2016_vertebrates (Mathelier et al., 2016).

 HOCOMOCO databases: HOCOMOCOv10_HUMAN, HOCOMOCOv10_MOUSE (Kulakovskiy et al., 2016).

 CIS-BP databases: Homo_sapiens, Mus_musculus (Weirauch et al., 2014).  EUKARYOTE wei2010 human and mouse databases included with MEME version

4.12.0.

2.6.1. TF binding enrichment in the unique MM enhancers near OEMM protein coding genes

Motif enrichment was performed on MM enhancers near OEMM protein coding genes, by getting the unique regions from all the interactions and merging contiguous regions using Bedtools. These regions were then inputted and processed as specified in Motif enrichment, section 2.6.

2.6.2. TF binding enrichment in the SMM enhancers regulating protein coding OESMM genes

Motif enrichment was performed on the SMM enhancers regulating protein coding OESMM genes, by getting only the unique regions from all the interactions and merging contiguous regions using Bedtools. These regions were then input and processed as specified in the start of the Motif enrichment, section 2.6.

A threshold of E-value < 0.05 was used for de novo motifs which were enriched in the regions for each enhancer set and subsequent similarity of TFs motifs with these de novo motifs. Only unique motifs - corresponding TFs combinations were reported. In cases where a motif for a TF was found for multiple species, only the human version was shown, if the human version was not present, only the mouse version was shown if it was available. The results using DREME and MEME were combined for each enhancer set. In turn, the results for each enhancer set were combined into a final table, showing whether each enriched motif in at least one MM subgroup enhancer was enriched in the other MM subgroup enhancers.

2.6.3. Gene expression comparison for TF genes binding to SMM enhancers regulating protein coding OESMM genes for the different conditions

Using the TF binding enrichments in the regulatory SMM enhancers (see TF binding

enrichment in the SMM enhancers regulating protein coding OESMM genes, section 2.6.2). A list of 425 TF genes was obtained by including genes having motif enrichment in at least one subgroup for regulatory SMM enhancers (38 TFs) and then overlapping and getting unique gene ids (Ensembl id) from:

 TF genes found in the HOCOMOCOv10 HUMAN database (641 TFs).  Genes with measured RNA-seq data in this study (57992 genes).

The 425 TF genes are referred to as annotated TF genes. To obtain Ensembl ids from gene names, the human gene symbol to Ensembl id conversions with the R package org.Hs.eg.db version 3.7.0 (Carlson, 2018) and a curated list for human genes (National Center for Biotechnology Information, 2019) were used.

For this list, the rLog with batch effects removed of the gene expression were obtained as previously specified on the sections 2.2.3 and 2.2.4 respectively for quantified genes (section 2.3.2) for all quality samples (Table 2-1, including the PC, MM CL and all MM samples). Since there are comparisons between PC, MM and MM CLs and given that MM CLs subgroups may not be comparable to MM subgroups, it was decided to not account for subgroup effect in this calculation.

Gene expression was obtained for each TF list and all combinations of:  Each group of samples PC, MM, MM CLs.

 Two groups from the annotated genes: one with the TF genes with significant binding in MM any subgroup and another one with no enrichment in MM subgroup.

For each combination, the mean regularized gene expression was calculated for the corresponding samples and genes. Then a two-sample one tail Kolmogorov-Smirnov test within each group of samples (MM, PC and MM CLs) was performed under the null hypothesis that the distribution of gene means was not greater for TFs with significant motif enrichment in any MM subgroup compared with TFs without significant motif enrichment in any MM subgroup.

100,000 permutation tests were performed through random sampling of a number of annotated genes equal to the number of TFs with a motif significantly enriched in any MM subgroup from all annotated genes. For MM, PC and MM CLs, for each sample the average expression of the selected TFs was calculated. The mean of all gene expression means (average expression of the selected TFs for the random sample of TFs) was compared to that of MM subgroup enriched annotated genes in any MM subgroup. The number of tests having the random sample average TF gene average expression higher or equal to that of TF motif enriched annotated genes in any MM subgroup was obtained and p-values were calculated. Similarly, as done with each combination of PC, MM and MM CLs samples, another set of combinations was performed for:

 Each MM subgroup.

 Two groups from the annotated genes: one with the TF genes with significant binding in the particular MM and another one with no enrichment in the particular MM subgroup.

For each combination, the mean regularized gene expression was calculated for the corresponding samples and genes. Then a two-sample one tail Kolmogorov-Smirnov test within each group of samples (MM subgroups) was performed under the null hypothesis that the distribution of gene means was not greater for TFs with significant motif enrichment in the particular MM subgroup compared with TFs without significant motif enrichment in the particular MM subgroup.

For each subgroup, 100,000 permutation tests were performed through random sampling of all annotated genes, each sample’s number of elements was equal to the number of TF motif enriched annotated genes in that particular MM subgroup. The average of all gene expression averages (average expression of the selected TFs for the random sample) was compared to that of MM subgroup enriched annotated genes in that MM subgroup. The number of tests having the random sample average TF gene average expression higher or equal to that of TF

motif enriched annotated TF genes in that MM subgroup was obtained and p-values were calculated.

Documento similar