Capitulo IV. Análisis de datos
Anexo 8. Terremotos, un mundo de posibilidades Videos, experimentos y preguntas
Proteomics as a relatively new ‘post-genomic’ science focuses on the large scale determination of the functional protein network in the cell. We applied mass spectrometry based proteomics to determine proteomes and phosphoproteomes in different cell types including tumor cells and liver cells and of various organisms ranging from Escherichia coli to human. Using SILAC, we investigated protein and phosphorylation changes in-vivo upon different treatments including phosphatase inhibition and growth factor stimulation. The determination of thousands of proteins that contain posttranslationally phosphorylated residues demands description, storage, management and recovery of the obtained data. For this purpose, we created the phosphorylation site database PHOSIDA (http://www.phosida.com) (Chapter 4). Its purpose is not only to make the obtained large- scale data public to the scientific community, but also to mine the data and to derive general patterns relating to phosphorylation events. By quantitative proteomics, we show that regulation through posttranslational modifications takes place on the site level rather than the level of the entire phosphoprotein. For example, many proteins contained phosphorylation sites that were differently regulated upon epidermal growth factor stimulation. This demonstrates the necessity to establish methods that extend the common approach of matching spectra to peptide sequences. We created a probability based algorithm to detect phosphorylation events on the site level (Chapter 3). Using a specified cutoff with respect to the localization probability of phosphorylation sites (p > 0.75), we tested the accuracy of our method using manually verified high confidence data of a previous phosphoproteomic study. We found that more than 90% of determined phosphorylation sites were correctly localized within the peptide sequence.
We extended the common proteomics workflow ranging from cell preparation to matching the measured spectra to protein sequences by the application of the ‘Knowledge Discovery in Databases’ (KDD) process to extract knowledge from the obtained large-scale data. For example, we found that only a small subset of phosphorylation sites was regulated upon growth factor stimulation. All quantitative phosphoproteomic studies showed that regulation through phosphorylation was most apparent for tyrosine residues. Our data sets suggest that the distribution of pS, pT, and pY is around 85%, 13%, and 2% on average. We also observed
146
that the number of phosphorylation events in prokaryotic cells is considerably different from the one observed in eukaryotes. In fly, for example, we determined more than 10,000 in-vivo phosphorylation sites on even very low abundant proteins including kinases and transcription factors. In comparison, we did not detected more than 100 phosphorylation events in any prokaryotic cell.
The comparison of our phosphoproteomic datasets with large-scale data from other studies, which were also integrated into PHOSIDA, underlined the novelty of our high accuracy data. Overall, around 80% of determined phosphorylation sites of each study were novel. Thus the determination of phosphoproteomes is far from being complete.
Using statistical tests that are integrated into the PHOSIDA environment, we found that phosphorylation events are distributed over all cellular compartments. Some compartments such as mitochondria, however, were underrepresented, whereas phosphorylation events in the nucleus were overrepresented. On the basis of integrated secondary structure and solvent accessibility predictions, we found that phosphorylation sites were predominantly located in loops and hinges on the surface of the protein. We also found evidence for significantly overrepresented consensus sequences that surround eukaryotic phosphorylation sites and make up kinase motifs. In contrast, we could not derive any significant motif from prokaryotic phosphosites. Besides mining methods that derive general patterns regarding function, cell compartment localization, structural constraints, consensus sequences and further categories, we investigated the evolution of phosphorylation (Chapter 9). The high conservation of phosphorylation throughout higher eukaryotes on the protein level as well as on the site level underlines the functional impact of phosphorylated proteins, which play key roles in signalling and therefore have to be preserved in evolution. In this regard, the yeast phosphoproteome presents an outlier, as yeast phosphorylation sites were not significantly more conserved than their non-phosphorylated counterparts. This observation is in agreement with the fact that many kinases evolved after the speciation event that separated yeast from higher eukaryotes. In addition, a non-negligable proportion of amino acids that are phosphorylated in human, but not conserved in mouse, point to background phosphorylation that does not have any functional impact on the underlying system and therefore no selective pressure.
Furthermore, the PHOSIDA knowledge discovery pipeline also includes a phosphorylation site predictor on the basis of a support vector machine (Chapter 7). The accuracy of predicting phosphorylated serines on the basis of the raw sequence was higher than 90% for each investigated eukaryotic organism.
The inclusion of various high confidence large scale data obtained from high accuracy quantitative phosphoproteomic studies along with a phosphorylation site predictor make PHOSIDA a rich environment to the biologist wishing to analyze phosphorylation events of proteins of interest. Moreover, the automated analysis pipeline based on the KDD process enables us to derive various patterns relating to phosphorylation.
We also constructed a proteome database, termed ‘Max-Planck Unified Proteome Database’ (MAPU) that includes proteomes of different organelles, tissues and cell types (Chapter 5). Obtained proteomic data were also mapped to the genome (Chapter 8). The reassignment of identified peptide sequences to corresponding genes allows not only the assignment of important protein features including phosphorylation to the coding genome sequences but also the experimental validation of predicted genes. Using the DAS technology, we linked our proteome database with the genome database EnsEMBL. Finally, the update and extension of the sex bias database SEBIDA was a further intent of my PhD study (Chapter 6).
We intend to extend the phosphorylation site database by the inclusion of other posttranslational modifications such as acetylation, for example. It will be interesting, whether the general constraints observed in phosphorylation events can also be found in other posttranslational modifications using the KDD process. In addition, we wish to establish the first machine learning approach that is capable of predicting acetylation events, on the basis of the raw sequence.
Another goal is to integrate the evolutionary annotation provided by the EnsEMBL Compara database into the PHOSIDA web application. This will enable the web users to study the evolutionary conservation of any given phosphorylated protein throughout 36 eukaryotes. Currently, the evolutionary section of the PHOSIDA online application is restricted to phylogenetic information throughout seven eukaryotes on the basis of our self-coded pipeline. Furthermore, we intend to link our proteome databases with other online environments such as PRIDE and Peptide Atlas, in order to establish a broad proteomic data network.
148