• No se han encontrado resultados

3.   MARCO METODOLÓGICO 37

3.1   Método de investigación según su propósito 37

Gene Expression

Gene expression (GE) is the process by which information from a gene is used in the synthesis of a functional gene product (e.g. protein). The measurement of expression is typically done by detecting messenger RNA (mRNA). Studies of gene expression profile have revealed its power in predicting disease outcome and selecting therapies for individual patients (Van’t Veer et al., 2002). For such a high dimensional data set (16615 features), PCA is applied to study important variation components. In Figure 4.4, the diagonal plots display the 1-dimensional distributions of the gene expression data block onto the first 4 principal component (PC) directions i.e. scores using the same format as in Figure 2.1. The off-diagonal plots show the 2-dimensional projections onto the subspaces generated by each pair of these 4 PC directions. Each symbol represents a patient and is colored by subtypes red for Luminal A, magenta for Luminal B, cyan for HER2 and blue for Basal-like.

From the diagonal plots, the first PC shows a clear subtype difference between Basal- like, HER2 and Luminal. The second PC presents a separation between Luminal A and the other subtypes. The third PC contains little subtype information. The fourth shows some but not strong evidence of separation between Basal-like and HER2. This strong connection between gene expression variations and class differences is mainly due to the fact that these subtypes are determined based on gene expression.

Copy Number Variation

The copy number variation (CN) are a form of structural variation of the two copies of a genome. It has been well known that differences in the DNA sequence of genomes have important impacts to personal traits. However, some recent studies have shown that copy number data also play an important role in characterizing individual risk of cancers and drug responses, e.g. Sebat et al. (2007); Xu et al. (2008). In Figure 4.5,

PC 1 LumA LumB Her2 Basal -50 0 50 PC 4 PC 2 PC 3 -100 -50 0 50 PC 1 PC 4 -100 0 100 PC 2 -100 0 100 PC 3

Figure 4.4: The first 4 PC projections of the gene expression data block. The first PC presents strong evidence of subtype differences between Basal-like, and the union of the others. The second PC shows a separation between Luminal A and the other subtypes.

the first 4 PC projections of the copy number data block are displayed similarly as in Figure 4.4. The first two largest variation PCs show little correlation with subtypes differences. The third PC presents strong evidence of differences between Basal-like and the other subtypes and the fourth PC shows some differences between Luminal A and the others.

Reverse Phase Protein Array

Reverse phase protein (micro)arrays (RPPA) is a new, sensitive, high-throughput technology for obtaining protein micro-arrays which provide quantitative profiling of disease associated proteins (Charboneau et al., 2002). A broad assessment of quanti- tative protein changes in diseased and healthy tissue can be offered by RPPA data. This protein profiling has the potential for detecting meaningful protein and pathway interactions of known proteins (Tibes et al., 2006). Figure 4.6 suggests that the first PC contains useful information for distinguishing the Basal-like from the Luminial B. The second PC presents the differences between Basal-like and Luminal which can be similarly found in the PCs of gene expression and copy number. The fourth displays

PC 1 LumA LumB Her2 Basal -40 -20 0 20 40 PC 4 PC 2 PC 3 -80 -60 -40 -20 0 20 40 PC 1 PC 4 -40 -20 0 20 40 60 PC 2 -40 -20 0 20 40 PC 3

Figure 4.5: The first 4 PC projections of the copy number data block. The third PC presents a strong evidence of differences between Basal-like and the other subtypes and the fourth PC shows some differences between Luminal A and the others.

separation between HER2 and the other subtypes which is not clearly presented in the first 4 PCs of gene expression and copy number.

Gene Mutation

Gene mutation is a permanent alteration of the nucleotide sequence of the genome, which might result in different types of change in sequences and thus alter the product of a gene, or prevent the gene from functioning properly or completely. Mutations in certain genes, described as high penetrance, are often associated with high risk of developing some types of cancer e.g. breast cancer (Tung et al., 2015). Figure 4.7 visualizes the first 4 PC projections of the Mutation data block. It can be seen that each PC tends to be driven by several influential data points and is not useful for revealing their associations with subtype differences.

Such observation is due to the low frequency of mutations of most patients, which can be observed in the left panel of Figure 4.8. Each dot corresponds to one patient, with colors based on breast cancer subtypes as in Figure 4.7. Less than 10 patients carry more than 5% mutations among the 18256 genes and these patients are the samples

PC 1 LumA LumB Her2 Basal -2 0 2 4 6 PC 4 PC 2 PC 3 -5 0 5 10 PC 1 PC 4 -5 0 5 PC 2 -5 0 5 PC 3

Figure 4.6: The first 4 PC projections of the RPPA data block. The second PC presents differences between Basal-like and Luminal. The fourth PC displays separation between Her2- like and the other subtypes.

PC 1 LumA LumB Her2 Basal -20 -10 0 10 PC 4 PC 2 PC 3 0 10 20 PC 1 PC 4 -20 0 20 PC 2 -10 0 10 20 PC 3

Figure 4.7: The first 4 PC projections of Mutation data block. Each PC is driven by several influential data points and is not useful for revealing their associations with subtype differences.

driving the largest 4 variations presented in Figure 4.7. To reveal additional insights of this gene mutation data, a bar plot in the right panel of Figure 4.8 presents the top 25 genes with highest chance of having mutations. The height of each bar represents the percentage of patients having a mutation in the corresponding gene with the name labeled in the bar. As can be seen, TP53 and PIK3CA are the major players which are known for greatly increasing the risk of developing breast cancer (Network et al., 2012).

0 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 Percentage of Mutations

Percentage of Mutations per Patient

0 0.05 0.1 0.15 0.2 0.25 0.3 0.35 Percentage of Mutations

Percentage of Mutations per Feature

TP53 PIK3CA TTN CROCCP2 GATA3 CDH1 MAP3K1 MUC16 KMT2C NBPF1 Unknown MUC4 SYNE1 NEB PTEN DCAKD USH2A FLG NCOA3 DMD FRG1B MUC12 RYR2 HMCN1 MACF1

Figure 4.8: The left panel shows percentages of mutations among 18256 genes for each patient. The right panel presents the major set of genes having mutations.