• No se han encontrado resultados

ANÁLISIS DE LOS RAYOS Y SU EXPRESIÓN Primer Rayo De Voluntad O Poder.

In document Curso completo de Magia Negra (página 196-200)

Choosing accurate statistical measure when describing phenotype is essential for making inferences in life sciences. There is strong motivation behind applying statistical methods in biomedical research. Firstly, all components of living organisms by the definition are subject to variation. Secondly, accuracy and precision of any measurement depends on the chosen method. Moreover, in the majority of studies researchers use estimates that depend on sampling variation. Any observational data is based on a sampling of particular phenotype population. Gaining information whether the difference in the measurements is because of the sampling variation or underlying phenomenon is at the heart of statistics and should be a pre-requisite to any experimental design, followed by the hypotheses testing and drawing conclusions. The tendency in the current research is relying on the sophisticated, highly automated statistical software and laboratory equipment. The understanding of experimental design and the rationale behind it is often missing. Moreover, in that kind of scientific system of mass information, it became extremely easy to overlook details and replicate incorrect measurements. The need for improving the applicability of statistical methods has been also recently noticed by others (Leek, 2015). Interestingly, in 2005 an epidemiologist, John Ioannidis (Stanford University, California) concluded that most of the published results are false (Nuzzo, 2014). This article and the interview given in 2013 by Randy Schekman, the US biologist, a Nobel Prize winner in physiology and medicine, made me re-think the way I work as a researcher and which part of the research I would like to focus on (Nobel winner boycott science journals). My primary background is in medical physics, however, between 2009 and 2012 I actively participated in biomedical laboratory research. Thus, I found myself at the clash of two different scientific worlds: highly predictable, the precise world of measurements versus the highly unpredictable world of experiments involving living systems. As a person with extensive training in physics, mathematics, and statistics (years 2003 and 2008), I found it challenging to work in the MDM2/p53 laboratory. I could not find the balance between meeting all the deadlines and ensuring the quality of research at the same time, e.g. by the time I would have been satisfied with the quality of my measurements, all H1299 cells were probably dead due to the lack of the nutrients in the measurement equipment. In 2010, I continued my research work in the nanoengineering laboratory, in the team that was a multi-disciplinary one. Again, the differences in scientific approaches amongst researchers, based on their different backgrounds, became very clear to me: engineers did not often appreciate the complexity and unpredictability of the living

- 36 -

systems, while biologists tended to focus on the big picture more, than on data processing and mathematical description of the laboratory data collected.

As a beginning researcher working on the collaborative projects I often struggled with small sizes of biological samples, incorrect use of average values, common lack of understanding of what standard deviation means in practice, analysing scatterplots that had some data points missing, with the lack of access to the raw numerical data from some of my collaborators, assumption that the data is always distributed normally (reporting only mean values with standard deviation even when the data was not distributed normally), manipulating the shape of the histogram by changing the bin width and drawing conclusions based on it. By practicing science, I realized the impact of graphical visualization of the laboratory outcomes on attracting business partners, the power of choosing proper central tendency measure when describing observations, the importance of checking and understanding data distribution, as well as the importance of associating all quantitative data with qualitative data. The need of merging quantitative and qualitative data as well as the amount of data I was expected to manage and analyze led me to find better, more cost- and time- effective data management solutions. I found that the answers to the research questions were often there, hidden in the lab-books of my colleagues or seemingly outdated books and journal articles. Therefore, I also learned the importance of data capture from the paper resources and the proper use of data storage systems as well as running database queries.

All the issues mentioned above made me learn and evolve as a scientist, and I regard them as one of the biggest challenges in the collaborative research. However, in this thesis, I decided to focus on how to tackle bypassing the normality assumption of the data distribution while dealing with the data describing living systems. Additionally, by measuring and simulating a relatively simple organism, Neurospora crassa, I decided to show how using an alternative statistical approach can make the analysis of biological data a more accurate and precise task, thus, helping to make better, evidence-based decisions.

Although the presented methodology is focused on the specific organism, it is based on a well-established set of MATLAB procedures and is a universal method of generating simulation values by previous measurements. Surprisingly, this methodology seems not to be implemented in any automated software for life sciences applications, although it can be successfully applied when the normality assumption is not met and when any other parametric method does not reflect the distribution of the data collected. Interestingly, suggested approach can aid decision making for all living systems, including human

- 37 -

populations. It can also increase value and reduce waste in biomedical research by more precise reflection of biological variation compared to the standard parametric distributions. Data points considered as outliers, which are unwanted in the great majority parametric methods, are very often at the heart of life sciences and neglecting them to satisfy the theoretical conditions of parametric tests might lead to missing scientific discoveries. The method I propose here is a non-parametric one and allows to construct a custom distribution, based strictly on the measured data. Choosing a non-parametric methodology, in this case, allowed me for better management of outliers and non-typical observations. Usually, observations we make about the behavior of various organisms living on Earth are

used to produce the numbers and associated statistics. Here, I apply a reversed engineering approach: I use the numbers and various statistics associated with the same numerical data to “produce” certain behavior and visualize the outcomes to see how statistical descriptions match reality. Examples of it are shown in the Figure 13 below.

Neurospora crassa wt large field of view (link to the animation)

Neurospora crassa wt large field of view (link to the animation)

Neurospora crassa wt small field of view

(link to the animation)

Neurospora crassa wt small field of view (link to the animation)

Figure 11Stochastic simulation of Neurospora crassa at the time t~30min. Produced image is based on the 100x100 MATLAB numerical matrix. The stochastic components of the system are simulated through the kernel estimation of the inverses of the cumulative distribution functions describing key growth parameters of Neurospora crassa

Figure 12 Stochastic simulation of Neurospora crassa at the time t~2h 16 min. Produced image is based on the 300x300 MATLAB numerical matrix. The stochastic components of the system are simulated through the kernel estimation of the inverses of the cumulative distribution functions describing key growth parameters of Neurospora crassa.

- 38 -

Here, I use statistics to measure and simulate the behavior of the filamentous fungus,

Neurospora crassa bearing in mind common statistical pitfalls. I never use extreme values to “adjust” results to the story. Introducing results only as proportions can be very misleading

Figure 13 Examples of in silico Neurospora crassa wild type (wt) compared with its real-world equivalents.(1) A. in silico fungus at the time t~25 min, B. hyphae at the colony periphery, source:

http://ec.asm.org/content/3/2/348/F3.large.jpg(2) A. in silico fungus at the time t~29min, B. hyphae at the colony periphery, source: http://www.fungalcell.org/colony-organisation-and-development-image-gallery

(moved to the University of Manchester) (3) A. in silico fungus at the time t~41min, B. hyphae at the colony interior, source: http://ec.asm.org/content/3/2/348/F3.large.jpg, (4) A. in silico fungus at the time t~2h 01min, B. Neurospora crassa colony by Douglas Ivey, source: http://newsroom.ucr.edu/573, (5) in silico fungus at the time t ~ 2h 15min, B. Neurospora crassa laboratory culture, source:

http://www.fgsc.net/neurosimages/cultureimages/CULTUREwild_type_FGSC988.jpg (6) Comparison of simulation outcomes for various statistical approaches. Each of the simulation modes comprises 50 independent simulations (an overlay of 50 final frames from 50 independent simulations, whereat each simulation consists of 150 time steps). (Model #1) simulation values generated from the Gaussian distribution – this causes inadequate assumptions on the central tendency measure and distribution of the data, (Model #2) simulation values withdrawn from the actual data with replacement – this causes very precise prediction of the central tendency measure, but at the cost of preserving variability, also, accurate data distribution is not preserved, (Model #3) simulation values generated from the kernel distributions – this guarantees precise prediction of the central tendency measure, preserving both data distribution and variability

- 39 -

when reported without absolute values. Therefore, next to the relative frequency counts, I give the information on the number of the measurements, along with the minimum, quartile, median, and maximum values of the key growth parameters of Neurospora crassa. Moreover, I report findings in the most consistent way possible throughout the thesis, to avoid any bias and imprecision when making conclusions. I consciously choose what to count and how to count. The choice of the growth parameters is a continuation of the prior experimental research on filamentous fungi, run by my colleague, Dr. Marie Held. The research question I ask is whether using normal (Gauss) distributions for the statistical description of

Neurospora crassa is justified. I express numerical results as images to make them more intuitive and thus accessible for human judgment. I make a clear assumption that Neurospora crassa extends from the tip and that all generated branches extend in exactly the same way, from the tip. Furthermore, I keep in mind that there are different types of averages, therefore, in the results I report all possible central tendency measure including means, medians, and modes. I include outliers in the computer simulations. However, the simulations are set up in a way that the probability of generating outliers is very low, what reflects the situation in a real world. I use the laboratory sample provided by Professor Roger Lew (York University, Canada) and show that by sampling a population I still can achieve valid simulation results regarding the whole population of that organism (please, see Figure 13 for comparison of simulation outcomes with Neurospora crassa grown in a laboratory conditions). Firstly, I test the statistical model by using MATLAB matrix consisting of 100x100 cells. Subsequently, I increase the size of the matrix to see how efficiently that model can predict the overall pattern of the colony evolution in time. Finally, I show that running the same parametrical model using different assumptions about data distribution affects results. The subject of the study a general one. Using kernel estimation for predicting the behavior of live systems improves the precision, especially regarding preserving variability and outliers. Thus, it might lead to more accurate interventions. However, I would like to stress the necessity of linking the proposed statistical model with qualitative data. Considring qualitative and quantitative information in parallel protects the analysis from the confounding effect, e.g. adding certain substances during experiments with Neurospora crassa, deleting certain genes, or changing pressure during laboratory experiment can radically change the behavior of that organism (Held, 2011b, Lew, 2011)

- 40 -

In document Curso completo de Magia Negra (página 196-200)