The quantitative outcomes of this thesis are analysed with repeated measure linear models, pairwise comparisons of samples and confidence intervals.
In general, the presented experiments consist of severalfactors, which represent experimental changes that are controlled by the experimenter, for example, which colour is presented. A factor has severallevels, that is, multiple ways the factor is presented or modified. For the example of colours, this could mean three levels: green, blue and red. An experiment can have multiple factors that are crossed. That means all combinations of factors are presented. If we had another factors display with the levels CRT and LED, we would have a3×2factorial design, with overall six conditions. To determine how each factor or level influence the outcome of an experiment statistical methods are employed.
During the time the research of this thesis was conducted a paradigm shift regarding statistical approaches began. This is reflected in changing methods used to evaluate the experimental data. While the experiments in Part II are analysed using statistics based on null-hypothesis testing and p- values, the later experiments in Part III are analysed using “new” statistics, which focuses on interpreting effect size measures and comparison of
2.6. Statistics
confidence intervals.
The following descriptions should serve as a short overview of the statis- tical methodology and a guide on how to interpret the presented results. For a more detailed explanations and mathematical background more specialised literature on quantitative analysis is available, for example Field (2013), Henson (2006) or Cumming (2013).
2.6.1
Omnibus tests
Omnibus tests analyse the variance in a data set to detect whether a change in an experimental condition is accompanied by a change in the experimental measurements. In this thesis, ANOVA F-tests with repeated measures are used to test for changes in experimental conditions. This means that for each condition each participant provides a data point. The overall data is then analysed for changes between the conditions based on the individual changes between conditions. The outcome indicates whether there could be a change between conditions, but not which conditions differ. Thus, of more than two levels exist for an experimental manipulation, additional analysis of the data is required, for example, follow-up statistical tests or direct visual observation of the data.
The overview of the results of a repeated measures ANOVA will be given in a table, listing the main effects and interactions in each line. They contain the following parameters:
Degrees of freedom.dfis related to the number of groups that are used in a given test. Both of these values describe the data that was used to calculate the statistics, but not the outcome. For a test involving a single factor the first degree of freedom are calculated as
df=number of groups−1
For a test involving multiple factors they are computed as the product of all the degrees of freedom for all factors. The second degrees of freedom
e
calculated as
e
df=number of observations−1
+number of participants−1
+number of levels−1
F-test statistic(F). The F-value describes the ratio between variance in the data set and the variance that can be explained by the variation of the tested experimental factor. Thus, a high F-value indicates a high degree of influence of the experimental factor on observed outcome, while a low F-value indicates a small effect on the outcome.
P-value(p). This value describes the probability that the observed results would have occurred at random, assuming that the experimental changes had no effect on the outcome. A common threshold to call a result statistical significant is a p-value of0.05. This practice, however, is under recent scrutiny and in general other ways of determining importance (for example, effect size measures) are advised.
Partial eta squared (η2p). This value represents a normalised measure of effect size measure. It is based on the differences between means in the given test and normalised by the observed variance. Thus it has no meaning regarding the original measurement, but can be used to compare the strength of an effect between different measures, factors and experiments. A higher value indicates a stronger effect of the manipulation on the outcome.
2.6.2
Additional Effect Size Measures
As follow-up tests for the omnibus test the differences between conditions are reported as mean paired differences MDin measurement units and as
standardised mean changes ddiff(MD divided by the standard deviation).
MDis useful as it gives an impression of change between two variables in
colour units, while the standardised values are useful to make the effect sizes comparable regardless of the measure and experiment.
2.6. Statistics
2.6.3
Confidence intervals
95% confidence intervals (CI) are generated through the Bias-Corrected Accelerated Non-Parametric bootstrapping algorithm8. They allow us to interpret the results and their reliability without depending on p-values and significance testing, which are strongly argued against in the current statistical and psychological literature (see Cumming 2013 and Henson 2006 for more details).
3
CHAPTER
THREE
RELATED
WORK
The research presented in this thesis is situated in the area of gaze- contingent displays. This chapter describes previous classifications of GCDs and highlighting the current state of the art. It then proceeds to situate this thesis by surveying the existing literature and creating a classification that highlights GCD techniques that have received less attention. To do this, it describes GCDs according to their intent: improv- ing rendering performance, investigating properties of perception and supporting perception.