2. GEOMORFOLOGÍA Y RELIEVE
5.5.6. ACTIVIDADES DE INVESTIGACIÓN
In order to gain further insight into the research question and the hypothesis, a com- parison between the different types of search and their requisite outcomes was re- quired. As research area—and by extension training background—had the most statistical impact, these variables became the focus of further analysis. All search- related metrics (boundary objects, the number of searches, and the average time spent between searches) were divided into two groups: keyword and visualisation. As every search was logged in the system and the type recorded, and as only one type of search was conducted during a particular time frame, this data was easily generated by the system.
The first type of test conducted was a Wilcoxon signed rank test [303] to de- termine if there was a statistical significance in the search metrics. Using keyword search as the control and visualisation search as the experiment (for the purposes of the test), there was no statistical significance found in the number of boundary ob- jects created (p = .133) nor in the average time between searches (p = .098). There was, however, a statistical significance in the number of searches conducted. Visu- alisation searches saw a statistically significant increase in the number of searches generated (p=.008).
While no statistical significance existed when comparing the means of the two experiment groups, a statistical significance does exist when running a non-parametric ANOVA. When comparing the number of boundary objects generated during the keyword search across research areas, a Kruskal-Wallis H test reveals a statistical significance: χ2(4) = 13.961, p = .007. This significance becomes even more pro- nounced when examining the number of boundary objects generated during the vi- sualisation search: χ2(4) = 16.906, p =.002.10 In a pairwise comparison of the key- word search, the statistically significant comparisons are between Computer Science– History (p=.030) and Computer Science–Library Studies (p=.036). In the visuali- sation search, a pairwise comparison showed only a statistical significance between Computer Science–History (p=.013).
9As noted in the search preference analysis, study type is ignored due to its likely association with training.
5.3. Analysis of Results 135
When considering the grouping by training type (as noted above), the results were somewhat similar. Again, a Kruskal-Wallis H test reveals a statistical signif- icance in the number of boundary objects created across type of training: χ2(1) =
12.788, p<.0005. In this instance the keyword search was actually slightly more sta- tistically relevant than the visualisation search using a Kruskal-Wallis H test, which had a result of χ2(1) = 10.262, p = .001. While both had normal distributions, it should be noted that the visualisation search in this case contained a couple of mi- nor outliers in the data. As they were minor, they were not discarded.
Comparing the other search metrics studied in the Wilcoxon signed rank test, there appears to be no statistical significance in variance among research area. See Table 5.2 for results. However when looking at the variance of the same metrics by training type, statistically significant evidence arises for how individuals engage with search. With the exception of the amount of time spent on the keyword search, all search metrics showed a statistically significant variance. See Table5.3for results.
TABLE5.2: Kruskal-Wallis H tests for search metrics by research area.
Search Metric KW Result Sig.
Average time (keyword) χ2(4) =3.777 .437
Average time (visualisation) χ2(4) =7.675 .104
Number of searches (keyword) χ2(4) =7.406 .116
Number of searches (visualisation) χ2(4) =8.634 .079
TABLE5.3: Kruskal-Wallis H tests for search metrics by training back- ground.
Search Metric KW Result Sig.
Average time (keyword) χ2(1) =2.948 .086
Average time (visualisation) χ2(1) =6.376 .012
Number of searches (keyword) χ2(1) =6729 .009
Number of searches (visualisation) χ2(1) =4.682 .030
Two separate binomial regressions were run in order to determine if any of the above metrics had an impact on the user’s reflective experience (i.e. their choice of preferred search method). In both models, preferred method was treated as a dichotomous variable, where 0 represents keyword search as the preferred method and 1 indicates visualisation as the preferred method.
The first regression looked at the effect of the number of searches generated dur- ing each type of search and the average time spent between searches in each type of search (for a total 4 metrics). The preferred method was used as the dependent variable. Linearity of the continuous variables was confirmed via the logit to the de- pendent variable as assessed by a Box-Tidwell procedure [304], using a Bonferroni correction [305] applied to all terms, resulting in statistical significance at p<.0056. There were no standardized residuals that fell more than±2.5 standard deviations from the mean, thus there was no need for correction. The model fit was considered statistically significant: χ2(4) = 10.815, p = .029, and the model explained 35.9%
of variance (using Nagelkerke R2) and correctly classified 77.1% of cases. See Table
5.4for the values for sensitivity, specificity, and positive and negative predictive val- ues. The area under the ROC curve (see Figure5.10) was .803 (with a 95% CI of .654 to .951), which is considered an excellent level of discrimination according to Hos- mer, Lemeshow, and Sturdivant [306]. However, of the 5 predictor variables, only 1 proved to be statistically significant: average number of seconds between searches using the visualisation search (see Table5.5) with it having only a slight impact (odds of 1.014) of increasing the likelihood to select the visualisation search.
TABLE 5.4: Sensitivity, Specificity, and Predictive Values for the Bi- nommial Regression using Search Metrics.
Metric %
Sensitivity 81.0 Specificity 71.4 Positive Predictive Value 80.9 Negative Predictive Value 71.4
TABLE5.5: Logistic regression predicting the likelihood of selecting Visualisation as the preferred method based on number of keyword searches, number of visualisation searches, average time between keyword searches, and average time between visualisation searches.
Variable B SE Wald df p Odds Lower Upper
# of Searches (Kw) -.039 .040 .965 1 .326 .962 .890 1.039 # of Searches (Viz) .028 .027 1.074 1 .300 1.028 .975 1.084 Avg. Time (Kw) .001 .007 .018 1 .894 1.001 .988 1.014 Avg. Time (Viz) .014 .007 .969 1 .046 1.014 1.000 1.028 Constant -1.687 1.705 .979 1 .322 .185
FIGURE5.10: ROC curve for binomial regression of search metrics to predict preferred method
5.3. Analysis of Results 137
The second regression looked at the effect of the number of boundary objects created for each search (for a total 2 metrics). The preferred method was used as the dependent variable. Linearity of the continuous variables was confirmed via the logit to the dependent variable as assessed by a Box-Tidwell procedure [304], using a Bonferroni correction [305] applied to all terms, resulting in statistical significance at p < .01. There were no standardized residuals that fell more than±2.5 standard deviations from the mean, thus there was no need for correction. The model fit was not considered statistically significant: χ2(2) = 1.835, p = .399; however, a Hosmer and Lemeshow test did not result in a statistically significant result: χ2(7) =
10.045, p = .186, which suggests the model is still a good fit [306] (see Table5.6for the contingencies).
TABLE5.6: Contingency Table for Hosmer and Lemeshow Test.
Step Observed (Kw) Expected (Kw) Observed (Viz) Expected (Viz) Total
1 2 2.576 2 1.424 4 2 2 2.063 2 1.937 4 3 4 1.723 0 2.277 4 4 1 1.578 3 2.422 4 5 2 1.429 2 2.571 4 6 1 1.600 4 3.400 5 7 0 1.255 4 2.745 4 8 2 1.210 2 2.790 4 9 0 .565 2 1.435 2
Additionally, the model explained 6.9% of variance (using Nagelkerke R2) and correctly classified 62.9% of cases. However, the area under the ROC curve (see Figure5.11) was .675 (with a 95% CI of .485 to .865), which is considered a poor level of discrimination according to Hosmer, Lemeshow, and Sturdivant [306]. None of the variables were statistically significant in the prediction, thus this regression is not considered applicable to prediction of preferred method. As such, it is not relevant to further analysis.
FIGURE5.11: ROC curve for binomial regression of boundary objects to predict preferred method
Finally, two additional binomial regressions were run to determine if any of the same metrics referenced above could be used to predict whether the participant fell into the "inductive" training group or the deductive training group. The first regres- sion tested the number of boundary objects created between the two types of search (for a total of 2 covariates). Linearity of the continuous variables was confirmed via the logit to the dependent variable as assessed by a Box-Tidwell procedure, us- ing a Bonferroni correction applied to all terms, resulting in statistical significance at p < .01. There were no standardized residuals that fell more than ±2.5 stan- dard deviations from the mean, thus there was no need for correction. The model fit was considered statistically significant: χ2(2) =19.418, p < .005, and the model explained 56.8% of variance (using Nagelkerke R2) and correctly classified 70.6% of cases. See Table5.7 for the values for sensitivity, specificity, and positive and neg- ative predictive values. The area under the ROC curve (see Figure 5.12) was .884 (with a 95% CI of .777 to .991), which is considered an excellent level of discrimi- nation according to Hosmer, Lemeshow, and Sturdivant [306]. Of the 2 predictor variables, only the number of boundary objects created during the keyword search proved to be statistically significant (see Table5.8) with each new boundary object created having a likelihood to predict an inductive participant by 1.544.
TABLE 5.7: Sensitivity, Specificity, and Predictive Values for the Bi- nommial Regression using Created Boundary Objects.
Metric %
Sensitivity 70.6 Specificity 77.8 Positive Predictive Value 75 Negative Predictive Value 73.7
5.3. Analysis of Results 139
TABLE5.8: Logistic regression predicting the likelihood of determin- ing if the user is an inductive reasoner based on the number of bound-
ary objects created in either the keyword or visualisation search.
Variable B SE Wald df p Odds Lower Upper
Boundary Objects Created (Kw) .434 .183 5.652 1 .017 1.544 1.079 2.208 Boundary Objects Created (Viz) .194 .120 2.610 1 .106 1.214 .960 1.536 Constant -2.334 .841 7.710 1 .005 .097
FIGURE5.12: ROC curve for binomial regression of boundary objects to predict training type
The second regression tested the number of searches conducted and the average amount of time between searches across the two types of search (for a total of 4 co- variates). Linearity of the continuous variables was confirmed via the logit to the dependent variable as assessed by a Box-Tidwell procedure, using a Bonferroni cor- rection applied to all terms, resulting in statistical significance at p < .0056. There were no standardized residuals that fell more than±2.5 standard deviations from the mean, thus there was no need for correction. The model fit was considered sta- tistically significant: χ2(4) = 16.210, p = .003, and the model explained 49.4% of variance (using Nagelkerke R2) and correctly classified 76.5% of cases. See Table5.9
for the values for sensitivity, specificity, and positive and negative predictive val- ues. The area under the ROC curve (see Figure5.13) was .837 (with a 95% CI of .706 to .968), which is considered an excellent level of discrimination according to Hos- mer, Lemeshow, and Sturdivant [306]. Of the 4 predictor variables, only the average amount of time spent between visualisation searches proved to be statistically sig- nificant (see Table5.10) with each additional second having a likelihood to predict an inductive participant by .982.
TABLE 5.9: Sensitivity, Specificity, and Predictive Values for the Bi- nommial Regression using Search Metrics.
Metric %
Sensitivity 76.5 Specificity 72.2 Positive Predictive Value 83.3 Negative Predictive Value 76.5
TABLE 5.10: Logistic regression predicting the likelihood of deter- mining if the user is an inductive reasoner based on number of key- word searches, number of visualisation searches, average time be- tween keyword searches, and average time between visualisation
searches.
Variable B SE Wald df p Odds Lower Upper
# of Searches (Kw) .084 .045 3.55 1 .059 1.088 .997 1.187 # of Searches (Viz) -.048 .031 2.432 1 .119 .954 .898 1.012 Avg. Time (Kw) .002 .006 .061 1 . .804 1.002 .989 1.014 Avg. Time (Viz) -.018 .008 5.069 1 .024 .982 .967 .998 Constant 2.034 1.859 1.197 1 .274 7.646
FIGURE5.13: ROC curve for binomial regression of search metrics to predict training type
A further investigation of the impact of this data is conducted in the discussion section of this chapter.