1. Planteamiento del Problema de Investigación
1.4. Alcance
The aim of this study was to review how empirical research on web accessibility is conducted, to understand the changing state of the art in web accessibility and to determine whether the accessibility of websites has improved over the last two decades. This involved a comprehensive and systematic review of website accessibility evaluation studies that have been published over a 15-year period between 1999 and 2014.
The reviewing process indicated an unanticipated degree of variability and noise in the data with studies varying on just about every aspect of web accessibility evaluation. Studies vary according to the number and sampling of websites and the webpages they contain; the type and location of the evaluated websites; the choice of evaluation methods; the selection of tools, browsers, devices and ATs; the selection of guidelines, standards and heuristics against which websites are evaluated; the number and sampling of user groups; the documentation of methodology; and the formatting and
presentation of results. This considerable methodological variability made it extremely difficult to meta-analyse the data or draw reliable conclusions from it.
Nevertheless, despite this considerable noise, the usable, comparable data that could be extracted from 155 of the 397 publications selected for inclusion in this systematic review shows a persistent occurrence of accessibility problems that does not appear to be improving. Between 1999 and 2014, on average, less than a third of evaluated websites conformed to WCAG 1.0 Level A, the most basic level of web accessibility. The type or location of the websites evaluated appears to have little bearing on accessibility levels. Different types of websites (e.g. governmental, educational, commercial etc.) across a range of countries are consistently found to have serious accessibility problems. Similarly, the choice of evaluation method and tools has had little impact on the results: evaluations using automated tools, evaluations relying on the judgement of one or more expert evaluators, and user evaluations with disabled people consistently report low levels of web accessibility.
These results are consistent with other, much smaller, longitudinal studies of web accessibility that cover much shorter and older time periods. For example, Hackett et al. (2005) evaluated the accessibility of a random selection of 40 websites over a five-year period from 1997 to 2002. Having access to the original source code allowed them to determine that the websites became progressively inaccessible while increasing in
complexity. Loiacono et al. (2009) evaluated the accessibility of Fortune 100 websites between 2000 and 2005. They evaluated the homepage of each website against all three levels of WCAG 1.0 and established that, whilst the number of level A violations had decreased over time, the number of level AA and AAA violations had either remained the same or increased. Loiacono et al. surmise that an increased use of automated evaluation tools had allowed companies to correct the more obvious errors but at the expense of more thorough manual inspections. Hanson and Richards (2013) evaluated the accessibility of over 100 high-trafficked and government websites from the US and UK over a 14-year period from 1999 to 2012. They determined that despite some evidence of improvement (particularly for government websites), overall the websites exhibited generally low conformance with accessibility indicators. The authors speculate that any improvements in accessibility may be attributed to the development of good coding practices and SEO techniques rather than a growing awareness and adherence to WCAG guidelines.
To be able to properly compare web accessibility evaluation studies and to reliably benchmark the accessibility of websites, the accessibility community needs a consistent methodology with more complete methods of recording and reporting the data. In Europe, many member states established their own initiatives to analyse and monitor the accessibility of public sector websites, such as the Observatorio Español de Accesibilidad (Spanish Accessibility Observatory)36 and the Polska Akademia
Dostępności (Polish Academy of Accessibility)37. As part of a wider initiative to
monitor the accessibility of Information and Communication Technology (ICT) products and services across Europe, the MeAC (Measuring Progress of eAccessibility in Europe) project evaluated websites from the 27 member states as well as 4 relevant third countries (Norway, Australia, Canada and the US). The most recent project report, MeAC III (Kubische, Cullen, Dolphin, Laurin, & Cederbom, 2013) indicated that there was still considerable room for improvement in web accessibility levels across Europe. In response to the then proposed European Union Directive 2016/2102/EU on
making websites and mobile apps of public sector bodies more accessible, a study led by the Swedish accessibility consultancy Funka proposed a unified European web
accessibility monitoring methodology (Laurin et al., 2015). The methodology includes
36https://administracionelectronica.gob.es/pae_Home/pae_Estrategias/pae_Accesibilidad/pae_observa
torio_accesibilidad_eng.html
both automated and manual evaluation methods performed by experts, in conjunction with self-declaration by website owners and “end-user involvement” (essentially a feedback/complaint mechanism). It also recommends that the sample of pages be based on the recommendations made in WCAG-EM (Velleman et al., 2014). It is hoped that this will harmonise existing approaches to monitoring accessibility and make it much easier to measure and compare accessibility compliance across member states. While attempts to establish a unified accessibility monitoring methodology are laudable, they are currently limited to a specific geographic region (the EU) and to specific sectors within it (public sector websites). Extending the approach to incorporate a more global perspective and to monitor both public and private sector websites would better reflect the diversity of the web and provide a more accurate indication of the current state of web accessibility.
This systematic review is not without its limitations. According to Timmins and McCabe (2005), “the use of appropriate keywords is the cornerstone of an effective search” (p. 4). Given the focus of this systematic review, the initial search terms included variations on the phrase “web accessibility”. Although these search terms seemed appropriate, they returned not only a high number of publications to process but also a considerable number of irrelevant publications. Some publications were certainly relevant to web accessibility but did not feature a website evaluation. Many publications used the term “web accessibility” to refer to the performance or availability of websites rather than their accessibility to people with disabilities. Other publications had little to no relevance but had been returned due to the inclusion of a web
accessibility statement on the website hosting the publication. It is unclear how this limitation could have been avoided: increasing the specificity of the search terms (to include, for example, “accessibility testing” or “accessibility guidelines”) may have increased the relevance of the results but at the cost of excluding publications that do not include such terms. An inspection of the keywords chosen by authors of web accessibility evaluation studies also confirms the appropriateness of the search terms: though some include additional keywords, such as “user testing” or “people with disabilities”, the majority of publications use one or more of the initial search terms. Despite the use of appropriate search terms, additional checks of each author’s publication record revealed that the initial searches had missed a considerable number of publications (387 potentially relevant publications, later narrowed down to 131
relevant publications). The reasons for this disparity are unclear. One possibility is that the publications chosen by authors to publish their work were not indexed by any of the databases or publisher’s portals examined. This seems unlikely, however, given the breadth of sources initially examined, which returned publications from across a range of disciplines. In addition to searching databases and publisher’s portals, it may have been prudent to manually examine the proceedings of specific journals and conferences, such as Transactions on Accessible Computing (TACCESS) or the Web for All
Conference (W4A). Whether this would have returned any additional publications, however, is questionable, as the initial search results included several publications from such sources. Another possibility is that, at the time the initial searches were conducted (September 2014), publications had not yet been indexed by any of the databases or publisher’s portals. However, this seems only likely to apply to publications from the latter half of 2014, of which few emerged from the additional checks of each author’s publication record. The additional publications were published as far back as 1999, which suggests another cause for the disparity. Of course, another possibility is
inconsistencies in the selection and extraction of publications. However, attempts were made to cross-check the data during the study selection process.
Publication and selection biases pose a potential threat to the validity of systematic reviews. Publication bias is the tendency for only positive results to be published, whereas selection bias occurs when the sample of studies selected for analysis is not representative of the larger population of studies. Although no explicit attempt was made to control for publication bias in this study, no evidence of it was seen in the publications examined, which included both positive and negative results. Furthermore, there do not appear to be any financial, political or ideological interests that researchers or publishers of web accessibility studies may have for presenting a particular type of result. The inclusion of only peer-reviewed literature and exclusion of so-called “grey literature” and unpublished results may have resulted in selection bias and potentially threatened the generalisability of the work. For instance, a global report on web accessibility commissioned by the United Nations and conducted by the British
accessibility company, Nomensa, was considered to be grey literature and therefore did not meet the inclusion criteria for this systematic review. The United Nations report was, however, further developed by Thompson, Burgstahler, Moore, Gunderson and Hoyt (2007), which did meet the inclusion criteria for this review. Nevertheless, this threat was alleviated by the selection of high quality research from a broad range of
sources. Overall, the selection of publications in this systematic review is appropriate and representative of the web accessibility evaluation studies published, and these studies present a balanced overview of the state of web accessibility.
3.5. Conclusions
This chapter has presented a systematic review of studies that have evaluated web accessibility published over the 15-year period between 1999 and 2014. As established in Chapter 2, numerous initiatives over this period have sought to support, encourage and compel web developers to fulfil their responsibility to develop accessible websites. These include the formation of projects, working groups, and task forces; the definition of web accessibility standards, policies and legal imperatives; and the development of tools, guidelines and other resources. One would have expected such initiatives to herald an increase in levels of web accessibility, either through promoting greater knowledge and awareness of the topic or by creating a legal and moral imperative. Instead, the usable, comparable data that could be extracted from the 397 studies selected for inclusion in this systematic review shows a persistent occurrence of accessibility problems that does not appear to be improving. Between 1999 and 2014, on average, less than a third of evaluated websites conformed to WCAG 1.0 Level A, the most basic level of web accessibility.
The reasons for this surprising outcome are unclear. One explanation could be the sampling of websites used in the publications selected for review. Without a deeper analysis of the sampling strategies taken, it is impossible to determine the real-world applicability of the results. Even if this were possible, the proliferation and growing complexity of websites over the 15-year period may have simply outstripped the rate at which web developer knowledge, awareness and implementation of web accessibility has increased. Furthermore, it is impossible to determine whether developers of the evaluated websites were even exposed to WCAG and other web accessibility initiatives. More consistent methodology and complete methods of recording and reporting the data in web accessibility evaluation studies would have allowed a much deeper analysis, which may have resulted in a more reliable conclusion. Nevertheless, the outcome of this systematic review supports the notion that existing tools, guidelines and resources do not adequately support web developers and highlights an important research gap to be filled.