Selecting data that are both appropriate to the research question and the resources
available to the research – including personal expertise, time and financial resources
- is one stage of the research process. In this section, a summary of the major
advantages and disadvantages of working with secondary data as opposed to primary
data is presented. A summary of this information is shown in Table 4.2.
The first major advantage of using secondary data is that the data collection process
is often informed by expertise and professionalism not commonly available to
smaller scale studies. For example, UK government surveys use a complex sample
design and weighting system which provides researchers with more control over the
analysis and manipulation of the sample population. Smaller scale surveys may
employ convenience sample where generalisability is questionable.
The second major advantage of using secondary data is the breadth of data available
for research purposes. There are few researchers who would have the resources to
collect data from a representative sample of the population in the UK, let alone
repeatedly collect this data or follow these individuals over time. Sorensen et al.
(1996) suggested that sample size depending on subgroup analysis,
representativeness and risk of bias (e.g. from recall, non-response) is an important
advantage of secondary data. The UK government conducts a multitude of surveys
on large representative samples of the population at set intervals providing a rich
Table 4.2 Advantages and Disadvantages of Primary and Secondary Data.
Type Advantages Disadvantages
Secondary Data • Inexpensive.
• Easily accessible.
• In some cases is available immediately. • Clarify research problem.
• May provide required background information. • May contribute to creativity.
• May not be current (e.g. census data). • Possibility that data is unreliable.
• Collected by someone else for some other purpose. • May not be applicable/suitable for research needs. • Required data may be unavailable or difficult to
obtain.
Primary Data • Applicable to research question.
• Up-to-date as collected for immediate data needs.
• Expensive. • Time consuming.
• Not immediately available.
• Specialist training and experience required to design study and collect data.
is of particular importantance in epidemiology and public health where the focus is
largely on the health of populations as opposed to individuals.
The third major advantage of secondary data is economy – that is that someone else
has already collected the data requiring no further resources directed at this phase
of the research. Many secondary data sets are publicly available for little or no cost.
This means costs associated with the use of secondary data can be significantly lower
than primary data (Herron, 1989) in relation to salary expenses, transportation, etc.
There is also a time saving associated with secondary data. Castle (2003) highlighted
that the traditional approach of primary data can lead to the loss of valuable time
and resources. With the data already collected, often pre-cleaned to some extent by
professionals and stored electronically such as in the UK Data Archive, researchers
can devote more time to understanding the research process and analysing the data.
Alvarez, Canduela and Raeside (2012) argued that students’ use of secondary data
will allow a more fuller understanding of the survey process as well as providing
sufficient data for meaningful analysis. Alvarez et al. (2012) suggested that this
provides an enhanced educational experience beyond that offered by small-scale
primary surveys which often have questionable sampling frames and low response
rates. Preference can also be an advantage – secondary data analysis provides an
important opportunity for researchers to spend more time testing hypotheses using
existing data as opposed to writing grant applications for funding to collect primary
data. Thus, researchers can be more creative in how they address research questions
The first major disadvantage of using secondary data is in relation to the distinct
difference between primary and secondary data – the purpose it was collected for.
As previously mentioned, in secondary data analysis the data is being explored for a
different purpose than that of the current study and thus particular information
which the secondary researchers might be interested in may not have been collected.
For example, data may have been collected on a different geographical region or
variables of interest may have been categorised differently (e.g. age was categorised
as intervals rather than a continuous variable or race categorised as white or other
rather than several groups).
A second major disadvantage of using secondary data is that because the secondary
researcher was not involved in the planning and execution of the data collection
phase, the reliability, accuracy, and quality of the data cannot be determined. More
specifically the secondary researcher does not know how consistently the data
collection phase was conducted and the extent to which the data was affected by
problems such as low response rates and respondents’ misinterpretation of
questions. To mitigate this, secondary data, such as that provided by the UK
government often have extensive documentation about their data collection phase,
response rates, variable categorisation and other technical information readily
available in data archives. Using only trusted secondary data sets and interrogating
4.4.3 Issues around ethics, access and rules around data management