• No se han encontrado resultados

Antecedentes y estado de la cuestión

C. Significación personal: por parte del aprendiz depende de factores intrapersonales, interpersonales y formales de los contenidos de enseñanza,

3.3.2. Principio II: Diseño de instrucción

In this chapter we have introduced the main features of meta-analysis, distinguishing between IPD and AD, one-stage and two-stage models and giving an overview of the missing data problems usually encountered by the analysts.

Many other key concepts and techniques related to meta-analysis have been developed, includ- ing:

• Meta-regression (Stanley and Jarrell, 1989), ”an extension to subgroup analyses that allows the effect of continuous, as well as categorical, characteristics to be investigated, and in principle allows the effects of multiple factors to be investigated simultaneously” (Higgins and Green, 2008);

• Statistical methods to detect and adjust for the recurrent issue of publication bias (East- erbrook et al., 1991), that is bias arising from the non publication of small studies that failed to prove the efficacy of a particular treatment;

• Methods for multivariate meta-analysis (Mavridis and Salanti, 2013), and

• Network meta-analysis, i.e. meta-analysis comparing several treatment effects through both direct and indirect comparisons between studies (Caldwell, 2014).

However we decided not to pursue these further here because they are not directly relevant to our work.

Some people argue that meta-analyses should be performed only with data coming from RCTs, excluding observational studies; this is because RCTs are considered a much more valid design for causal inference. However, many reviews have shown how including data from observational studies may improve the inference and that RCT-only meta-analyses often give similar results to meta-analyses of observational studies (Shrier et al., 2007).

Before continuing with the presentation of the strategies for addressing the missing data issues in meta-analysis, Chapter 3 presents the main missing data methods in a general setting. We will come back to the IPD-MA setting in Chapter 4.

3

Missing Data Methods

This chapter gives an overview of the problems raised by missing data in clinical research and of the methods currently used to handle them. In Section 3.1 we outline the different missing data mechanisms, which are the key concepts underpinning the analysis of partially observed data. Then, in Section 3.2 we review the oldest and most straightforward methods for the analysis of partially observed data, discussing why we cannot always rely on them. In Section 3.3 we introduce Rubin’s Multiple Imputation, describing both the Joint Modelling approach (Subsection 3.3.1) and Full Conditional Specification (Subsection 3.3.2). Finally, in Section 3.4 we compare this two imputation methods, itemizing pros and cons of both.

3.1

Introduction

Since unfortunately most of the studies carried out in medical and social sciences are unable to collect all the planned data, missing data are a common issue in statistical analyses. Not only do they cause a loss of information, and hence power, but they may also cause bias in parameter estimates and hence potentially lead to misleading inferences. The first paper to deal systematically with these issues was probably (Rubin, 1976), which is considered a landmark paper in this area. In it, Rubin suggested three different kinds of missing data mechanisms: Missing At Random, Missing Completely At Random and Missing Not At Random. In order to explain the different mechanisms, we define some notation.

Imagine that we intended to collect data on n units; we can call Yi the vector of variables that we planned to observe for unit i. To reflect the missing data, we split this vector into two parts: YOi , the sub-vector of observed variables for that unit, and YMi , the sub-vector of missing variables. It is important to stress that these variables do exist, but have just not been

collected in our study. Finally, we can define a third vector, called Ri, which is a vector of binary variables such that:

Yi,j missing ⇒ Ri,j = 0, Yi,j observed ⇒ Ri,j = 1.

Given this notation, we can now define the three missing data mechanisms as follows:

• Missing Completely At Random (MCAR): Data are defined to be MCAR if the prob- ability of a value being missing is completely unrelated to both the observed and the underlying unseen values on each unit. Algebraically:

P(Ri|YOi , Y M

i ) = P(Ri).

• Missing At Random (MAR): Data are defined to be MAR if, given or conditional on the observed variables on a unit, the probability of a value being missing is completely unrelated to underlying unseen values on that unit. In formulae:

P(Ri|YOi , Y M

i ) = P(Ri|YOi ).

It is quite important here to stress that this does not mean that the probability of a variable being missing is completely independent from its value. Marginally, Ri depends on YMi , but given YOi this dependence is broken.

• Missing Not At Random (MNAR): Data are defined to be MNAR if the probability of a value being missing depends on the underlying missing values on that unit, even after conditioning on the observed variables. Algebraically:

P(Ri|Yi) = P(Ri|YOi , Y M i ).

We emphasize that these are just assumptions that we make about the reasons for the data being missing in the context of the analysis at hand, and not only they are un-testable, but they are not properties of the dataset itself. For this reason, even though we might often be tempted to rely on the MAR assumption, because of the simplifications that it brings to the analysis, we should always check the robustness of the results of our analysis to different assumptions with sensitivity analysis (Molenberghs et al., 2014, Part V).