Missing Data and Multiple Imputation: An Unbiased ApproachThe default method of dealing with missing data in statistical analyses is to only use the complete observations (complete case analysis), which can lead to unexpected bias when data do not meet the assumption of missing completely at random (MCAR). For the assumption of MCAR to be met, missingness cannot be related to either the observed or unobserved variables. A less stringent assumption, missing at random (MAR), requires that missingness not be associated with the value of the missing variable itself, but can be associated with the other observed variables. When data are truly MAR as opposed to MCAR, the default complete case analysis method can lead to biased results. There are statistical options available to adjust for data that are MAR, including multiple imputation (MI) which is consistent and efficient at estimating effects. Multiple imputation uses informing variables to determine statistical distributions for each piece of missing data. Then multiple datasets are created by randomly drawing on the distributions for each piece of missing data. Since MI is efficient, only a limited number, usually less than 20, of imputed datasets are required to get stable estimates. Each imputed dataset is analyzed using standard statistical techniques, and then results are combined to get overall estimates of effect. A simulation study will be demonstrated to show the results of using the default complete case analysis, and MI in a linear regression of MCAR and MAR simulated data. Further, MI was successfully applied to the association study of CO2 levels and headaches when initial analysis showed there may be an underlying association between missing CO2 levels and reported headaches. Through MI, we were able to show that there is a strong association between average CO2 levels and the risk of headaches. Each unit increase in CO2 (mmHg) resulted in a doubling in the odds of reported headaches.
Document ID
20140003850
Acquisition Source
Johnson Space Center
Document Type
Abstract
Authors
Foy, M. (Wyle Integrated Science and Engineering Group Houston, TX, United States)
VanBaalen, M. (NASA Johnson Space Center Houston, TX, United States)
Wear, M. (Wyle Integrated Science and Engineering Group Houston, TX, United States)
Mendez, C. (MEI Technologies, Inc. Houston, TX, United States)
Mason, S. (MEI Technologies, Inc. Houston, TX, United States)
Meyers, V. (NASA Johnson Space Center Houston, TX, United States)
Alexander, D. (NASA Johnson Space Center Houston, TX, United States)
Law, J. (NASA Johnson Space Center Houston, TX, United States)
Date Acquired
April 28, 2014
Publication Date
February 11, 2014
Subject Category
Statistics And ProbabilityNumerical Analysis
Report/Patent Number
JSC-CN-30027Report Number: JSC-CN-30027
Meeting Information
Meeting: NASA Human Research Program Investigators'' Workshop
Location: Galveston, TX
Country: United States
Start Date: February 11, 2014
Sponsors: Universities Space Research Association, National Space Biomedical Research Inst.