Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Missing data

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 31 - 60 of 61

Full-Text Articles in Statistics and Probability

Multiple Imputation Of Missing Data In Structural Equation Models With Mediators And Moderators Using Gradient Boosted Machine Learning, Robert J. Milletich Ii Oct 2016

Multiple Imputation Of Missing Data In Structural Equation Models With Mediators And Moderators Using Gradient Boosted Machine Learning, Robert J. Milletich Ii

Psychology Theses & Dissertations

Mediation and moderated mediation models are two commonly used models for indirect effects analysis. In practice, missing data is a pervasive problem in structural equation modeling with psychological data. Multiple imputation (MI) is one method used to estimate model parameters in the presence of missing data, while accounting for uncertainty due to the missing data. Unfortunately, commonly used MI methods are not equipped to handle categorical variables or nonlinear variables such as interactions. In this study, we introduce a general MI framework that uses the Bayesian bootstrap (BB) method to generate posterior inferences for indirect effects and gradient boosted machine …


Crtgeedr: An R Package For Doubly Robust Generalized Estimating Equations Estimations In Cluster Randomized Trials With Missing Data, Melanie Prague, Rui Wang, Victor De Gruttola Feb 2016

Crtgeedr: An R Package For Doubly Robust Generalized Estimating Equations Estimations In Cluster Randomized Trials With Missing Data, Melanie Prague, Rui Wang, Victor De Gruttola

Harvard University Biostatistics Working Paper Series

No abstract provided.


Correction Of Verication Bias Using Log-Linear Models For A Single Binaryscale Diagnostic Tests, Haresh Rochani, Hani M. Samawi, Robert L. Vogel, Jingjing Yin Dec 2015

Correction Of Verication Bias Using Log-Linear Models For A Single Binaryscale Diagnostic Tests, Haresh Rochani, Hani M. Samawi, Robert L. Vogel, Jingjing Yin

Biostatistics: Faculty Publications

In diagnostic medicine, the test that determines the true disease status without an error is referred to as the gold standard. Even when a gold standard exists, it is extremely difficult to verify each patient due to the issues of costeffectiveness and invasive nature of the procedures. In practice some of the patients with test results are not selected for verification of the disease status which results in verification bias for diagnostic tests. The ability of the diagnostic test to correctly identify the patients with and without the disease can be evaluated by measures such as sensitivity, specificity and predictive …


The Effects Of A Planned Missingness Design On Examinee Motivation And Psychometric Quality, Matthew S. Swain May 2015

The Effects Of A Planned Missingness Design On Examinee Motivation And Psychometric Quality, Matthew S. Swain

Dissertations, 2014-2019

Assessment practitioners in higher education face increasing demands to collect assessment and accountability data to make important inferences about student learning and institutional quality. The validity of these high-stakes decisions is jeopardized, particularly in low-stakes testing contexts, when examinees do not expend sufficient motivation to perform well on the test. This study introduced planned missingness as a potential solution. In planned missingness designs, data on all items are collected but each examinee only completes a subset of items, thus increasing data collection efficiency, reducing examinee burden, and potentially increasing data quality. The current scientific reasoning test served as the Long …


Integrating Data Transformation In Principal Components Analysis, Mehdi Maadooliat, Jianhua Z. Huang, Jianhua Hu Mar 2015

Integrating Data Transformation In Principal Components Analysis, Mehdi Maadooliat, Jianhua Z. Huang, Jianhua Hu

Mathematics, Statistics and Computer Science Faculty Research and Publications

Principal component analysis (PCA) is a popular dimension-reduction method to reduce the complexity and obtain the informative aspects of high-dimensional datasets. When the data distribution is skewed, data transformation is commonly used prior to applying PCA. Such transformation is usually obtained from previous studies, prior knowledge, or trial-and-error. In this work, we develop a model-based method that integrates data transformation in PCA and finds an appropriate data transformation using the maximum profile likelihood. Extensions of the method to handle functional data and missing values are also developed. Several numerical algorithms are provided for efficient computation. The proposed method is illustrated …


Some General Guidelines For Choosing Missing Data Handling Methods In Educational Research, Jehanzeb R. Cheema Nov 2014

Some General Guidelines For Choosing Missing Data Handling Methods In Educational Research, Jehanzeb R. Cheema

Journal of Modern Applied Statistical Methods

The effect of a number of factors, such as the choice of analytical method, the handling method for missing data, sample size, and proportion of missing data, were examined to evaluate the effect of missing data treatment on accuracy of estimation. A methodological approach involving simulated data was adopted. One outcome of the statistical analyses undertaken in this study is the formulation of easy-to-implement guidelines for educational researchers that allows one to choose one of the following factors when all others are given: sample size, proportion of missing data in the sample, method of analysis, and missing data handling method.


Phylogenetic Linkage Among Hiv-Infected Village Residents In Botswana: Estimation Of Clustering Rates In The Presence Of Missing Data, Nicole Bohme Carnegie, Rui Wang, Vladimir Novitsky, Victor G. Degruttola Jun 2013

Phylogenetic Linkage Among Hiv-Infected Village Residents In Botswana: Estimation Of Clustering Rates In The Presence Of Missing Data, Nicole Bohme Carnegie, Rui Wang, Vladimir Novitsky, Victor G. Degruttola

Harvard University Biostatistics Working Paper Series

No abstract provided.


Jmasm 32: Multiple Imputation Of Missing Multilevel, Longitudinal Data: A Case When Practical Considerations Trump Best Practices?, Jennifer E. V. Lloyd, Jelena Obradović, Richard M. Carpiano, Frosso Motti-Stefanidi May 2013

Jmasm 32: Multiple Imputation Of Missing Multilevel, Longitudinal Data: A Case When Practical Considerations Trump Best Practices?, Jennifer E. V. Lloyd, Jelena Obradović, Richard M. Carpiano, Frosso Motti-Stefanidi

Journal of Modern Applied Statistical Methods

A pedagogical tool is presented for applied researchers dealing with incomplete multilevel, longitudinal data. It explains why such data pose special challenges regarding missingness. Syntax created to perform a multiply-imputed growth modeling procedure in Stata Version 11 (StataCorp, 2009) is also described.


Targeted Estimation Of Variable Importance Measures With Interval-Censored Outcomes, Stephanie Sapp, Mark J. Van Der Laan, Kimberly Page Feb 2013

Targeted Estimation Of Variable Importance Measures With Interval-Censored Outcomes, Stephanie Sapp, Mark J. Van Der Laan, Kimberly Page

U.C. Berkeley Division of Biostatistics Working Paper Series

In most experimental and observational studies, participants are not followed in continuous time. Instead, data is collected about participants only at certain monitoring times. These monitoring times are random, and often participant specific. As a result, outcomes are only known up to random time intervals, resulting in interval-censored data. In contrast, when estimating variable importance measures on interval-censored outcomes, practitioners often ignore the presence of interval-censoring, and instead treat the data as continuous or right-censored, applying ad-hoc approaches to mask the true interval-censoring. In this paper, we describe Targeted Minimum Loss-based Estimation methods tailored for estimation of variable importance measures …


In Praise Of Simplicity Not Mathematistry! Ten Simple Powerful Ideas For The Statistical Scientist, Roderick J. Little Jan 2013

In Praise Of Simplicity Not Mathematistry! Ten Simple Powerful Ideas For The Statistical Scientist, Roderick J. Little

The University of Michigan Department of Biostatistics Working Paper Series

Ronald Fisher was by all accounts a first-rate mathematician, but he saw himself as a scientist, not a mathematician, and he railed against what George Box called (in his Fisher lecture) "mathematistry". Mathematics is the indispensable foundation for statistics, but our subject is constantly under assault by people who want to turn statistics into a branch of mathematics, making the subject as impenetrable to non-mathematicians as possible. Valuing simplicity, I describe ten simple and powerful ideas that have influenced my thinking about statistics, in my areas of research interest: missing data, causal inference, survey sampling, and statistical modeling in general. …


A Cautionary Note On Generalized Linear Models For Covariance Of Unbalanced Longitudinal Data, Jianhua Z. Huang, Min Chen, Mehdi Maadooliat, Mohsen Pourahmadi Mar 2012

A Cautionary Note On Generalized Linear Models For Covariance Of Unbalanced Longitudinal Data, Jianhua Z. Huang, Min Chen, Mehdi Maadooliat, Mohsen Pourahmadi

Mathematics, Statistics and Computer Science Faculty Research and Publications

Missing data in longitudinal studies can create enormous challenges in data analysis when coupled with the positive-definiteness constraint on a covariance matrix. For complete balanced data, the Cholesky decomposition of a covariance matrix makes it possible to remove the positive-definiteness constraint and use a generalized linear model setup to jointly model the mean and covariance using covariates (Pourahmadi, 2000). However, this approach may not be directly applicable when the longitudinal data are unbalanced, as coherent regression models for the dependence across all times and subjects may not exist. Within the existing generalized linear model framework, we show how to overcome …


Proxy Pattern-Mixture Analysis For A Binary Variable Subject To Nonresponse., Rebecca H. Andridge, Roderick J. Little Nov 2011

Proxy Pattern-Mixture Analysis For A Binary Variable Subject To Nonresponse., Rebecca H. Andridge, Roderick J. Little

The University of Michigan Department of Biostatistics Working Paper Series

We consider assessment of the impact of nonresponse for a binary survey

variable Y subject to nonresponse, when there is a set of covariates

observed for nonrespondents and respondents. To reduce dimensionality and

for simplicity we reduce the covariates to a continuous proxy variable X

that has the highest correlation with Y, estimated from a probit

regression analysis of respondent data. We extend our previously proposed

proxy-pattern mixture analysis (PPMA) for continuous outcomes to the binary

outcome using a latent variable approach. The method does not assume data

are missing at random, and creates a framework for sensitivity analyses.

Maximum …


A Multivariate Variable Model With Possibility Of Missing Data On A Stochastic Process, Ehsan B. Samani Jun 2011

A Multivariate Variable Model With Possibility Of Missing Data On A Stochastic Process, Ehsan B. Samani

Applications and Applied Mathematics: An International Journal (AAM)

A joint model for multivariate responses with potentially non-random missing values on a stochastic process is proposed. A full likelihood-based approach that allows yielding maximum likelihood estimates of the model parameters is used. Sensitivity of the results to the assumptions is also investigated. A common way to investigate whether perturbations of model components influence key results of the analysis is to compare the results derived from the original and perturbed models using a general index of sensitivity (ISNI). The approach is illustrated by analyzing a finance data set.


The Performance Of Multiple Imputation For Likert-Type Items With Missing Data, Walter Leite, S. Natasha Beretvas May 2010

The Performance Of Multiple Imputation For Likert-Type Items With Missing Data, Walter Leite, S. Natasha Beretvas

Journal of Modern Applied Statistical Methods

The performance of multiple imputation (MI) for missing data in Likert-type items assuming multivariate normality was assessed using simulation methods. MI was robust to violations of continuity and normality. With 30% of missing data, MAR conditions resulted in negatively biased correlations. With 50% missingness, all results were negatively biased.


An Evaluation Of Multiple Imputation For Meta-Analytic Structural Equation Modeling, Carolyn F. Furlow, S. Natasha Beretvas May 2010

An Evaluation Of Multiple Imputation For Meta-Analytic Structural Equation Modeling, Carolyn F. Furlow, S. Natasha Beretvas

Journal of Modern Applied Statistical Methods

A simulation study was used to evaluate multiple imputation (MI) to handle MCAR correlations in the first step of meta-analytic structural equation modeling: the synthesis of the correlation matrix and the test of homogeneity. No substantial parameter bias resulted from using MI. Although some SE bias was found for meta-analyses involving smaller numbers of studies, the homogeneity test was never rejected when using MI.


Applying Multiple Imputation With Geostatistical Models To Account For Item Nonresponse In Environmental Data, Breda Munoz, Virginia M. Lesser, Ruben A. Smith May 2010

Applying Multiple Imputation With Geostatistical Models To Account For Item Nonresponse In Environmental Data, Breda Munoz, Virginia M. Lesser, Ruben A. Smith

Journal of Modern Applied Statistical Methods

Methods proposed to solve the missing data problem in estimation procedures should consider the type of missing data, the missing data mechanism, the sampling design and the availability of auxiliary variables correlated with the process of interest. This article explores the use of geostatistical models with multiple imputation to deal with missing data in environmental surveys. The method is applied to the analysis of data generated from a probability survey to estimate Coho salmon abundance in streams located in western Oregon watersheds.


Multiple Imputation For The Comparison Of Two Screening Tests In Two-Phase Alzheimer Studies, Ofer Harel, Xiao-Hua Zhou Sep 2006

Multiple Imputation For The Comparison Of Two Screening Tests In Two-Phase Alzheimer Studies, Ofer Harel, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Two-phase designs are common in epidemiological studies of dementia, and especially in Alzheimer research. In the first phase, all subjects are screened using a common screening test(s), while in the second phase, only a subset of these subjects is tested using a more definitive verification assessment, i.e. golden standard test. When comparing the accuracy of two screening tests in a two-phase study of dementia, inferences are commonly made using only the verified sample. It is well documented that in that case, there is a risk for bias, called verification bias. When the two screening tests have only two values (e.g. …


A Comparison For Longitudinal Data Missing Due To Truncation, Rong Liu Jan 2006

A Comparison For Longitudinal Data Missing Due To Truncation, Rong Liu

Theses and Dissertations

Many longitudinal clinical studies suffer from patient dropout. Often the dropout is nonignorable and the missing mechanism needs to be incorporated in the analysis. The methods handling missing data make various assumptions about the missing mechanism, and their utility in practice depends on whether these assumptions apply in a specific application. Ramakrishnan and Wang (2005) proposed a method (MDT) to handle nonignorable missing data, where missing is due to the observations exceeding an unobserved threshold. Assuming that the observations arise from a truncated normal distribution, they suggested an EM algorithm to simplify the estimation.In this dissertation the EM algorithm is …


Autologous Stem Cell Transplant: Factors Predicting The Yield Of Cd34+ Cells, Elizabeth Anne Lawson Dec 2005

Autologous Stem Cell Transplant: Factors Predicting The Yield Of Cd34+ Cells, Elizabeth Anne Lawson

Theses and Dissertations

Stem cell transplant is often considered the last hope for the survival for many cancer patients. The CD34+ cell content of a collection of stem cells has appeared as the most reliable indicator of the quantity of desired cells in a peripheral blood stem cell harvest and is used as a surrogate measure of the sample quality. Factors predicting the yield of CD34+ cells in a collection are not yet fully understood. Throughout the literature, there has been conflicting evidence with regards to age, gender, disease status, and prior radiation. In addition to the factors that have already been explored, …


Inference For P(Y, Vee Ming Ng Nov 2005

Inference For P(Y, Vee Ming Ng

Journal of Modern Applied Statistical Methods

Some tests and confidence bounds for the reliability parameter R=P(Y


Multiple Imputation For Correcting Verification Bias, Ofer Harel, Xiao-Hua Zhou May 2005

Multiple Imputation For Correcting Verification Bias, Ofer Harel, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

In the case in which all subjects are screened using a common test, and only a subset of these subjects are tested using a golden standard test, it is well documented that there is a risk for bias, called verification bias. When the test has only two levels (e.g. positive and negative) and we are trying to estimate the sensitivity and specificity of the test, one is actually constructing a confidence interval for a binomial proportion. Since it is well documented that this estimation is not trivial even with complete data, we adopt Multiple imputation (MI) framework for verification bias …


Assessing Treatment Effects In Randomized Longitudinal Two-Group Designs With Missing Observations, James Algina, H. J. Keselman Nov 2004

Assessing Treatment Effects In Randomized Longitudinal Two-Group Designs With Missing Observations, James Algina, H. J. Keselman

Journal of Modern Applied Statistical Methods

SAS’s PROC MIXED can be problematic when analyzing data from randomized longitudinal two-group designs when observations are missing over time. Overall (1996, 1999) and colleagues found a number of procedures that are effective in controlling the number of false positives (Type I errors) and are yet sensitive (powerful) to detect treatment effects. Two favorable methods incorporate time in study and baseline scores to model the missing data mechanism; one method was a single-stage PROC MIXED ANCOVA solution and the other was a two-stage endpoint analysis using the change scores as dependent scores. Because the twostage approach can lack sensitivity to …


Modeling Incomplete Longitudinal Data, Hakan Demirtas Nov 2004

Modeling Incomplete Longitudinal Data, Hakan Demirtas

Journal of Modern Applied Statistical Methods

This article presents a review of popular parametric, semiparametric and ad-hoc approaches for analyzing incomplete longitudinal data.


Non-Parametric Estimation Of Roc Curves In The Absence Of A Gold Standard, Xiao-Hua Zhou, Pete Castelluccio, Chuan Zhou Jul 2004

Non-Parametric Estimation Of Roc Curves In The Absence Of A Gold Standard, Xiao-Hua Zhou, Pete Castelluccio, Chuan Zhou

UW Biostatistics Working Paper Series

In evaluation of diagnostic accuracy of tests, a gold standard on the disease status is required. However, in many complex diseases, it is impossible or unethical to obtain such the gold standard. If an imperfect standard is used as if it were a gold standard, the estimated accuracy of the tests would be biased. This type of bias is called imperfect gold standard bias. In this paper we develop a maximum likelihood (ML) method for estimating ROC curves and their areas of ordinal-scale tests in the absence of a gold standard. Our simulation study shows the proposed estimates for the …


Does Weighting For Nonresponse Increase The Variance Of Survey Means?, Rod Little, Sonya L. Vartivarian Apr 2004

Does Weighting For Nonresponse Increase The Variance Of Survey Means?, Rod Little, Sonya L. Vartivarian

The University of Michigan Department of Biostatistics Working Paper Series

Nonresponse weighting is a common method for handling unit nonresponse in surveys. A widespread view is that the weighting method is aimed at reducing nonresponse bias, at the expense of an increase in variance. Hence, the efficacy of weighting adjustments becomes a bias-variance trade-off. This note suggests that this view is an oversimplification -- nonresponse weighting can in fact lead to a reduction in variance as well as bias. A covariate for a weighting adjustment must have two characteristics to reduce nonresponse bias - it needs to be related to the probability of response, and it needs to be related …


Modern Statistical Methods For Handling Missing Repeated Measurements In Obesity Trial Data: Beyond Locf, Gary L. Gadbury, C. S. Coffey, D. B. Allison Aug 2003

Modern Statistical Methods For Handling Missing Repeated Measurements In Obesity Trial Data: Beyond Locf, Gary L. Gadbury, C. S. Coffey, D. B. Allison

Mathematics and Statistics Faculty Research & Creative Works

This paper brings together some modern statistical methods to address the problem of missing data in obesity trials with repeated measurements. Such missing data occur when subjects miss one or more follow-up visits or drop out early from an obesity trial. a common approach to dealing with missing data because of dropout is 'last observation carried forward' (LOCF). This method, although intuitively appealing, requires restrictive assumptions to produce valid statistical conclusions. We review the need for obesity trials, the assumptions that must be made regarding missing data in such trials, and some modern statistical methods for analyzing data containing missing …


Mixtures Of Varying Coefficient Models For Longitudinal Data With Discrete Or Continuous Non-Ignorable Dropout, Joseph W. Hogan, Xihong Lin, Benjamin A. Herman May 2003

Mixtures Of Varying Coefficient Models For Longitudinal Data With Discrete Or Continuous Non-Ignorable Dropout, Joseph W. Hogan, Xihong Lin, Benjamin A. Herman

The University of Michigan Department of Biostatistics Working Paper Series

The analysis of longitudinal repeated measures data is frequently complicated by missing data due to informative dropout. We describe a mixture model for joint distribution for longitudinal repeated measures, where the dropout distribution may be continuous and the dependence between response and dropout is semiparametric. Specifically, we assume that responses follow a varying coefficient random effects model conditional on dropout time, where the regression coefficients depend on dropout time through unspecified nonparametric functions that are estimated using step functions when dropout time is discrete (e.g., for panel data) and using smoothing splines when dropout time is continuous. Inference under the …


Analyzing Group By Time Effects In Longitudinal Two-Group Randomized Trial Designs With Missing Data, James Algina, H. J. Keselman, Abdul R. Othman May 2003

Analyzing Group By Time Effects In Longitudinal Two-Group Randomized Trial Designs With Missing Data, James Algina, H. J. Keselman, Abdul R. Othman

Journal of Modern Applied Statistical Methods

We investigated bias, sampling variability, Type I error and power of nine approaches for testing the group by time interaction in a repeated measures design under three types of missing data mechanisms. One procedure due to Overall, Ahn, Shivakumar, and Kalburgi (1999) performed reasonably well over a range of conditions.


Chronic Disease Data And Analysis: Current State Of The Field, Ralph D'Agostino Sr., Lisa M. Sullivan Nov 2002

Chronic Disease Data And Analysis: Current State Of The Field, Ralph D'Agostino Sr., Lisa M. Sullivan

Journal of Modern Applied Statistical Methods

Chronic disease usually spans years of a person’s lifetime and includes a disease free period, a preclinical, or latent period, where there are few overt signs of disease, a clinical period where the disease manifests and is eventually diagnosed, and a follow-up period where the disease might progress steadily or remain stable. It is often of interest to investigate the relationship between risk factors measured at a point in time (usually during the disease free or preclinical period), and the development of disease at some future point (e.g., 10 years later). We outline some popular designs for the identification of …


Program For Missing Data In The Multivariate Normal Distribution, Chi-Ping Lu Jan 1975

Program For Missing Data In The Multivariate Normal Distribution, Chi-Ping Lu

All Graduate Plan B and other Reports, Spring 1920 to Spring 2023

Missing data can often cause many problems in research work. Therefore for carrying out analysis, some procedure for obtaining estimates in the presence of missing data should be applied. Various theories and techniques have been developed for different types of problems.

Analysis of the Multivariate Normal Distribution with missing data is one of the areas studied. It has been discussed earlier by Wilkes (1932), Lord (1955), Edgett (1956) and Hartley (1958). They have established some basic concepts and an outline in the way of estimation.

In the last ten years, A. A. Afifi and R. M. Elasfoff also have contributed …