Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Missing data

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 61

Full-Text Articles in Statistics and Probability

Variance Shrinkage In Dunnett-Type Multiple Comparisons With Missing Data, Md Habibullah Jan 2026

Variance Shrinkage In Dunnett-Type Multiple Comparisons With Missing Data, Md Habibullah

UNF Graduate Theses and Dissertations

Dunnett’s procedure is widely used for comparing multiple treatments with a control, but its application becomes challenging in the presence of missing data and multiple com- parisons. An improved Dunnett-type procedure addresses this by using multiple imputation under Rubin’s framework and constructing unified confidence intervals based on a multi- variate t distribution, allowing valid simultaneous inference while controlling the family-wise error rate (FWER). This work further extends the method by incorporating shrinkage-based variance estimation. Specifically, individual group variances are shrunk toward a common value to improve stability. This approach is particularly effective when group variances are similar or moderately different, …


Statistical Inference For Noisy Matrix Completion Incorporating Auxiliary Information, Shujie Ma, Po-Yao Niu, Yichong Zhang, Yinchu Zhu Apr 2025

Statistical Inference For Noisy Matrix Completion Incorporating Auxiliary Information, Shujie Ma, Po-Yao Niu, Yichong Zhang, Yinchu Zhu

Research Collection School Of Economics

This article investigates statistical inference for noisy matrix completion in a semi-supervised model when auxiliary covariates are available. The model consists of two parts. One part is a low-rank matrix induced by unobserved latent factors; the other part models the effects of the observed covariates through a coefficient matrix which is composed of high-dimensional column vectors. We model the observational pattern of the responses through a logistic regression of the covariates, and allow its probability to go to zero as the sample size increases. We apply an iterative least squares (LS) estimation approach in our considered context. The iterative LS …


Complex Missing Data Problems In Education Surveys, Thomas Wesley Robertson Jan 2025

Complex Missing Data Problems In Education Surveys, Thomas Wesley Robertson

Electronic Theses & Dissertations (2024 - present)

Missing data are a nearly universal problem in human subjects research, including in education. However, reporting and addressing missing data is an issue, despite guidelines from the APA style guide and the What Works Clearinghouse, as well as guidance from prominent statisticians on the best methods to use. Prior research conducted in 2004 and 2014 found that in the field of education, most studies do not report or address missing data. In addition, no study has looked specifically at how missing data are reported and addressed in complex surveys. The current study has two main objectives: first, to determine if …


Improvement And Evaluation Of Multiple Imputation By Heckman's One-Step Mle For Binary Mnar Outcomes And Various Types Of Mar Covariates, Xin W. Shore Sep 2024

Improvement And Evaluation Of Multiple Imputation By Heckman's One-Step Mle For Binary Mnar Outcomes And Various Types Of Mar Covariates, Xin W. Shore

Mathematics & Statistics ETDs

Missing data is inevitable in clinical epidemiology. It becomes one of the major challenges in the analyses and can potentially undermine the validity of results and conclusions. Although methods for handling missing data with mechanisms of missing completely at random (MCAR) or missing at random (MAR) have been widely researched, methods adapted for the missing not at random (MNAR) mechanism are less studied. Galimard et al. (2018) have derived a method to use multiple imputation by Heckman's One-Step ML Estimation for binary MNAR outcome and continuous MAR covariates (MIHEml). This dissertation focuses on updating MIHEml in terms of …


Evaluation Of Imputation Methods Focusing On Categorical Outcomes, Nadia Bernardo Mendoza Jan 2024

Evaluation Of Imputation Methods Focusing On Categorical Outcomes, Nadia Bernardo Mendoza

CGU Theses & Dissertations

In general, standard statistical analysis models typically rely on completely observed cases, excluding incomplete rows from the dataset. This approach poses particular challenges when the objective is to predict a rare outcome, especially when some of the ob servations with the rare outcome are incomplete. In such cases, the available information to support the model in predicting this event is reduces. Theoretically correct models may pre dict all instances in the majority class achieving high accuracy, but fail in predicting the rare cases, which are often the most interesting ones. Therefore, it is crucial to make the most of all …


Addressing The Analytical And Computational Challenges Using Machine Learning In Biomedical Research, Yizhuo Wang Dec 2023

Addressing The Analytical And Computational Challenges Using Machine Learning In Biomedical Research, Yizhuo Wang

Dissertations and Theses (Open Access)

In the contemporary healthcare field, professionals are confronted with an ever-growing volume of clinical data stored in electronic health records, alongside the genomic data stemming from laboratory experiments. As a response to this deluge of data, the application of machine learning (ML) techniques is gaining popularity since ML techniques have demonstrated an exceptional proficiency in processing big data and deciphering complex nonlinear patterns that are intrinsic to biomedical research.

My research leverages ML's capabilities to address the computational challenges spanning diverse areas, including adaptive clinical trial designs, survival analysis, and high-dimensional genetic data analysis. Specifically, Chapter 2 focused on the …


Weighted Mean Difference Statistics For Paired Data In The Presence Of Missing Values, Yuntong Li, Brent J. Shelton, William St Clair, Heidi L. Weiss, John L. Villano, Arnold Stromberg, Chi Wang, Li Chen Aug 2023

Weighted Mean Difference Statistics For Paired Data In The Presence Of Missing Values, Yuntong Li, Brent J. Shelton, William St Clair, Heidi L. Weiss, John L. Villano, Arnold Stromberg, Chi Wang, Li Chen

Markey Cancer Center Faculty Publications

Missing data is a common issue in many biomedical studies. Under a paired design, some subjects may have missing values in either one or both of the conditions due to loss of follow-up, insufficient biological samples, etc. Such partially paired data complicate statistical comparison of the distribution of the variable of interest between the two conditions. In this article, we propose a general class of test statistics based on the difference in weighted sample means without imposing any distributional or model assumption. An optimal weight is derived from this class of tests. Simulation studies show that our proposed test with …


A Multistate Competing Risks Framework For Preconception Prediction Of Pregnancy Outcomes, Kaitlyn Cook, Neil J. Perkins, Enrique Schisterman, Sebastien Haneuse Dec 2022

A Multistate Competing Risks Framework For Preconception Prediction Of Pregnancy Outcomes, Kaitlyn Cook, Neil J. Perkins, Enrique Schisterman, Sebastien Haneuse

Statistical and Data Sciences: Faculty Publications

Background: Preconception pregnancy risk profiles—characterizing the likelihood that a pregnancy attempt results in a full-term birth, preterm birth, clinical pregnancy loss, or failure to conceive—can provide critical information during the early stages of a pregnancy attempt, when obstetricians are best positioned to intervene to improve the chances of successful conception and full-term live birth. Yet the task of constructing and validating risk assessment tools for this earlier intervention window is complicated by several statistical features: the final outcome of the pregnancy attempt is multinomial in nature, and it summarizes the results of two intermediate stages, conception and gestation, whose outcomes …


Statistical Challenges And Methods For Missing And Imbalanced Data, Rose Adjei Dec 2022

Statistical Challenges And Methods For Missing And Imbalanced Data, Rose Adjei

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Missing data remains a prevalent issue in every area of research. The impact of missing data, if not carefully handled, can be detrimental to any statistical analysis. Some statistical challenges associated with missing data include, loss of information, reduced statistical power and non-generalizability of findings in a study. It is therefore crucial that researchers pay close and particular attention when dealing with missing data. This multi-paper dissertation provides insight into missing data across different fields of study and addresses some of the above mentioned challenges of missing data through simulation studies and application to real datasets. The first paper of …


Using The Fraction Of Missing Information (Fmi) In Selecting Auxiliary Variables To Impute Missingness In Confirmatory Factor Analysis (Cfa), Dareen Taha Alzahrani Jan 2022

Using The Fraction Of Missing Information (Fmi) In Selecting Auxiliary Variables To Impute Missingness In Confirmatory Factor Analysis (Cfa), Dareen Taha Alzahrani

Electronic Theses and Dissertations

This study aimed to investigate the effectiveness of using the fraction of missing information (FMI) to select auxiliary variables in imputing missing data in confirmatory factor analysis (CFA). This was done by conducting two studies (a simulation study and an empirical study). A Monte Carlo simulation technique was used to compare the performance and the effect of the restrictive strategy based on FMI and the inclusive strategy on parameter estimate bias and parameter estimate efficiency. The missing data mechanisms, missing data proportion, correlation strength between the analysis variables and auxiliary variables, and the inclusive and restrictive strategies were assessed in …


Conditional And Marginal Imputation Models For Multilevel Data, Gang Liu Aug 2021

Conditional And Marginal Imputation Models For Multilevel Data, Gang Liu

Legacy Theses & Dissertations (2009 - 2024)

This dissertation study extends sequential hierarchical regression imputation (SHRIMP) methods to multilevel datasets with three levels of nesting and proposes a marginal method based on marginalized multilevel model (MMM) framework. Specifically, the proposed model consists of two levels such that the first level relates the marginal mean of responses with covariates through a generalized regression model and the second level includes subject specific random effects within the same generalized regression model. To draw the inference on the population-averaged or subject-specified coefficients, the hierarchical regression and/or MMM is applied as the imputation and estimation models. We employ Markov Chain Monte Carlo …


Performance Comparison Of Imputation Methods For Mixed Data Missing At Random With Small And Large Sample Data Set With Different Variability, Kyei Afari Aug 2021

Performance Comparison Of Imputation Methods For Mixed Data Missing At Random With Small And Large Sample Data Set With Different Variability, Kyei Afari

Electronic Theses and Dissertations

One of the concerns in the field of statistics is the presence of missing data, which leads to bias in parameter estimation and inaccurate results. However, the multiple imputation procedure is a remedy for handling missing data. This study looked at the best multiple imputation methods used to handle mixed variable datasets with different sample sizes and variability along with different levels of missingness. The study employed the predictive mean matching, classification and regression trees, and the random forest imputation methods. For each dataset, the multiple regression parameter estimates for the complete datasets were compared to the multiple regression parameter …


Compare And Contrast Maximum Likelihood Method And Inverse Probability Weighting Method In Missing Data Analysis, Scott Sun May 2021

Compare And Contrast Maximum Likelihood Method And Inverse Probability Weighting Method In Missing Data Analysis, Scott Sun

Mathematical Sciences Technical Reports (MSTR)

Data can be lost for different reasons, but sometimes the missingness is a part of the data collection process. Unbiased and efficient estimation of the parameters governing the response mean model requires the missing data to be appropriately addressed. This paper compares and contrasts the Maximum Likelihood and Inverse Probability Weighting estimators in an Outcome-Dependendent Sampling design that deliberately generates incomplete observations. WE demonstrate the comparison through numerical simulations under varied conditions: different coefficient of determination, and whether or not the mean model is misspecified.


Performance Comparison Of Multiple Imputation Methods For Quantitative Variables For Small And Large Data With Differing Variability, Vincent Onyame May 2021

Performance Comparison Of Multiple Imputation Methods For Quantitative Variables For Small And Large Data With Differing Variability, Vincent Onyame

Electronic Theses and Dissertations

Missing data continues to be one of the main problems in data analysis as it reduces sample representativeness and consequently, causes biased estimates. Multiple imputation methods have been established as an effective method of handling missing data. In this study, we examined multiple imputation methods for quantitative variables on twelve data sets with varied sizes and variability that were pseudo generated from an original data. The multiple imputation methods examined are the predictive mean matching, Bayesian linear regression and linear regression, non-Bayesian in the MICE (Multiple Imputation Chain Equation) package in the statistical software, R. The parameter estimates generated from …


Imputation, Modelling And Optimal Sampling Design For Digital Camera Data In Recreational Fisheries Monitoring, Ebenezer Afrifa-Yamoah Jan 2021

Imputation, Modelling And Optimal Sampling Design For Digital Camera Data In Recreational Fisheries Monitoring, Ebenezer Afrifa-Yamoah

Theses: Doctorates and Masters

Digital camera monitoring has evolved as an active application-oriented scheme to help address questions in areas such as fisheries, ecology, computer vision, artificial intelligence, and criminology. In recreational fisheries research, digital camera monitoring has become a viable option for probability-based survey methods, and is also used for corroborative and validation purposes. In comparison to onsite surveys (e.g. boat ramp surveys), digital cameras provide a cost-effective method of monitoring boating activity and fishing effort, including night-time fishing activities. However, there are challenges in the use of digital camera monitoring that need to be resolved. Notably, missing data problems and the cost …


Almost All Missing Data Are Mnar, Thomas R. Knapp Sep 2020

Almost All Missing Data Are Mnar, Thomas R. Knapp

Journal of Modern Applied Statistical Methods

Rubin (1976, and elsewhere) claimed that there are three kinds of “missingness”: missing completely at random; missing at random; and missing not at random. He gave examples of each. The article that now follows takes an opposing view by arguing that almost all missing data are missing not at random.


Multiple Imputation Using Influential Exponential Tilting In Case Of Non-Ignorable Missing Data, Kavita Gohil Jan 2020

Multiple Imputation Using Influential Exponential Tilting In Case Of Non-Ignorable Missing Data, Kavita Gohil

College of Graduate Studies: Theses & Dissertations

Modern research strategies rely predominantly on three steps, data collection, data analysis, and inference. In research, if the data is not collected as designed, researchers may face challenges of having incomplete data, especially when it is non-ignorable. These situations affect the subsequent steps of evaluation and make them difficult to perform. Inference with incomplete data is a challenging task in data analysis and clinical trials when missing data related to the condition under the study. Moreover, results obtained from incomplete data are prone to biases. Parameter estimation with non-ignorable missing data is even more challenging to handle and extract useful …


Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li Jan 2020

Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li

Theses and Dissertations--Statistics

Comparing the distribution of biomarker measurements between two groups under either an unpaired or paired design is a common goal in many biomarker studies. However, analyzing biomarker data is sometimes challenging because the data may not be normally distributed and contain a large fraction of zero values or missing values. Although several statistical methods have been proposed, they either require data normality assumption, or are inefficient. We proposed a novel two-part semiparametric method for data under an unpaired setting and a nonparametric method for data under a paired setting. The semiparametric method considers a two-part model, a logistic regression for …


Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui Jan 2020

Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui

Theses and Dissertations--Statistics

In this dissertation, we investigate three distinct but interrelated problems for nonparametric analysis of clustered data and multivariate data in pre-post factorial design.

In the first project, we propose a nonparametric approach for one-sample clustered data in pre-post intervention design. In particular, we consider the situation where for some clusters all members are only observed at either pre or post intervention but not both. This type of clustered data is referred to us as partially complete clustered data. Unlike most of its parametric counterparts, we do not assume specific models for data distributions, intra-cluster dependence structure or variability, in effect …


Evaluation Of Modern Missing Data Handling Methods For Coefficient Alpha, Katerina Matysova Dec 2019

Evaluation Of Modern Missing Data Handling Methods For Coefficient Alpha, Katerina Matysova

College of Education and Human Sciences: Dissertations, Theses, and Student Research

When assessing a certain characteristic or trait using a multiple item measure, quality of that measure can be assessed by examining the reliability. To avoid multiple time points, reliability can be represented by internal consistency, which is most commonly calculated using Cronbach’s coefficient alpha. Almost every time human participants are involved in research, there is missing data involved. Missing data means that even though complete data were expected to be collected, some data are missing. Missing data can follow different patterns as well as be the result of different mechanisms. One traditional way to deal with missing data is listwise …


The Estimation Of Missing Values In Rectangular Lattice Designs, Emmanuel Ogochukwu Ossai, Abimibola Victoria Oladugba Sep 2019

The Estimation Of Missing Values In Rectangular Lattice Designs, Emmanuel Ogochukwu Ossai, Abimibola Victoria Oladugba

Journal of Modern Applied Statistical Methods

Algebraic expressions for estimating missing data when one or more observation(s) are missing in Rectangular lattice designs with repetition were derived using the method of minimizing the residual sum of squares. Results showed that the estimated value(s) were significantly approximate to that of the actual value(s).


Comparison Of Imputation Methods For Mixed Data Missing At Random, Kaitlyn Heidt May 2019

Comparison Of Imputation Methods For Mixed Data Missing At Random, Kaitlyn Heidt

Electronic Theses and Dissertations

A statistician's job is to produce statistical models. When these models are precise and unbiased, we can relate them to new data appropriately. However, when data sets have missing values, assumptions to statistical methods are violated and produce biased results. The statistician's objective is to implement methods that produce unbiased and accurate results. Research in missing data is becoming popular as modern methods that produce unbiased and accurate results are emerging, such as MICE in R, a statistical software. Using real data, we compare four common imputation methods, in the MICE package in R, at different levels of missingness. The …


Fixed Choice Design And Augmented Fixed Choice Design For Network Data With Missing Observations, Miles Q. Ott, Matthew T. Harrison, Krista J. Gile, Nancy P. Barnett, Joseph W. Hogan Jan 2019

Fixed Choice Design And Augmented Fixed Choice Design For Network Data With Missing Observations, Miles Q. Ott, Matthew T. Harrison, Krista J. Gile, Nancy P. Barnett, Joseph W. Hogan

Statistical and Data Sciences: Faculty Publications

The statistical analysis of social networks is increasingly used to understand social processes and patterns. The association between social relationships and individual behaviors is of particular interest to sociologists, psychologists, and public health researchers. Several recent network studies make use of the fixed choice design (FCD), which induces missing edges in the network data. Because of the complex dependence structure inherent in networks, missing data can pose very difficult problems for valid statistical inference. In this article, we introduce novel methods for accounting for the FCD censoring and introduce a new survey design, which we call the augmented fixed choice …


Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao Jan 2019

Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao

Theses and Dissertations

In longitudinal studies, outcomes are measured repeatedly over time, but in reality clinical studies are full of missing data points of monotone and non-monotone nature. Often this missingness is related to the unobserved data so that it is non-ignorable. In such context, pattern-mixture model (PMM) is one popular tool to analyze the joint distribution of outcome and missingness patterns. Then the unobserved outcomes are imputed using the distribution of observed outcomes, conditioned on missing patterns. However, the existing methods suffer from model identification issues if data is sparse in specific missing patterns, which is very likely to happen with a …


Forecasting Crashes, Credit Card Default, And Imputation Analysis On Missing Values By The Use Of Neural Networks, Jazmin Quezada Jan 2019

Forecasting Crashes, Credit Card Default, And Imputation Analysis On Missing Values By The Use Of Neural Networks, Jazmin Quezada

Open Access Theses & Dissertations

A neural network is a system of hardware and/or software patterned after the operation of neurons in the human brain. Neural networks,- also called Artificial Neural Networks - are a variety of deep learning technology, which also falls under the umbrella of artificial intelligence, or AI. Recent studies shows that Artificial Neural Network has the highest coefficient of determination (i.e. measure to assess how well a model explains and predicts future outcomes.) in comparison to the K-nearest neighbor classifiers, logistic regression, discriminant analysis, naive Bayesian classifier, and classification trees. In this work, the theoretical description of the neural network methodology …


Handling Missing Data In Single-Case Studies, Chao-Ying Joanne Peng, Li-Ting Chen Jun 2018

Handling Missing Data In Single-Case Studies, Chao-Ying Joanne Peng, Li-Ting Chen

Journal of Modern Applied Statistical Methods

Multiple imputation is illustrated for dealing with missing data in a published SCED study. Results were compared to those obtained from available data. Merits and issues of implementation are discussed. Recommendations are offered on primal/advanced readings, statistical software, and future research.


Missing Data In Longitudinal Surveys: A Comparison Of Performance Of Modern Techniques, Paola Zaninotto, Amanda Sacker Dec 2017

Missing Data In Longitudinal Surveys: A Comparison Of Performance Of Modern Techniques, Paola Zaninotto, Amanda Sacker

Journal of Modern Applied Statistical Methods

Using a simulation study, the performance of complete case analysis, full information maximum likelihood, multivariate normal imputation, multiple imputation by chained equations and two-fold fully conditional specification to handle missing data were compared in longitudinal surveys with continuous and binary outcomes, missing covariates, and an interaction term.


Impact Of Home Visit Capacity On Genetic Association Studies Of Late-Onset Alzheimer's Disease, David W. Fardo, Laura E. Gibbons, Shubhabrata Mukherjee, M. Maria Glymour, Wayne Mccormick, Susan M. Mccurry, James D. Bowen, Eric B. Larson, Paul K. Crane Aug 2017

Impact Of Home Visit Capacity On Genetic Association Studies Of Late-Onset Alzheimer's Disease, David W. Fardo, Laura E. Gibbons, Shubhabrata Mukherjee, M. Maria Glymour, Wayne Mccormick, Susan M. Mccurry, James D. Bowen, Eric B. Larson, Paul K. Crane

Biostatistics Faculty Publications

INTRODUCTION—Findings for genetic correlates of late-onset Alzheimer's disease (LOAD) in studies that rely solely on clinic visits may differ from those with capacity to follow participants unable to attend clinic visits.

METHODS—We evaluated previously identified LOAD-risk single nucleotide variants in the prospective Adult Changes in Thought study, comparing hazard ratios (HRs) estimated using the full data set of both in-home and clinic visits (n = 1697) to HRs estimated using only data that were obtained from clinic visits (n = 1308). Models were adjusted for age, sex, principal components to account for ancestry, and additional health indicators.

RESULTS …


Multiple Ratio Imputation By The Emb Algorithm: Theory And Simulation, Masayoshi Takahashi May 2017

Multiple Ratio Imputation By The Emb Algorithm: Theory And Simulation, Masayoshi Takahashi

Journal of Modern Applied Statistical Methods

Although multiple imputation is the gold standard of treating missing data, single ratio imputation is often used in practice. Based on Monte Carlo simulation, the Expectation-Maximization with Bootstrapping (EMB) algorithm to create multiple ratio imputation is used to fill in the gap between theory and practice.


Jmasm44: Implementing Multiple Ratio Imputation By The Emb Algorithm (R), Masayoshi Takahashi May 2017

Jmasm44: Implementing Multiple Ratio Imputation By The Emb Algorithm (R), Masayoshi Takahashi

Journal of Modern Applied Statistical Methods

Although single ratio imputation is often used to deal with missing values in practice, there is a paucity of discussion regarding multiple ratio imputation. Code in the R statistical environment is presented to execute multiple ratio imputation by the Expectation-Maximization with Bootstrapping (EMB) algorithm.