Open Access. Powered by Scholars. Published by Universities.®
- Institution
- Keyword
-
- Causal inference (5)
- Counterfactual (5)
- Diagnostic tests (4)
- Prediction (4)
- Sensitivity (4)
-
- Classification (3)
- Confounding (3)
- Double robust estimation (3)
- G-computation estimation (3)
- Specificity (3)
- Bootstrap (2)
- Bootstrap method (2)
- Case-control sampling (2)
- Cross-validation (2)
- Current status data (2)
- Disease screening (2)
- Genetics (2)
- Health expenditures (2)
- Inverse probability of treatment weighted estimation (2)
- Linear regression (2)
- Log-normal (2)
- Longitudinal data (2)
- Multiple testing (2)
- Nonparametric maximum likelihood estimation (2)
- Point-of-service health plan (2)
- Q-Q plots (2)
- Referral to specialists (2)
- Regression splines (2)
- Skewed data (2)
- Skewed distributions (2)
- Publication Year
- Publication
-
- Harvard University Biostatistics Working Paper Series (21)
- U.C. Berkeley Division of Biostatistics Working Paper Series (20)
- UW Biostatistics Working Paper Series (15)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (7)
- The University of Michigan Department of Biostatistics Working Paper Series (4)
-
- CHIP Documents (1)
- COBRA Preprint Series (1)
- Dissertations and Theses (Open Access) (1)
- Dissertations, Theses, and Capstone Projects (1)
- EVMS School of Health Professions Faculty Publications (1)
- Electronic Theses & Dissertations (2024 - present) (1)
- Electronic Theses and Dissertations (1)
- Loma Linda University Electronic Theses, Dissertations & Projects (1)
- Pomona Senior Theses (1)
- SMU Data Science Review (1)
- Theses and Dissertations (1)
- Theses, Dissertations and Capstones (1)
- Western Research Forum (1)
- Publication Type
Articles 1 - 30 of 80
Full-Text Articles in Statistical Theory
A Predictive Coding Account Of Spatial Working Memory Following Prophylactic Levetiracetam Administration Prior To Traumatic Brain Injury, Omeima Mutwali
A Predictive Coding Account Of Spatial Working Memory Following Prophylactic Levetiracetam Administration Prior To Traumatic Brain Injury, Omeima Mutwali
Dissertations, Theses, and Capstone Projects
Traumatic brain injury (TBI) symptom prevention and remediation is an important area of research that would benefit vulnerable groups, including active-duty and veteran soldiers. These patients can sustain penetrative forces in fields of combat or in training, which result in focal lesions that trigger inflammatory and degenerative processes in the brain. Both primary and secondary injuries are associated with changes to cognition, behavior and affective state. This disease poses increased risk of epileptogenesis, as well. Given these outcomes, prior research has evaluated levetiracetam (LEV) as a prophylactic treatment for seizures, cognitive deficits and negative emotionality. LEV acts as a presynaptic …
Comparative Machine Learning Models For Disease Risk Prediction, Mercy Mawusi Agbley
Comparative Machine Learning Models For Disease Risk Prediction, Mercy Mawusi Agbley
Theses, Dissertations and Capstones
Accurate prediction of disease outcomes is crucial for improving clinical decision-making and enabling early intervention. This study compares the performance of various statistical and machine learning models for clinical risk prediction using two healthcare datasets: diabetic retinopathy and heart disease. The models assessed include Logistic Regression, LASSO, k-Nearest Neighbors (KNN), Support Vector Machines (SVM), Neural Networks, Random Forests, Gradient Boosting Machines (GBM), and a stacked ensemble model. Prior to modeling, datasets were split into train and test sets. Standardization was applied to numeric features whilst categorical features were one-hot encoded. These transformations were later applied to the test set. Principal …
First-Generation Medical School Applicants: A Quantitative Study Designed To Identify Areas Of Educational Support, Bethsabe Romero, Amanda K. Burbage
First-Generation Medical School Applicants: A Quantitative Study Designed To Identify Areas Of Educational Support, Bethsabe Romero, Amanda K. Burbage
EVMS School of Health Professions Faculty Publications
First-generation (First Gen) students are unique medical school applicants. Due to their lived experience, they approach patient care by prioritizing trust, comfort and understanding. They have proven ability to overcome obstacles and were found to be more resilient than their continuing generation (Cont Gen) peers. Despite these notable attributes, they face unique challenges in gaining medical school acceptance. There are very few quantitative studies examining this student subpopulation, and our study identifies characteristics of first-generation medical school applicants while highlighting areas of needed support. This cross-sectional study used deidentified Application and Matriculating Student Questionnaire survey data that was obtained from …
“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King
“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King
Pomona Senior Theses
The work of this thesis is twofold — first, qualitatively characterizing the confluence between the British eugenics and statistics movements in the late 19th and early 20th centuries, and second, quantitatively analyzing the effect of this foundation on pedagogical materials in the growing field of statistics between 1880 and 1970. Towards the first goal, the history of the method of least squares, state statistics, and positive and negative eugenics are outlined, followed by a close reading of the foundational texts authored by Francis Galton and Karl Pearson that introduced linear regression. Towards the latter goal, English-language statistics textbooks published between …
Theoretical Foundations And Applied Performance Of Periodicity-Aware Imputation: Variable Bandpass Block Bootstrap Methods For Incomplete Time Series, Asmaa Ahmad
Electronic Theses & Dissertations (2024 - present)
Time series data are prevalent across a wide range of disciplines, including health surveillance, public policy, and environmental monitoring. In the presence of underlying cyclical patterns, the integrity of time series analysis depends critically on the ability to detect, model, and impute structured missing data without compromising the temporal structure. This dissertation introduces and validates a novel imputation framework that integrates the Variable Bandpass Periodic Block Bootstrap (VBPBB) into multiple imputation procedures, improving the accuracy, robustness, and interpretability of time series models under high rates of missingness and noise. The overarching goal of this dissertation was to develop and evaluate …
Latent Classification Of Multi-Model Non-Homogeneous Continuous-Time Markov Chains, Joonha Chang
Latent Classification Of Multi-Model Non-Homogeneous Continuous-Time Markov Chains, Joonha Chang
Dissertations and Theses (Open Access)
The continuous-time Markov chain (CTMC) model and latent clustering models are commonly used to study longitudinal measures of categorical outcomes. Because of its simple but powerful Markovian property, CTMC models have been widely used in medical and public health researches. Due to limitations in the standard CTMC model, there have been some studies on non-homogeneous continuous-time Markov chain (NH-CTMC) that utilized time-dependent rates, but the progresses have been limited. NH-CTMC can be more powerful than CTMC by its default nature of time-dependent rate that can be fitted to a wider range of applications in medical studies. In this study, we …
Bayesian Methods For Graphical Models With Neighborhood Selection., Sagnik Bhadury
Bayesian Methods For Graphical Models With Neighborhood Selection., Sagnik Bhadury
Electronic Theses and Dissertations
Graphical models determine associations between variables through the notion of conditional independence. Gaussian graphical models are a widely used class of such models, where the relationships are formalized by non-null entries of the precision matrix. However, in high-dimensional cases, covariance estimates are typically unstable. Moreover, it is natural to expect only a few significant associations to be present in many realistic applications. This necessitates the injection of sparsity techniques into the estimation method. Classical frequentist methods, like GLASSO, use penalization techniques for this purpose. Fully Bayesian methods, on the contrary, are slow because they require iteratively sampling over a quadratic …
Accounting For The Uncertainty Due To Chemicals Below The Detection Limit In Mixture Analysis, Paul M. Hargarten
Accounting For The Uncertainty Due To Chemicals Below The Detection Limit In Mixture Analysis, Paul M. Hargarten
Theses and Dissertations
Humans are exposed to multiple chemicals every day. Epidemiological studies have shown that chemical mixtures are associated with cancers, allergies, neurodevelopmental disorders, and other adverse health effects. To assess these associations, investigators are increasingly using chemical mixture approaches like weighted quantile sum (WQS) regression. In these studies, the research objectives are to determine whether a mixture of correlated chemicals is associated with an adverse health outcome and to identify the important chemicals. However, as experimental equipment measures each exposure to a chemical-specific detection limit, the exposures are unknown between zero and the detection limit. Indeed, the number of exposures below …
Interpreting Patient Reported Outcomes In Orthopaedic Surgery: A Systematic Review, Shgufta Docter, Zina Fathalla, Michael Lukacs, Michaela Khan, Morgan Jennings, Shu-Hsuan Liu, Dong Zi, Dianne Bryant
Interpreting Patient Reported Outcomes In Orthopaedic Surgery: A Systematic Review, Shgufta Docter, Zina Fathalla, Michael Lukacs, Michaela Khan, Morgan Jennings, Shu-Hsuan Liu, Dong Zi, Dianne Bryant
Western Research Forum
Background: Reporting methods of patient reported outcome measures (PROMs) vary in orthopaedic surgery literature. While most studies report statistical significance, the interpretation of results would be improved if authors reported confidence intervals (CIs), the minimally clinically important difference (MCID), and number needed to treat (NNT).
Objective: To assess the quality and interpretability of reporting the results of PROMs. To evaluate reporting, we will assess the proportion of studies that reported (1) 95% CIs, (2) MCID, and (3) NNT. To evaluate interpretation, we will assess the proportion of studies that discussed results using the MCID or the effect sizes and how …
Overcoming Small Data Limitations In Heart Disease Prediction By Using Surrogate Data, Alfeo Sabay, Laurie Harris, Vivek Bejugama, Karen Jaceldo-Siegl
Overcoming Small Data Limitations In Heart Disease Prediction By Using Surrogate Data, Alfeo Sabay, Laurie Harris, Vivek Bejugama, Karen Jaceldo-Siegl
SMU Data Science Review
In this paper, we present a heart disease prediction use case showing how synthetic data can be used to address privacy concerns and overcome constraints inherent in small medical research data sets. While advanced machine learning algorithms, such as neural networks models, can be implemented to improve prediction accuracy, these require very large data sets which are often not available in medical or clinical research. We examine the use of surrogate data sets comprised of synthetic observations for modeling heart disease prediction. We generate surrogate data, based on the characteristics of original observations, and compare prediction accuracy results achieved from …
Evaluation Of Progress Towards The Unaids 90-90-90 Hiv Care Cascade: A Description Of Statistical Methods Used In An Interim Analysis Of The Intervention Communities In The Search Study, Laura Balzer, Joshua Schwab, Mark J. Van Der Laan, Maya L. Petersen
Evaluation Of Progress Towards The Unaids 90-90-90 Hiv Care Cascade: A Description Of Statistical Methods Used In An Interim Analysis Of The Intervention Communities In The Search Study, Laura Balzer, Joshua Schwab, Mark J. Van Der Laan, Maya L. Petersen
U.C. Berkeley Division of Biostatistics Working Paper Series
WHO guidelines call for universal antiretroviral treatment, and UNAIDS has set a global target to virally suppress most HIV-positive individuals. Accurate estimates of population-level coverage at each step of the HIV care cascade (testing, treatment, and viral suppression) are needed to assess the effectiveness of "test and treat" strategies implemented to achieve this goal. The data available to inform such estimates, however, are susceptible to informative missingness: the number of HIV-positive individuals in a population is unknown; individuals tested for HIV may not be representative of those whom a testing intervention fails to reach, and HIV-positive individuals with a viral …
Conditional Screening For Ultra-High Dimensional Covariates With Survival Outcomes, Hyokyoung Grace Hong, Jian Kang, Yi Li
Conditional Screening For Ultra-High Dimensional Covariates With Survival Outcomes, Hyokyoung Grace Hong, Jian Kang, Yi Li
The University of Michigan Department of Biostatistics Working Paper Series
Identifying important biomarkers that are predictive for cancer patients' prognosis is key in gaining better insights into the biological influences on the disease and has become a critical component of precision medicine. The emergence of large-scale biomedical survival studies, which typically involve excessive number of biomarkers, has brought high demand in designing efficient screening tools for selecting predictive biomarkers. The vast amount of biomarkers defies any existing variable selection methods via regularization. The recently developed variable screening methods, though powerful in many practical setting, fail to incorporate prior information on the importance of each biomarker and are less powerful in …
Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret
Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret
UW Biostatistics Working Paper Series
We have frequently implemented crossover studies to evaluate new therapeutic interventions for genital herpes simplex virus infection. The outcome measured to assess the efficacy of interventions on herpes disease severity is the viral shedding rate, defined as the frequency of detection of HSV on the genital skin and mucosa. We performed a simulation study to ascertain whether our standard model, which we have used previously, was appropriately considering all the necessary features of the shedding data to provide correct inference. We simulated shedding data under our standard, validated assumptions and assessed the ability of 5 different models to reproduce the …
Adaptive Pair-Matching In The Search Trial And Estimation Of The Intervention Effect, Laura Balzer, Maya L. Petersen, Mark J. Van Der Laan
Adaptive Pair-Matching In The Search Trial And Estimation Of The Intervention Effect, Laura Balzer, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In randomized trials, pair-matching is an intuitive design strategy to protect study validity and to potentially increase study power. In a common design, candidate units are identified, and their baseline characteristics used to create the best n/2 matched pairs. Within the resulting pairs, the intervention is randomized, and the outcomes measured at the end of follow-up. We consider this design to be adaptive, because the construction of the matched pairs depends on the baseline covariates of all candidate units. As consequence, the observed data cannot be considered as n/2 independent, identically distributed (i.i.d.) pairs of units, as current practice assumes. …
Meta-Analysis Of Social-Personality Psychological Research, Blair T. Johnson, Alice H. Eagly
Meta-Analysis Of Social-Personality Psychological Research, Blair T. Johnson, Alice H. Eagly
CHIP Documents
This publication provides a contemporary treatment of the subject of meta-analysis in relation to social-personality psychology. Meta-analysis literally refers to the statistical pooling of the results of independent studies on a given subject, although in practice it refers as well to other steps of research synthesis, including defining the question under investigation, gathering all available research reports, coding of information about the studies and their effects, and interpretation/dissemination of results. Discussed as well are the hallmarks of high-quality meta-analyses.
Estimating Effects On Rare Outcomes: Knowledge Is Power, Laura B. Balzer, Mark J. Van Der Laan
Estimating Effects On Rare Outcomes: Knowledge Is Power, Laura B. Balzer, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Many of the secondary outcomes in observational studies and randomized trials are rare. Methods for estimating causal effects and associations with rare outcomes, however, are limited, and this represents a missed opportunity for investigation. In this article, we construct a new targeted minimum loss-based estimator (TMLE) for the effect of an exposure or treatment on a rare outcome. We focus on the causal risk difference and statistical models incorporating bounds on the conditional risk of the outcome, given the exposure and covariates. By construction, the proposed estimator constrains the predicted outcomes to respect this model knowledge. Theoretically, this bounding provides …
Assessing Association For Bivariate Survival Data With Interval Sampling: A Copula Model Approach With Application To Aids Study, Hong Zhu, Mei-Cheng Wang
Assessing Association For Bivariate Survival Data With Interval Sampling: A Copula Model Approach With Application To Aids Study, Hong Zhu, Mei-Cheng Wang
Johns Hopkins University, Dept. of Biostatistics Working Papers
In disease surveillance systems or registries, bivariate survival data are typically collected under interval sampling. It refers to a situation when entry into a registry is at the time of the first failure event (e.g., HIV infection) within a calendar time interval, the time of the initiating event (e.g., birth) is retrospectively identified for all the cases in the registry, and subsequently the second failure event (e.g., death) is observed during the follow-up. Sampling bias is induced due to the selection process that the data are collected conditioning on the first failure event occurs within a time interval. Consequently, the …
A Regularization Corrected Score Method For Nonlinear Regression Models With Covariate Error, David M. Zucker, Malka Gorfine, Yi Li, Donna Spiegelman
A Regularization Corrected Score Method For Nonlinear Regression Models With Covariate Error, David M. Zucker, Malka Gorfine, Yi Li, Donna Spiegelman
Harvard University Biostatistics Working Paper Series
No abstract provided.
Variable Importance Analysis With The Multipim R Package, Stephan J. Ritter, Nicholas P. Jewell, Alan E. Hubbard
Variable Importance Analysis With The Multipim R Package, Stephan J. Ritter, Nicholas P. Jewell, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
We describe the R package multiPIM, including statistical background, functionality and user options. The package is for variable importance analysis, and is meant primarily for analyzing data from exploratory epidemiological studies, though it could certainly be applied in other areas as well. The approach taken to variable importance comes from the causal inference field, and is different from approaches taken in other R packages. By default, multiPIM uses a double robust targeted maximum likelihood estimator (TMLE) of a parameter akin to the attributable risk. Several regression methods/machine learning algorithms are available for estimating the nuisance parameters of the models, including …
Minimum Description Length And Empirical Bayes Methods Of Identifying Snps Associated With Disease, Ye Yang, David R. Bickel
Minimum Description Length And Empirical Bayes Methods Of Identifying Snps Associated With Disease, Ye Yang, David R. Bickel
COBRA Preprint Series
The goal of determining which of hundreds of thousands of SNPs are associated with disease poses one of the most challenging multiple testing problems. Using the empirical Bayes approach, the local false discovery rate (LFDR) estimated using popular semiparametric models has enjoyed success in simultaneous inference. However, the estimated LFDR can be biased because the semiparametric approach tends to overestimate the proportion of the non-associated single nucleotide polymorphisms (SNPs). One of the negative consequences is that, like conventional p-values, such LFDR estimates cannot quantify the amount of information in the data that favors the null hypothesis of no disease-association.
We …
Landmark Prediction Of Survival, Layla Parast, Tianxi Cai
Landmark Prediction Of Survival, Layla Parast, Tianxi Cai
Harvard University Biostatistics Working Paper Series
No abstract provided.
Nonparametric Regression With Missing Outcomes Using Weighted Kernel Estimating Equations, Lu Wang, Andrea Rotnitzky, Xihong Lin
Nonparametric Regression With Missing Outcomes Using Weighted Kernel Estimating Equations, Lu Wang, Andrea Rotnitzky, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
Survival Analysis With Error-Prone Time-Varying Covariates: A Risk Set Calibration Approach, Xiaomei Liao, David M. Zucker, Yi Li, Donna Spiegelman
Survival Analysis With Error-Prone Time-Varying Covariates: A Risk Set Calibration Approach, Xiaomei Liao, David M. Zucker, Yi Li, Donna Spiegelman
Harvard University Biostatistics Working Paper Series
No abstract provided.
Causal Inference For Nested Case-Control Studies Using Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan
Causal Inference For Nested Case-Control Studies Using Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
A nested case-control study is conducted within a well-defined cohort arising out of a population of interest. This design is often used in epidemiology to reduce the costs associated with collecting data on the full cohort; however, the case control sample within the cohort is a biased sample. Methods for analyzing case-control studies have largely focused on logistic regression models that provide conditional and not marginal causal estimates of the odds ratio. We previously developed a Case-Control Weighted Targeted Maximum Likelihood Estimation (TMLE) procedure for case-control study designs, which relies on the prevalence probability q0. We propose the use of …
Comparing Risk Scoring Systems Beyond The Roc Paradigm In Survival Analysis, Hajime Uno, Lu Tian, Tianxi Cai, Isaac S. Kohane, L. J. Wei
Comparing Risk Scoring Systems Beyond The Roc Paradigm In Survival Analysis, Hajime Uno, Lu Tian, Tianxi Cai, Isaac S. Kohane, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Spatial Cluster Detection For Repeatedly Measured Outcomes While Accounting For Residential History, Andrea J. Cook, Diane Gold, Yi Li
Spatial Cluster Detection For Repeatedly Measured Outcomes While Accounting For Residential History, Andrea J. Cook, Diane Gold, Yi Li
Harvard University Biostatistics Working Paper Series
No abstract provided.
Spatial Cluster Detection For Weighted Outcomes Using Cumulative Geographic Residuals, Andrea J. Cook, Yi Li, David Arterburn, Ram C. Tiwari
Spatial Cluster Detection For Weighted Outcomes Using Cumulative Geographic Residuals, Andrea J. Cook, Yi Li, David Arterburn, Ram C. Tiwari
Harvard University Biostatistics Working Paper Series
No abstract provided.
The Importance Of Scale For Spatial-Confounding Bias And Precision Of Spatial Regression Estimators, Christopher J. Paciorek
The Importance Of Scale For Spatial-Confounding Bias And Precision Of Spatial Regression Estimators, Christopher J. Paciorek
Harvard University Biostatistics Working Paper Series
Increasingly, regression models are used when residuals are spatially correlated. Prominent examples include studies in environmental epidemiology to understand the chronic health effects of pollutants. I consider the effects of residual spatial structure on the bias and precision of regression coefficients, developing a simple framework in which to understand the key issues and derive informative analytic results. When the spatial residual is induced by an unmeasured confounder, regression models with spatial random effects and closely-related models such as kriging and penalized splines are biased, even when the residual variance components are known. Analytic and simulation results show how the bias …
Analysis Of Randomized Comparative Clinical Trial Data For Personalized Treatment Selections, Tianxi Cai, Lu Tian, Peggy H. Wong, L. J. Wei
Analysis Of Randomized Comparative Clinical Trial Data For Personalized Treatment Selections, Tianxi Cai, Lu Tian, Peggy H. Wong, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Group Comparison Of Eigenvalues And Eigenvectors Of Diffusion Tensors, Armin Schwartzman, Robert F. Dougherty, Jonathan E. Taylor
Group Comparison Of Eigenvalues And Eigenvectors Of Diffusion Tensors, Armin Schwartzman, Robert F. Dougherty, Jonathan E. Taylor
Harvard University Biostatistics Working Paper Series
No abstract provided.