Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (264)
- Medicine and Health Sciences (217)
- Public Health (216)
- Epidemiology (210)
- Life Sciences (24)
-
- Biology (11)
- Public Health Education and Promotion (9)
- Exercise Science (8)
- Kinesiology (8)
- Social and Behavioral Sciences (8)
- Applied Statistics (5)
- Health Services Administration (5)
- Mathematics (5)
- Statistical Models (4)
- Arts and Humanities (3)
- Chemistry (3)
- Computer Sciences (3)
- Engineering (3)
- Social Work (3)
- Biochemistry (2)
- Biochemistry, Biophysics, and Structural Biology (2)
- Business (2)
- Categorical Data Analysis (2)
- Data Science (2)
- Design of Experiments and Sample Surveys (2)
- Pharmacy and Pharmaceutical Sciences (2)
- Sociology (2)
- Sports Studies (2)
- Keyword
-
- Dietary inflammatory index (35)
- Inflammation (27)
- Diet (14)
- Obesity (14)
- Statistics (14)
-
- Physical Sciences and Mathematics, Statistics and Probability (12)
- Bayesian (11)
- Nutrition (11)
- Colorectal cancer (10)
- Pregnancy (10)
- Risk (10)
- South Carolina (10)
- Epidemiology (9)
- Mental health (9)
- United States (9)
- Cancer (8)
- Dementia (8)
- HIV (8)
- Mortality (8)
- Diet quality (7)
- NHANES (7)
- Physical activity (7)
- Risk factors (7)
- Biomarkers (6)
- C-reactive protein (6)
- Children (6)
- Metabolic syndrome (6)
- Physical Sciences and Mathematics, Statistics (6)
- Physical Sciences and Mathematics, Statistics and Probability, Biostatistics (6)
- Semiparametric (6)
- Publication Year
- Publication
- Publication Type
Articles 331 - 360 of 376
Full-Text Articles in Statistics and Probability
Powerful Association Test Combining Rare Variant And Gene Expression Using Family Data From Genetic Analysis Workshop 19, Yen Yi Ho, Weihua Guan, Michael O'Connell, Saonli Basu
Powerful Association Test Combining Rare Variant And Gene Expression Using Family Data From Genetic Analysis Workshop 19, Yen Yi Ho, Weihua Guan, Michael O'Connell, Saonli Basu
Faculty Publications
Background: Genetic association studies aim to test for disease or trait association with genetic variants, either throughout the human genome or in regions of interest. However, for most diseases and traits, the combined effects of associated genetic variants explain only a small proportion of the genetic variation. This "missing heritability" may be a result of the small effects of common variants considered in the genetic association studies. Rare variants may also play an important role in understanding the missing heritability of complex traits. Method: We propose a novel weight-adjustment approach to combine gene expression into rare variant analysis. Results from …
Sample Size Calculation For Ph Mixture Cure Model, Yihong Zhan
Sample Size Calculation For Ph Mixture Cure Model, Yihong Zhan
Theses and Dissertations
With the development of advanced medical technology, a significant proportion of patients can be cured of many chronic diseases. Because a substantial fraction of patients have censored information, the standard survival model, such as the proportional hazards (PH) model cannot capture the cured information of patients. Thus PH mixture cure model is developed to handle the survival data with potential cured information. A corresponding sample size formula based on log rank test has been proposed by Wang et al. (2012) and the probability of death in their formula is only contributed by the control arm. However, to calculate the sample …
Parametric Reversed Hazards Model For Left Censored Data With Application To Hiv, Farahnaz Islam
Parametric Reversed Hazards Model For Left Censored Data With Application To Hiv, Farahnaz Islam
Theses and Dissertations
Left censoring is generally a rare type of censoring in time-to-event data, however there are some fields such as HIV related studies where it commonly occurs. Currently, there is no clear recommendation in the literature on the optimal model and distribution to analyze left-censored data. Recommendations can help researchers apply more accurate models for this type of censoring. This study derives the Parametric Reversed Hazards (PRH) Model for a variety of distributions which may be appropriate for left censored data. The performance of these derived PRH models to analyze HIV viral load data are compared using extensive simulations and a …
Registration And Clustering Of Functional Observations, Zizhen Wu
Registration And Clustering Of Functional Observations, Zizhen Wu
Theses and Dissertations
As an important exploratory analysis, curves of similar shape are often classified into groups, which we call clustering of functional data. Phase variations or time distortions are often encountered in the biological processes, such as growth patterns or gene profiles. As a result of time distortion, curves of similar shape may not be aligned. Regular clustering methods for functional data usually ignore the presence of phase variations, which may result in low clustering accuracy. However, it is difficult to account for phase variation without knowing the cluster structure.
In this dissertation, we first propose a Bayesian method that simultaneously clusters …
Modern Estimation Problems In Group Testing, Md Shamim Sarker
Modern Estimation Problems In Group Testing, Md Shamim Sarker
Theses and Dissertations
In the simplest form of group testing, pools are formed by compositing a fixed number of individual specimens (e.g., blood, urine, swab, etc.) and then the pools are tested for a binary characteristic, such as presence or absence of a disease. Group testing is commonly used to screen for a variety of sexually transmitted diseases in epidemiological applications where the main goal is to increase testing efficiency. In this dissertation, we study three estimation problems that are motivated by real-life applications. We propose new methods to model group testing data for both single and multiple infections. In the first problem, …
Semiparametric Joint Dynamic Modeling Of A Longitudinal Marker, Recurrent Competing Risks, And A Terminal Event, Piaomu Liu
Theses and Dissertations
The joint modeling framework has found extensive applications in cancer and other biomedical research. For example, recent initiatives and developments in precision medicine call for appropriate prognostic tools to assist individualized or personalized approaches in cancer diagnosis and treatment. Data generated by clinical trials and medical research often include correlated longitudinal marker measurements and time- to-event information, which are possibly a recurrent event, competing risks, and a survival outcome. Primary interests of joint modeling include the association between the longitudinal marker measurements and time-to-event data, as well as predictions of survival probabilities of new observational units from the same population. …
Spatio-Temporal Analysis Of The Occupational Fatal Victimization Of Law Enforcement Officers In The Us, Xueyi Xing
Spatio-Temporal Analysis Of The Occupational Fatal Victimization Of Law Enforcement Officers In The Us, Xueyi Xing
Theses and Dissertations
The models with constant coefficients of the covariates across space and time are commonly used in spatio-temporal analyses. However, the associations between risk factors and the outcome could have locally differential temporal trends in many cases. In this study, a Bayesian latent cluster modeling strategy is employed to identify potential spatial clusters in which locally specific sets of temporally varying coefficients of covariates are allowed. A state-level panel data of police officers occupational fatal victimization for the years 1979-2010 is used. To accommodate overdisperson and excess zeros, a negative binomial model and zero-inflated Poisson/negative binomial models are also utilized. A …
Regression Models For Count Data Based On The Double Poisson Distribution, Rebecca Wardrop
Regression Models For Count Data Based On The Double Poisson Distribution, Rebecca Wardrop
Theses and Dissertations
This paper explores the double Poisson distribution. The probability mass function and the difficulties associated with derivative-based optimization for this distribution are discussed. Stata software developed for estimation of double Poisson regression is detailed. Simulations are used to test the software. Data which are over-, under-, and equidispersed relative to the Poisson are generated and the software is utilized to estimate a regression model, a zero-inflated model, and a marginalized zero-inflated model all based on the double Poisson distribution. The estimated power of the test for φ = 1 for the double Poisson models are compared to the power of …
Score Test Derivations And Implementations For Bivariate Probability Mass And Density Functions With An Application To Copula Functions, Roy Bower
Theses and Dissertations
This dissertation is comprised and grounded in statistical theory with an application to solving real world problems. In particular, the development and implementation of multiple score tests under a variety of scenarios are derived, applied, and interpreted. In chapter 2, I propose a score test for independence of the marginals based on Lakshminarayana’s bivariate Poisson distribution. Each marginal distribution of the bivariate model is a univariate Poisson distribution, and the parameters of the bivariate distribution can be estimated using maximum likelihood methods. The simulation study shows that the score test maintains size close to the nominal level. To assess the …
Some Issues In Markov Chain Monte Carlo Estimation For Item Response Theory, Han Kil Lee
Some Issues In Markov Chain Monte Carlo Estimation For Item Response Theory, Han Kil Lee
Theses and Dissertations
Both the marginalized Bayesian modal estimation (MBME) and Metropolis-Hasting within Gibbs (MH/Gibbs) are the popular estimation methods for Item Response Theory (IRT). However, predictions from MBME and MH/Gibbs are not directly comparable because of two problems. First, the examinees with the same response pattern do not produce the same ability estimates from MH/Gibbs while MBME provides identical estimates. This problem can be handled by updating each response pattern instead of updating each examinee. Second, standard errors from MBME are smaller than standard error estimates from MH/Gibbs. This pattern occurs because of two speculated reasons; correlation between item parameter estimations and …
Semiparametric Estimation Methods For Complex Accelerated Failure Time Model, Yinding Wang
Semiparametric Estimation Methods For Complex Accelerated Failure Time Model, Yinding Wang
Theses and Dissertations
The proportional hazards (PH) model and the accelerated failure time (AFT) model are the two most popular survival models in fitting the right-censored data. The AFT model is a useful alternative to the PH model, particularly when the PH assumption is not satisfied. Usually, the linear association is assumed with logarithm of survival time in the AFT model. However, the nonlinear association may exist in practice. The first project aims to handle the nonlinear component in the AFT model, which is called the semiparametric additive partial accelerated failure time (AP-AFT) model. Two estimation methods based on the rank-smooth method and …
Bayesian Ensemble Of Regression Trees For Multinomial Probit And Quantile Regression, Bereket P. Kindo
Bayesian Ensemble Of Regression Trees For Multinomial Probit And Quantile Regression, Bereket P. Kindo
Theses and Dissertations
This dissertation proposes multinomial probit Bayesian additive regression trees (MPBART), ordered multiclass Bayesian additive classification trees (O-MBACT) and Bayesian quantile additive regression trees (BayesQArt) as extensions of BART - Bayesian additive regression trees for tackling multinomial choice, multiclass classification, ordinal regression and quantile regression problems. The proposed models exhibit very good predictive performances. In particular, ranking among the top performing procedures when non-linear relationships exist between the response and the predictors. The proposed procedures can readily be applied on data sets with the number of predictors larger than the number of observations.
MPBART is sufficiently flexible to allow inclusion of …
Frailty Probit Models For Clustered Interval-Censored Failure Time Data, Haifeng Wu
Frailty Probit Models For Clustered Interval-Censored Failure Time Data, Haifeng Wu
Theses and Dissertations
Survival analysis is an important branch of statistics that deals with time to event data or survival data. An important feature of such data is that the survival time of interest is usually not completely known but is censored due to the design of the study or an early dropout. In this dissertation we focus on studying clustered interval-censored data, a special type of survival data. Interval-censored data arise in many epidemiological, social science, and medical studies, in which subjects are examined at periodical follow-up visits. The survival (or failure) time of interest is never exactly observed but is known …
Association Between Nutritional Awareness And Diet Quality: Evidence From The Observation Of Cardiovascular Risk Factors In Luxembourg (Oriscav-Lux) Study, Ala'a Alkerwi, Nicolas Sauvageot, Leoné Malan, Nitin Shivappa, James R. Hébert
Association Between Nutritional Awareness And Diet Quality: Evidence From The Observation Of Cardiovascular Risk Factors In Luxembourg (Oriscav-Lux) Study, Ala'a Alkerwi, Nicolas Sauvageot, Leoné Malan, Nitin Shivappa, James R. Hébert
Faculty Publications
This study examined the association between nutritional awareness and diet quality, as indicated by energy density, dietary diversity and adequacy to achieve dietary recommendations, while considering the potentially important role of socioeconomic status (SES). Data were derived from 1351 subjects, aged 18–69 years and enrolled in the ORISCAV-LUX study. Energy density score (EDS), dietary diversity score (DDS) and Recommendation Compliance Index (RCI) were calculated based on data derived from a food frequency questionnaire. Nutritional awareness was defined as self-perception of the importance assigned to eating balanced meals, and classified as high, moderate, or of little importance. Initially, a General Linear …
The Candidate Cancer Gene Database: A Database Of Cancer Driver Genes From Forward Genetic Screens In Mice, Kenneth L. Abbott, Erik T. Nyre, Juan Abrahante, Yen Yi Ho, Rachel Isaksson Vogel, Timothy K. Starr
The Candidate Cancer Gene Database: A Database Of Cancer Driver Genes From Forward Genetic Screens In Mice, Kenneth L. Abbott, Erik T. Nyre, Juan Abrahante, Yen Yi Ho, Rachel Isaksson Vogel, Timothy K. Starr
Faculty Publications
Identification of cancer driver gene mutations is crucial for advancing cancer therapeutics. Due to the overwhelming number of passenger mutations in the human tumor genome, it is difficult to pinpoint causative driver genes. Using transposon mutagenesis in mice many laboratories have conducted forward genetic screens and identified thousands of candidate driver genes that are highly relevant to human cancer. Unfortunately, this information is difficult to access and utilize because it is scattered across multiple publications using different mouse genome builds and strength metrics. To improve access to these findings and facilitate metaanalyses, we developed the Candidate Cancer Gene Database (CCGD, …
Bayesian Semi- And Non-Parametric Analysis For Spatially Correlated Survival Data, Haiming Zhou
Bayesian Semi- And Non-Parametric Analysis For Spatially Correlated Survival Data, Haiming Zhou
Theses and Dissertations
Flexible incorporation of both geographical patterning and risk effects in cancer survival models is becoming increasingly important, due in part to the recent availability of large cancer registries. The analysis of spatial survival data is challenged by the presence of spatial dependence and censoring for survival times. Accurately modeling the risk factors and geographical pattern that explain the differences in survival is particularly of interest. Within this dissertation, the first chapter reviews commonlyused baseline priors, semiparametric and nonparametric Bayesian survival models and recent approaches for accommodating spatial dependence, both conditional and marginal. The last three chapters contribute three flexible survival …
Methods For Identifying Regions Of Brain Activation Using Fmri Meta-Data, Meredith A. Ray
Methods For Identifying Regions Of Brain Activation Using Fmri Meta-Data, Meredith A. Ray
Theses and Dissertations
Functional neuroimaging is a relatively young discipline within the neurosciences that has led to significant advances in our understanding of the human brain and progress in neuroscientific research related to public health. Accurately identifying activated regions in the brain showing a strong association with an outcome of interest is crucial in terms of disease prediction and prevention. Functional magnetic resonance imaging (fMRI) is the most widely used method for this type of study as it has the ability to measure and identify the location of changes in tissue perfusion, blood oxygenation, and blood volume. In practice, the three-dimensional brain locations …
Semiparametric Regression Analysis Of Bivariate Interval-Censored Data, Naichen Wang
Semiparametric Regression Analysis Of Bivariate Interval-Censored Data, Naichen Wang
Theses and Dissertations
Survival analysis is a long-lasting and popular research area and has numerous applications in all fields such as social science, engineering, economics, industry, and public health. Interval-censored data are a special type of survival data, in which the survival time of interest is never exactly observed but is known to fall within some observed interval. Interval-censored data arise commonly in real-life studies, in which subjects are examined at periodical or irregular follow-up visits. In this dissertation, we develop efficient statistical approaches for regression analysis of bivariate intervalcensored data, in which the two survival times of interest are correlated and both …
Simulation Based Evaluation Of Multiscale Small Area Health Models, Purbasha Dasgupta
Simulation Based Evaluation Of Multiscale Small Area Health Models, Purbasha Dasgupta
Theses and Dissertations
The effects of scale on the analysis of spatial data, often referred to as the modifiable areal unit problem in spatial studies, is one of the issues often encountered in small area health models. These spatial effects of scale are also seen in the areas of disease mapping where data are usually available in counts. Often there is a need to consider the different scales of aggregation that exist within count data, since inferences based on analyses can vary if we change the definition of the unit of analysis. This thesis provides a framework that describes the distribution of relative …
An Efficient Algorithm To Explore Liquid Association On A Genome-Wide Scale, Tina Gunderson, Yen Yi Ho
An Efficient Algorithm To Explore Liquid Association On A Genome-Wide Scale, Tina Gunderson, Yen Yi Ho
Faculty Publications
Background: The growing wealth of public available gene expression data has made the systemic studies of how genes interact in a cell become more feasible. Liquid association (LA) describes the extent to which coexpression of two genes may vary based on the expression level of a third gene (the controller gene). However, genome-wide application has been difficult and resource-intensive. We propose a new screening algorithm for more efficient processing of LA estimation on a genome-wide scale and apply its use to a data set. Results: On a test subset of the data, the fast screening algorithm achieved >99.8 agreement with …
Non- And Semi-Parametric Bayesian Inference With Recurrent Events And Coherent Systems Data, A. K. M. Fazlur Rahman
Non- And Semi-Parametric Bayesian Inference With Recurrent Events And Coherent Systems Data, A. K. M. Fazlur Rahman
Theses and Dissertations
This dissertation deals with non- and semi-parametric Bayesian inference of gap-time distribution with recurrent event data and simultaneous inference of component and system reliabilities of coherent systems data. Recurrent event data arise from a wide variety of studies/fields such as clinical trials, epidemiology, public health, biomedicine (e.g. repeated heart attack, repeated tumor occurrences of a cancer patient). In Chapter 2 we develop nonparametric Bayes and empirical Bayes estimators of the survivor function \bar{F} = 1 - F, of the gap-time distribution by assigning a Dirichlet process prior on F. We develop a closed form estimator of \bar{F} as well as …
Ranking World Class Chess Players Using Only Results From Head-To-Head Games, Sterling Swygert
Ranking World Class Chess Players Using Only Results From Head-To-Head Games, Sterling Swygert
Senior Theses
This honors thesis explores a method of ranking the world’s top ten chess grand- masters using only the outcomes of games containing only players in that very set. This method allows for players in a single era to be quickly ranked via algorithmic and numerical means, including very specific information, from a statistical stand- point. Furthermore, unlike the rating systems that are commonly used, the Elo and the Glicko systems, this method is Classicist in its statistical approach, rather than Bayesian. Finally, this ranking method also differs from others as it limits the infor- mation to games between the individuals …
Socioeconomic And Demographic Factors Modify The Association Between Informal Caregiving And Health In The Sandwich Generation, Elizabeth K. Do, Steven A. Cohen, Monique J. Brown Ph.D., Mph
Socioeconomic And Demographic Factors Modify The Association Between Informal Caregiving And Health In The Sandwich Generation, Elizabeth K. Do, Steven A. Cohen, Monique J. Brown Ph.D., Mph
Faculty Publications
Background
Nearly 50 million Americans provide informal care to an older relative or friend. Many are members of the “sandwich generation”, providing care for elderly parents and children simultaneously. Although evidence suggests that the negative health consequences of caregiving are more severe for sandwiched caregivers, little is known about how these associations vary by sociodemographic factors.
Methods
We abstracted data from the Behavioral Risk Factor Surveillance System to determine how the association between caregiving and health varies by sociodemographic factors, using ordinal logistic regression with interaction terms and stratification by number of children, income, and race/ethnicity.
Results
The association between …
Oldtimers & Newcomers In Collective Action, Stefanie R. Chamberlain
Oldtimers & Newcomers In Collective Action, Stefanie R. Chamberlain
Theses and Dissertations
Most work on groups facing collective action assumes that group membership is static, or fixed. Yet static membership is rare, with members joining and leaving groups. In this thesis, I propose to explore how the presence of newcomers to groups affects group coordination. Past research has shown an overall negative effect of newcomers on group contributions. The proposed thesis attempts to further establish the effect by determining whether newcomers, oldtimers, or both are responsible for the declining cooperation in groups. While the empirical component is focused solely on establishing who is responsible for driving down cooperation rates in dynamic groups, …
Modular Network Construction Using Eqtl Data: An Analysis Of Computational Costs And Benefits, Yen Yi Ho, Leslie M. Cope, Giovanni Parmigiani
Modular Network Construction Using Eqtl Data: An Analysis Of Computational Costs And Benefits, Yen Yi Ho, Leslie M. Cope, Giovanni Parmigiani
Faculty Publications
Background: In this paper, we consider analytic methods for the integrated analysis of genomic DNA variation and mRNA expression (also named as eQTL data), to discover genetic networks that are associated with a complex trait of interest. Our focus is the systematic evaluation of the trade-off between network size and network search efficiency in the construction of these networks. Results: We developed a modular approach to network construction, building from smaller networks to larger ones, thereby reducing the search space while including more variables in the analysis. The goal is achieving a lower computational cost while maintaining high confidence in …
Applications Of Bayesian Nonparametrics To Reliability And Survival Data, Li Li
Applications Of Bayesian Nonparametrics To Reliability And Survival Data, Li Li
Theses and Dissertations
Reliability and survival data are widely encountered across many common settings. Subjects under investigation often include machines, bioassays, patients, etc.; their reliability or survival distribution, and its association with covariate processes, are commonly of interest. Within this dissertation, the first two chapters focus on reliability data where repairable systems fail and get interventions, e.g. repairs in the event process. It begins with a nonparametric test for the commonly assumed ''good as old'' assumption for minimal repair models and then a semi-parametric regression model is introduced for reliability data using Kijima's effective age. The third chapter focuses on survival data observed …
Methods For Clustering Mixed Data, Jeanmarie L. Hendrickson
Methods For Clustering Mixed Data, Jeanmarie L. Hendrickson
Theses and Dissertations
We give a brief introduction to cluster analysis and then propose and discuss a few methods for clustering mixed data. In particular, a model-based clustering method for mixed data based on Everitt's (1988) work is described, and we use a simulated annealing method to estimate the parameters for Everitt's model. A penalized log likelihood with the simulated annealing method is proposed as a remedy for the parameter estimates being drawn to extremes. Everitt's approach and the proposed method are compared based on their performance in clustering simulated data. We then use the penalized log likelihood method on a heart disease …
Bayesian Analysis Of Continuous Curve Functions, Wen Cheng
Bayesian Analysis Of Continuous Curve Functions, Wen Cheng
Theses and Dissertations
We consider Bayesian analysis of continuous curve functions in 1D, 2D and 3D spaces. A fundamental feature of the analysis is that it is invariant under a simultaneous warping/re-parameterization of all target curves, as well as translation, rotation and scale of each individual if necessary. We introduce Bayesian models based on a special curve representation named Square Root Velocity Function (SRVF) introduced by Srivastava et al. (2011, IEEE PAMI). A Gaussian process model for the SRVFs of curves is proposed, and suitable prior models such as the Dirichlet distribution are employed for modeling the warping function as a cumulative distribution …
Association Between Adverse Childhood Experiences And Diagnosis Of Cancer, Monique J. Brown, Leroy R. Thacker, Steven A. Cohen
Association Between Adverse Childhood Experiences And Diagnosis Of Cancer, Monique J. Brown, Leroy R. Thacker, Steven A. Cohen
Faculty Publications
Objective: Adverse childhood experiences (ACEs) are linked to multiple adverse health outcomes. This study examined the association between ACEs and cancer diagnosis.
Methods: Data from the 2010 Behavioral Risk Factor Surveillance System (BRFSS) survey were used. The BRFSS is the largest ongoing telephone health survey, conducted in all US states, the District of Columbia, Puerto Rico, Guam and the U.S. Virgin Islands, and provides data on a variety of health issues among the non-institutionalized adult population. Principal component analysis (PCA) was used to derive components for ACEs. Multivariable logistic regression models were used to provide adjusted odds ratios (OR) and …
Estimation And Q-Matrix Validation For Diagnostic Classification Models, Yuling Feng
Estimation And Q-Matrix Validation For Diagnostic Classification Models, Yuling Feng
Theses and Dissertations
Diagnostic classification models (DCMs) are structured latent class models widely discussed in the field of psychometrics. They model subjects' underlying attribute patterns and classify subjects into unobservable groups based on their mastery of attributes required to answer the items correctly. The effective implementation of DCMs depends on correct specification of a Q-matrix which is a binary matrix linking attribute patterns to items. Current literature on assessing the appropriateness of Q-matrix specifications has focused on validation methods for the deterministic-input, noisy-and-gate (DINA) model. The goal of the study is to develop general Q-matrix validation methods that can be applied to a …