Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (16)
- Life Sciences (15)
- Statistical Models (15)
- Bioinformatics (11)
- Statistical Methodology (11)
-
- Medicine and Health Sciences (7)
- Multivariate Analysis (6)
- Statistical Theory (5)
- Survival Analysis (5)
- Biochemistry (4)
- Biochemistry, Biophysics, and Structural Biology (4)
- Genetics and Genomics (4)
- Longitudinal Data Analysis and Time Series (4)
- Microarrays (4)
- Categorical Data Analysis (3)
- Chemistry (3)
- Clinical Trials (3)
- Computational Biology (3)
- Computer Sciences (3)
- Data Science (3)
- Mathematics (3)
- Physics (3)
- Public Health (3)
- Vital and Health Statistics (3)
- Analytical Chemistry (2)
- Applied Mathematics (2)
- Biological and Chemical Physics (2)
- Institution
- Keyword
-
- Causal inference (8)
- Bayesian (4)
- Propensity score (4)
- Average treatment effect (3)
- Observational studies (3)
-
- Propensity scores (3)
- Statistics (3)
- ATE (2)
- Bioinformatics (2)
- Breast cancer (2)
- Differential expression (2)
- Dimension Reduction (2)
- Factor model (2)
- Features Extraction (2)
- Gene expression (2)
- Machine learning (2)
- Mass spectrometry (2)
- Mathematical statistics (2)
- Metabolomics (2)
- Observational data (2)
- Ordinal outcome (2)
- RNA-seq (2)
- Regression analysis (2)
- ScRNA-seq (2)
- Sensitivity (2)
- Shrinkage (2)
- Simulation (2)
- Sojourn time (2)
- Survival Analysis (2)
- Variable selection (2)
Articles 31 - 60 of 61
Full-Text Articles in Biostatistics
Marginal Methods And Software For Clustered Data With Cluster- And Group-Size Informativeness., Mary Elizabeth Gregg
Marginal Methods And Software For Clustered Data With Cluster- And Group-Size Informativeness., Mary Elizabeth Gregg
Electronic Theses and Dissertations
Clustered data result when observations have some natural organizational association. In such data, cluster size is defined as the number of observations belonging to a cluster. A phenomenon termed informative cluster size (ICS) occurs when observation outcomes vary in a systematic way related to the cluster size. An additional form of informativeness, termed informative within-cluster group size (IWCGS), arises when the distribution of group-defining categorical covariates within clusters similarly carries information related to outcomes. Standard methods for the marginal analysis of clustered data can produce biased estimates and inference when data have informativeness. A reweighting methodology has been developed that …
Novel Bayesian Methodology For The Analysis Of Single-Cell Rna Sequencing Data., Michael Sekula
Novel Bayesian Methodology For The Analysis Of Single-Cell Rna Sequencing Data., Michael Sekula
Electronic Theses and Dissertations
With single-cell RNA sequencing (scRNA-seq) technology, researchers are able to gain a better understanding of health and disease through the analysis of gene expression data at the cellular-level; however, scRNA-seq data tend to have high proportions of zero values, increased cell-to-cell variability, and overdispersion due to abnormally large expression counts, which create new statistical problems that need to be addressed. This dissertation includes three research projects that propose Bayesian methodology suitable for scRNA-seq analysis. In the first project, a hurdle model for identifying differentially expressed genes across cell types in scRNA-seq data is presented. This model incorporates a correlated random …
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Electronic Theses and Dissertations
Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …
Statistical Methods For Estimating And Testing Treatment Effect For Multiple Treatment Groups In Observational Studies., Xiaofang Yan
Statistical Methods For Estimating And Testing Treatment Effect For Multiple Treatment Groups In Observational Studies., Xiaofang Yan
Electronic Theses and Dissertations
Note: Abstract would not save due to an issue with some of the characters.
Identifying Risk Factors Related To Premature Birth Through Binary Logistic And Proportional Odds Ordinal Logistic Regression, Clayton Elwood
Identifying Risk Factors Related To Premature Birth Through Binary Logistic And Proportional Odds Ordinal Logistic Regression, Clayton Elwood
Electronic Theses and Dissertations
Premature birth has been identified as the single greatest cause of death worldwide in children under the age of five. This thesis will implement binary logistic regression and proportional odds ordinal logistic regression to predict different levels of premature birth and identify associated risk factors. The models will be built from the Center for Disease Control and Prevention's 2014 Vital Statistics Natality Birth Data containing nearly 4 million live births within the United States. Odds ratios and confidence intervals on risk factors were produced utilizing binary logistic regression.
Novel Bayesian Methodology In Multivariate Problems., Debamita Kundu
Novel Bayesian Methodology In Multivariate Problems., Debamita Kundu
Electronic Theses and Dissertations
This dissertation involves developing novel Bayesian methodology for multivariate problems. In particular, it focuses on two contexts: shrinkage based variable selection in multivariate regression and simultaneous covariance estimation of multiple groups. Both these projects are centered around fully Bayesian inference schemes based on hierarchical modeling to capture context-specific features of the data and the development of computationally efficient estimation algorithm. Variable selection over a potentially large set of covariates in a linear model is quite popular. In the Bayesian context, common prior choices can lead to a posterior expectation of the regression coefficients that is a sparse (or nearly sparse) …
Innate Immunity, The Hepatic Extracellular Matrix, And Liver Injury: Mathematical Modeling Of Metastatic Potential And Tumor Development In Alcoholic Liver Disease., Shanice V. Hudson
Innate Immunity, The Hepatic Extracellular Matrix, And Liver Injury: Mathematical Modeling Of Metastatic Potential And Tumor Development In Alcoholic Liver Disease., Shanice V. Hudson
Electronic Theses and Dissertations
The overarching goals of the current work are to fill key gaps in the current understanding of alcohol consumption and the risk of metastasis to the liver. Considering the evidence this research group has compiled confirming that the hepatic matrisome responds dynamically to injury, an altered extracellular matrix (ECM) profile appears to be a key feature of pre-fibrotic inflammatory injury in the liver. This group has demonstrated that the hepatic ECM responds dynamically to alcohol exposure, in particular, sensitizing the liver to LPS-induced inflammatory damage. Although the study of alcohol in its role as a contributing factor to oncogenesis and …
Bayesian Analytical Approaches For Metabolomics : A Novel Method For Molecular Structure-Informed Metabolite Interaction Modeling, A Novel Diagnostic Model For Differentiating Myocardial Infarction Type, And Approaches For Compound Identification Given Mass Spectrometry Data., Patrick J. Trainor
Electronic Theses and Dissertations
Metabolomics, the study of small molecules in biological systems, has enjoyed great success in enabling researchers to examine disease-associated metabolic dysregulation and has been utilized for the discovery biomarkers of disease and phenotypic states. In spite of recent technological advances in the analytical platforms utilized in metabolomics and the proliferation of tools for the analysis of metabolomics data, significant challenges in metabolomics data analyses remain. In this dissertation, we present three of these challenges and Bayesian methodological solutions for each. In the first part we develop a new methodology to serve a basis for making higher order inferences in metabolomics, …
Generalized Spatiotemporal Modeling And Causal Inference For Assessing Treatment Effects For Multiple Groups For Ordinal Outcome., Soutik Ghosal
Generalized Spatiotemporal Modeling And Causal Inference For Assessing Treatment Effects For Multiple Groups For Ordinal Outcome., Soutik Ghosal
Electronic Theses and Dissertations
This dissertation consists of three projects and can be categorized in two broad research areas: generalized spatiotemporal modeling and causal inference based on observational data. In the first project, I introduce a Bayesian hierarchical mixed effect hurdle model with a nested random effect structure to model the count for primary care providers and understand their spatial and temporal variation. This study further enables us to identify the health professional shortage areas and the possible impacting factors. In the second project, I have unified popular parametric and nonparametric propensity score-based methods to assess the treatment effect of multiple groups for ordinal …
Improving The Detection Limit Of Tau Aggregates For Use With Biological Samples, Emily Rickman Hager
Improving The Detection Limit Of Tau Aggregates For Use With Biological Samples, Emily Rickman Hager
Electronic Theses and Dissertations
The protein Tau is found in neurofibrillary tangles in Alzheimer's disease and over 20 other neurodegenerative diseases. An assay has been developed to detect minute amounts of fibrils from human brain tissue. This assay subjects brain tissue extract and recombinant Tau to several rounds of sonication and incubation. Incubation allows recombinant Tau to add itself to the ends of the existing fibrils in brain tissue extract. Sonication breaks the existing fibrils in the brain tissue extract offering more ends for Tau to add onto. Cycles of sonication and incubation have been shown to allow for amplification of Tau fibrils from …
Functional Data Analysis Methods For Predicting Disease Status., Sarah Kendrick
Functional Data Analysis Methods For Predicting Disease Status., Sarah Kendrick
Electronic Theses and Dissertations
Introduction: Differential scanning calorimetry (DSC) is used to determine thermally-induced conformational changes of biomolecules within a blood plasma sample. Recent research has indicated that DSC curves (or thermograms) may have different characteristics based on disease status and, thus, may be useful as a monitoring and diagnostic tool for some diseases. Since thermograms are curves measured over a range of temperature values, they are often considered as functional data. In this dissertation we propose and apply functional data analysis (FDA) techniques to analyze DSC data from the Lupus Family Registry and Repository (LFRR). The aim is to develop FDA methods to …
Sample Size Calculations And Normalization Methods For Rna-Seq Data., Xiaohong Li
Sample Size Calculations And Normalization Methods For Rna-Seq Data., Xiaohong Li
Electronic Theses and Dissertations
High-throughput RNA sequencing (RNA-seq) has become the preferred choice for transcriptomics and gene expression studies. With the rapid growth of RNA-seq applications, sample size calculation methods for RNA-seq experiment design and data normalization methods for DEG analysis are important issues to be explored and discussed. The underlying theme of this dissertation is to develop novel sample size calculation methods in RNA-seq experiment design using test statistics. I have also proposed two novel normalization methods for analysis of RNA-seq data. In chapter one, I present the test statistical methods including Wald’s test, log-transformed Wald’s test and likelihood ratio test statistics for …
Bayesian Approach On Short Time-Course Data Of Protein Phosphorylation, Casual Inference For Ordinal Outcome And Causal Analysis Of Dietary And Physical Activity In T2dm Using Nhanes Data., You Wu
Electronic Theses and Dissertations
This dissertation contains three different projects in proteomics and causal inferences. In the first project, I apply a Bayesian hierarchical model to assess the stability of phosphorylated proteins under short-time cold ischemia. This study provides inference on the stability of these phosphorylated proteins, which is valuable when using these proteins as biomarkers for a disease. in the second project, I perform a comparative study of different confounding-adjusted to estimate the treatment effect when the outcome variable is ordinal using observational data. The adjusted U-statistics method is compared with other methods such as ordinal logistic regression, propensity score based stratification and …
Likelihood-Based Methods For Analysis Of Copy Number Variation Using Next Generation Sequencing Data., Udika Iroshini Bandara
Likelihood-Based Methods For Analysis Of Copy Number Variation Using Next Generation Sequencing Data., Udika Iroshini Bandara
Electronic Theses and Dissertations
A Copy Number Variation (CNV) detection problem is considered using Circular Binary Segmentation (CBS) procedures, including newly developed procedures based on likelihood ratio tests with the parametric bootstrap for models based on discrete distributions for count data (Poisson and negative binomial) and a widely-used DNAcopy package. Results from the literature concerning maximum likelihood estimation for the negative binomial distribution are reviewed. The Newton-Raphson method is used to find the root of the derivative of the profile log likelihood function when applicable, and it is proven that this method converges to the true Maximum Likeihood Estimate (MLE), if the starting point …
Estimation Of The Three Key Parameters And The Lead Time Distribution In Lung Cancer Screening., Ruiqi Liu
Estimation Of The Three Key Parameters And The Lead Time Distribution In Lung Cancer Screening., Ruiqi Liu
Electronic Theses and Dissertations
This dissertation contains three research projects on cancer screening probability modeling. Cancer screening is the primary technique for early detection. The goal of screening is to catch the disease early before clinical symptoms appear. In these projects, the three key parameters and lead time distribution were estimated to provide a statistical point of view on the effectiveness of cancer screening programs. In the first project, cancer screening probability model was used to analyze the computed tomography (CT) scan group in the National Lung Screening Trial (NLST) data. Three key parameters were estimated using Bayesian approach and Markov Chain Monte Carlo …
Multilevel Models For Longitudinal Data, Aastha Khatiwada
Multilevel Models For Longitudinal Data, Aastha Khatiwada
Electronic Theses and Dissertations
Longitudinal data arise when individuals are measured several times during an ob- servation period and thus the data for each individual are not independent. There are several ways of analyzing longitudinal data when different treatments are com- pared. Multilevel models are used to analyze data that are clustered in some way. In this work, multilevel models are used to analyze longitudinal data from a case study. Results from other more commonly used methods are compared to multilevel models. Also, comparison in output between two software, SAS and R, is done. Finally a method consisting of fitting individual models for each …
Propensity Score Based Methods For Estimating The Treatment Effects Based On Observational Studies., Younathan Abdia
Propensity Score Based Methods For Estimating The Treatment Effects Based On Observational Studies., Younathan Abdia
Electronic Theses and Dissertations
This dissertation consists of two interconnected research projects. The first project was a study of propensity scores based statistical methods for estimating the average treatment effect (ATE) and the average treatment effect among treated (ATT) when there are two treatment groups. The ATE is defined as the mean of the individual causal effects in the whole population, while ATT is defined as the treatment effect for the treated population. Propensity score based statistical methods, such as matching, regression, stratification, inverse probability weighting (IPW), and doubly robust (DR) methods were used to estimate the ATE and ATT. Simulation studies and case …
Some Contributions To Nonparametric And Semiparametric Inference For Clustered And Multistate Data., Sandipan Dutta
Some Contributions To Nonparametric And Semiparametric Inference For Clustered And Multistate Data., Sandipan Dutta
Electronic Theses and Dissertations
This dissertation is composed of research projects that involve methods which can be broadly classified as either nonparametric or semiparametric. Chapter 1 provides an introduction of the problems addressed in these projects, a brief review of the related works that have done so far, and an outline of the methods developed in this dissertation. Chapter 2 describes in details the first project which aims at developing a rank-sum test for clustered data where an outcome from group in a cluster is associated with the number of observations belonging to that group in that cluster. Chapter 3 proposes the use of …
A Log Rank Test For Clustered Data Under Informative Within-Cluster Group Size., Mary Elizabeth Gregg
A Log Rank Test For Clustered Data Under Informative Within-Cluster Group Size., Mary Elizabeth Gregg
Electronic Theses and Dissertations
The log rank test is a popular nonparametric test for comparing the marginal survival distribution of two groups. When data are organized within clusters and the size of clusters or the distribution of group membership within a cluster is related to an outcome of interest, traditional methods of data analysis can be biased. In this thesis, we develop a within-cluster group weighted log rank test to compare marginal survival time distributions between groups from clustered data, correcting for cluster size and intra-cluster group size informativeness. The performance of this new test is compared with the unweighted and cluster-weighted log rank …
Propensity Score Methods : A Simulation And Case Study Involving Breast Cancer Patients., John Craycroft
Propensity Score Methods : A Simulation And Case Study Involving Breast Cancer Patients., John Craycroft
Electronic Theses and Dissertations
Observational data presents unique challenges for analysis that are not encountered with experimental data resulting from carefully designed randomized controlled trials. Selection bias and unbalanced treatment assignments can obscure estimations of treatment effects, making the process of causal inference from observational data highly problematic. In 1983, Paul Rosenbaum and Donald Rubin formalized an approach for analyzing observational data that adjusts treatment effect estimates for the set of non-treatment variables that are measured at baseline. The propensity score is the conditional probability of assignment to a treatment group given the covariates. Using this score, one may balance the covariates across treatment …
Semi-Parametric Methods For Personalized Treatment Selection And Multi-State Models., Chathura K. Siriwardhana
Semi-Parametric Methods For Personalized Treatment Selection And Multi-State Models., Chathura K. Siriwardhana
Electronic Theses and Dissertations
This dissertation contains three research projects on personalized medicine and a project on multi-state modelling. The idea behind personalized medicine is selecting the best treatment that maximizes interested clinical outcomes of an individual based on his or her genetic and genomic information. We propose a method for treatment assignment based on individual covariate information for a patient. Our method covers more than two treatments and it can be applied with a broad set of models and it has very desirable large sample properties. An empirical study using simulations and a real data analysis show the applicability of the proposed procedure. …
Inference For A Zero-Inflated Conway-Maxwell-Poisson Regression For Clustered Count Data., Hyoyoung Choo-Wosoba
Inference For A Zero-Inflated Conway-Maxwell-Poisson Regression For Clustered Count Data., Hyoyoung Choo-Wosoba
Electronic Theses and Dissertations
This dissertation is directed toward developing a statistical methodology with applications of the Conway-Maxwell-Poisson (CMP) distribution (Conway, R. W., and Maxwell, W. L., 1962) to count data. The count data for this dissertation exhibit three different characteristics: clustering, zero inflation, and dispersion. Clustering suggests that observations within clusters are correlated, and the zero inflation phenomenon occurs when the data exhibit excessive zero counts. Dispersion implies that the mean is greater/smaller than the variance unlike a Poisson distribution. The dissertation starts with an introduction of inference for a zero-inflated clustered count data in the first chapter. Then, it presents novel methodologies …
Novel Applications Of And Extensions To Linear Regression Methods For The Biomedical And Materials Sciences., Joe Bible
Electronic Theses and Dissertations
In this work we present three topics, each of which centered on either the application or modification of various linear regression methods. Our work with respect to the “Materials Genome” project while undermined by oversimplification and data integrity issues in its early stages, provides a sound platform from which the project can proceed successfully. Building upon a growing body of knowledge around the use of Weighted Generalized Estimating Equations (WGEE), our second investigation proposes an extension to that framework intended to address the inherent bias present in the analysis of clustered longitudinal data with potentially informative cluster sizes and temporal …
Optcluster : An R Package For Determining The Optimal Clustering Algorithm And Optimal Number Of Clusters., Michael N. Sekula
Optcluster : An R Package For Determining The Optimal Clustering Algorithm And Optimal Number Of Clusters., Michael N. Sekula
Electronic Theses and Dissertations
Determining the best clustering algorithm and ideal number of clusters for a particular dataset is a fundamental difficulty in unsupervised clustering analysis. In biological research, data generated from Next Generation Sequencing technology and microarray gene expression data are becoming more and more common, so new tools and resources are needed to group such high dimensional data using clustering analysis. Different clustering algorithms can group data very differently. Therefore, there is a need to determine the best groupings in a given dataset using the most suitable clustering algorithm for that data. This paper presents the R package optCluster as an efficient …
Summary Of Survival Analysis With Sas Procedures., Derek Duane Childers 1990-
Summary Of Survival Analysis With Sas Procedures., Derek Duane Childers 1990-
Electronic Theses and Dissertations
The research conducted for this thesis was performed to summarize some of the most commonly used survival analysis techniques as well as to create one macro that will provide the solutions for these techniques. Some of the techniques that this thesis focuses on are survival and hazard functions, mean and median survival times, life table, log rank test, proportional hazards/model building, and competing risk. To further analyze these survival analysis techniques I will use the Bone Marrow Transplantation for Leukemia dataset. This trial consists of either acute myelocytic leukemia (AML 99 patients) or acute lymphoblastic leukemia (ALL 38 patients). There …
Assessing The Social And Ecological Factors That Influence Childhood Overweight And Obesity, Katie Callahan
Assessing The Social And Ecological Factors That Influence Childhood Overweight And Obesity, Katie Callahan
Electronic Theses and Dissertations
The prevalence of childhood overweight and obesity is increasing at an alarming rate in the United States. Currently more than 1 in 3 children aged 2-19 are overweight or obese. This is of major concern because childhood overweight and obesity leads to chronic conditions such as type II diabetes and tracks into adulthood, where more severe adverse health outcomes arise. In this study I used the premise of the social ecological model (SEM) to analyze the common levels that a child is exposed to daily; the intrapersonal level, the interpersonal level, the school level, and the community level to better …
Penalized Regressions For Variable Selection Model, Single Index Model And An Analysis Of Mass Spectrometry Data., Yubing Wan
Electronic Theses and Dissertations
The focus of this dissertation is to develop statistical methods, under the framework of penalized regressions, to handle three different problems. The first research topic is to address missing data problem for variable selection models including elastic net (ENet) method and sparse partial least squares (SPLS). I proposed a multiple imputation (MI) based weighted ENet (MI-WENet) method based on the stacked MI data and a weighting scheme for each observation. Numerical simulations were implemented to examine the performance of the MIWENet method, and compare it with competing alternatives. I then applied the MI-WENet method to examine the predictors for the …
Patient Rule Induction Method For Subgroup Identification Given Censored Data., Patrick James Trainor
Patient Rule Induction Method For Subgroup Identification Given Censored Data., Patrick James Trainor
Electronic Theses and Dissertations
The identification of subgroups in clinical studies is an important aspect of personalized medicine. In order to develop tailored therapeutics, the factors that characterize subgroups with differential prognosis, response to treatment, and incidence of adverse events or toxicities must be elucidated. We present a generalization of a statistical learning algorithm, Patient Rule Induction Method (PRIM), that is well suited for this task given a right-censored time-to-event outcome measure. This algorithm works to recursively partition a covariate space into mutually exclusive boxes that can be utilized to define subgroups. Conceptually the algorithm is similar to classification and regression trees but rather …
Statistical Methods For Assessing Treatment Effects For Observational Studies., Kristopher C. Gardner 1984-
Statistical Methods For Assessing Treatment Effects For Observational Studies., Kristopher C. Gardner 1984-
Electronic Theses and Dissertations
Though randomized clinical (RCTs) trials are the gold standard for comparing treatments, they are often infeasible or exclude clinically important subjects, or generally represent an idealized medical setting rather than real practice. Observational data provide an opportunity to study practice-based evidence, but also present challenges for analysis. Traditional statistical methods which are suitable for RCTs may be inadequate for the observational studies. In this project, four of the most popular statistical methods for observational studies: ANCOVA, propensity score matching, regression with the propensity score as a covariate, and instrumental variables (IV) are investigated through application to MarketScan insurance claims data. …
Compound Identification Using Penalized Linear Regression., Ruiqi Liu
Compound Identification Using Penalized Linear Regression., Ruiqi Liu
Electronic Theses and Dissertations
In this study, we propose a new method for compound identification using penalized linear regression. Compound identification is often achieved by matching the experimental mass spectra to the mass spectra stored in a reference library based on mass spectral similarity. In the context of the linear regression, the response variable is an experimental mass spectrum (i.e., query) and all the compounds in the reference library are the independent variables. However, the number of compounds in the reference library is much larger than the range of m/z values so that the data become high dimensional data with suffering from singularity. For …