Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (108)
- Applied Statistics (39)
- Life Sciences (38)
- Medicine and Health Sciences (33)
- Statistical Models (29)
-
- Statistical Methodology (23)
- Applied Mathematics (19)
- Social and Behavioral Sciences (14)
- Longitudinal Data Analysis and Time Series (13)
- Genetics and Genomics (12)
- Data Science (9)
- Design of Experiments and Sample Surveys (9)
- Multivariate Analysis (9)
- Public Health (9)
- Mathematics (8)
- Bioinformatics (7)
- Microarrays (7)
- Probability (7)
- Psychology (7)
- Computer Sciences (6)
- Genomics (6)
- Other Applied Mathematics (6)
- Categorical Data Analysis (5)
- Ecology and Evolutionary Biology (5)
- Genetics (5)
- Ordinary Differential Equations and Applied Dynamics (5)
- Dynamic Systems (4)
- Engineering (4)
- Keyword
-
- Epidemiology (8)
- Biostatistics (6)
- Neuroscience (6)
- Bayesian (5)
- Medicine (5)
-
- Clinical trial (4)
- Environment (4)
- Genomics (4)
- Bayesian Hierarchical Models (3)
- Gene expression (3)
- Longitudinal data (3)
- Machine learning (3)
- Mixed integer programming (3)
- Other (3)
- Power (3)
- Simulation (3)
- Support vector machine (3)
- Adaptive design (2)
- Algorithm (2)
- Chemicals (2)
- Classification (2)
- Cluster analysis (2)
- Covariance (2)
- DNA Methylation (2)
- Data analysis (2)
- Dimensionality reduction (2)
- Dropout (2)
- Evolution (2)
- Forensic Science (2)
- Genetics (2)
- Publication Year
- Publication
- Publication Type
Articles 121 - 150 of 176
Full-Text Articles in Statistics and Probability
Data Files To Accompany "The Support Vector Machine And Mixed Integer Linear Programming: Ramp Loss Svm With L1-Norm Regularization", Eric J. Hess, J. Paul Brooks
Data Files To Accompany "The Support Vector Machine And Mixed Integer Linear Programming: Ramp Loss Svm With L1-Norm Regularization", Eric J. Hess, J. Paul Brooks
Statistical Sciences and Operations Research Data
These files accompany, "The Support Vector Machine and Mixed Integer Linear Programming: Ramp Loss SVM with L1-Norm Regularization" by Eric J. Hess and J. Paul Brooks, presented at the 2015 INFORMS Computing Society Conference, Operations Research and Computing: Algorithms and Software for Analytics, Richmond, Virginia January 11-13, 2015.
The files contain instances of optimization problems that are described in the paper and for which results are reported. The files are in CPLEX LP format. The naming convention of the files is as follows: ndBTj0F.lp, where is the number of samples, is the number of attributes, and refers to …
Dynamic Bayesian Approaches To The Statistical Calibration Problem, Derick Lorenzo Rivers
Dynamic Bayesian Approaches To The Statistical Calibration Problem, Derick Lorenzo Rivers
Theses and Dissertations
The problem of statistical calibration of a measuring instrument can be framed both in a statistical context as well as in an engineering context. In the first, the problem is dealt with by distinguishing between the "classical" approach and the "inverse" regression approach. Both of these models are static models and are used to estimate "exact" measurements from measurements that are affected by error. In the engineering context, the variables of interest are considered to be taken at the time at which you observe the measurement. The Bayesian time series analysis method of Dynamic Linear Models (DLM) can be used …
Incorporating Dependence Boundaries In Simulating Associated Discrete Data, Mary E. Haynes
Incorporating Dependence Boundaries In Simulating Associated Discrete Data, Mary E. Haynes
Theses and Dissertations
In the study of associated discrete variables, limitations on the range of the possible association measures (Pearson correlation, odds ratio, etc.) arise from the form of the joint probability function between the variables. These limitations are known as the Fréchet bounds. The bounds for cases involving associated binary variables are explored in the context of simulating datasets with a desired correlation and set of marginal probabilities. A new method for creating such datasets is compared to an existing method that uses the multivariate probit. A method for simulating associated binary variables using a desired odds ratio and known marginal probabilities …
Methods For Integrative Analysis Of Genomic Data, Paul Manser
Methods For Integrative Analysis Of Genomic Data, Paul Manser
Theses and Dissertations
In recent years, the development of new genomic technologies has allowed for the investigation of many regulatory epigenetic marks besides expression levels, on a genome-wide scale. As the price for these technologies continues to decrease, study sizes will not only increase, but several different assays are beginning to be used for the same samples. It is therefore desirable to develop statistical methods to integrate multiple data types that can handle the increased computational burden of incorporating large data sets. Furthermore, it is important to develop sound quality control and normalization methods as technical errors can compound when integrating multiple genomic …
Latent Variable Models Given Incompletely Observed Surrogate Outcomes And Covariates, Chunfeng Ren
Latent Variable Models Given Incompletely Observed Surrogate Outcomes And Covariates, Chunfeng Ren
Theses and Dissertations
Latent variable models (LVMs) are commonly used in the scenario where the outcome of the main interest is an unobservable measure, associated with multiple observed surrogate outcomes, and affected by potential risk factors. This thesis develops an approach of efficient handling missing surrogate outcomes and covariates in two- and three-level latent variable models. However, corresponding statistical methodologies and computational software are lacking efficiently analyzing the LVMs given surrogate outcomes and covariates subject to missingness in the LVMs. We analyze the two-level LVMs for longitudinal data from the National Growth of Health Study where surrogate outcomes and covariates are subject to …
The Total Picture: Multiple Chemical Exposures To Pregnant Women In The Us – An Nhanes Study Of Data From 2003 Through 2010, Teri Cabana
Theses and Dissertations
INTRODUCTION: Chemical exposures to US pregnant women have been shown to have adverse health impacts on both mother and fetus. A prior paper revealed that US pregnant women in 2003-2004 had widespread exposure to multiple chemicals. The goal of this research is to examine how environmental chemical exposures to US pregnant women have changed from 2003 to 2010 and to look further at the extent of simultaneous exposure to multiple chemicals in US pregnant women using biomonitoring data available through NHANES (the National Health and Nutritional Examination Survey). METHODS: Using available NHANES data from the following cycles (2003-2004, 2005-2006, 2007-2008, …
Regularization Methods For Predicting An Ordinal Response Using Longitudinal High-Dimensional Genomic Data, Jiayi Hou
Theses and Dissertations
Ordinal scales are commonly used to measure health status and disease related outcomes in hospital settings as well as in translational medical research. Notable examples include cancer staging, which is a five-category ordinal scale indicating tumor size, node involvement, and likelihood of metastasizing. Glasgow Coma Scale (GCS), which gives a reliable and objective assessment of conscious status of a patient, is an ordinal scaled measure. In addition, repeated measurements are common in clinical practice for tracking and monitoring the progression of complex diseases. Classical ordinal modeling methods based on the likelihood approach have contributed to the analysis of data in …
Review And Extension For The O’Brien Fleming Multiple Testing Procedure, Hanan Hammouri
Review And Extension For The O’Brien Fleming Multiple Testing Procedure, Hanan Hammouri
Theses and Dissertations
O'Brien and Fleming (1979) proposed a straightforward and useful multiple testing procedure (group sequential testing procedure) for comparing two treatments in clinical trials where subject responses are dichotomous (e.g. success and failure). O'Brien and Fleming stated that their group sequential testing procedure has the same Type I error rate and power as that of a fixed one-stage chi-square test, but gives the opportunity to terminate the trial early when one treatment is clearly performing better than the other. We studied and tested the O'Brien and Fleming procedure specifically by correcting the originally proposed critical values. Furthermore, we updated the O’Brien …
Response Adaptive Design Using Auxiliary And Primary Outcomes, Shuxian Sinks
Response Adaptive Design Using Auxiliary And Primary Outcomes, Shuxian Sinks
Theses and Dissertations
Response adaptive designs intend to allocate more patients to better treatments without undermining the validity and the integrity of the trial. The immediacy of the primary response (e.g. deaths, remission) determines the efficiency of the response adaptive design, which often requires outcomes to be quickly or immediately observed. This presents difficulties for survival studies, which may require long durations to observe the primary endpoint. Therefore, we introduce auxiliary endpoints to assist the adaptation with the primary endpoint, where an auxiliary endpoint is generally defined as any measurement that is positively associated with the primary endpoint. Our proposed design (referred to …
The Estimation And Evaluation Of Optimal Thresholds For Two Sequential Testing Strategies, Amber R. Wilk
The Estimation And Evaluation Of Optimal Thresholds For Two Sequential Testing Strategies, Amber R. Wilk
Theses and Dissertations
Many continuous medical tests often rely on a threshold for diagnosis. There are two sequential testing strategies of interest: Believe the Positive (BP) and Believe the Negative (BN). BP classifies a patient positive if either the first test is greater than a threshold θ1 or negative on the first test and greater than θ2 on the second test. BN classifies a patient positive if the first test is greater than a threshold θ3 and greater than θ4 on the second test. Threshold pairs θ = (θ1, θ2) or (θ3, θ4), depending on strategy, are defined as optimal if they maximized …
Choosing The Cut Point For A Restricted Mean In Survival Analysis, A Data Driven Method, Emily H. Sheldon
Choosing The Cut Point For A Restricted Mean In Survival Analysis, A Data Driven Method, Emily H. Sheldon
Theses and Dissertations
Survival Analysis generally uses the median survival time as a common summary statistic. While the median possesses the desirable characteristic of being unbiased, there are times when it is not the best statistic to describe the data at hand. Royston and Parmar (2011) provide an argument that the restricted mean survival time should be the summary statistic used when the proportional hazards assumption is in doubt. Work in Restricted Means dates back to 1949 when J.O. Irwin developed a calculation for the standard error of the restricted mean using Greenwood’s formula. Since then the development of the restricted mean has …
Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese
Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese
Theses and Dissertations
Batch effects are due to probe-specific systematic variation between groups of samples (batches) resulting from experimental features that are not of biological interest. Principal components analysis (PCA) is commonly used as a visual tool to determine whether batch effects exist after applying a global normalization method. However, PCA yields linear combinations of the variables that contribute maximum variance and thus will not necessarily detect batch effects if they are not the largest source of variability in the data. We present an extension of principal components analysis to quantify the existence of batch effects, called guided PCA (gPCA). We describe a …
Characterization Of A Weighted Quantile Score Approach For Highly Correlated Data In Risk Analysis Scenarios, Caroline Carrico
Characterization Of A Weighted Quantile Score Approach For Highly Correlated Data In Risk Analysis Scenarios, Caroline Carrico
Theses and Dissertations
In risk evaluation, the effect of mixtures of environmental chemicals on a common adverse outcome is of interest. However, due to the high dimensionality and inherent correlations among chemicals that occur together, the traditional methods (e.g. ordinary or logistic regression) are unsuitable. We extend and characterize a weighted quantile score (WQS) approach to estimating an index for a set of highly correlated components. In the case with environmental chemicals, we use the WQS to identify “bad actors” and estimate body burden. The accuracy of the WQS was evaluated through extensive simulation studies in terms of validity (ability of the WQS …
Accounting For Model Uncertainty In Linear Mixed-Effects Models, Adam Sima
Accounting For Model Uncertainty In Linear Mixed-Effects Models, Adam Sima
Theses and Dissertations
Standard statistical decision-making tools, such as inference, confidence intervals and forecasting, are contingent on the assumption that the statistical model used in the analysis is the true model. In linear mixed-effect models, ignoring model uncertainty results in an underestimation of the residual variance, contributing to hypothesis tests that demonstrate larger than nominal Type-I errors and confidence intervals with smaller than nominal coverage probabilities. A novel utilization of the generalized degrees of freedom developed by Zhang et al. (2012) is used to adjust the estimate of the residual variance for model uncertainty. Additionally, the general global linear approximation is extended to …
A Comparison Of Methods Of Analysis To Control For Confounding In A Cohort Study Of A Dietary Intervention, Esinhart Hali
A Comparison Of Methods Of Analysis To Control For Confounding In A Cohort Study Of A Dietary Intervention, Esinhart Hali
Theses and Dissertations
Comparing samples from different populations can be biased by confounding. There are several statistical methods that can be used to control for confounding. These include; multiple linear regression, propensity score matching, propensity score/logit of propensity score as a single covariate in a linear regression model, stratified analysis using propensity score quintiles, weighted analysis using propensity scores or trimmed scores. The data were from two studies of a dietary intervention (FIBERR and RNP). The outcome variable was change from baseline to one month for eight outcome measures; fat, fiber, and fruits/ vegetables behavior, fat, fiber, and fruits/vegetables intentions, fat and fruits/vegetables …
The Effect Of Baseline Cluster Stratification On The Power Of Pre-Post Analysis, Fengjiao Hu
The Effect Of Baseline Cluster Stratification On The Power Of Pre-Post Analysis, Fengjiao Hu
Theses and Dissertations
The purpose of study is to check whether the power of detecting the effect of intervention versus control in a pre- and post-study can be increased by using a stratified randomized controlled design. A stratified randomized controlled design with two study arms and two time points, where strata are determined by clustering on baseline outcomes of the primary measure, is considered. A modified hierarchical clustering algorithm is developed which guarantees optimality as well as requiring each cluster to have at least one subject per study arm. The power is calculated based on simulated bivariate normal distributed primary measures with mixture …
Does Pair-Matching On Ordered Baseline Measures Increase Power: A Simulation Study, Yan Jin
Does Pair-Matching On Ordered Baseline Measures Increase Power: A Simulation Study, Yan Jin
Theses and Dissertations
It has been shown that pair-matching on an ordered baseline with normally distributed measures reduces the variance of the estimated treatment effect (Park and Johnson, 2006). The main objective of this study is to examine if pair-matching improves the power when the distribution is a mixture of two normal distributions. Multiple scenarios with a combination of different sample sizes and parameters are simulated. The power curves are provided for three cases, with and without matching, as follows: analysis of post-intervention data only, adding baseline as a covariate, and classic pre-post comparison. The study shows that the additional variance reduction provided …
Unbiased Estimation For The Contextual Effect Of Duration Of Adolescent Height Growth On Adulthood Obesity And Health Outcomes Via Hierarchical Linear And Nonlinear Models, Robert Carrico
Theses and Dissertations
This dissertation has multiple aims in studying hierarchical linear models in biomedical data analysis. In Chapter 1, the novel idea of studying the durations of adolescent growth spurts as a predictor of adulthood obesity is defined, established, and illustrated. The concept of contextual effects modeling is introduced in this first section as we study secular trend of adulthood obesity and how this trend is mitigated by the durations of individual adolescent growth spurts and the secular average length of adolescent growth spurts. It is found that individuals with longer periods of fast height growth in adolescence are more prone to …
Statistical Methods For Normalization And Analysis Of High-Throughput Genomic Data, Tobias Guennel
Statistical Methods For Normalization And Analysis Of High-Throughput Genomic Data, Tobias Guennel
Theses and Dissertations
High-throughput genomic datasets obtained from microarray or sequencing studies have revolutionized the field of molecular biology over the last decade. The complexity of these new technologies also poses new challenges to statisticians to separate biological relevant information from technical noise. Two methods are introduced that address important issues with normalization of array comparative genomic hybridization (aCGH) microarrays and the analysis of RNA sequencing (RNA-Seq) studies. Many studies investigating copy number aberrations at the DNA level for cancer and genetic studies use comparative genomic hybridization (CGH) on oligo arrays. However, aCGH data often suffer from low signal to noise ratios resulting …
Hypothesis Testing And Power Calculations For Taxonomic-Based Human Microbiome Data, P. S. Larossa, J. Paul Brooks, Elena Deych, Edward L. Boone, David J. Edwards, Qin Wang, Erica Sodergren, George Weinstock, William D. Shannon
Hypothesis Testing And Power Calculations For Taxonomic-Based Human Microbiome Data, P. S. Larossa, J. Paul Brooks, Elena Deych, Edward L. Boone, David J. Edwards, Qin Wang, Erica Sodergren, George Weinstock, William D. Shannon
Statistical Sciences and Operations Research Publications
This paper presents new biostatistical methods for the analysis of microbiome data based on a fully parametric approach using all the data. The Dirichlet-multinomial distribution allows the analyst to calculate power and sample sizes for experimental design, perform tests of hypotheses (e.g., compare microbiomes across groups), and to estimate parameters describing microbiome properties. The use of a fully parametric model for these data has the benefit over alternative non-parametric approaches such as bootstrapping and permutation testing, in that this model is able to retain more information contained in the data. This paper details the statistical approaches for several tests of …
Inferential Methods For High-Throughput Methylation Data, Maria Capparuccini
Inferential Methods For High-Throughput Methylation Data, Maria Capparuccini
Theses and Dissertations
The role of abnormal DNA methylation in the progression of disease is a growing area of research that relies upon the establishment of sound statistical methods. The common method for declaring there is differential methylation between two groups at a given CpG site, as summarized by the difference between proportions methylated db=b1-b2, has been through use of a Filtered Two Sample t-test, using the recommended filter of 0.17 (Bibikova et al., 2006b). In this dissertation, we performed a re-analysis of the data used in recommending the threshold by fitting a mixed-effects ANOVA model. It was determined that the 0.17 filter …
Power And Sample Size For Three-Level Cluster Designs, Tina Cunningham
Power And Sample Size For Three-Level Cluster Designs, Tina Cunningham
Theses and Dissertations
Over the past few decades, Cluster Randomized Trials (CRT) have become a design of choice in many research areas. One of the most critical issues in planning a CRT is to ensure that the study design is sensitive enough to capture the intervention effect. The assessment of power and sample size in such studies is often faced with many challenges due to several methodological difficulties. While studies on power and sample size for cluster designs with one and two levels are abundant, the evaluation of required sample size for three-level designs has been generally overlooked. First, the nesting effect introduces …
Stereotype Logit Models For High Dimensional Data, Andre Williams
Stereotype Logit Models For High Dimensional Data, Andre Williams
Theses and Dissertations
Gene expression studies are of growing importance in the field of medicine. In fact, subtypes within the same disease have been shown to have differing gene expression profiles (Golub et al., 1999). Often, researchers are interested in differentiating a disease by a categorical classification indicative of disease progression. For example, it may be of interest to identify genes that are associated with progression and to accurately predict the state of progression using gene expression data. One challenge when modeling microarray gene expression data is that there are more genes (variables) than there are observations. In addition, the genes usually demonstrate …
An Inferential Framework For Network Hypothesis Tests: With Applications To Biological Networks, Phillip Yates
An Inferential Framework For Network Hypothesis Tests: With Applications To Biological Networks, Phillip Yates
Theses and Dissertations
The analysis of weighted co-expression gene sets is gaining momentum in systems biology. In addition to substantial research directed toward inferring co-expression networks on the basis of microarray/high-throughput sequencing data, inferential methods are being developed to compare gene networks across one or more phenotypes. Common gene set hypothesis testing procedures are mostly confined to comparing average gene/node transcription levels between one or more groups and make limited use of additional network features, e.g., edges induced by significant partial correlations. Ignoring the gene set architecture disregards relevant network topological comparisons and can result in familiar n<
An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall
An Empirical Approach To Evaluating Sufficient Similarity: Utilization Of Euclidean Distance As A Similarity Measure, Scott Marshall
Theses and Dissertations
Individuals are exposed to chemical mixtures while carrying out everyday tasks, with unknown risk associated with exposure. Given the number of resulting mixtures it is not economically feasible to identify or characterize all possible mixtures. When complete dose-response data are not available on a (candidate) mixture of concern, EPA guidelines define a similar mixture based on chemical composition, component proportions and expert biological judgment (EPA, 1986, 2000). Current work in this literature is by Feder et al. (2009), evaluating sufficient similarity in exposure to disinfection by-products of water purification using multivariate statistical techniques and traditional hypothesis testing. The work of …
Cost And Accuracy Comparisons In Medical Testing Using Sequential Testing Strategies, Anwar Ahmed
Cost And Accuracy Comparisons In Medical Testing Using Sequential Testing Strategies, Anwar Ahmed
Theses and Dissertations
The practice of sequential testing is followed by the evaluation of accuracy, but often not by the evaluation of cost. This research described and compared three sequential testing strategies: believe the negative (BN), believe the positive (BP) and believe the extreme (BE), the latter being a less-examined strategy. All three strategies were used to combine results of two medical tests to diagnose a disease or medical condition. Descriptions of these strategies were provided in terms of accuracy (using the maximum receiver operating curve or MROC) and cost of testing (defined as the proportion of subjects who need 2 tests to …
A Numerical Method For Estimating The Variance Of Age At Maximum Growth Rate In Growth Models, Semhar Ogbagaber
A Numerical Method For Estimating The Variance Of Age At Maximum Growth Rate In Growth Models, Semhar Ogbagaber
Theses and Dissertations
Most studies on maturation and body composition using the Fels Longitudinal data mention peak height velocity (PHV) as an important outcome measure. The PHV is often derived from growth models such as the triple logistic model fitted to the stature (height) data. The age at PHV is sometimes ordinalized to designate an individual as an early, average or late maturer. In theory, age at PHV is the age at which the rate of growth reaches the maximum. Theoretically, for a well behaved growth function, this could be obtained by setting the second derivative of the growth function to zero and …
Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi
Bayesian And Frequentist Approaches For The Analysis Of Multiple Endpoints Data Resulting From Exposure To Multiple Health Stressors., Epiphanie Nyirabahizi
Theses and Dissertations
In risk analysis, Benchmark dose (BMD)methodology is used to quantify the risk associated with exposure to stressors such as environmental chemicals. It consists of fitting a mathematical model to the exposure data and the BMD is the dose expected to result in a pre-specified response or benchmark response (BMR). Most available exposure data are from single chemical exposure, but living objects are exposed to multiple sources of hazards. Furthermore, in some studies, researchers may observe multiple endpoints on one subject. Statistical approaches to address multiple endpoints problem can be partitioned into a dimension reduction group and a dimension preservative group. …
Is Screening Cargo Containers For Smuggled Nuclear Threats Worthwhile?, Jason R. W. Merrick, Laura A. Mclay
Is Screening Cargo Containers For Smuggled Nuclear Threats Worthwhile?, Jason R. W. Merrick, Laura A. Mclay
Statistical Sciences and Operations Research Publications
In recent years, Customs and Border Protection (CBP) has installed radiation sensors to screen cargo containers entering theUnited States. They are concerned that terrorists could use containers to smuggle radiological material into the country and carry out attacks with dirty bombs or a nuclear device. Recent studies have questioned the value of improving this screening system with new sensor technology. The cost of delays caused by frequent false alarms outweighs any reduction in the probability of an attack in an expected cost analysis. We extend existing methodology in three ways to demonstrate how additional factors affect the value of screening …
Nonlinear Models In Multivariate Population Bioequivalence Testing, Bassam Dahman
Nonlinear Models In Multivariate Population Bioequivalence Testing, Bassam Dahman
Theses and Dissertations
In this dissertation a methodology is proposed for simultaneously evaluating the population bioequivalence (PBE) of a generic drug to a pre-licensed drug, or the bioequivalence of two formulations of a drug using multiple correlated pharmacokinetic metrics. The univariate criterion that is accepted by the food and drug administration (FDA) for testing population bioequivalence is generalized. Very few approaches for testing multivariate extensions of PBE have appeared in the literature. One method uses the trace of the covariance matrix as a measure of total variability, and another uses a pooled variance instead of the reference variance. The former ignores the correlation …