Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (162)
- Applied Statistics (91)
- Engineering (86)
- Statistical Models (59)
- Operations Research, Systems Engineering and Industrial Engineering (47)
-
- Design of Experiments and Sample Surveys (41)
- Operational Research (35)
- Statistical Methodology (34)
- Social and Behavioral Sciences (32)
- Life Sciences (29)
- Applied Mathematics (22)
- Medicine and Health Sciences (22)
- Business (20)
- Longitudinal Data Analysis and Time Series (19)
- Data Science (17)
- Multivariate Analysis (17)
- Computer Sciences (16)
- Mathematics (14)
- Survival Analysis (14)
- Environmental Sciences (13)
- Electrical and Computer Engineering (12)
- Aviation (11)
- Psychology (11)
- Oceanography and Atmospheric Sciences and Meteorology (10)
- Arts and Humanities (9)
- Bioinformatics (9)
- Categorical Data Analysis (9)
- Genetics and Genomics (9)
- Institution
- Keyword
-
- Bayesian (21)
- Statistics (20)
- Physical Sciences and Mathematics, Statistics and Probability (12)
- Machine learning (11)
- Simulation (11)
-
- Biostatistics (9)
- Classification (9)
- Survival analysis (8)
- Design of experiments (7)
- Semiparametric (7)
- Bayesian Hierarchical Models (6)
- Bayesian methods (6)
- Goodness-of-fit tests (6)
- Longitudinal data (6)
- Markov chain Monte Carlo (6)
- Maximum likelihood (6)
- Physical Sciences and Mathematics, Statistics (6)
- Physical Sciences and Mathematics, Statistics and Probability, Biostatistics (6)
- Variable selection (6)
- #antcenter (5)
- Clinical trial (5)
- Clustering (5)
- EM algorithm (5)
- Experimental design (5)
- Gene expression (5)
- Gibbs sampling (5)
- MCMC (5)
- Neural networks (5)
- Optimization (5)
- Parameter estimation (5)
- Publication Year
Articles 331 - 360 of 565
Full-Text Articles in Statistics and Probability
Characterization Of A Weighted Quantile Score Approach For Highly Correlated Data In Risk Analysis Scenarios, Caroline Carrico
Characterization Of A Weighted Quantile Score Approach For Highly Correlated Data In Risk Analysis Scenarios, Caroline Carrico
Theses and Dissertations
In risk evaluation, the effect of mixtures of environmental chemicals on a common adverse outcome is of interest. However, due to the high dimensionality and inherent correlations among chemicals that occur together, the traditional methods (e.g. ordinary or logistic regression) are unsuitable. We extend and characterize a weighted quantile score (WQS) approach to estimating an index for a set of highly correlated components. In the case with environmental chemicals, we use the WQS to identify “bad actors” and estimate body burden. The accuracy of the WQS was evaluated through extensive simulation studies in terms of validity (ability of the WQS …
Extensions For The Epsilon Skew Exponential Power Distribution Family, Ahmad N. Flaih
Extensions For The Epsilon Skew Exponential Power Distribution Family, Ahmad N. Flaih
Theses and Dissertations
Practical applied statistics reveals that the analysis of many real data that exhibit both fat-tailless and skewness shows significant departures from normality assumptions. In these circumstances, the adoption of more flexible models that cope with near normal data may be appropriate in place of adopting the robust approach, semi-parametric or nonparametric models, and Box-Cox transformation. An alternative approach is to consider using the Epsilon Skew Exponential Power (ESEP) distribution developed by Elsalloukh [12] and [15] which embeds the normal, fat-tailedness, and skewness distributions as special cases. We develop models based on using the ESEP family, in particular, the ESEP receiver …
Accounting For Model Uncertainty In Linear Mixed-Effects Models, Adam Sima
Accounting For Model Uncertainty In Linear Mixed-Effects Models, Adam Sima
Theses and Dissertations
Standard statistical decision-making tools, such as inference, confidence intervals and forecasting, are contingent on the assumption that the statistical model used in the analysis is the true model. In linear mixed-effect models, ignoring model uncertainty results in an underestimation of the residual variance, contributing to hypothesis tests that demonstrate larger than nominal Type-I errors and confidence intervals with smaller than nominal coverage probabilities. A novel utilization of the generalized degrees of freedom developed by Zhang et al. (2012) is used to adjust the estimate of the residual variance for model uncertainty. Additionally, the general global linear approximation is extended to …
Estimation And Q-Matrix Validation For Diagnostic Classification Models, Yuling Feng
Estimation And Q-Matrix Validation For Diagnostic Classification Models, Yuling Feng
Theses and Dissertations
Diagnostic classification models (DCMs) are structured latent class models widely discussed in the field of psychometrics. They model subjects' underlying attribute patterns and classify subjects into unobservable groups based on their mastery of attributes required to answer the items correctly. The effective implementation of DCMs depends on correct specification of a Q-matrix which is a binary matrix linking attribute patterns to items. Current literature on assessing the appropriateness of Q-matrix specifications has focused on validation methods for the deterministic-input, noisy-and-gate (DINA) model. The goal of the study is to develop general Q-matrix validation methods that can be applied to a …
Heaped Data In Count Models, Tammy Harris
Heaped Data In Count Models, Tammy Harris
Theses and Dissertations
Heaped data result when subjects who recall the frequency of events prefer for reporting from a limited set of rounded responses or preferred digits over reporting exact counts. These rounded responses and digit preferences (also referred to as data coarsening) could be characterized by reported frequencies (or counts) favoring multiples of 20, reporting counts ending with 0 or 5, or a preference for reporting an even number over an odd number or vice versa. This mixture of values is a type of measurement error (pattern of misreporting) that can lead to biased estimation and imprecision in discrete quantitative data. Sometimes …
Models And Software Development For Interval-Censored Data, Chun Pan
Models And Software Development For Interval-Censored Data, Chun Pan
Theses and Dissertations
Interval-censored time-to-event data occur naturally in studies of diseases where the symptoms are not directly observable, and periodic clinical examinations are required for detection. Due to the lack of well-established procedures, interval-censored data have been conventionally treated as right-censored data, however, this introduces bias at the first place. This dissertation focuses on methodological research and software development for interval-censored data. Specifically, it consists of three projects. The first project is to create an R package for regression analysis and survival curve estimation of interval-censored data based on several published papers by our research team. In the second project, a Bayesian …
The Complete Plus-Minus: A Case Study Of The Columbus Blue Jackets, Nathan Spagnola
The Complete Plus-Minus: A Case Study Of The Columbus Blue Jackets, Nathan Spagnola
Theses and Dissertations
A new hockey statistic termed the Complete Plus-Minus (CPM) was created to calculate the abilities of hockey players in the National Hockey League (NHL). This new statistic was used to analyze the Columbus Blue Jackets for the 2011-2012 season. The CPM for the Blue Jackets was created using two logistic regressions that modeled a goal being scored for and against the Blue Jackets. Whether a goal was scored for or against the team were the responses, while events on the ice were the predictors in the model. It was found that the team's poor performance was due to a weak …
A New Method For The Comparison Of Survival Distributions, Jaymie Shanahan
A New Method For The Comparison Of Survival Distributions, Jaymie Shanahan
Theses and Dissertations
The assessment of overall homogeneity of time-to-event curves is a key element in survival analysis in biomedical research. The currently commonly used testing methods, e.g. log-rank test, Wilcoxon test, and Kolmogorov-Smirnov test, may have a significant loss of statistical testing power under certain circumstances. In this thesis we replicate a testing method (Lin & Xu, 2009) that is robust for the comparison of the overall homogeneity of survival curves based on the absolute difference of the area under the survival curves using normal approximation by Greenwood's formula, and propose a new weight component to their test statistic. The weight component …
Protein Identification Using Bayesian Stochastic Search, Christina Nicole Lewis
Protein Identification Using Bayesian Stochastic Search, Christina Nicole Lewis
Theses and Dissertations
Current methods for protein identification in tandem mass spectrometry (MS/MS) involve database searches or de novo peptide sequencing, with database searches being the standard method. With database searches, issues arise when the species is not in the database. Shortcomings of de novo peptide sequencing and database searches include chemical noise, overly complex fragments, and incomplete b and y ion sequences. Here we present a Bayesian approach to identifying peptides. Our model uses prior information about the average relative abundances of bond cleavages and the prior probability of any particular amino acid sequence. The proposed likelihood function is composed of two …
Modeling Mixed Unfolding/Monotone Dichotomous Item Exams, Na Yang
Modeling Mixed Unfolding/Monotone Dichotomous Item Exams, Na Yang
Theses and Dissertations
Item response theory (IRT) is widely applied to analyze educational and psychological assessments. Readily available IRT implementations allow for two common types of models: monotone models used for dominance scales (Guttman 1950; Rasch 1960/1980; Birnbaum 1968; Mokken 1971) and unfolding models used for proximity scales (Coombs, 1964; Andrich, 1996; Roberts, Donoghue and Laughlin, 2000).
When an exam contains items following both types of models, there is currently no method to distinguish the item types, estimate their characteristics, or estimate the examinee characteristics. Thus, there is no existing methodology to simultaneously analyze items like ``At a minimum, I am in favor …
Permutation Testing For Covariance Matrices, With Applications In Shape Analysis, Blake Cassidy Hill
Permutation Testing For Covariance Matrices, With Applications In Shape Analysis, Blake Cassidy Hill
Theses and Dissertations
In many applications, it is of interest to compare covariance structures. In this work, we propose hypothesis tests for comparing covariance matrices for data in different groups, especially in shape analysis. The main motivation for the work is comparing covariance matrices of the size and shapes of damaged versus undamaged DNA molecules. A practical motivation behind analyzing the differences between these DNA covariance matrices is to compare the variation between the two groups during situations where the molecules are repairing. The testing methods proposed in this dissertation consist of three types of permutation testing methods for differences in covariance structures. …
Advanced Methodology Developments In Mixture Cure Models, Chao Cai
Advanced Methodology Developments In Mixture Cure Models, Chao Cai
Theses and Dissertations
Modern medical treatments have substantially improved cure rates for many chronic diseases and have generated increasing interest in appropriate statistical models to handle survival data with non-negligible cure fractions. The mixture cure models are designed to model such data set, which assume that studied population is a mixture of being cured and uncured. In this dissertation, I will develop two programs named smcure and NPHMC in R. The first program aims to facilitate estimating two popular mixture cure models: the proportional hazards (PH) mixture cure model and accelerated failure time (AFT) mixture cure model. The second program focuses on designing …
A Comparison Of Methods Of Analysis To Control For Confounding In A Cohort Study Of A Dietary Intervention, Esinhart Hali
A Comparison Of Methods Of Analysis To Control For Confounding In A Cohort Study Of A Dietary Intervention, Esinhart Hali
Theses and Dissertations
Comparing samples from different populations can be biased by confounding. There are several statistical methods that can be used to control for confounding. These include; multiple linear regression, propensity score matching, propensity score/logit of propensity score as a single covariate in a linear regression model, stratified analysis using propensity score quintiles, weighted analysis using propensity scores or trimmed scores. The data were from two studies of a dietary intervention (FIBERR and RNP). The outcome variable was change from baseline to one month for eight outcome measures; fat, fiber, and fruits/ vegetables behavior, fat, fiber, and fruits/vegetables intentions, fat and fruits/vegetables …
An Integrated Screening And Optimization Strategy, Nathaniel Jackson Rohbock
An Integrated Screening And Optimization Strategy, Nathaniel Jackson Rohbock
Theses and Dissertations
Within statistical methods, design of experiments (DOE) is well suited to make good inference from a minimal amount of data. Two types of designs within DOE are screening designs and optimization designs. Traditionally, these approaches have been necessarily separated by a gap between the objectives of each design and the methods available. Despite being so separated, in practice these designs are frequently connected by sequential experimentation. In fact, from the genesis of a project, the experimentor often knows that both designs will be necessary to accomplish his objectives. Due to advances in the understanding of experimental designs with complex aliasing …
The Effect Of Baseline Cluster Stratification On The Power Of Pre-Post Analysis, Fengjiao Hu
The Effect Of Baseline Cluster Stratification On The Power Of Pre-Post Analysis, Fengjiao Hu
Theses and Dissertations
The purpose of study is to check whether the power of detecting the effect of intervention versus control in a pre- and post-study can be increased by using a stratified randomized controlled design. A stratified randomized controlled design with two study arms and two time points, where strata are determined by clustering on baseline outcomes of the primary measure, is considered. A modified hierarchical clustering algorithm is developed which guarantees optimality as well as requiring each cluster to have at least one subject per study arm. The power is calculated based on simulated bivariate normal distributed primary measures with mixture …
Does Pair-Matching On Ordered Baseline Measures Increase Power: A Simulation Study, Yan Jin
Does Pair-Matching On Ordered Baseline Measures Increase Power: A Simulation Study, Yan Jin
Theses and Dissertations
It has been shown that pair-matching on an ordered baseline with normally distributed measures reduces the variance of the estimated treatment effect (Park and Johnson, 2006). The main objective of this study is to examine if pair-matching improves the power when the distribution is a mixture of two normal distributions. Multiple scenarios with a combination of different sample sizes and parameters are simulated. The power curves are provided for three cases, with and without matching, as follows: analysis of post-intervention data only, adding baseline as a covariate, and classic pre-post comparison. The study shows that the additional variance reduction provided …
An Applied Investigation Of Gaussian Markov Random Fields, Jessica Lyn Olsen
An Applied Investigation Of Gaussian Markov Random Fields, Jessica Lyn Olsen
Theses and Dissertations
Recently, Bayesian methods have become the essence of modern statistics, specifically, the ability to incorporate hierarchical models. In particular, correlated data, such as the data found in spatial and temporal applications, have benefited greatly from the development and application of Bayesian statistics. One particular application of Bayesian modeling is Gaussian Markov Random Fields. These methods have proven to be very useful in providing a framework for correlated data. I will demonstrate the power of GMRFs by applying this method to two sets of data; a set of temporal data involving car accidents in the UK and a set of spatial …
Xprime-Em: Eliciting Expert Prior Information For Motif Exploration Using The Expectation-Maximization Algorithm, Wei Zhou
Theses and Dissertations
Understanding the possible mechanisms of gene transcription regulation is a primary challenge for current molecular biologists. Identifying transcription factor binding sites (TFBSs), also called DNA motifs, is an important step in understanding these mechanisms. Furthermore, many human diseases are attributed to mutations in TFBSs, which makes identifying those DNA motifs significant for disease treatment. Uncertainty and variations in specific nucleotides of TFBSs present difficulties for DNA motif searching. In this project, we present an algorithm, XPRIME-EM (Eliciting EXpert PRior Information for Motif Exploration using the Expectation-Maximization Algorithm), which can discover known and de novo (unknown) DNA motifs simultaneously from a …
Estimation Of The Effects Of Parental Measures On Child Aggression Using Structural Equation Modeling, Jordan Daniel Pyper
Estimation Of The Effects Of Parental Measures On Child Aggression Using Structural Equation Modeling, Jordan Daniel Pyper
Theses and Dissertations
A child's parents are the primary source of knowledge and learned behaviors for developing children, and the benefits or repercussions of certain parental practices can be long lasting. Although parenting practices affect behavioral outcomes for children, families tend to be diverse in their circumstances and needs. Research attempting to ascertain cause and effect relationships between parental influences and child behavior can be difficult due to the complex nature of family dynamics and the intricacies of real life. Structural equation modeling (SEM) is an appropriate method for this research as it is able to account for the complicated nature of child-parent …
Unbiased Estimation For The Contextual Effect Of Duration Of Adolescent Height Growth On Adulthood Obesity And Health Outcomes Via Hierarchical Linear And Nonlinear Models, Robert Carrico
Theses and Dissertations
This dissertation has multiple aims in studying hierarchical linear models in biomedical data analysis. In Chapter 1, the novel idea of studying the durations of adolescent growth spurts as a predictor of adulthood obesity is defined, established, and illustrated. The concept of contextual effects modeling is introduced in this first section as we study secular trend of adulthood obesity and how this trend is mitigated by the durations of individual adolescent growth spurts and the secular average length of adolescent growth spurts. It is found that individuals with longer periods of fast height growth in adolescence are more prone to …
Support Vector Machines For Classification And Imputation, Spencer David Rogers
Support Vector Machines For Classification And Imputation, Spencer David Rogers
Theses and Dissertations
Support vector machines (SVMs) are a powerful tool for classification problems. SVMs have only been developed in the last 20 years with the availability of cheap and abundant computing power. SVMs are a non-statistical approach and make no assumptions about the distribution of the data. Here support vector machines are applied to a classic data set from the machine learning literature and the out-of-sample misclassification rates are compared to other classification methods. Finally, an algorithm for using support vector machines to address the difficulty in imputing missing categorical data is proposed and its performance is demonstrated under three different scenarios …
Species Identification And Strain Attribution With Unassembled Sequencing Data, Owen Eric Francis
Species Identification And Strain Attribution With Unassembled Sequencing Data, Owen Eric Francis
Theses and Dissertations
Emerging sequencing approaches have revolutionized the way we can collect DNA sequence data for applications in bioforensics and biosurveillance. In this research, we present an approach to construct a database of known biological agents and use this database to develop a statistical framework to analyze raw reads from next-generation sequence data for species identification and strain attribution. Our method capitalizes on a Bayesian statistical framework that accommodates information on sequence quality, mapping quality and provides posterior probabilities of matches to a known database of target genomes. Importantly, our approach also incorporates the possibility that multiple species can be present in …
Hitters Vs. Pitchers: A Comparison Of Fantasy Baseball Player Performances Using Hierarchical Bayesian Models, Scott D. Huddleston
Hitters Vs. Pitchers: A Comparison Of Fantasy Baseball Player Performances Using Hierarchical Bayesian Models, Scott D. Huddleston
Theses and Dissertations
In recent years, fantasy baseball has seen an explosion in popularity. Major League Baseball, with its long, storied history and the enormous quantity of data available, naturally lends itself to the modern-day recreational activity known as fantasy baseball. Fantasy baseball is a game in which participants manage an imaginary roster of real players and compete against one another using those players' real-life statistics to score points. Early forms of fantasy baseball began in the early 1960s, but beginning in the 1990s, the sport was revolutionized due to the advent of powerful computers and the Internet. The data used in this …
Using An Experimental Mixture Design To Identify Experimental Regions With High Probability Of Creating A Homogeneous Monolithic Column Capable Of Flow, Charles C. Willden
Using An Experimental Mixture Design To Identify Experimental Regions With High Probability Of Creating A Homogeneous Monolithic Column Capable Of Flow, Charles C. Willden
Theses and Dissertations
Graduate students in the Brigham Young University Chemistry Department are working to develop a filtering device that can be used to separate substances into their constituent parts. The device consists of a monomer and water mixture that is polymerized into a monolith inside of a capillary. The ideal monolith is completely solid with interconnected pores that are small enough to cause the constituent parts to pass through the capillary at different rates, effectively separating the substance. Although the end objective is to minimize pore sizes, it is necessary to first identify an experimental region where any combination of input variables …
A Dempster-Shafer Method For Multi-Sensor Fusion, Bethany G. Foley
A Dempster-Shafer Method For Multi-Sensor Fusion, Bethany G. Foley
Theses and Dissertations
The Dempster-Shafer Theory, a generalization of the Bayesian theory, is based on the idea of belief and as such can handle ignorance. When all of the required information is available, many data fusion methods provide a solid approach. Yet, most do not have a good way of dealing with ignorance. In the absence of information, these methods must then make assumptions about the sensor data. However, the real data may not fit well within the assumed model. Consequently, the results are often unsatisfactory and inconsistent. The Dempster-Shafer Theory is not hindered by incomplete models or by the lack of prior …
Using Multiattribute Utility Copulas In Support Of Uav Search And Destroy Operations, Beau A. Nunnally
Using Multiattribute Utility Copulas In Support Of Uav Search And Destroy Operations, Beau A. Nunnally
Theses and Dissertations
The multiattribute utility copula is an emerging form of utility function used by decision analysts to study decisions with dependent attributes. Failure to properly address attribute dependence may cause errors in selecting the optimal policy. This research examines two scenarios of interest to the modern warfighter. The first scenario employs a utility copula to determine the type, quantity, and altitude of UAVs to be sent to strike a stationary target. The second scenario employs a utility copula to examine the impact of attribute dependence on the optimal routing of UAVs in a contested operational environment when performing a search and …
Computer Aided Multi-Data Fusion Dismount Modeling, Juan L. Morales
Computer Aided Multi-Data Fusion Dismount Modeling, Juan L. Morales
Theses and Dissertations
Recent research efforts strive to address the growing need for dismount surveillance, dismount tracking and characterization. Current work in this area utilizes hyperspectral and multispectral imaging systems to exploit spectral properties in order to detect areas of exposed skin and clothing characteristics. Because of the large bandwidth and high resolution, hyperspectral imaging systems pose great ability to characterize and detect dismounts. A multi-data dismount modeling system where the development and manipulation of dismount models is a necessity. This thesis demonstrates a computer aided multi-data fused dismount model, which facilitates studies of dismount detection, characterization and identification. The system is created …
Covariance Analysis Of Vision Aided Navigation By Bootstrapping, Andrew L. Relyea
Covariance Analysis Of Vision Aided Navigation By Bootstrapping, Andrew L. Relyea
Theses and Dissertations
Inertial Navigation System (INS) aiding using bearing measurements taken over time of stationary ground features is investigated. A cross country flight, in two and three dimensional space, is considered, as well as a vertical drop in three dimensional space. The objective is to quantify the temporal development of the uncertainty in the navigation states of an aircraft INS which is aided by taking bearing measurements of ground objects which have been geolocated using ownship position. It is shown that during wings level flight at constant speed and a fixed altitude, an aircraft that tracks ground objects and over time sequentially …
Bayesian Pollution Source Apportionment Incorporating Multiple Simultaneous Measurements, Jonathan Casey Christensen
Bayesian Pollution Source Apportionment Incorporating Multiple Simultaneous Measurements, Jonathan Casey Christensen
Theses and Dissertations
We describe a method to estimate pollution profiles and contribution levels for distinct prominent pollution sources in a region based on daily pollutant concentration measurements from multiple measurement stations over a period of time. In an extension of existing work, we will estimate common source profiles but distinct contribution levels based on measurements from each station. In addition, we will explore the possibility of extending existing work to allow adjustments for synoptic regimes—large scale weather patterns which may effect the amount of pollution measured from individual sources as well as for particular pollutants. For both extensions we propose Bayesian methods …
Predicting Maximal Oxygen Consumption (Vo2max) Levels In Adolescents, Brent A. Shepherd
Predicting Maximal Oxygen Consumption (Vo2max) Levels In Adolescents, Brent A. Shepherd
Theses and Dissertations
Maximal oxygen consumption (VO2max) is considered by many to be the best overall measure of an individual's cardiovascular health. Collecting the measurement, however, requires subjecting an individual to prolonged periods of intense exercise until their maximal level, the point at which their body uses no additional oxygen from the air despite increased exercise intensity, is reached. Collecting VO2max data also requires expensive equipment and great subject discomfort to get accurate results. Because of this inherent difficulty, it is often avoided despite its usefulness. In this research, we propose a set of Bayesian hierarchical models to predict VO2max levels in adolescents, …