Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (38)
- Life Sciences (17)
- Statistical Models (17)
- Statistical Methodology (13)
- Biostatistics (12)
-
- Engineering (12)
- Categorical Data Analysis (8)
- Industrial Engineering (8)
- Operations Research, Systems Engineering and Industrial Engineering (8)
- Social and Behavioral Sciences (7)
- Education (6)
- Medicine and Health Sciences (6)
- Multivariate Analysis (5)
- Probability (5)
- Statistical Theory (5)
- Educational Assessment, Evaluation, and Research (4)
- Geography (4)
- Longitudinal Data Analysis and Time Series (4)
- Operational Research (4)
- Applied Mathematics (3)
- Cancer Biology (3)
- Cell and Developmental Biology (3)
- Genetics and Genomics (3)
- Medical Specialties (3)
- Plant Sciences (3)
- Remote Sensing (3)
- Vital and Health Statistics (3)
- Agriculture (2)
- Keyword
-
- Pure sciences (14)
- Bayesian (6)
- Statistics (6)
- Applied sciences (4)
- Bayesian Statistics (3)
-
- Differential Item Functioning (3)
- Horseshoe (3)
- Bayesian Analysis (2)
- Biological sciences (2)
- Clinical trials (2)
- Conditional autoregressive prior (2)
- Distance Correlation (2)
- Item response theory (2)
- Landsat time series (2)
- MCMC (2)
- Measurement invariance (2)
- Ovarian cancer (2)
- Poisson (2)
- Psychometrics (2)
- Quality control (2)
- Regression (2)
- Statistical Learning (2)
- Survival analysis (2)
- Variable selection (2)
- 5-fluorouracil (1)
- Aberrant responding (1)
- Adaptation (1)
- Aggregate failure-time data (1)
- Analytics (1)
- Appraisals (1)
Articles 31 - 60 of 84
Full-Text Articles in Statistics and Probability
Aberrant Responding With Underlying Dominance And Unfolding Response Processes: Examining Model Fit And Performance Of Person-Fit Statistics, Jennifer A. Reimers
Aberrant Responding With Underlying Dominance And Unfolding Response Processes: Examining Model Fit And Performance Of Person-Fit Statistics, Jennifer A. Reimers
Graduate Theses and Dissertations
Researchers have recognized that respondents may not answer items in a way that accurately reflects their attitude or trait level being measured. The resulting response data that deviates from what would be expected has been shown to have significant effects on the psychometric properties of a scale and analytical results. However, many studies that have investigated the detection of aberrant data and its effects have done so using dominance item response theory (IRT) models. It is unknown whether the impacts of aberrant data and the methodology used to identify aberrant responding when using dominance IRT models apply similarly when scales …
Data-Driven Statin Initiation Evaluation And Optimization For Prediabetes Population, Muhenned A. Abdulsahib
Data-Driven Statin Initiation Evaluation And Optimization For Prediabetes Population, Muhenned A. Abdulsahib
Graduate Theses and Dissertations
This dissertation develops quantitative models to support medical decision making of statininitiation considering the uncertainty in disease progression for prediabetes patients. A mathematical model is built to help medical decision-makers take action of statin initiation under uncertainty in future prediabetes progressions. The association between cholesterol drug use, such as statin, and elevating glucose level attracted considerable amounts of attention in the literature. Statin effects on glucose vary with respect to different levels of glucose. The first chapter of this dissertation introduces the problem and an overview of the tools that will be used to solve it. In the second chapter …
Statistical Modeling, Learning And Computing For Stochastic Dynamics Of Complex Systems, Mohammadmahdi Hajiha
Statistical Modeling, Learning And Computing For Stochastic Dynamics Of Complex Systems, Mohammadmahdi Hajiha
Graduate Theses and Dissertations
With the recent advances in sensor technology, it is much easier to collect and store streams of system operational and environmental (SOE) data. These data can be used as input to model the underlying behavior of complex engineered systems and phenomenons if appropriate algorithms with well-defined assumptions are developed. This dissertation is comprised of the research work to show the applicability of SOE data when fed into proposed tailored algorithms. The first purposes of these algorithms are to estimate and analyze the reliability of a system as elaborated in Chapter 2. This chapter provides the derivation of closed-form expressions that …
Evaluating The Efficiency Of Markov Chain Monte Carlo Algorithms, Thuy Scanlon
Evaluating The Efficiency Of Markov Chain Monte Carlo Algorithms, Thuy Scanlon
Graduate Theses and Dissertations
Markov chain Monte Carlo (MCMC) is a simulation technique that produces a Markov chain designed to converge to a stationary distribution. In Bayesian statistics, MCMC is used to obtain samples from a posterior distribution for inference. To ensure the accuracy of estimates using MCMC samples, the convergence to the stationary distribution of an MCMC algorithm has to be checked. As computation time is a resource, optimizing the efficiency of an MCMC algorithm in terms of effective sample size (ESS) per time unit is an important goal for statisticians. In this paper, we use simulation studies to demonstrate how the Gibbs …
Privacy-Preserving Cloud-Assisted Data Analytics, Wei Bao
Privacy-Preserving Cloud-Assisted Data Analytics, Wei Bao
Graduate Theses and Dissertations
Nowadays industries are collecting a massive and exponentially growing amount of data that can be utilized to extract useful insights for improving various aspects of our life. Data analytics (e.g., via the use of machine learning) has been extensively applied to make important decisions in various real world applications. However, it is challenging for resource-limited clients to analyze their data in an efficient way when its scale is large. Additionally, the data resources are increasingly distributed among different owners. Nonetheless, users' data may contain private information that needs to be protected.
Cloud computing has become more and more popular in …
Statistical Modeling For High-Dimensional Compositional Data With Applications To The Human Microbiome, Thy Dao
Graduate Theses and Dissertations
Compositional data refer to the data that lie on a simplex, which are common in many scientific domains such as genomics, geology, and economics. As the components in a composition must sum to one, traditional tests based on unconstrained data become inappropriate, and new statistical methods are needed to analyze this special type of data. This dissertation is motivated by some statistical problems arising in the analysis of compositional data. In particular, we focus on the high-dimensional and over-dispersed setting, where the dimensionality of compositions is greater than the sample size and the dispersion parameter is moderate or large. In …
Knowledge Discovery From Complex Event Time Data With Covariates, Samira Karimi
Knowledge Discovery From Complex Event Time Data With Covariates, Samira Karimi
Graduate Theses and Dissertations
In particular engineering applications, such as reliability engineering, complex types of data are encountered which require novel methods of statistical analysis. Handling covariates properly while managing the missing values is a challenging task. These type of issues happen frequently in reliability data analysis. Specifically, accelerated life testing (ALT) data are usually conducted by exposing test units of a product to severer-than-normal conditions to expedite the failure process. The resulting lifetime and/or censoring data are often modeled by a probability distribution along with a life-stress relationship. However, if the probability distribution and life-stress relationship selected cannot adequately describe the underlying failure …
Comparative Evaluation Of Statistical Dependence Measures, Eman Abdel Rahman Ibrahim
Comparative Evaluation Of Statistical Dependence Measures, Eman Abdel Rahman Ibrahim
Graduate Theses and Dissertations
Measuring and testing dependence between random variables is of great importance in many scientific fields. In the case of linearly correlated variables, Pearson’s correlation coefficient is a commonly used measure of the correlation strength. In the case of nonlinear correlation, several innovative measures have been proposed, such as distance-based correlation, rank-based correlations, and information theory-based correlation. This thesis focuses on the statistical comparison of several important correlations, including Spearman’s correlation, mutual information, maximal information coefficient, biweight midcorrelation, distance correlation, and copula correlation, under various simulation settings such as correlative patterns and the level of random noise. Furthermore, we apply those …
Gene Set Testing By Distance Correlation, Sho-Hsien Su
Gene Set Testing By Distance Correlation, Sho-Hsien Su
Graduate Theses and Dissertations
Pathways are the functional building blocks of complex diseases such as cancers. Pathway-level studies may provide insights on some important biological processes. Gene set test is an important tool to study the differential expression of a gene set between two groups, e.g., cancer vs normal. The differential expression of a gene set could be due to the difference in mean, variability, or both. However, most existing gene set tests only target the mean difference but overlook other types of differential expression. In this thesis, we propose to use the recently developed distance correlation for gene set testing. To assess the …
Development Of An Effect Size To Classify The Magnitude Of Dif In Dichotomous And Polytomous Items, James D. Weese
Development Of An Effect Size To Classify The Magnitude Of Dif In Dichotomous And Polytomous Items, James D. Weese
Graduate Theses and Dissertations
A standardized effect size for the SIBTEST/POLYSIBTEST procedure is proposed, allowing for Differential Item Functioning (DIF) to be classified with a single set of DIF heuristics regardless of whether data are dichotomous or polytomous. This proposed standardized effect size accounts for both variability in responses and whether participants are included in the SIBTEST/POLYSIBTEST calculations. First, a new set of unstandardized effect size heuristics are established for dichotomous data that are more aligned with Educational Testing Service (ETS) standards using two and three parameter logistic (2PL and 3PL) models. Second, a standardized effect size is proposed and compared to other DIF …
Conditional Distance Correlation Test For Gene Expression Level, Dna Methylation Level And Copy Number, Shanshan Zhang
Conditional Distance Correlation Test For Gene Expression Level, Dna Methylation Level And Copy Number, Shanshan Zhang
Graduate Theses and Dissertations
Over the past years, efforts have been devoted to the genome-wide analysis of genetic and epigenetic profiles to better understand the underlying biological mechanisms of complex diseases such as cancer. It is of great importance to unravel the complex dependence structure between biological factors, and many conditional dependence tests have been developed to meet this need. The traditional partial correlation method can only capture the linear partial correlation, but not the nonlinear correlation. To overcome this limitation, we propose to use the innovative conditional distance correlation (CDC), which measures the conditional dependence between random vectors and detect nonlinear relations. In …
Quantifying The Simultaneous Effect Of Socio-Economic Predictors And Build Environment On Spatial Crime Trends, Alfieri Daniel Ek
Quantifying The Simultaneous Effect Of Socio-Economic Predictors And Build Environment On Spatial Crime Trends, Alfieri Daniel Ek
Graduate Theses and Dissertations
Proper allocation of law enforcement agencies falls under the umbrella of risk terrainmodeling (Caplan et al., 2011, 2015; Drawve, 2016) that primarily focuses on crime prediction and prevention by spatially aggregating response and predictor variables of interest. Although mental health incidents demand resource allocation from law enforcement agencies and the city, relatively less emphasis has been placed on building spatial models for mental health incidents events. Analyzing spatial mental health events in Little Rock, AR over 2015 to 2018, we found evidence of spatial heterogeneity via Moran’s I statistic. A spatial modeling framework is then built using generalized linear models, …
Assessing Differential Item Functioning In The Perceived Stress Scale, Nana Amma Berko Asamoah
Assessing Differential Item Functioning In The Perceived Stress Scale, Nana Amma Berko Asamoah
Graduate Theses and Dissertations
When an item on a test functions differently for subgroups of respondents with respect to an exogenous variable (or covariate) after conditioning on the latent variable of interest, the item is said to exhibit Differential Item Functioning (DIF). The 10-item Perceived Stress Scale (PSS10) is administered to respondents via MTurk to quantify “perceived stress” and identify if items on the scale function differently for specific subgroups defined by age, sex, race, marital status, number of children, employment status and social media usage.
The purpose of this study was to compare traditional DIF detection approaches (Mantel-Haenszel, logistic regression, likelihood ratio test …
Learning Networks With Categorical Data Using Distance Correlation, And A Novel Graph-Based Multivariate Test, Jian Tinker
Learning Networks With Categorical Data Using Distance Correlation, And A Novel Graph-Based Multivariate Test, Jian Tinker
Graduate Theses and Dissertations
We study the use of distance correlation for statistical inference on categorical data, especially the induction of probability networks. Szekely et al. first defined distance correlation for continuous variables in [42], and Zhang translated the concept into the categorical setting in [57] by defining dCor(X,Y) for categorical variables X = (x1,...,xI) and Y = (y1,...,yJ) where P(X=xi)=[pi]i and P(Y=yi)=[pi]j with the formula [Please open the document]
Part I of the dissertation covers the background we need to understand this formula, and prepares us to analyze the properties and performance of its applications.
Part II then presents the main results of …
Effect Of Predictor Dependence On Variable Selection For Linear And Log-Linear Regression, Apu Chandra Das
Effect Of Predictor Dependence On Variable Selection For Linear And Log-Linear Regression, Apu Chandra Das
Graduate Theses and Dissertations
We propose a Bayesian approach to the Dirichlet-Multinomial (DM) regression model, which uses horseshoe, Laplace, and horseshoe plus priors for shrinkage and selection. The Dirichlet-Multinomial model can be used to find the significant association between a set of available covariates and taxa for a microbiome sample. We incorporate the covariates in a log-linear regression framework. We design a simulation study to make a comparison among the performance of the three shrinkage priors in terms of estimation accuracy and the ability to detect true signals. Our results have clearly separated the performance of the three priors and indicated that the horseshoe …
Models For Data Analysis In Accelerated Reliability Growth, Cesar Alexander Ruiz Torres
Models For Data Analysis In Accelerated Reliability Growth, Cesar Alexander Ruiz Torres
Graduate Theses and Dissertations
This work develops new methodologies for analyzing accelerated testing data in the context of a reliability growth program for a complex multi-component system. Each component has multiple failure modes and the growth program consists of multiple test-fix stages with corrective actions applied at the end of each stage. The first group of methods considers time-to-failure data and test covariates for predicting the final reliability of the system. The time-to-failure of each failure mode is assumed to follow a Weibull distribution with rate parameter proportional to an acceleration factor. Acceleration factors are specific to each failure mode and test covariates. We …
Structural Analysis Of The Multifunctional Spoiie Regulatory Protein Of Clostridioides Difficile., Blythe Emily Bunkers
Structural Analysis Of The Multifunctional Spoiie Regulatory Protein Of Clostridioides Difficile., Blythe Emily Bunkers
Graduate Theses and Dissertations
Clostridioides (formally Clostridium) difficile is a medically relevant pathogen pertinent to infectious disease research. C. difficile is distinctly known for its ability to produce two toxins, enterotoxin A and cytotoxin B, and the propensity to colonize the mammalian gastrointestinal tract. It is known that metabolism is tightly correlated with sporulation in endospore producers such as C. difficile, but an interesting and novel regulatory relationship found by the Ivey lab has yet to be understood. The relationship explored in this study is observed between the sporulation factor, SpoIIE, which represses expression of an ABC peptide transporter, app. In this study, two …
Measuring Sexual Excitation And Sexual Inhibition In A Dutch-Speaking Sample, Malachi Willis
Measuring Sexual Excitation And Sexual Inhibition In A Dutch-Speaking Sample, Malachi Willis
Graduate Theses and Dissertations
Background: Individual differences in sexual excitation and sexual inhibition are important predictors of sexual functioning. Psychometric instruments for these aspects of sexual response were originally developed separately for men (Sexual Inhibition /Sexual Excitation Scales [SIS/SES]) and women (Sexual Excitation/Sexual Inhibition Inventory for Women [SESII-W]). These measures were then adapted to function similarly in samples comprising both men and women (Sexual Inhibition/Sexual Excitation Scales-Short Form [SIS/SES-SF] and Sexual Excitation/Sexual Inhibition Inventory for Women and Men [SESII-W/M], respectively). No published study to our knowledge has administered the SIS/SES and SESII-W/M questionnaires to a sample of both women and men. In the present …
Detecting Differentially Co-Expressed Gene Modules Via The Edge-Count Test, Anne Gratius Lin
Detecting Differentially Co-Expressed Gene Modules Via The Edge-Count Test, Anne Gratius Lin
Graduate Theses and Dissertations
Background
Gene expression profiling by microarray has been used to uncover molecular variations in many different diseases. Complementary to conventional differential expression analysis, differential co-expression analysis can identify gene markers from the systematic and granular level. There are three aspects for differential co-expression network analysis, including the network global topological comparison, differential co-expression cluster identification, and differential co-expressed genes and gene pair identification. To date, most of the methods available still rely on Pearson’s correlation coefficient despite its nonlinear insensitivity.
Results
Here we present an approach that is robust to nonlinearity by using the edge-count test for differential co-expression analysis. …
Spatio-Temporal Prediction Of Arkansas Gubernatorial Election, Michael Harris
Spatio-Temporal Prediction Of Arkansas Gubernatorial Election, Michael Harris
Graduate Theses and Dissertations
Our goal is to create spatio-temporal models for predicting future gubernatorial elections. For a concrete example of how well our models work we use past data to predict the 2018 Arkansas gubernatorial election and use the existing 2018 election data to check our models predictive accuracy. Gubernatorial election data was collected from the Arkansas Secretary of State website while related covariate data was collected from the website for the Federal Reserve Bank of St. Louis. The data we collect is on the county level. For predictive purposes we fit multiple models to the data using Markov chain Monte Carlo and …
Probabilistic Models For Order-Picking Operations With Multiple In-The-Aisle Pick Positions, Jingming Liu
Probabilistic Models For Order-Picking Operations With Multiple In-The-Aisle Pick Positions, Jingming Liu
Graduate Theses and Dissertations
The development of probability density functions (pdfs) for travel time of a narrow aisle lift truck (NALT) and an automated storage and retrieval (AS/R) machine is the focus of the dissertation. The multiple in-the-aisle pick positions (MIAPP) order picking system can be modeled as an M/G/1 queueing problem in which storage and retrieval requests are the customers and the vehicle (NALT or AS/R machine) is the server. Service time is the sum of travel time and the deterministic time to pick up and deposit a pallet (TPD).
Our first contribution is the development of travel time pdfs for retrieval operations …
Spatio-Temporal Analysis Of Tree Ring Chronology And Precipitation, Ruizhe Yin
Spatio-Temporal Analysis Of Tree Ring Chronology And Precipitation, Ruizhe Yin
Graduate Theses and Dissertations
Tree ring chronology data is known to reflect regional climate due to the strong impact of rainfall and temperature. Therefore, tree ring data can be used to reconstruct historical climate in order to understand how climate changed in the past and make prediction about the future behavior of the climate. For simplicity, this research only considers the influence of precipitation on tree ring growth within the New England area. A total of 94 measurement sites are used to record tree ring width over 881 years and corresponding precipitation data are given at some locations for 121 years. We developed a …
Effect Of Cross-Validation On The Output Of Multiple Testing Procedures, Josh Dallas Price
Effect Of Cross-Validation On The Output Of Multiple Testing Procedures, Josh Dallas Price
Graduate Theses and Dissertations
High dimensional data with sparsity is routinely observed in many scientific disciplines. Filtering out the signals embedded in noise is a canonical problem in such situations requiring multiple testing. The Benjamini--Hochberg procedure using False Discovery Rate control is the gold standard in large scale multiple testing. In Majumder et al. (2009) an internally cross-validated form of the procedure is used to avoid a costly replicate study and the complications that arise from population selection in such studies (i.e. extraneous variables). I implement this procedure and run extensive simulation studies under increasing levels of dependence among parameters and different data generating …
A Bayesian Framework For Estimating Seismic Wave Arrival Time, Hua Zhong
A Bayesian Framework For Estimating Seismic Wave Arrival Time, Hua Zhong
Graduate Theses and Dissertations
Because earthquakes have a large impact on human society, statistical methods for better studying earthquakes are required. One characteristic of earthquakes is the arrival time of seismic waves at a seismic signal sensor. Once we can estimate the earthquake arrival time accurately, the earthquake location can be triangulated, and assistance can be sent to that area correctly. This study presents a Bayesian framework to predict the arrival time of seismic waves with associated uncertainty. We use a change point framework to model the different conditions before and after the seismic wave arrives. To evaluate the performance of the model, we …
Comparing Elo, Glicko, Irt, And Bayesian Irt Statistical Models For Educational And Gaming Data, Breanna Morrison
Comparing Elo, Glicko, Irt, And Bayesian Irt Statistical Models For Educational And Gaming Data, Breanna Morrison
Graduate Theses and Dissertations
Statistical models used for estimating skill or ability levels often vary by field, however their underlying mathematical models can be very similar. Differences in the underlying models can be due to the need to accommodate data with different underlying formats and structure. As the models from varying fields increase in complexity, their ability to be applied to different types of data may have the ability to increase. Models that are applied to educational or psychological data have advanced to accommodate a wide range of data formats, including increased estimation accuracy with sparsely populated data matrices. Conversely, the field of online …
Advanced Statistics In Arkansas Sports Reporting, Andrew Lee Epperson
Advanced Statistics In Arkansas Sports Reporting, Andrew Lee Epperson
Graduate Theses and Dissertations
This study seeks to analyze how Arkansas’ sports journalists are adapting to the recent surge in available advanced statistics that are being used by certain national news organizations. Using in-depth qualitative research that includes in-depth interviews with a number of individuals in the print, broadcast, and athletics side of sports coverage, we discover how journalists and coaches use these next-generation analytics, what they fundamentally mean for the evolution of each respective path, and why so few Arkansas reporters and writers use them at the time of this paper’s defense. We see how budgets and deadlines restrict the use of these …
A Hidden Markov Factor Analysis Framework For Seizure Detection In Epilepsy Patients, Mahboubeh Madadi
A Hidden Markov Factor Analysis Framework For Seizure Detection In Epilepsy Patients, Mahboubeh Madadi
Graduate Theses and Dissertations
Approximately 1% of the world population suffers from epilepsy. Continuous long-term electroencephalographic (EEG) monitoring is the gold-standard for recording epileptic seizures and assisting in the diagnosis and treatment of patients with epilepsy. Detection of seizure from the recorded EEG is a laborious, time consuming and expensive task. In this study, we propose an automated seizure detection framework to assist electroencephalographers and physicians with identification of seizures in recorded EEG signals. In addition, an automated seizure detection algorithm can be used for treatment through automatic intervention during the seizure activity and on time triggering of the injection of a radiotracer to …
A Generative Statistical Approach For Data Classification In A Biologically Inspired Design Tool, Marvin Manuel Arroyo Rujano
A Generative Statistical Approach For Data Classification In A Biologically Inspired Design Tool, Marvin Manuel Arroyo Rujano
Graduate Theses and Dissertations
The objective of the research this thesis describes is to find a way to classify text-based descriptions of biological adaption to support Biologically Inspired design. Biologically inspired design is a fairly new field with ongoing research. There are different tools to assist designers and biologists in bio-inspired design. Some of the most common are BioTRIZ and AskNature. In recent years, more tools have been proposed to aid and make research in the field easier, for example, the Biologically Inspired Adaptive System Design (BIASD) tool. This tool was designed with the goal of helping designers in early design stages generate more …
Sequential Inference For Hidden Markov Models, Michael Ellis
Sequential Inference For Hidden Markov Models, Michael Ellis
Graduate Theses and Dissertations
In many applications data are collected sequentially in time with very short time intervals between observations. If one is interested in using new observations as they arrive in time then non-sequential Bayesian inference methods, such as Markov Chain Monte Carlo (MCMC) sampling, can be too slow. Increasingly, state space models are being used to model nonlinear and non-Gaussian systems. The structure of state space models allows for sequential Bayesian inference so that an approximation to the posterior distribution of interest can be updated as new observations arrive. In special cases, the exact posterior distribution can be updated through conjugate Bayesian …
Quantitative Microbial Risk Assessment For Parts, Ground, And Msc Poultry Product Including Intervention Analysis And Exploration Of Enterobacteriaceae As An Indicator Organism In Poultry Processing, Leigh Ann Parette
Graduate Theses and Dissertations
Samples collected at five different large bird poultry processing facilities over a period of 7 months from prescald to post debone locations were enumerated for Enterobacteriaceae, Salmonella spp., and Campylobacter spp. and the results were used to create Quantitative Microbial Risk Analyses (QMRA) models for parts, ground, and mechanically separated chicken (MSC) products. Sensitivity analyses indicated the points in the process at which reductions would be most advantageous to the endpoint and simulation models were run to test reductions required to meet the current USDA performance standards.
These data were analyzed to determine the reductions from one node (location) to …