Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (3)
- Business (2)
- Business Analytics (2)
- Business Intelligence (2)
- Data Science (2)
-
- Probability (2)
- Statistical Models (2)
- Analysis (1)
- Bioinformatics (1)
- Biostatistics (1)
- Categorical Data Analysis (1)
- Computational Biology (1)
- Computer Engineering (1)
- Data Storage Systems (1)
- Digital Communications and Networking (1)
- Engineering (1)
- Genetics and Genomics (1)
- Life Sciences (1)
- Longitudinal Data Analysis and Time Series (1)
- Mathematics (1)
- Microarrays (1)
- Operations and Supply Chain Management (1)
- Other Business (1)
- Other Statistics and Probability (1)
- Service Learning (1)
- Social and Behavioral Sciences (1)
- Sociology (1)
- Keyword
-
- Distance Correlation (2)
- Blockchain (1)
- Blockchain solutions (1)
- Categorical Data (1)
- Communication (1)
-
- Consumer-Packaged Goods (1)
- Data (1)
- Data analysis (1)
- Differential Item Functioning (1)
- Dirichlet-Multinomial regression (1)
- Gene Set Test (1)
- Gibbs sampler (1)
- Graph-Based Multivariate Test (1)
- Horseshoe (1)
- Horseshoe plus (1)
- Inventory Management (1)
- Item Response Theory (1)
- Laplace (1)
- Learning Networks (1)
- MCMC (1)
- MCMC algorithm (1)
- Machine Learning (1)
- Markov Chain Monte Carlo (1)
- Metropolis-Hasting (1)
- Microbiome data (1)
- Model-based Recursive Partitioning (1)
- Monte Carlo Markov Chain (1)
- Multivariate Independent (1)
- Outlier (1)
- Overdispersion (1)
Articles 1 - 7 of 7
Full-Text Articles in Multivariate Analysis
Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald
Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald
Data Science Undergraduate Honors Theses
Coca-Cola is a popular soft drink brand with sales occurring in every Walmart store across the world, which generates large quantities of data and requires a robust supply chain system. However, the company does not currently have a sophisticated, automated, and/or prescriptive system for detecting where, when, and why inventory outages occur and applying preventative measures to avoid loss of revenue from the absence of inventory on store shelves. This thesis proposes and applies a novel, prescriptive system for this purpose. An inventory outage can be seen as a ‘negative’ statistical outlier in a time series of inventory for an …
How Blockchain Solutions Enable Better Decision Making Through Blockchain Analytics, Sammy Ter Haar
How Blockchain Solutions Enable Better Decision Making Through Blockchain Analytics, Sammy Ter Haar
Information Systems Undergraduate Honors Theses
Since the founding of computers, data scientists have been able to engineer devices that increase individuals’ opportunities to communicate with each other. In the 1990s, the internet took over with many people not understanding its utility. Flash forward 30 years, and we cannot live without our connection to the internet. The internet of information is what we called early adopters with individuals posting blogs for others to read, this was known as Web 1.0. As we progress, platforms became social allowing individuals in different areas to communicate and engage with each other, this was known as Web 2.0. As Dr. …
Evaluating The Efficiency Of Markov Chain Monte Carlo Algorithms, Thuy Scanlon
Evaluating The Efficiency Of Markov Chain Monte Carlo Algorithms, Thuy Scanlon
Graduate Theses and Dissertations
Markov chain Monte Carlo (MCMC) is a simulation technique that produces a Markov chain designed to converge to a stationary distribution. In Bayesian statistics, MCMC is used to obtain samples from a posterior distribution for inference. To ensure the accuracy of estimates using MCMC samples, the convergence to the stationary distribution of an MCMC algorithm has to be checked. As computation time is a resource, optimizing the efficiency of an MCMC algorithm in terms of effective sample size (ESS) per time unit is an important goal for statisticians. In this paper, we use simulation studies to demonstrate how the Gibbs …
Gene Set Testing By Distance Correlation, Sho-Hsien Su
Gene Set Testing By Distance Correlation, Sho-Hsien Su
Graduate Theses and Dissertations
Pathways are the functional building blocks of complex diseases such as cancers. Pathway-level studies may provide insights on some important biological processes. Gene set test is an important tool to study the differential expression of a gene set between two groups, e.g., cancer vs normal. The differential expression of a gene set could be due to the difference in mean, variability, or both. However, most existing gene set tests only target the mean difference but overlook other types of differential expression. In this thesis, we propose to use the recently developed distance correlation for gene set testing. To assess the …
Assessing Differential Item Functioning In The Perceived Stress Scale, Nana Amma Berko Asamoah
Assessing Differential Item Functioning In The Perceived Stress Scale, Nana Amma Berko Asamoah
Graduate Theses and Dissertations
When an item on a test functions differently for subgroups of respondents with respect to an exogenous variable (or covariate) after conditioning on the latent variable of interest, the item is said to exhibit Differential Item Functioning (DIF). The 10-item Perceived Stress Scale (PSS10) is administered to respondents via MTurk to quantify “perceived stress” and identify if items on the scale function differently for specific subgroups defined by age, sex, race, marital status, number of children, employment status and social media usage.
The purpose of this study was to compare traditional DIF detection approaches (Mantel-Haenszel, logistic regression, likelihood ratio test …
Learning Networks With Categorical Data Using Distance Correlation, And A Novel Graph-Based Multivariate Test, Jian Tinker
Learning Networks With Categorical Data Using Distance Correlation, And A Novel Graph-Based Multivariate Test, Jian Tinker
Graduate Theses and Dissertations
We study the use of distance correlation for statistical inference on categorical data, especially the induction of probability networks. Szekely et al. first defined distance correlation for continuous variables in [42], and Zhang translated the concept into the categorical setting in [57] by defining dCor(X,Y) for categorical variables X = (x1,...,xI) and Y = (y1,...,yJ) where P(X=xi)=[pi]i and P(Y=yi)=[pi]j with the formula [Please open the document]
Part I of the dissertation covers the background we need to understand this formula, and prepares us to analyze the properties and performance of its applications.
Part II then presents the main results of …
Effect Of Predictor Dependence On Variable Selection For Linear And Log-Linear Regression, Apu Chandra Das
Effect Of Predictor Dependence On Variable Selection For Linear And Log-Linear Regression, Apu Chandra Das
Graduate Theses and Dissertations
We propose a Bayesian approach to the Dirichlet-Multinomial (DM) regression model, which uses horseshoe, Laplace, and horseshoe plus priors for shrinkage and selection. The Dirichlet-Multinomial model can be used to find the significant association between a set of available covariates and taxa for a microbiome sample. We incorporate the covariates in a log-linear regression framework. We design a simulation study to make a comparison among the performance of the three shrinkage priors in terms of estimation accuracy and the ability to detect true signals. Our results have clearly separated the performance of the three priors and indicated that the horseshoe …