Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (3)
- Engineering (3)
- Data Science (2)
- Industrial Engineering (2)
- Operations Research, Systems Engineering and Industrial Engineering (2)
-
- Biostatistics (1)
- Computer Engineering (1)
- Computer Sciences (1)
- Data Storage Systems (1)
- Databases and Information Systems (1)
- Digital Communications and Networking (1)
- Education (1)
- Educational Assessment, Evaluation, and Research (1)
- Educational Psychology (1)
- Industrial Technology (1)
- Information Security (1)
- Multivariate Analysis (1)
- Operational Research (1)
- Probability (1)
- Programming Languages and Compilers (1)
- Statistical Methodology (1)
- Statistical Theory (1)
- Systems Architecture (1)
- Vital and Health Statistics (1)
- Keyword
-
- Aggregate failure-time data (1)
- Applied sciences (1)
- Bayesian method (1)
- Bayesian nonparametric (1)
- Big data (1)
-
- Boosting Trees (1)
- Categorical Data (1)
- Chosen plaintext attack (CPA) (1)
- Cloud-Assisted Systems (1)
- Compositional data (1)
- Constructed-response (1)
- Data privacy (1)
- Distance Correlation (1)
- Distance-based (1)
- Ensemble Learning (1)
- Fair exchange (1)
- Feature Selection (1)
- Geometric Programming (GP) (1)
- Gradient Boosted Trees (1)
- Gradient projection method (GPM) (1)
- Graph-Based Multivariate Test (1)
- Heterogeneous architectures (1)
- Hierarchical rater model (1)
- High dimensional compositional data (1)
- Item response theory (1)
- Learning Networks (1)
- Logistic regression (1)
- Machine learning (1)
- Mathematical optimization (1)
- Maximum likelihood estimation (1)
Articles 1 - 8 of 8
Full-Text Articles in Categorical Data Analysis
Mle And Eap Methods For Estimating Ability Scores For Data Of Varying Sample Size And Item Length, Sahar Taji
Mle And Eap Methods For Estimating Ability Scores For Data Of Varying Sample Size And Item Length, Sahar Taji
Graduate Theses and Dissertations
In this research, the performance of two popular estimators, Maximum Likelihood Estimator(MLE) and Bayesian Expected a Posteriori (EAP) is studied and compared in estimating the latent ability score in an Item Response Theory (IRT) model. The 2-Parameter Logistic (2PL) IRT model which is characterized by difficulty and discrimination item parameters is used to estimate the latent ability scores. Several datasets are generated for variety of sample size and item length values. The Monte-Carlo simulation is used to analyze the performance of the estimators. Results show that MLE produces reliable results with low root mean square error (RMSE) across all datasets. …
Ensemble Tree-Based Machine Learning For Imaging Data, Reza Iranzad
Ensemble Tree-Based Machine Learning For Imaging Data, Reza Iranzad
Graduate Theses and Dissertations
In particular medical imaging data, such as positron emission tomography (PET), computed tomography (CT), and fluorescence intravital microscopy (IVM), have become prevalent for use in a wide variety of applications, from diagnostic purposes, tracking diseases' progress, and monitoring the effectiveness of treatments to decision-making processes. The detailed information generated by medical imaging has enabled physicians to provide more comprehensive care. Although numerous machine learning algorithms, especially those used for imaging data, have been developed, dealing with unique structures in imaging data remained a big challenge. In this dissertation, we are proposing novel statistical tree-based methods with more efficient and more …
Posterior Predictive Model Checking Of The Hierarchical Rater Model, Nnamdi Chika Ezike
Posterior Predictive Model Checking Of The Hierarchical Rater Model, Nnamdi Chika Ezike
Graduate Theses and Dissertations
Fitting wrongly specified models to observed data may lead to invalid inferences about the model parameters of interest. The current study investigated the performance of the posterior predictive model checking (PPMC) approach in detecting model-data misfit of the hierarchical rater model (HRM). The HRM is a rater-mediated model that incorporates components of the polytomous item response theory (IRT) model, such as the partial credit model (PCM) and generalized partial credit model (GPCM), at the second level of the hierarchy, to model examinees’ responses to performance assessments. To date, the HRM has not been rigorously evaluated using PPMC techniques. Monte Carlo …
Privacy-Preserving Cloud-Assisted Data Analytics, Wei Bao
Privacy-Preserving Cloud-Assisted Data Analytics, Wei Bao
Graduate Theses and Dissertations
Nowadays industries are collecting a massive and exponentially growing amount of data that can be utilized to extract useful insights for improving various aspects of our life. Data analytics (e.g., via the use of machine learning) has been extensively applied to make important decisions in various real world applications. However, it is challenging for resource-limited clients to analyze their data in an efficient way when its scale is large. Additionally, the data resources are increasingly distributed among different owners. Nonetheless, users' data may contain private information that needs to be protected.
Cloud computing has become more and more popular in …
Statistical Modeling For High-Dimensional Compositional Data With Applications To The Human Microbiome, Thy Dao
Graduate Theses and Dissertations
Compositional data refer to the data that lie on a simplex, which are common in many scientific domains such as genomics, geology, and economics. As the components in a composition must sum to one, traditional tests based on unconstrained data become inappropriate, and new statistical methods are needed to analyze this special type of data. This dissertation is motivated by some statistical problems arising in the analysis of compositional data. In particular, we focus on the high-dimensional and over-dispersed setting, where the dimensionality of compositions is greater than the sample size and the dispersion parameter is moderate or large. In …
Knowledge Discovery From Complex Event Time Data With Covariates, Samira Karimi
Knowledge Discovery From Complex Event Time Data With Covariates, Samira Karimi
Graduate Theses and Dissertations
In particular engineering applications, such as reliability engineering, complex types of data are encountered which require novel methods of statistical analysis. Handling covariates properly while managing the missing values is a challenging task. These type of issues happen frequently in reliability data analysis. Specifically, accelerated life testing (ALT) data are usually conducted by exposing test units of a product to severer-than-normal conditions to expedite the failure process. The resulting lifetime and/or censoring data are often modeled by a probability distribution along with a life-stress relationship. However, if the probability distribution and life-stress relationship selected cannot adequately describe the underlying failure …
Learning Networks With Categorical Data Using Distance Correlation, And A Novel Graph-Based Multivariate Test, Jian Tinker
Learning Networks With Categorical Data Using Distance Correlation, And A Novel Graph-Based Multivariate Test, Jian Tinker
Graduate Theses and Dissertations
We study the use of distance correlation for statistical inference on categorical data, especially the induction of probability networks. Szekely et al. first defined distance correlation for continuous variables in [42], and Zhang translated the concept into the categorical setting in [57] by defining dCor(X,Y) for categorical variables X = (x1,...,xI) and Y = (y1,...,yJ) where P(X=xi)=[pi]i and P(Y=yi)=[pi]j with the formula [Please open the document]
Part I of the dissertation covers the background we need to understand this formula, and prepares us to analyze the properties and performance of its applications.
Part II then presents the main results of …
Probabilistic Graphical Modeling On Big Data, Ming-Hua Chung
Probabilistic Graphical Modeling On Big Data, Ming-Hua Chung
Graduate Theses and Dissertations
The rise of Big Data in recent years brings many challenges to modern statistical analysis and modeling. In toxicogenomics, the advancement of high-throughput screening technologies facilitates the generation of massive amount of biological data, a big data phenomena in biomedical science. Yet, researchers still heavily rely on key word search and/or literature review to navigate the databases and analyses are often done in rather small-scale. As a result, the rich information of a database has not been fully utilized, particularly for the information embedded in the interactive nature between data points that are largely ignored and buried. For the past …