Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,820 Full-Text Articles 23,917 Authors 9,922,835 Downloads 282 Institutions

All Articles in Statistics and Probability

Faceted Search

12,820 full-text articles. Page 202 of 487.

An Investigation Of Chi-Square And Entropy Based Methods Of Item-Fit Using Item Level Contamination In Item Response Theory, William R. Dardick, Brandi A. Weiss 2020 The George Washington University

An Investigation Of Chi-Square And Entropy Based Methods Of Item-Fit Using Item Level Contamination In Item Response Theory, William R. Dardick, Brandi A. Weiss

Journal of Modern Applied Statistical Methods

New variants of entropy as measures of item-fit in item response theory are investigated. Monte Carlo simulation(s) examine aberrant conditions of item-level misfit to evaluate relative (compare EMRj, X2, G2, S-X2, and PV-Q1) and absolute (Type I error and empirical power) performance. EMRj has utility in discovering misfit.


Hereditarily Irreducible Maps, Hussam Abobaker, Włodzimierz J. Charatonik 2020 Missouri University of Science and Technology

Hereditarily Irreducible Maps, Hussam Abobaker, Włodzimierz J. Charatonik

Mathematics and Statistics Faculty Research & Creative Works

A map f:X→Y from a continuum X onto a continuum Y is said to be hereditarily irreducible, if f(A)⊊f(B) for any subcontinua A and B such that A⊊B. We investigate properties of hereditarily irreducible maps between continua. Special attention is given to maps between graphs and maps from the interval.


Video Game Genre Classification Based On Deep Learning, Yuhang Jiang 2020 Western Kentucky University

Video Game Genre Classification Based On Deep Learning, Yuhang Jiang

Masters Theses & Specialist Projects

Video games have played a more and more important role in our life. While the genre classification is a deeply explored research subject by leveraging the strength of deep learning, the automatic video game genre classification has drawn little attention in academia. In this study, we compiled a large dataset of 50,000 video games, consisting of the video game covers, game descriptions and the genre information. We explored three approaches for genre classification using deep learning techniques. First, we developed five image-based models utilizing pre-trained computer vision models such as MobileNet, ResNet50 and Inception, based on the game covers. Second, …


A Monte Carlo Analysis Of Ordinary Least Squares Versus Equal Weights, James Brewer Ayres 2020 Western Kentucky University

A Monte Carlo Analysis Of Ordinary Least Squares Versus Equal Weights, James Brewer Ayres

Masters Theses & Specialist Projects

Equal weights are an alternative weighting procedure to the optimal weights offered by ordinary least squares regression analysis. Also called units weights, equal weights are formed by standardizing scores on the predictor variables and averaging these standardized scores to create a composite score. Research is limited regarding the conditions under which equal weights result in cross-validated 𝑅𝑅2 values that meet or exceed optimal weights. In this study, I explored the effect of various predictor-criterion correlations, predictor intercorrelations, and sample sizes to determine the relative performance of equal and optimal weighting schemes upon cross-validation. Results indicated that optimally weighted predictors explained …


A Data Analytic Framework For Physical Fatigue Management Using Wearable Sensors, Zahra Sedighi Maman, Ying-Ju Chen, Amir Baghdadi, Seamus Lombardo, Lora A. Cavuoto, Fadel M. Megahed 2020 Adelphi University

A Data Analytic Framework For Physical Fatigue Management Using Wearable Sensors, Zahra Sedighi Maman, Ying-Ju Chen, Amir Baghdadi, Seamus Lombardo, Lora A. Cavuoto, Fadel M. Megahed

Mathematics Faculty Publications

The use of expert systems in optimizing and transforming human performance has been limited in practice due to the lack of understanding of how an individual's performance deteriorates with fatigue accumulation, which can vary based on both the worker and the workplace conditions. As a first step toward realizing the human-centered approach to artificial intelligence and expert systems, this paper lays the foundation for a data analytic approach to managing fatigue in physically-demanding workplaces. The proposed framework capitalizes on continuously collected human performance data from wearable sensor technologies, and is centered around four distinct phases of fatigue: (a) detection, where …


A Two-Stage Machine Learning Framework To Predict Heart Transplantation Survival Probabilities Over Time With A Monotonic Probability Constraint, Hamidreza Ahady Dolatsaraa, Ying-Ju (Tessa) Chen, Christy Evans, Ashish Gupta, Fadel M. Megahed 2020 Clark University

A Two-Stage Machine Learning Framework To Predict Heart Transplantation Survival Probabilities Over Time With A Monotonic Probability Constraint, Hamidreza Ahady Dolatsaraa, Ying-Ju (Tessa) Chen, Christy Evans, Ashish Gupta, Fadel M. Megahed

Mathematics Faculty Publications

The overarching goal of this paper is to develop a modeling framework that can be used to obtain personalized, data-driven and monotonically constrained probability curves. This research is motivated by the important problem of improving the predictions for organ transplantation outcomes, which can inform updates made to organ allocation protocols, post-transplantation care pathways, and clinical resource utilization. In pursuit of our overarching goal and motivating problem, we propose a novel two-stage machine learning-based framework for obtaining monotonic probabilities over time. The first stage uses the standard approach of using independent machine learning models to predict transplantation outcomes for each time-period …


Estimation And Inference Under Model Uncertainty, Yizheng Wei 2020 University of South Carolina

Estimation And Inference Under Model Uncertainty, Yizheng Wei

Theses and Dissertations

Chapter 1 of this dissertation proposes a consistent and locally efficient estimator to estimate the model parameters for a logistic mixed effect model with random slopes. Our approach relaxes two typical assumptions: the random effects being normally distributed, and the covariates and random effects being independent of each other. Adhering to these assumptions is particularly difficult in health studies where in many cases we have limited resources to design experiments and gather data in long-term studies, while new findings from other fields might emerge, suggesting the violation of such assumptions. So it is crucial if we could have an estimator …


Categorical And Fuzzy Ensemble-Based Algorithms For Cluster Analysis, Bridget Nicole Manning 2020 University of South Carolina

Categorical And Fuzzy Ensemble-Based Algorithms For Cluster Analysis, Bridget Nicole Manning

Theses and Dissertations

This dissertation focuses on improving multivariate methods of cluster analysis. In Chapter 3 we discuss methods relevant to the categorical clustering of tertiary data while Chapter 4 considers the clustering of quantitative data using ensemble algorithms. Lastly, in Chapter 5, future research plans are discussed to investigate the clustering of spatial binary data.

Cluster analysis is an unsupervised methodology whose results may be influenced by the types of variables recorded on observations. When dealing with the clustering of categorical data, solutions produced may not accurately reflect the structure of the process that generated them. Increased variability within the latent structure …


Incorporation And Measurement Of Uncertainty In Clustered And Spatial Data, Yuan Hong 2020 University of South Carolina

Incorporation And Measurement Of Uncertainty In Clustered And Spatial Data, Yuan Hong

Theses and Dissertations

Analyzing population representative datasets for local estimation and predictions over time is important for monitoring related public health issues, however, there are many statistical challenges associated with such analyses. Mixed effect models are one of the common options which can incorporate time and spatial effect in the model and related inference is well established.

In the first part of this dissertation, to estimate area-level prevalence using individuallevel data, small area estimation (SAE) with post-stratified mixed effect models were used where sampling weights were also incorporated into it. However, if poststratification which requires more computation effort can improve estimation accuracy is …


Contrasting Cumulative Risk And Multiple Individual Risk Models Of The Relationship Between Adverse Childhood Experiences (Aces) And Adult Health Outcomes, Marianna LaNoue, Brandon George, Deborah L Helitzer, Scott W Keith 2020 Thomas Jefferson University

Contrasting Cumulative Risk And Multiple Individual Risk Models Of The Relationship Between Adverse Childhood Experiences (Aces) And Adult Health Outcomes, Marianna Lanoue, Brandon George, Deborah L Helitzer, Scott W Keith

College of Population Health Faculty Papers

BACKGROUND: A very large body of research documents relationships between self-reported Adverse Childhood Experiences (srACEs) and adult health outcomes. Despite multiple assessment tools that use the same or similar questions, there is a great deal of inconsistency in the operationalization of self-reported childhood adversity for use as a predictor variable. Alternative conceptual models are rarely used and very limited evidence directly contrasts conceptual models to each other. Also, while a cumulative numeric 'ACE Score' is normative, there are differences in the way it is calculated and used in statistical models. We investigated differences in model fit and performance between the …


An Explainable And Statistically Validated Ensemble Clustering Model Applied To The Identification Of Traumatic Brain Injury Subgroups, Dacosta Yeboah, Louis Steinmeister, Daniel B. Hier, Bassam Hadi, Donald C. Wunsch, Gayla R. Olbricht, Tayo Obafemi-Ajayi 2020 Mercy Hospital Neurosurgery

An Explainable And Statistically Validated Ensemble Clustering Model Applied To The Identification Of Traumatic Brain Injury Subgroups, Dacosta Yeboah, Louis Steinmeister, Daniel B. Hier, Bassam Hadi, Donald C. Wunsch, Gayla R. Olbricht, Tayo Obafemi-Ajayi

Electrical and Computer Engineering Faculty Research & Creative Works

We present a framework for an explainable and statistically validated ensemble clustering model applied to Traumatic Brain Injury (TBI). The objective of our analysis is to identify patient injury severity subgroups and key phenotypes that delineate these subgroups using varied clinical and computed tomography data. Explainable and statistically-validated models are essential because a data-driven identification of subgroups is an inherently multidisciplinary undertaking. In our case, this procedure yielded six distinct patient subgroups with respect to mechanism of injury, severity of presentation, anatomy, psychometric, and functional outcome. This framework for ensemble cluster analysis fully integrates statistical methods at several stages of …


Logistic Regression Under Sparse Data Conditions, David A. Walker, Thomas J. Smith 2020 Northern Illinois University

Logistic Regression Under Sparse Data Conditions, David A. Walker, Thomas J. Smith

Journal of Modern Applied Statistical Methods

The impact of sparse data conditions was examined among one or more predictor variables in logistic regression and assessed the effectiveness of the Firth (1993) procedure in reducing potential parameter estimation bias. Results indicated sparseness in binary predictors introduces bias that is substantial with small sample sizes, and the Firth procedure can effectively correct this bias.


Estimating A Multilevel Model With Complex Survey Data: Demonstration Using Timss, Julie Lorah 2020 Indiana University Bloomington

Estimating A Multilevel Model With Complex Survey Data: Demonstration Using Timss, Julie Lorah

Journal of Modern Applied Statistical Methods

Analysis of complex survey data is demonstrated for the multilevel model. Description of specific aspects of analysis, including plausible values, sampling weights, and replicate weights is provided. Following this, example TIMSS data and models are described and results are presented.


Cost Effectiveness Of Sample Pooling To Test For Sars-Cov-2, Baha Abdalhamid, Christopher Richard Bilder, Jodi Louise Garrett, Peter Charles Iwen 2020 University of Nebraska Medical Center

Cost Effectiveness Of Sample Pooling To Test For Sars-Cov-2, Baha Abdalhamid, Christopher Richard Bilder, Jodi Louise Garrett, Peter Charles Iwen

Department of Statistics: Faculty Publications

No abstract provided.


Use Of Advanced Statistical Techniques To Predict All-Cause Mortality In The Systolic Blood Pressure Intervention Trial, William Kostis, Javier Cabrera, Chun Pang Lin, John Kostis, Jennifer Wellings, Stavros Zinonos, Jeanne Dobrzynski, Daniel Blickstein 2020 Rutgers Robert Wood Johnson Medical School

Use Of Advanced Statistical Techniques To Predict All-Cause Mortality In The Systolic Blood Pressure Intervention Trial, William Kostis, Javier Cabrera, Chun Pang Lin, John Kostis, Jennifer Wellings, Stavros Zinonos, Jeanne Dobrzynski, Daniel Blickstein

Department of Medicine Faculty Papers

Background: The Systolic Blood Pressure Intervention Trial (SPRINT) was conducted in patients with hypertension and additional risk for cardiovascular disease who were randomized to the intensive blood pressure group targeting systolic blood pressure (SBP) less than 120 mm Hg and to the standard group where the target was less than 140 mm Hg. Analyses were done in the matched group of participants with the same gender, same age (±2 years) and same SBP (±3 mm Hg) at three months of treatment regardless of initial randomization to intensive or standard group (shaded area in Figure 1). Methods and results: During 3.26 …


Concomitant Of Order Statistics From New Bivariate Gompertz Distribution, Sumit Kumar, M. J. S. Khan, Surinder Kumar 2020 Babasaheb Bhimrao Ambedkar University

Concomitant Of Order Statistics From New Bivariate Gompertz Distribution, Sumit Kumar, M. J. S. Khan, Surinder Kumar

Journal of Modern Applied Statistical Methods

For the new bivariate Gompertz distribution, the expression for probability density function (pdf) of rth order statistics and pdf of concomitant arising from rth order statistics are derived. The properties of concomitant arising from the corresponding order statistics are used to derive these results. The exact expression for moment generating function (mgf) of concomitant of order rth statistics is derived. Also, the mean of concomitant arising from rth order statistics is computed using the mgf of concomitant of rth order statistics, and the exact expression for joint density of concomitant of two non-adjacent order statistics …


Time Series Analysis Of Offshore Buoy Light Detection And Ranging (Lidar) Windspeed Data, Aditya Garapati, Charles J. Henderson, Carl Walenciak, Brian T. Waite 2020 Southern Methodist University

Time Series Analysis Of Offshore Buoy Light Detection And Ranging (Lidar) Windspeed Data, Aditya Garapati, Charles J. Henderson, Carl Walenciak, Brian T. Waite

SMU Data Science Review

In this paper, modeling techniques for the forecasting of wind speed using historical values observed by Light Detection and Ranging (LIDAR) sensors in an offshore context are described. Both univariate time series and multivariate time series modeling techniques leveraging meteorological data collected simultaneously with the LIDAR data are evaluated for potential contributions to predictive ability. Accurate and timely ability to predict wind values is essential to the effective integration of wind power into existing power grid systems. It allows for both the management of rapid ramp-up / down of base production capacity due to highly variable wind power inputs and …


Teaching Computational Machine Learning (Without Statistics), Katherine M. Kinnaird 2020 Smith College

Teaching Computational Machine Learning (Without Statistics), Katherine M. Kinnaird

Statistical and Data Sciences: Faculty Publications

This paper presents an undergraduate machine learning course that emphasizes algorithmic understanding and programming skills while assuming no statistical training. Emphasizing the development of good habits of mind, this course trains students to be independent machine learning practitioners through an iterative, cyclical framework for teaching concepts while adding increasing depth and nuance. Beginning with unsupervised learning, this course is sequenced as a series of machine learning ideas and concepts with specific algorithms acting as concrete examples. This paper also details course organization including evaluation practices and logistics.


The Effect Of Corporate Governance Mechanisms And Their Interactions On Earnings Quality, Nasrin Azar 2020 Universiti Malaya

The Effect Of Corporate Governance Mechanisms And Their Interactions On Earnings Quality, Nasrin Azar

Student Works (2020-2029)

Corporate governance (CG) mechanisms play an essential role in improving financial reporting quality, especially earnings quality. Due to corporate failures around the world, there has been a renewed interest in the effect of CG on earnings quality. The primary objectives of this thesis are to (1) investigate the impact of CG characteristics on earnings quality, and (2) examine whether the interactions between CG mechanisms influence earnings quality. Based on the agency and resource dependence theories, this study develops and examines ten main hypotheses (23 sub-hypotheses) to achieve these objectives. This study identifies CG mechanisms such as the board of directors, …


Simple Unequal Allocation Procedure For Ranked Set Sampling With Skew Distributions, Dinesh Bhoj, Girish Chandra 2020 Rutgers University

Simple Unequal Allocation Procedure For Ranked Set Sampling With Skew Distributions, Dinesh Bhoj, Girish Chandra

Journal of Modern Applied Statistical Methods

A practical unbalanced Ranked Set Sampling (RSS) model is proposed to estimate the population mean of positively skewed distributions. The gains in the relative precisions of the population mean based on the proposed model for chosen distributions are uniformly higher than those based on balanced RSS and the t-model proposed in Kaur et al. (1997). The relative precisions of the simple unequal allocation model are, with one exception, better than (s, t)-model which is better than t-model. The relative precision of the proposed model is very close or equal to the optimal Neyman allocation model.


Digital Commons powered by bepress