Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Models (13)
- Statistical Methodology (11)
- Education (10)
- Social and Behavioral Sciences (10)
- Design of Experiments and Sample Surveys (8)
-
- Applied Statistics (7)
- Educational Assessment, Evaluation, and Research (7)
- Mathematics (6)
- Public Affairs, Public Policy and Public Administration (5)
- Social Statistics (5)
- Life Sciences (4)
- Medicine and Health Sciences (4)
- Public Health (4)
- Applied Mathematics (3)
- Biostatistics (3)
- Computer Sciences (3)
- Engineering (3)
- Epidemiology (3)
- Policy Design, Analysis, and Evaluation (3)
- Statistical Theory (3)
- Data Science (2)
- Environmental Sciences (2)
- Probability (2)
- Science and Mathematics Education (2)
- Artificial Intelligence and Robotics (1)
- Business (1)
- Categorical Data Analysis (1)
- Chemical Engineering (1)
- Keyword
-
- Evaluation (3)
- Odds ratio (3)
- Survival analysis (3)
- Bayesian (2)
- Design parameters (2)
-
- Equivalence test (2)
- Experimental design (2)
- Monte Carlo simulation (2)
- Poisson regression (2)
- Statistics (2)
- Stochastic block models (2)
- Variable selection (2)
- Zero-inflated (2)
- Adaptive test (1)
- Additive outliers (1)
- Adjusting for covariates (1)
- Alcoholism (1)
- Amazon product review data (1)
- Asymptotic normality (1)
- Autoregressive (1)
- BRBM GLM (1)
- Bayesian HBR (1)
- Bias correction (1)
- Bifurcating (1)
- Binary data (1)
- Bipartite networks (1)
- Bipartite sbm (1)
- Bivariate (1)
- Bootstrap (1)
- Brain image (1)
- Publication Year
- Publication
- Publication Type
Articles 1 - 30 of 119
Full-Text Articles in Statistics and Probability
Community Detection In Bipartite Networks Using Bipartite Stochastic Block Models With Node-Level Covariates, Geraldine Elaine Percival
Community Detection In Bipartite Networks Using Bipartite Stochastic Block Models With Node-Level Covariates, Geraldine Elaine Percival
Dissertations
In this age, monumental webs of data demands for perpetual cultivation of ways to untangle these webs of information. One of the many curiosities is how to systematically group entities. When clustering, one avenue to take is ascertaining the interconnectedness between the data points, and gauging their influence to each other. This perspective is programmed to model the relationships between the data presented as a network. Many of the methods being used today are algorithm-based, which may pose limitations in understanding and explaining the uncertainty revolving around the data. Hence, it is proposed to steer towards a model-based approach that …
The Heterogeneity Of Healthcare And Insurance In The Midwest, William Robb
The Heterogeneity Of Healthcare And Insurance In The Midwest, William Robb
Honors Theses
Background
Healthcare varies in each state from labor and capital to the laws governing the state and the finances they run under. United States citizens have criticized healthcare and call for change as the United States lacks a universal healthcare system. Even with a universal healthcare system, there are systematic issues and a financial downturn that need to be addressed even if it were to be adopted. Half of the counties in the U.S. see a shortage in labor and hospitals leading to a desert for both cases. Funding through the government is unstable and lacks clear and meaningful payments …
Identifying Draft-Worthy Pitchers: A Statistical Analysis Of Northwoods League Performance Metrics, Stephen Tyrpak
Identifying Draft-Worthy Pitchers: A Statistical Analysis Of Northwoods League Performance Metrics, Stephen Tyrpak
Honors Theses
The Northwoods League (NWL) is a collegiate summer wooden-bat baseball league that provides college athletes with an opportunity to maintain and enhance their skills during the offseason while competing in game-like conditions. Although prior research has explored player performance in other summer leagues, such as the Cape Cod Baseball League, academic investigation into predictors of draft outcomes in the NWL remains limited. This study addresses this gap by evaluating whether pitch-tracking metrics collected via Trackman systems can predict a pitcher’s likelihood of selection in the Major League Baseball (MLB) draft and by identifying which performance characteristics are most strongly associated …
Testing For Broad Alternatives In Stratified Contingency Tables, Nan Mi
Testing For Broad Alternatives In Stratified Contingency Tables, Nan Mi
Dissertations
In medical and social sciences fields, data are measured in terms of discrete categories. The primary question of interest involves the relationship between a set of factors and a set of response variables under studies. Moreover, the distribution of the response variables may be influenced by another set of variables called confounders. The data from such studies are summarized in 3-way tables. The hypothesis we are interested in can be expressed in terms of "no partial association" between the sub-populations and the response levels.
The methods for testing the association or independence in a 2x2 contingency table have been developed, …
Estimating Climate Risk In Financial Markets, Olanrewaju Oluwadamilare Olaniyan
Estimating Climate Risk In Financial Markets, Olanrewaju Oluwadamilare Olaniyan
Masters Theses
The growing impact of climate change on financial markets necessitates a rigorous approach to climate risk assessment. This thesis examines methods for quantifying climate-related financial risks, with a focus on distinguishing climate risk from broader market movements (represented by S&P 500). Using a factor model, we isolate climate risk factors to better understand sector-specific volatility. The insurance sector is used as a proxy for climate risk exposure, given its sensitivity to climate-related losses and regulatory changes. We apply Extreme Value Theory (EVT); the Block Maxima Method and the Peaks Over Threshold Method, to identify excess risk patterns in financial portfolios. …
Nonparametric Finite Mixture Of Ising Graphical Models, Manal Hamadi Alloqmani
Nonparametric Finite Mixture Of Ising Graphical Models, Manal Hamadi Alloqmani
Dissertations
Statistical applications in fields such as bioinformatics, genomics, speech processing, image processing, and communications often involve large-scale models in which thousands or millions of random variables are linked in complex ways. Graphical models provide a general methodology for approaching these problems, and indeed many of the models developed by researchers in these applied fields are instances of the general graphical model formalism. This formalism gives a nice framework for capturing complex dependencies among the random variables and building a large-scale model for high-dimensional data. Recently, high-dimensional data are more assumed to come from one population and follow a parametric or …
Quasi – Monte Carlo Estimation For Functional Generalized Linear Mixed Models, Ruvini Jayamaha
Quasi – Monte Carlo Estimation For Functional Generalized Linear Mixed Models, Ruvini Jayamaha
Waldo Library Student Exhibits
Functional Data Analysis (FDA) is a topic of growing interest in the Statistics community. The data in FDA are smooth curves or surfaces in time or space which can be conceptualized as functions.
We propose a Functional Generalized Linear Mixed Model (FGLMM) to fit EEG data and estimate the parameters using Quasi-Monte Carlo Method.
This proposed model deals with non-Gaussian scalar response, functional predictor, and random effects. We relax the assumption of link and variance functions.
Estimating And Applying Parameters Necessary To Plan Cluster Randomized Trials (Crts) And Multisite Cluster Randomized Trials (Mscrts), Dea Mulolli
Dissertations
Cluster randomized trials (CRTs) are commonly used to study the effectiveness of educational interventions. During the design phase of a study, it is critical for researchers to ensure their studies are adequately powered to detect meaningful treatment effects, including both main and moderator effects. Designing CRTs with adequate power to detect main and moderator effects requires accurate estimates of design parameters. This research aims to advance the literature on design parameters for power analyses, specifically focusing on empirical estimates of intraclass correlations (ICCs). The work consists of three research papers that examine the role of including the teacher level in …
A Contingency Table Alternative To Poisson Regression In Comparing The Frequency Distributions Of Two Populations, Sandra Tay
Dissertations
When testing the conditional independence between a binary outcome and a binary treatment indicator, conditioned on a categorical variable with k levels, typically represented by a K × 2 frequency table, researchers often turn to Poisson regression and the Cochran-Mantel-Haenszel (CMH) test. However, a common challenge encountered in these analyses is the presence of treatment effect heterogeneity. Introducing an interaction term between treatment indicators and effect modifiers in log-linear regression offers potential solutions, yet the equidispersion assumption of Poisson regression remains problematic. On the other hand, the CMH test assumes similar treatment effects across all strata, disregarding potential variations among …
An Experimental Study Of Supervised Machine Learning Techniques For Minor Class Prediction Utilizing Kernel Density Estimation: Factors Impacting Model Performance, Abdullah Mana Alfarwan
An Experimental Study Of Supervised Machine Learning Techniques For Minor Class Prediction Utilizing Kernel Density Estimation: Factors Impacting Model Performance, Abdullah Mana Alfarwan
Dissertations
This dissertation examined classification outcome differences among four popular individual supervised machine learning (ISML) models (logistic regression, decision tree, support vector machine, and multilayer perceptron) when predicting minor class membership within imbalanced datasets. The study context and the theoretical population sampled focus on one aspect of the larger problem of student retention and dropout prediction in higher education (HE): identification.
This study differs from current literature by implementing an experimental design approach with simulated student data that closely mirrors HE situational and student data. Specifically, this study tested the predictive ability of the four ISML classification models (CLS) under experimentally …
Alternative Adjacency Matrices And Spatial Analysis, Jaeseong Hwang
Alternative Adjacency Matrices And Spatial Analysis, Jaeseong Hwang
Dissertations
Spatial analysis is essential for comprehending the spatial distribution of diseases and various phenomena across geographic regions. This study investigates the utilization of alternative adjacency matrices in spatial analysis, with a specific focus on implementing Poisson regression models. This study intricately explores the methodology behind constructing alternative weight matrices, specifying weight matrices, and comparing the performance of Poisson models using five different weight matrices.
The popular Poisson model model is described, and five different definitions of weight matrices are defined, which are the following: binary weight matrix, inverse distance weight matrix using Euclidean distance, Graph distance matrix, Path matrix, and …
Quasi-Monte Carlo Estimation For Functional Generalized Linear Mixed Models., Ruvini Kumari Jayamaha Hitihamilage
Quasi-Monte Carlo Estimation For Functional Generalized Linear Mixed Models., Ruvini Kumari Jayamaha Hitihamilage
Dissertations
Functional Data Analysis (FDA) is a topic of growing interest in the statistics community and is applied in a wide range of fields such as Anthropology, Epidemiology, Meteorology, Neurology and Engineering. The data in FDA are smooth curves or surfaces in time or space which can be conceptualized as functions. Because of the smooth nature of the data and the measurements are highly correlated, making the classical methods such as univariate or multivariate analysis are infeasible for such data. Functional data Analysis (FDA) deals with these kinds of more detailed, complex, and structured data.
In this dissertation, we propose a …
Establishing Practical Equivalence Of Factor Loadings In Multigroup Confirmatory Factor Analysis, Christopher Edward Shank
Establishing Practical Equivalence Of Factor Loadings In Multigroup Confirmatory Factor Analysis, Christopher Edward Shank
Dissertations
This dissertation compares the performance of equivalence test (EQT) and null hypothesis test (NHT) procedures for identifying invariant and noninvariant factor loadings under a range of experimental manipulations. EQT is the statistically appropriate approach when the research goal is to find evidence of group similarity rather than group difference; despite this, the conventional approach to measurement invariance analysis relies upon NHT. EQT has proved effective for invariance detection using global model-data fit statistics in simulated and real-world data (Counsell et al., 2020) but its use in partial measurement invariance (PMI) analysis for evaluation of factor loading differences between groups has …
Development Of An App For The Kalamazoo Nature Center, Ernest Au
Development Of An App For The Kalamazoo Nature Center, Ernest Au
Honors Theses
Kalamazoo Nature Center (KNC), which has been recognized by its peers as one of the top nature centers in the country, is home to over 14 miles of hiking trails winding through woods, wetlands, and prairies. There are numerous places/plots in KNC that have an interesting and impressive history besides being home to a variety of animals and hundreds of wildflowers and other plant life. To improve the visitor’s experience at KNC, we will design a software app via the senior capstone project at the department of Computer Science at WMU. As the first step towards establishing a reference model …
Nonparametric Tests For Replicated Latin Squares, Joseph Yang
Nonparametric Tests For Replicated Latin Squares, Joseph Yang
Dissertations
Two classes of nonparametric procedures for a replicated Latin square design that test for both general and increasing alternatives are developed. The two classes of procedures are similar in the sense that both transform the data so that existing well-known tests for randomized complete block designs can be utilized. On the other hand, the two classes differ in the way that the data is transformed - one class essentially aggregates the data while the other class aligns the data. Within these contexts, the exact distributions and asymptotic distributions are discussed, when applicable. The exact distributions are easily computed using the …
Evaluating The Performance Of Estimators In Sem And Irt With Ordinal Variables, Bo Klauth
Evaluating The Performance Of Estimators In Sem And Irt With Ordinal Variables, Bo Klauth
Dissertations
In conducting confirmatory factor analysis with ordered response items, the literature suggests that when the number of responses is five and item skewness (IS) is approximately normal, researchers can employ maximum likelihood with robust standard errors (MLR). However, MLR can yield biased factor loadings (FL) and FL standard errors (FLSE) when the variables are ordinal. Other estimators are available. Unweighted least squares and weighted least squares with adjusted mean and variance (ULSMV and WLSMV) are known as the estimators for CFA with ordinal variables (CFA-OV). Another estimator, marginal maximum likelihood (MML), is used in the item response theory (IRT), specifically …
Functional Generalized Linear Mixed Models, Harmony Luce
Functional Generalized Linear Mixed Models, Harmony Luce
Dissertations
With the advancements in data collection technologies, researchers in various fields such as epidemiology, chemometrics, and environmental science face the challenges of obtaining useful information from more detailed, complex, and intricately-structured data. Since the existing methods often are not suitable for such data, new statistical methods are developed to accommodate the complicated data structures.
As a part of such efforts, this dissertation proposes Functional Generalized Linear Mixed Model (FGLMM), which extends classical generalized linear mixed models to include functional covariates. Functional Data Analysis (FDA) is a rapidly developing area of statistics for data which can be naturally viewed as smooth …
Learning Finite Mixture Of Ising Graphical Models, Chong Gu
Learning Finite Mixture Of Ising Graphical Models, Chong Gu
Dissertations
The Ising model is valuable in examining complex interactions within a system, but its estimation is challenging. In this work, we proposed penalized likelihood procedures to infer conditional dependence structure when observed data come from heterogeneous resources in high-dimensional setting. The proposed method can be efficiently implemented by taking advantage of coordinate-ascent, minorization–maximization principles and EM algorithm. A BIC-type criterion will be utilized for the selection of the tuning parameter in the penalized likelihood approaches. The effectiveness of the proposed method is supported by simulation studies and a real-world example.
Statistical Clustering Of Networks With Additional Information, Paul Atandoh
Statistical Clustering Of Networks With Additional Information, Paul Atandoh
Dissertations
As the online market grows rapidly, many companies and researchers are interested in analyzing product review dataset which includes ratings and text review data. In the first project, we mainly focus on analyzing the text review data. In the current literature, it is common to use only text analysis tools to analyze review dataset. But in our work, we propose a method that utilizes both a text analysis method such as topic modeling and a statistical network model to build network among individuals and find interesting communities. We introduce a promising framework that incorporates topic modeling technique to define the …
High-Dimensional Variable Selection Via Knockoffs Using Gradient Boosting, Amr Essam Mohamed
High-Dimensional Variable Selection Via Knockoffs Using Gradient Boosting, Amr Essam Mohamed
Dissertations
As data continue to grow rapidly in size and complexity, efficient and effective statistical methods are needed to detect the important variables/features. Variable selection is one of the most crucial problems in statistical applications. This problem arises when one wants to model the relationship between the response and the predictors. The goal is to reduce the number of variables to a minimal set of explanatory variables that are truly associated with the response of interest to improve the model accuracy. Effectively choosing the true influential variables and controlling the False Discovery Rate (FDR) without sacrificing power has been a challenge …
Active Vs Passive Investing, Garret Buchheit
Active Vs Passive Investing, Garret Buchheit
Honors Theses
With the increased popularity of passive investing, the long-term investment success of active management is being questioned more frequently. For this reason, this research seeks to find whether actively managed funds produce sufficient returns that cover the fees and management costs associated with them. A comparative analysis was made with 5401 actively managed U.S. mutual funds and several common market indices over three, five, and ten-year time spans ranging from 2012 to 2021. Additionally, an analysis was made comparing active and passive management in the recessionary period of 2007 to 2009. Finally, analysis was conducted on annual holdings turnover rates …
Mixture Of Functional Graphical Models, Qihai Liu
Mixture Of Functional Graphical Models, Qihai Liu
Dissertations
With the development of data collection technologies that use powerful monitoring devices and computational tools, many scientific fields are now obtaining more detailed and more complicatedly structured data, e.g., functional data. This leads to increasing challenges of extracting information from the large complex data. Making use of these data to gain insight into complex phenomena requires characterizing the relationships among a large number of functional variables. Functional data analysis (FDA) is a rapidly developing area of statistics for data which can be naturally viewed as a smooth curve or function. It is a method that changes the frame of data …
Estimation Of Odds Ratio In 2 X 2 Contingency Tables With Small Cell Counts, Guohao Zhu
Estimation Of Odds Ratio In 2 X 2 Contingency Tables With Small Cell Counts, Guohao Zhu
Dissertations
This study is focusing on properties of estimators of odds ratio or its logarithm in case of 2x2 tables with small counts. The odds ratio represents the odds that an outcome of interest will occur given a particular exposure, compared to the odds of the outcome occurring in the absence of that exposure. Both parameters are often used to quantify the strength of association of two binary variables and are common measurements reported in case-control, cohort, and cross-sectional studies.
Because of their wide applicability, both parameters, odds ratio, and its logarithm, have been intensively studied in the literature. However, most …
On Simes’S Second Conjecture: An Extended Single-Step Simes Test Procedure For Multiple Testing, Matthew G. Hudson
On Simes’S Second Conjecture: An Extended Single-Step Simes Test Procedure For Multiple Testing, Matthew G. Hudson
Dissertations
One of the major concerns with multiple tests of significance is controlling the family wise error rate. Various methods have been developed to ensure that the false positive rate be maintained at some prespecified level. One of the most well know being the Bonferroni procedure. Simes presented an improved Bonferroni procedure for testing the global hypothesis that is more powerful and less conservative, especially with positively correlated tests. While Simes’s procedure is more powerful, it does not allow for making inferences on the individual hypotheses. However, the Simes procedure has since become the foundation of many p-value based multiple testing …
Statistical Properties And Applications Of Press Statistic, Ida Marie Alcantara
Statistical Properties And Applications Of Press Statistic, Ida Marie Alcantara
Dissertations
The most popularly used statistic R2 has a fundamental weakness in model building: it favors adding more predictors to the model because R2 can only increase. In effect, the additional predictors start fitting the noise in data. Other criterion in selecting a regression model such as R2 adj , AIC, SBC, and Mallow’s Cp does not guarantee the model selected will also make better prediction of future values. To avoid this, data scientists withhold a percentage of the data for validation purposes. The PRESS statistic does something similar by withholding each observation in calculating its own …
Elementary/Middle School Pre-Service Teachers’ Understanding Of Variability And The Use Of Dynamical Statistical Software, Yaomingxin Lu
Elementary/Middle School Pre-Service Teachers’ Understanding Of Variability And The Use Of Dynamical Statistical Software, Yaomingxin Lu
Research and Creative Activities Poster Day
A primary purpose of the study was to examine the effects of using dynamical statistical software (DSS) on prospective teachers’ (PSTs) understanding of statistical concepts, especially variability. Data were collected from PSTs enrolled in a probability and statistics course designed for prospective K-8 teachers. After initial analysis of the data using coding and classification schemes, we found the need to develop a more targeted framework to analyze students’ different levels of understanding. The variability framework (Garfield & Ben-Zvi, 2005) and the Structure of Observed Learning Outcomes (SOLO) taxonomy were then used in combination to develop a revised framework in order …
Statistical Models For Correlated Data, Xiaomeng Niu
Statistical Models For Correlated Data, Xiaomeng Niu
Dissertations
Correlated data arise frequently in many studies where multiple response variables or repeatedly measured responses within subjects are correlated. My dissertation topic lies broadly in developing various statistical methodologies for correlated types of data such as longitudinal data, clustered data, and multivariate data.
Multiple response variables might be relevant within subjects. A univariate procedure fitting each response separately does not take into account the correlation among responses. To improve estimation efficiency for the regression parameter, this study proposes two estimation procedures by accommodating correlations among the response variables. The proposed procedures do not require knowledge of the true correlation structure …
Denoising Large Neuroimage Mri Data Using Spatial Random Effect Models, Leonard Chukuma Johnson
Denoising Large Neuroimage Mri Data Using Spatial Random Effect Models, Leonard Chukuma Johnson
Dissertations
Spatial smoothing in Magnetic Resonance image (MRI) involves applying a filter to remove high frequency information and consequently improves signal-to-noise ratio that can greatly aid neurosurgeons in pre-surgical planning stages of tumor resection. This immensely reduces the time spent on Electrical stimulation mapping (ESM) prior to surgery. MRI's three-dimensional data provides voxel intensities with complex spatial relationship. The standard de facto spatial smoothing method, Gaussian Kernel smoothing, is satisfactory since a uniform smoothing is done for the whole brain. Secondly, the kernel smoothing technique assumes normality for the voxel intensity, but there is ample evidence in current research that indicates …
Statistical Properties Of Population Stability Index, Bilal Yurdakul
Statistical Properties Of Population Stability Index, Bilal Yurdakul
Dissertations
Population stability is an important concept in model management. It is crucial to monitor whether the current population has changed from the population used during development of a model. For example, has the distribution of credit scores changed, and is the existing credit score model still valid? Population change may occur for many reasons–change in the economic environment, strategic change in the business, policy changes within the company, or changes in regulatory environment.
The population stability index (PSI) is a statistic that measures how much a variable has shifted over time, and is used to monitor applicability of a statistical …
Statistical And Clinical Equivalence Of Measurements, Puntipa Wanitjirattikal
Statistical And Clinical Equivalence Of Measurements, Puntipa Wanitjirattikal
Dissertations
This study proposes a test for statistical equivalence of two measurements. Typically, a new measurement process Υ is compared to an existing or standard measurement process Χ. We are assuming that Χ and Υ are measurements on the same scale. The paired t-test may be used to check for significant difference between (Χ, Υ) pairs. However, the paired t-test is intended to detect shift-type relationships of the form Υ=Χ+δ1 and may have low power for scale-type relations of the form Υ=γΧ.
We propose a test that has reasonable power to …