Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Models (18)
- Social and Behavioral Sciences (13)
- Statistical Methodology (13)
- Education (10)
- Design of Experiments and Sample Surveys (9)
-
- Computer Sciences (8)
- Mathematics (8)
- Applied Statistics (7)
- Data Science (7)
- Educational Assessment, Evaluation, and Research (7)
- Applied Mathematics (6)
- Environmental Sciences (6)
- Life Sciences (6)
- Engineering (5)
- Medicine and Health Sciences (5)
- Social Statistics (5)
- Biostatistics (4)
- Categorical Data Analysis (4)
- Multivariate Analysis (4)
- Public Affairs, Public Policy and Public Administration (4)
- Statistical Theory (4)
- Artificial Intelligence and Robotics (3)
- Business (3)
- Longitudinal Data Analysis and Time Series (3)
- Probability (3)
- Biometry (2)
- Chemical Engineering (2)
- Earth Sciences (2)
- Institution
- Keyword
-
- Machine learning (6)
- Survival analysis (4)
- Evaluation (3)
- Odds ratio (3)
- Archimedean copula (2)
-
- Bayesian (2)
- Classification (2)
- Deep learning (2)
- Design parameters (2)
- Equivalence test (2)
- Experimental design (2)
- Monte Carlo simulation (2)
- Poisson regression (2)
- Statistics (2)
- Stochastic block models (2)
- Variable selection (2)
- Zero-inflated (2)
- Adaptive test (1)
- Additive outliers (1)
- Adjusting for covariates (1)
- Alcoholism (1)
- Amazon product review data (1)
- Artificial intelligence (1)
- Asymmetric (1)
- Asymptotic normality (1)
- Atrial fibrillation in Hispanics (1)
- Attitudes about disability (1)
- Autoregressive (1)
- BRBM GLM (1)
- Bathymetry (1)
Articles 1 - 30 of 121
Full-Text Articles in Statistics and Probability
Geometric Convergence And State-Space Decompositions For Stochastic Gradient Descent Markov Chains, Philip Zaleski
Geometric Convergence And State-Space Decompositions For Stochastic Gradient Descent Markov Chains, Philip Zaleski
Dissertations
No abstract provided.
Machine Learning For Predictive Energy And Emissions Modeling Of Vehicles And Power Grids In The United States, S M Tanvir Faysal Alam Chowdhoury
Machine Learning For Predictive Energy And Emissions Modeling Of Vehicles And Power Grids In The United States, S M Tanvir Faysal Alam Chowdhoury
Dissertations
The environmental benefits of electric vehicle (EV) adoption depend on more than replacing internal combustion engine vehicles with electric powertrains. EV adoption reshapes electricity demand, interacts with regional generation mixes, and influences travel behavior and congestion, creating a coupled transportation-energy system in which vehicle and power-plant emissions must be evaluated together. This dissertation develops machine-learning frameworks for predicting energy consumption and emissions from vehicles and power grids under rising EV adoption. The first component forecasts grid emissions from EV charging. Using simulation data from NREL's Cambium database, a Prophet-based time-series framework predicts carbon dioxide, nitrous oxide, and methane emission rates …
Community Detection In Bipartite Networks Using Bipartite Stochastic Block Models With Node-Level Covariates, Geraldine Elaine Percival
Community Detection In Bipartite Networks Using Bipartite Stochastic Block Models With Node-Level Covariates, Geraldine Elaine Percival
Dissertations
In this age, monumental webs of data demands for perpetual cultivation of ways to untangle these webs of information. One of the many curiosities is how to systematically group entities. When clustering, one avenue to take is ascertaining the interconnectedness between the data points, and gauging their influence to each other. This perspective is programmed to model the relationships between the data presented as a network. Many of the methods being used today are algorithm-based, which may pose limitations in understanding and explaining the uncertainty revolving around the data. Hence, it is proposed to steer towards a model-based approach that …
Variable Importance, Knockoff Filters, And Improving False Discovery And False Negative Rates, Nicholas Ehlman
Variable Importance, Knockoff Filters, And Improving False Discovery And False Negative Rates, Nicholas Ehlman
Dissertations
Tree ensemble methods such as Random Forests and Boosted Trees have introduced a range of variable importance statistics, offering powerful tools for feature selection. The advent of knockoff filters marked a significant advancement by combining the use of these variable importance statistics with the ability to control the False Discovery Rate (FDR). However, achieving a low FDR frequently comes at the cost of a high False Negative Rate (FNR), limiting the power of such approaches. In this work, we propose a novel method for leveraging knockoff variables to keep both FDR and FNR low. While this method does not have …
Testing For Broad Alternatives In Stratified Contingency Tables, Nan Mi
Testing For Broad Alternatives In Stratified Contingency Tables, Nan Mi
Dissertations
In medical and social sciences fields, data are measured in terms of discrete categories. The primary question of interest involves the relationship between a set of factors and a set of response variables under studies. Moreover, the distribution of the response variables may be influenced by another set of variables called confounders. The data from such studies are summarized in 3-way tables. The hypothesis we are interested in can be expressed in terms of "no partial association" between the sub-populations and the response levels.
The methods for testing the association or independence in a 2x2 contingency table have been developed, …
Nonparametric Finite Mixture Of Ising Graphical Models, Manal Hamadi Alloqmani
Nonparametric Finite Mixture Of Ising Graphical Models, Manal Hamadi Alloqmani
Dissertations
Statistical applications in fields such as bioinformatics, genomics, speech processing, image processing, and communications often involve large-scale models in which thousands or millions of random variables are linked in complex ways. Graphical models provide a general methodology for approaching these problems, and indeed many of the models developed by researchers in these applied fields are instances of the general graphical model formalism. This formalism gives a nice framework for capturing complex dependencies among the random variables and building a large-scale model for high-dimensional data. Recently, high-dimensional data are more assumed to come from one population and follow a parametric or …
Determination Of Electrochemical Parameters For Predicting Reaction Mechanism And Algorithmic Approaches To Pain Assessment, Huize Xue
Dissertations
This dissertation introduces novel advancements in electrochemical kinetics and pain assessment, structured into two main parts. The first part focuses on the comprehensive analysis of the kinetic and mechanistic aspects of electrochemical reactions, utilizing a combination of experimental techniques and simulation methods. A new software tool, Envismetrics, was developed using Python to facilitate the analysis of complex electrochemical data, including cyclic voltammetry (CV), chronoamperometry (CA), and hydrodynamic voltammetry (HDV). The software was rigorously tested and validated with well-characterized redox systems such as the ferricyanide/ferrocyanide couple, dimethylamine borane (DMAB), and Per- and Polyfluoroalkyl Substances (PFAS). It was successfully used to determine …
Machine Learning Methods For Pattern Recognition Analysis Of Genomic And Molecular Data, Kuang Du
Machine Learning Methods For Pattern Recognition Analysis Of Genomic And Molecular Data, Kuang Du
Dissertations
While immune therapies achieve remarkable success in treating various cancers, only a subset of patients achieves a durable clinical response, and many exhibit innate or acquired resistance. Precision medicine aims to tailor treatments to individual patients based on specific biological markers, ensuring that each patient receives the therapy most likely to be effective. Predictive biomarkers and gene signatures offer potential for more personalized treatment strategies by identifying patients likely to benefit. Recent studies suggest that gene signatures, comprising sets of genes, hold predictive value for certain clinical variables. Typically derived from biological expert knowledge, these signatures demonstrate substantial predictive potential, …
Multi-Label Classification Using Conformal Prediction, Chhavi Tyagi
Multi-Label Classification Using Conformal Prediction, Chhavi Tyagi
Dissertations
In many machine learning applications, such as image tagging, document classi-fication, and medical diagnosis, a data instance can be associated with multiple classes in parallel so that each instance is associated with multiple response variables simultaneously defining multi-label classification. Standard multi-label classification methods that provide point predictions have been developed. They lack in quantifying the uncertainty of predictions. These methods also lack in accounting for label dependencies and are very computationally expensive. This dissertation develops two methods of multi-label classification using conformal prediction that quantify the uncertainty of predictions. Chapter 1 introduces notations and tools that have been used in …
Large Deviation Theory In Stochastic Processes: Applications To Biological Modeling, Moshe C. Silverstein
Large Deviation Theory In Stochastic Processes: Applications To Biological Modeling, Moshe C. Silverstein
Dissertations
This dissertation delves into developing and applying stochastic models to analyze complex biological systems. It leverages Large Deviation Theory (LDT) to gain insights into these systems, focusing on two key examples: neural networks and calcium signaling dynamics. Traditional deterministic methods frequently fail to capture biological processes' randomness and inherent variability. Meanwhile, many stochastic approaches struggle to be mathematically tractable or provide accessible insights. The approach introduced in this study provides rigorous mathematical frameworks to enhance understanding of these stochastic behaviors while remaining tractable and insightful.
A stochastic model for a random biological neural network is constructed that addresses the dependencies …
Estimating And Applying Parameters Necessary To Plan Cluster Randomized Trials (Crts) And Multisite Cluster Randomized Trials (Mscrts), Dea Mulolli
Dissertations
Cluster randomized trials (CRTs) are commonly used to study the effectiveness of educational interventions. During the design phase of a study, it is critical for researchers to ensure their studies are adequately powered to detect meaningful treatment effects, including both main and moderator effects. Designing CRTs with adequate power to detect main and moderator effects requires accurate estimates of design parameters. This research aims to advance the literature on design parameters for power analyses, specifically focusing on empirical estimates of intraclass correlations (ICCs). The work consists of three research papers that examine the role of including the teacher level in …
A Contingency Table Alternative To Poisson Regression In Comparing The Frequency Distributions Of Two Populations, Sandra Tay
Dissertations
When testing the conditional independence between a binary outcome and a binary treatment indicator, conditioned on a categorical variable with k levels, typically represented by a K × 2 frequency table, researchers often turn to Poisson regression and the Cochran-Mantel-Haenszel (CMH) test. However, a common challenge encountered in these analyses is the presence of treatment effect heterogeneity. Introducing an interaction term between treatment indicators and effect modifiers in log-linear regression offers potential solutions, yet the equidispersion assumption of Poisson regression remains problematic. On the other hand, the CMH test assumes similar treatment effects across all strata, disregarding potential variations among …
An Experimental Study Of Supervised Machine Learning Techniques For Minor Class Prediction Utilizing Kernel Density Estimation: Factors Impacting Model Performance, Abdullah Mana Alfarwan
An Experimental Study Of Supervised Machine Learning Techniques For Minor Class Prediction Utilizing Kernel Density Estimation: Factors Impacting Model Performance, Abdullah Mana Alfarwan
Dissertations
This dissertation examined classification outcome differences among four popular individual supervised machine learning (ISML) models (logistic regression, decision tree, support vector machine, and multilayer perceptron) when predicting minor class membership within imbalanced datasets. The study context and the theoretical population sampled focus on one aspect of the larger problem of student retention and dropout prediction in higher education (HE): identification.
This study differs from current literature by implementing an experimental design approach with simulated student data that closely mirrors HE situational and student data. Specifically, this study tested the predictive ability of the four ISML classification models (CLS) under experimentally …
Alternative Adjacency Matrices And Spatial Analysis, Jaeseong Hwang
Alternative Adjacency Matrices And Spatial Analysis, Jaeseong Hwang
Dissertations
Spatial analysis is essential for comprehending the spatial distribution of diseases and various phenomena across geographic regions. This study investigates the utilization of alternative adjacency matrices in spatial analysis, with a specific focus on implementing Poisson regression models. This study intricately explores the methodology behind constructing alternative weight matrices, specifying weight matrices, and comparing the performance of Poisson models using five different weight matrices.
The popular Poisson model model is described, and five different definitions of weight matrices are defined, which are the following: binary weight matrix, inverse distance weight matrix using Euclidean distance, Graph distance matrix, Path matrix, and …
Quasi-Monte Carlo Estimation For Functional Generalized Linear Mixed Models., Ruvini Kumari Jayamaha Hitihamilage
Quasi-Monte Carlo Estimation For Functional Generalized Linear Mixed Models., Ruvini Kumari Jayamaha Hitihamilage
Dissertations
Functional Data Analysis (FDA) is a topic of growing interest in the statistics community and is applied in a wide range of fields such as Anthropology, Epidemiology, Meteorology, Neurology and Engineering. The data in FDA are smooth curves or surfaces in time or space which can be conceptualized as functions. Because of the smooth nature of the data and the measurements are highly correlated, making the classical methods such as univariate or multivariate analysis are infeasible for such data. Functional data Analysis (FDA) deals with these kinds of more detailed, complex, and structured data.
In this dissertation, we propose a …
Establishing Practical Equivalence Of Factor Loadings In Multigroup Confirmatory Factor Analysis, Christopher Edward Shank
Establishing Practical Equivalence Of Factor Loadings In Multigroup Confirmatory Factor Analysis, Christopher Edward Shank
Dissertations
This dissertation compares the performance of equivalence test (EQT) and null hypothesis test (NHT) procedures for identifying invariant and noninvariant factor loadings under a range of experimental manipulations. EQT is the statistically appropriate approach when the research goal is to find evidence of group similarity rather than group difference; despite this, the conventional approach to measurement invariance analysis relies upon NHT. EQT has proved effective for invariance detection using global model-data fit statistics in simulated and real-world data (Counsell et al., 2020) but its use in partial measurement invariance (PMI) analysis for evaluation of factor loading differences between groups has …
Atrial Fibrillation Management In Hispanic Adults, Tania Borja
Atrial Fibrillation Management In Hispanic Adults, Tania Borja
Dissertations
Background: Research has found atrial fibrillation (AF) to be the primary or a contributing cause of death on 183,321 death certificates, and an underlying cause of death for 26,535 Americans in 2019. Findings indicate an increased AF diagnosis in White people compared to racial and ethnic minorities, contrasting widespread findings of increased prevalence of cardiovascular disease and ischemic strokes in minorities. Significant disparities—by race and socioeconomic status in disease distribution and access to testing and lifesaving treatments—have been documented, specifically associated with social determinants of health (SDOH); i.e., the conditions in which people are born, grow, live, work, and age. …
Nonparametric Tests For Replicated Latin Squares, Joseph Yang
Nonparametric Tests For Replicated Latin Squares, Joseph Yang
Dissertations
Two classes of nonparametric procedures for a replicated Latin square design that test for both general and increasing alternatives are developed. The two classes of procedures are similar in the sense that both transform the data so that existing well-known tests for randomized complete block designs can be utilized. On the other hand, the two classes differ in the way that the data is transformed - one class essentially aggregates the data while the other class aligns the data. Within these contexts, the exact distributions and asymptotic distributions are discussed, when applicable. The exact distributions are easily computed using the …
Evaluating The Performance Of Estimators In Sem And Irt With Ordinal Variables, Bo Klauth
Evaluating The Performance Of Estimators In Sem And Irt With Ordinal Variables, Bo Klauth
Dissertations
In conducting confirmatory factor analysis with ordered response items, the literature suggests that when the number of responses is five and item skewness (IS) is approximately normal, researchers can employ maximum likelihood with robust standard errors (MLR). However, MLR can yield biased factor loadings (FL) and FL standard errors (FLSE) when the variables are ordinal. Other estimators are available. Unweighted least squares and weighted least squares with adjusted mean and variance (ULSMV and WLSMV) are known as the estimators for CFA with ordinal variables (CFA-OV). Another estimator, marginal maximum likelihood (MML), is used in the item response theory (IRT), specifically …
Functional Generalized Linear Mixed Models, Harmony Luce
Functional Generalized Linear Mixed Models, Harmony Luce
Dissertations
With the advancements in data collection technologies, researchers in various fields such as epidemiology, chemometrics, and environmental science face the challenges of obtaining useful information from more detailed, complex, and intricately-structured data. Since the existing methods often are not suitable for such data, new statistical methods are developed to accommodate the complicated data structures.
As a part of such efforts, this dissertation proposes Functional Generalized Linear Mixed Model (FGLMM), which extends classical generalized linear mixed models to include functional covariates. Functional Data Analysis (FDA) is a rapidly developing area of statistics for data which can be naturally viewed as smooth …
Learning Finite Mixture Of Ising Graphical Models, Chong Gu
Learning Finite Mixture Of Ising Graphical Models, Chong Gu
Dissertations
The Ising model is valuable in examining complex interactions within a system, but its estimation is challenging. In this work, we proposed penalized likelihood procedures to infer conditional dependence structure when observed data come from heterogeneous resources in high-dimensional setting. The proposed method can be efficiently implemented by taking advantage of coordinate-ascent, minorization–maximization principles and EM algorithm. A BIC-type criterion will be utilized for the selection of the tuning parameter in the penalized likelihood approaches. The effectiveness of the proposed method is supported by simulation studies and a real-world example.
Special Education: Inclusion And Exclusion In The K-12 U.S. Educational System, Erik Brault
Special Education: Inclusion And Exclusion In The K-12 U.S. Educational System, Erik Brault
Dissertations
The U.S. Department of Education defines students with disabilities as those having a physical or mental impairment that substantially limits one or more life activities. Previous research has found that students with disabilities placed in inclusive environments perform better academically and socially compared to students with disabilities who are placed in segregated environments. Yet, we know that inclusion in K-12 general education classrooms across the country is not consistently implemented.
The purpose of this study was to better understand the effects, if any, of general education high school teachers’ personal and professional experiences and knowledge on their attitudes toward educating …
Utilizing New Technologies To Measure Therapy Effectiveness For Mental And Physical Health, Jonathan Ossie
Utilizing New Technologies To Measure Therapy Effectiveness For Mental And Physical Health, Jonathan Ossie
Dissertations
Mental health is quickly becoming a major policy concern, with recent data reporting increasing and disproportionately worse mental health outcomes, including anxiety, depression, increased substance abuse, and elevated suicidal ideation. One specific population that is especially high risk for these issues is the military community because military conflict, deployment stressors, and combat exposure contribute to the risk of mental health problems.
Although several pharmacological approaches have been employed to combat this epidemic, their efficacy is mixed at best, which has led to novel nonpharmacological approaches. One such approach is Operation Surf, a nonprofit that provides nature-based programs advocating the restorative …
Statistical Clustering Of Networks With Additional Information, Paul Atandoh
Statistical Clustering Of Networks With Additional Information, Paul Atandoh
Dissertations
As the online market grows rapidly, many companies and researchers are interested in analyzing product review dataset which includes ratings and text review data. In the first project, we mainly focus on analyzing the text review data. In the current literature, it is common to use only text analysis tools to analyze review dataset. But in our work, we propose a method that utilizes both a text analysis method such as topic modeling and a statistical network model to build network among individuals and find interesting communities. We introduce a promising framework that incorporates topic modeling technique to define the …
High-Dimensional Variable Selection Via Knockoffs Using Gradient Boosting, Amr Essam Mohamed
High-Dimensional Variable Selection Via Knockoffs Using Gradient Boosting, Amr Essam Mohamed
Dissertations
As data continue to grow rapidly in size and complexity, efficient and effective statistical methods are needed to detect the important variables/features. Variable selection is one of the most crucial problems in statistical applications. This problem arises when one wants to model the relationship between the response and the predictors. The goal is to reduce the number of variables to a minimal set of explanatory variables that are truly associated with the response of interest to improve the model accuracy. Effectively choosing the true influential variables and controlling the False Discovery Rate (FDR) without sacrificing power has been a challenge …
Mixture Of Functional Graphical Models, Qihai Liu
Mixture Of Functional Graphical Models, Qihai Liu
Dissertations
With the development of data collection technologies that use powerful monitoring devices and computational tools, many scientific fields are now obtaining more detailed and more complicatedly structured data, e.g., functional data. This leads to increasing challenges of extracting information from the large complex data. Making use of these data to gain insight into complex phenomena requires characterizing the relationships among a large number of functional variables. Functional data analysis (FDA) is a rapidly developing area of statistics for data which can be naturally viewed as a smooth curve or function. It is a method that changes the frame of data …
Parameter Estimation And Inference Of Spatial Autoregressive Model By Stochastic Gradient Descent, Gan Luan
Parameter Estimation And Inference Of Spatial Autoregressive Model By Stochastic Gradient Descent, Gan Luan
Dissertations
Stochastic gradient descent (SGD) is a popular iterative method for model parameter estimation in large-scale data and online learning settings since it goes through the data in only one pass. While SGD has been well studied for independent data, its application to spatially-correlated data largely remains unexplored. This dissertation develops SGD-based parameter estimation and statistical inference algorithms for the spatial autoregressive (SAR) model, a common model for spatial lattice data.
This research contains three parts. (I) The first part concerns SGD estimation and inference for the SAR mean regression model. A new SGD algorithm based on maximum likelihood estimator (MLE) …
Dependent Censoring In Survival Analysis, Zhongcheng Lin
Dependent Censoring In Survival Analysis, Zhongcheng Lin
Dissertations
This dissertation mainly consists of two parts. In the first part, some properties of bivariate Archimedean Copulas formed by two time-to-event random variables are discussed under the setting of left censoring, where these two variables are subject to one left-censored independent variable respectively. Some distributional results for their joint cdf under different censoring patterns are presented. Those results are expected to be useful in both model fitting and checking procedures for Archimedean copula models with bivariate left-censored data. As an application of the theoretical results that are obtained, a moment estimator of the dependence parameter in Archimedean copula models is …
Estimation Of Odds Ratio In 2 X 2 Contingency Tables With Small Cell Counts, Guohao Zhu
Estimation Of Odds Ratio In 2 X 2 Contingency Tables With Small Cell Counts, Guohao Zhu
Dissertations
This study is focusing on properties of estimators of odds ratio or its logarithm in case of 2x2 tables with small counts. The odds ratio represents the odds that an outcome of interest will occur given a particular exposure, compared to the odds of the outcome occurring in the absence of that exposure. Both parameters are often used to quantify the strength of association of two binary variables and are common measurements reported in case-control, cohort, and cross-sectional studies.
Because of their wide applicability, both parameters, odds ratio, and its logarithm, have been intensively studied in the literature. However, most …
Novel Statistical Modeling Methods For Traffic Video Analysis, Hang Shi
Novel Statistical Modeling Methods For Traffic Video Analysis, Hang Shi
Dissertations
Video analysis is an active and rapidly expanding research area in computer vision and artificial intelligence due to its broad applications in modern society. Many methods have been proposed to analyze the videos, but many challenging factors remain untackled. In this dissertation, four statistical modeling methods are proposed to address some challenging traffic video analysis problems under adverse illumination and weather conditions.
First, a new foreground detection method is presented to detect the foreground objects in videos. A novel Global Foreground Modeling (GFM) method, which estimates a global probability density function for the foreground and applies the Bayes decision rule …