Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (122)
- Statistical Models (73)
- Social and Behavioral Sciences (60)
- Mathematics (58)
- Institutional and Historical (52)
-
- Statistical Methodology (50)
- Education (41)
- Categorical Data Analysis (40)
- Biostatistics (35)
- Other Statistics and Probability (30)
- Data Science (27)
- Computer Sciences (25)
- Statistical Theory (25)
- Medicine and Health Sciences (22)
- Higher Education (21)
- Applied Mathematics (19)
- Probability (18)
- Engineering (17)
- Design of Experiments and Sample Surveys (16)
- Longitudinal Data Analysis and Time Series (15)
- Business (14)
- Economics (14)
- Life Sciences (14)
- Multivariate Analysis (12)
- Science and Mathematics Education (12)
- Curriculum and Instruction (11)
- Educational Assessment, Evaluation, and Research (11)
- Other Mathematics (11)
- Institution
-
- Wright State University (47)
- Southern Methodist University (37)
- California Polytechnic State University, San Luis Obispo (22)
- Claremont Colleges (16)
- Utah State University (16)
-
- University of South Carolina (14)
- Central Bank of Nigeria (11)
- Nova Southeastern University (11)
- City University of New York (CUNY) (9)
- Embry-Riddle Aeronautical University (9)
- GALILEO, University System of Georgia (9)
- The University of Akron (9)
- University of Arkansas, Fayetteville (9)
- University of South Florida (9)
- Brigham Young University (8)
- Wayne State University (8)
- University of Nebraska - Lincoln (7)
- Ursinus College (7)
- Western Kentucky University (7)
- Air Force Institute of Technology (5)
- East Tennessee State University (5)
- Minnesota State University, Mankato (5)
- University of Central Florida (5)
- University of North Dakota (5)
- Bridgewater State University (3)
- Chapman University (3)
- Georgia Southern University (3)
- University of Connecticut (3)
- University of Denver (3)
- University of New Hampshire (3)
- Publication Year
- Publication
-
- Wright State University Student Fact Books (43)
- Statistical Science Theses and Dissertations (32)
- Theses and Dissertations (20)
- Electronic Theses and Dissertations (14)
- Statistics (14)
-
- Economic and Financial Review (10)
- Williams Honors College, Honors Research Projects (9)
- Mathematics Grants Collections (8)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (7)
- Journal of Humanistic Mathematics (7)
- Senior Theses (7)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (6)
- DataScan (6)
- Graduate Theses and Dissertations (6)
- Journal of Modern Applied Statistical Methods (6)
- Numeracy (6)
- Master's Theses (5)
- NovaFacts (5)
- Open Educational Resources (5)
- Pomona Faculty Publications and Research (4)
- Statistics and Probability (4)
- College of Graduate Studies: Theses & Dissertations (3)
- Essential Studies UNDergraduate Showcase (3)
- Faculty Publications (3)
- Honors Program Theses and Projects (3)
- Honors Projects (3)
- Honors Scholar Theses (3)
- Honors Theses and Capstones (3)
- Publications (3)
- SMU Data Science Review (3)
- Publication Type
- File Type
Articles 1 - 30 of 412
Full-Text Articles in Statistics and Probability
Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang
Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang
Statistical Science Theses and Dissertations
This dissertation addresses two distinct topics related to count time series analysis and topological medical image analysis, respectively. The first part of the dissertation comprises an application of a count time series model to analysis of US monthly sex trafficking data and development of a new model for multivariate count data that exhibits serial dependence and overdispersion. By imposing a family of multivariate mixed Poisson distributions on the count random vector, the proposed model can accommodate a broad range of overdispersion as well as positive contemporaneous correlations. For maximum likelihood estimation, a computationally feasible EM-type algorithm is derived based on …
Exact X2 Statistic Critical Values For Dice Fairness Testing, Warren Campbell
Exact X2 Statistic Critical Values For Dice Fairness Testing, Warren Campbell
SEAS Faculty Publications
Two questions were addressed: 1) What are the exact values of the c2 statistic critical values? 2. Does a dice tower offer improved fairness of dice rolls? Critical values of the statistic asymptotically approach those given by the continuous chi-square distribution, but the exact distribution is discrete. The exact distributions only asymptotically approach the chi-square distribution, and the convergence is slow (1/number of rolls). Th exact distributions are a function of the number of rolls, the chi-square distribution is not a function of the number of rolls. Exact values of c2 at the 90, 95, and 99 percent …
Statistical Methods In Research, Horahenage Dixon Vimalajeewa
Statistical Methods In Research, Horahenage Dixon Vimalajeewa
UNL Faculty Course Portfolios
This course portfolio documents the design, delivery, assessment, student-learning evidence, and reflective evaluation of STAT 801A-700: Statistical Methods in Research, an online asynchronous service course designed to introduce students from diverse disciplinary backgrounds to foundational statistical reasoning and applied data analysis. The course is a non-calculus-based introduction to statistical methods used to answer research questions, with emphasis on collecting, organizing, describing, analyzing, and drawing conclusions from data. The course also emphasizes applications relevant to biology, agriculture, and other research-oriented fields. This is an online distance course aimed to develop students’ understanding of basic probability and statistical concepts, recognize the importance …
Multi-Level Variable Selection Using A Bart-Enhanced Mixed-Effects Framework, Keming Zhang, Yaoyao Li, Jungang Zou, Sijian Wang, Bernadette A. Fausto, Liangyuan Hu
Multi-Level Variable Selection Using A Bart-Enhanced Mixed-Effects Framework, Keming Zhang, Yaoyao Li, Jungang Zou, Sijian Wang, Bernadette A. Fausto, Liangyuan Hu
College of Health Professions Faculty Papers
Selecting important individual- and cluster-level predictors has become increasingly critical in healthcare research, where data often exhibit hierarchical structures due to collection from multiple clusters. Mixed-effects models, which account for within-cluster correlation and between-cluster heterogeneity, are a natural approach for multilevel variable selection. However, currently available variable selection methods for multilevel data are predominantly based on mixed-effects models that impose restrictive parametric assumptions, potentially limiting their utility when the underlying relationships are nonlinear or involve interactions. While nonparametric methods have shown promise for variable selection in non-clustered data, they have been much less studied in the multilevel setting. Moreover, nonparametric …
An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis
An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis
Civil and Environmental Engineering Theses and Dissertations
Urban areas are increasingly exposed to natural hazards while accommodating a growing share of the global population, yet a consistent science-based framework for quantifying urban and community resilience remains lacking. This dissertation develops a physics-based analytical framework grounded in statistical mechanics and the quantitative theory of Brownian motion. A city is conceptualized as a complex medium in which citizens move analogously to Brownian particles within a viscoelastic environment, influenced by socioeconomic interactions and infrastructure functionality.
A central premise is that urban resilience, interpreted as engineering resilience (an outcome), can be quantified through a single metric: the mean-square displacement MSD=⟨r²(t)⟩, of …
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Statistical Science Theses and Dissertations
Impact evaluations of regional development programs often require estimating counterfactual outcomes for a small number of treated regions using survey-based areal data. In practice, evaluators typically rely on two-group quasi-experimental methods such as propensity score matching (PSM) and Difference-in-Differences (DiD). These approaches perform poorly when only a few regions receive treatment, and when the set of observed covariates is limited or only partially relevant. Moreover, they typically do not explicitly exploit the spatial and temporal dependence present in survey-based areal data such as in ACS (American Community Survey). This dissertation develops a family of Bayesian spatial predictive models for directly …
A Statistical Analysis Of Current And Future Hurricane Activity In The North Indian Ocean, Basil Lund
A Statistical Analysis Of Current And Future Hurricane Activity In The North Indian Ocean, Basil Lund
2026 Symposium
A hurricane is defined as a tropical storm with winds sustained at 74 mph or greater. I examined major (category 3 and above) hurricane activity over the North Indian Ocean from the years 1972-2019 as reported by Colorado State University Hurricane Forecast Archive. Using RStudio, I conducted a binomial analysis of the CSU dataset to calculate probabilities of zero to ten years with one or more major North Indian Ocean hurricanes in the next decade. I conducted a geometric analysis to determine probabilities associated with waiting periods for the next year with a major hurricane, as well as a Poisson …
Survival Patterns Among Adult And Pediatric Bone Cancer Patients, Ethan Estes
Survival Patterns Among Adult And Pediatric Bone Cancer Patients, Ethan Estes
Mathematical Sciences Undergraduate Honors Theses
Recently noted, Huang et al. (2023), machine learning (ML) models, while offering great advantages over traditional statistical predictive modeling methods, are less explored in the analysis of survival and other similar time-to-event predictive data modeling. ML methods such as neural networks offer a great deal of promise but need to be further explored to investigate their comparative power in predicting survival outcomes. Focusing specifically on survival analysis in adult and pediatric bone cancer patients, traditional methods, like shown in Emmert-Streib and Dehmer (2019), will be shown with machine learning models using methods in Hothorn, Hornik, and Zeileis (2006). In this …
Deep Learning Frameworks For Biological Data Integration And Generation, Alexa Beachum
Deep Learning Frameworks For Biological Data Integration And Generation, Alexa Beachum
Statistical Science Theses and Dissertations
Data integration represents a key area of research for analyzing the rapidly growing volume of high-dimensional biological data across sources, stages, and modalities. To model and understand these complex, often non-linear relationships, deep learning has become an increasingly powerful tool. Here, we present two novel deep learning frameworks that address distinct but complementary integration challenges. The first framework aligns single-cell omics data across temporal stages, and the second bridges imaging and omics modalities to generate patient-level molecular profiles.
In Chapter 1, we briefly summarize existing approaches---both statistical and deep learning-based---for single-cell omics data integration and discuss their limitations for handling …
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
SPARK Symposium Presentations
Ulnar Collateral Ligament (UCL) reconstruction, commonly referred to as Tommy John Surgery, has seen a significant rise among Major League Baseball (MLB) pitchers, prompting growing interest in identifying the mechanical and performance-based factors that contribute to injury risk. While previous studies have examined these relationships using traditional frequentist approaches separately, this study combines multiple different model techniques to present a broad framework for finding significant predictors of UCL Surgery. These models include Lasso and Ridge Regression, Principal Component Regression (PCR) , Partial Least Squares Regression (PLS) , Random Forest, Multiple Linear Regression, and a Bayesian Statistical Model. Using these models, …
Statistics For Social Scientists Who Are A Little Afraid Of Them, Kristie L. Campana
Statistics For Social Scientists Who Are A Little Afraid Of Them, Kristie L. Campana
MSU Authors Collection
This is a free textbook aimed at helping advanced undergraduate/beginner graduate students navigate statistics in the social sciences. All resources for this book are released under a creative commons license CC BY-SA 4.0. which means you are able to share, copy, and distribute the material in any medium or format, and that you can adapt and transform upon this content for any purpose. However, this can occur only under the following terms:
You must give appropriate attributions to the text, provide a link to the license, and indicate if any alterations to the text have been made. This can be …
Predicting Criminal Behavior In Major Us Cities, Madison A. Price
Predicting Criminal Behavior In Major Us Cities, Madison A. Price
SPARK Symposium Presentations
In recent years, especially post pandemic, there has been a decrease in crime in the United States. Unfortunately, the country’s violent crime rates are still significantly higher compared to similar high-income countries, so what predicts crime in major American cities? There is tons of research to support the idea that demographics can offer some insight into predicting crime. There are countless online resources that seek to identify major crime centrals in the United States (Petrino, 2025). In the late 1990s, researchers noticed that crime rates in cities had a downward slope due to an important contributor: demographic change (Fox & …
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Williams Honors College, Honors Research Projects
Unplanned 30-day hospital readmission remains a fundamental challenge in US healthcare, associated with increased risk to patient recovery and representing an estimated $52.4 billion in annual expenses (Beauvais et al., 2022). While the rigorously validated LACE index serves as the clinical standard for readmission modeling, its linear structure and four explanatory variables lack the complexity to capture the high-dimensional and interactive nature of patient risk. This study utilizes an admission granularity level cohort of the MIMIC-IV database to develop and compare machine learning architectures against the baseline LACE index. Due to the imbalanced prevalence of readmission, the penalized logistic regression, …
From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins
From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins
Honors Undergraduate Theses
With their deserts, castles, and ghost houses, the environments of Super Mario games are colorful, whimsical, and charming, but why are they so compelling, and what happens when our analysis of these environments extends beyond individual levels to expansive game worlds? Drawing on Cresswell’s theory of place (2014) and recent work on musical place-building in Mario Kart 8 (Heazlewood-Dale, 2024), I propose a spectrum between localized and globalized scale in games. As game environments become increasingly globalized, the music may be similarly altered to account for this shift in scale. Consequently, players may then encounter a broader, less musically congruent …
Statistics 103a Instructor Guide, Elizabeth R. Wentworth
Statistics 103a Instructor Guide, Elizabeth R. Wentworth
Open Educational Resources
This set of slides contains reading, original videos, activities and instructions for instructors to run a complete 12 week course in any modality. These resources can be used to supplement in-person instruction or can be used for either a hybrid or asynchronous course.
Inferential Statistics For Industrial Organizational Psychologists: A Practical Guide For Testing Hypotheses Using R, Caitlin Lapine
Inferential Statistics For Industrial Organizational Psychologists: A Practical Guide For Testing Hypotheses Using R, Caitlin Lapine
Open Touro Created
2026
This text aims to provide a practical guide for students in industrial organizational psychology or related fields to complete inferential statistics using R open-source programming language. It provides information about when to use particular statistical analyses and how to perform those with statistical software.
Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens
Statistical Science Theses and Dissertations
Electronic Health Records (EHR) contain a wealth of structured and unstructured patient data that can be leveraged for computable phenotyping, the process of algorithmically identifying patient cohorts with specific diseases or conditions. Traditional rule-based phenotyping approaches, while interpretable, often struggle with scalability, portability across institutions, and effective use of unstructured clinical narratives. Recent advances in large language models (LLMs) present new opportunities for synthesizing complex free-text information into concise, clinically meaningful representations. However, integrating LLMs into phenotyping workflows requires careful design to maintain transparency, interpretability, and measurable uncertainty—features essential for clinical adoption and downstream applications such as decision support.
We …
Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong
Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong
Statistical Science Theses and Dissertations
Recurrent event data frequently arise in clinical studies where individuals experience repeated, possibly related, events over time. These data are often accompanied by sparse and irregular longitudinal measurements, creating challenges for traditional joint modeling approaches that struggle to account for time-dependent associations and within-subject correlations. We propose FRAILTY (Functional Regression with AutoRegressIve fraiLTY), a novel two-step framework that integrates functional principal component analysis (PACE) with a dynamic frailty model featuring autoregressive structure. FRAILTY accommodates both scalar and functional predictors and captures within-subject dependence across recurrent events. To further extend its utility, we develop a multivariate joint modeling framework that simultaneously …
Empirical Evaluation Of Bayes Error Rate Bounds In Binary Classification, Riley May
Empirical Evaluation Of Bayes Error Rate Bounds In Binary Classification, Riley May
All Graduate Theses and Dissertations, Fall 2023 to Present
Classification tasks are fundamental in statistical machine learning. In classification tasks, a general goal is to build or select a model that can correctly classify data with as few errors as possible. However, for a particular dataset, the minimal number of errors achievable is seldom zero since overlap in the data makes errors unavoidable. As a result, it is often difficult for machine learning practitioners and data scientists to know whether classification errors can be reduced through further refinement. A potential solution to this lies in the Bayes error rate (BER). The BER is the lowest error rate achievable for …
Advancing Statistical Methods For Multivariate And Network Meta-Analysis, Yifei Wang
Advancing Statistical Methods For Multivariate And Network Meta-Analysis, Yifei Wang
Statistical Science Theses and Dissertations
Multivariate meta-analysis (MMA) and network meta-analysis (NMA) are essential tools for synthesizing evidence across multiple correlated outcomes and treatments. However, these tools face practical challenges, including outcome reporting bias (ORB), unreported within-study correlations, and computational burden. ORB can distort effect estimates in MMA, while missing within-study correlations in multivariate NMA may lead to biased conclusions. To address these challenges, this dissertation introduces two novel statistical methods. For MMA, we propose SemiMMA, a semiparametric and scalable approach that treats ORB as a missing-not-at-random problem and combines inverse propensity weighting (IPW) with the generalized method of moments (GMM). For multivariate NMA, we …
Incorporating Propensity Score Weighting And Nonresposne Adjustments Into Complex Survey Data With Survival Outcomes, Xinrui Shi
Theses and Dissertations
Propensity score weighting (PSW) plays a key role in minimizing confounding in observational research, especially when estimating treatment effects for time-to-event outcomes. However, its integration into survey data with complex design – particularly data with multiple stage sampling and censoring – remains underexplored. One significant challenge in such settings is the presence of nonresponse, which can introduce additional bias and complicate the use of standard weight adjustments. Moreover, there has been limited study on how PS weights can be effectively combined with nonresponse weighting adjustments in complex survey data that include survival outcomes. This dissertation aims to extend current methodologies …
Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi
Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi
Open Educational Resources
Data analysis using standard statistical methods and relevant computer software. Emphasis on real-world data, interpretation, and misinterpretation of computer output.
This syllabus contains open source notebook about data analysis content.
Nonlinear Power Function Model Changepoint Detection., Jacob Steven Townson
Nonlinear Power Function Model Changepoint Detection., Jacob Steven Townson
Electronic Theses and Dissertations
Most work surrounding changepoint analysis focuses on linear models. This dissertation explores changepoint detection in nonlinear power function models, specifically focusing on models where the constant multiplier and power are the parameters to be estimated in addition to the changepoint parameter. The study assumes an asymptotic framework as the number of observations approaches infinity. The study explores various model fitting algorithms, and decides to employ the Newton-Raphson method for parameter estimation, with a custom implementation developed to optimize the process. The research first establishes the strong consistency of estimators for the model without a changepoint. Building on this result, consistency …
Latent Variable Dyadic Regression Models For Predicting Over/Under Bets In Sports Betting, Alexcia Trejo
Latent Variable Dyadic Regression Models For Predicting Over/Under Bets In Sports Betting, Alexcia Trejo
Graduate Theses and Dissertations
This thesis explores the use of latent factor models to uncover hidden structures in pair wise outcomes derived from Over/Under betting markets in sports betting. Specifically, we implement and evaluate the Eigen model, a latent space model that represents dyadic data using node-specific vectors whose inner product govern edge probabilities. By modeling relationships between teams as adjacency matrices of binary outcomes, we investigate the extent to which the Eigen model captures both homophily, the tendency of similar teams to yield consistent betting results, and stochastic equivalence, where different teams exhibit indistinguishable patterns of Over/Under outcomes. A Bayesian formulation of the …
A Text Mining And Sentiment Analysis Of Valuable Cie Texts Using R, Eric Sugarman, Ethan Turber-Ortiz, Hannah Quinn
A Text Mining And Sentiment Analysis Of Valuable Cie Texts Using R, Eric Sugarman, Ethan Turber-Ortiz, Hannah Quinn
Mathematics, Computer Science & Statistics Presentations
The purpose of this project was to perform a sentiment analysis of three texts used in Ursinus College's Common Intellectual Experience (CIE) course: Between the World and Me by Ta-Nehisi Coates, The New Jim Crow by Michelle Alexander and Discourse on Method by Rene Descartes. Word count and word cloud analysis were also performed on the texts as well as term frequency and bigram analysis.
Cohens_D, Manish Rami
Cohens_D, Manish Rami
Software
This Python script calculates the effect size Cohen's d in a two group situation with known means and Standard Deviations.
Use this effect size if the sample size in your experiment is large and the two SDs are similar.
Repositioning The Game: Traditional Positions Vs Tracking-Based Archetypes In Nba Performance Models, Jacob Floyd
Repositioning The Game: Traditional Positions Vs Tracking-Based Archetypes In Nba Performance Models, Jacob Floyd
Senior Theses
Driven by the rise of advanced analytics and player tracking technologies, the NBA has transitioned away from traditional positional roles and toward more fluid player archetypes. This investigation uses principal component analysis and k-means clustering to group players based on season-long tracking data, creating new pseudo-positions that more accurately reflect modern playing styles. Predictive models were then built using both the classic position system and the newly generated clusters to forecast player scoring performance. Across every model comparison, both in terms of fit and predictive accuracy, the cluster-based system significantly outperformed the traditional position-based model. These results reinforce the idea …
Robust Spacecraft Autonomy For Deep Space Exploration In Special Euclidean Group Se(3), Matthew Wittal
Robust Spacecraft Autonomy For Deep Space Exploration In Special Euclidean Group Se(3), Matthew Wittal
Doctoral Dissertations and Master's Theses
Over the past half-century, humanity has gained extensive experience conducting manned spaceflight near Earth. Arguably, "near Earth" could even include the Moon — the most distant destination humans have reached. However, "near" in this work primarily refers low Earth orbit (LEO). One could argue that we have not truly left Earth since the Apollo, as spacecraft in some LEOs remain subject to atmospheric drag thus emphasizing their continued connection to Earth's immediate environment. Reflecting on this, it becomes clear that humanity has largely remained bound to Earth’s immediate vicinity since the Apollo missions reached the Moon. However, that is set …
Glass Delta, Manish Rami
Glass Delta, Manish Rami
Software
A Python script to calculate the effect size Glass' delta in a two group experiment with different standard deviation.
On The Gumbel-Weibull{Cauchy} Distribution, Jennifer D. Pippin
On The Gumbel-Weibull{Cauchy} Distribution, Jennifer D. Pippin
Theses, Dissertations and Capstones
Developing new statistical distributions and seeking higher flexibility in modeling different shapes of data remain a strong emphasis in research. The T-R{Y } framework, introduced in [3], utilizes three statistical distributions in order to generate a new distribution. Many research papers appeared in literature to develop distributions based on the T-R{Y } framework. In this thesis, a member of the T-R{Y } framework, namely the Gumbel-Weibull{Cauchy} (GWC), is introduced. Statistical properties of the GWC are studied, such as the quantile function, the hazard function, transformations, Shannon entropy, the …