Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies,
2026
Southern Methodist University
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Statistical Science Theses and Dissertations
Impact evaluations of regional development programs often require estimating counterfactual outcomes for a small number of treated regions using survey-based areal data. In practice, evaluators typically rely on two-group quasi-experimental methods such as propensity score matching (PSM) and Difference-in-Differences (DiD). These approaches perform poorly when only a few regions receive treatment, and when the set of observed covariates is limited or only partially relevant. Moreover, they typically do not explicitly exploit the spatial and temporal dependence present in survey-based areal data such as in ACS (American Community Survey). This dissertation develops a family of Bayesian spatial predictive models for directly …
Mitigating Parameter Identifiability Issues Through Model Calibration On The Data-Informed Active Subspace: An Example In Tumor Growth,
2026
Lafayette College
Mitigating Parameter Identifiability Issues Through Model Calibration On The Data-Informed Active Subspace: An Example In Tumor Growth, Allison L. Lewis, Rebecca A. Everett
Biology and Medicine Through Mathematics Conference
No abstract provided.
Using Effect Sizes, Confidence Intervals, And The Bayes Factor To Better Understand The T-Test, Analysis Of Variance, And Regression Results,
2026
Ball State University
Using Effect Sizes, Confidence Intervals, And The Bayes Factor To Better Understand The T-Test, Analysis Of Variance, And Regression Results, Holmes Finch
Perspectives on Early Childhood Psychology and Education
Null hypothesis testing is a widely used paradigm for assessing research hypotheses across the social sciences. Despite their ubiquity, researchers have discussed a number of problems and limitations to hypothesis testing and have suggested alternatives that might provide greater depth and explanation of research results. The purpose of this paper is to describe the use of several such alternatives and to show how they can be integrated with one another and with null hypothesis testing in order to provide a more holistic view of research hypotheses.
The Impatience Of Winning: An Analysis Of Time Discounting, Predictive Modeling, And The Nba Draft,
2026
Ursinus College
The Impatience Of Winning: An Analysis Of Time Discounting, Predictive Modeling, And The Nba Draft, Alec R. Plante
Business and Economics Honors Papers
This paper examines whether NBA draft decisions can be better explained by incorporating non-geometric time discounting into a model of general manager decision making. Using a dataset of 285 NBA draft prospects over a 12-year period, the impact of college statistics on Value Over Replacement Player (VORP) is determined, and these impact values are then used to create a “predicted” VORP for the first 4 seasons of each player’s career: a projection of what a general manager might think of a prospect’s future value given their college statistics. Following this, geometric and hyperbolic time discounting models are applied to estimate …
Do Dreams Reflect Our Culture? A Statistical Analysis On Dream Narratives,
2026
Northern Illinois University
Do Dreams Reflect Our Culture? A Statistical Analysis On Dream Narratives, Michal Kuderski
Honors Capstones
Dreams are often viewed as personal experiences, but they may also reflect cultural influences. This project investigates whether dream content varies across cultures by analyzing written dream reports from American, Japanese, and Peruvian college students using data from DreamBank.net. The study applies text analysis techniques to identify common themes and compares language patterns, including the use of ‘I’ and 'We,' to examine differences in self-focus. Statistical methods for count data are used to evaluate these patterns, along with resampling to address differences in sample size. Preliminary findings suggest that both dream themes and language use may vary by cultural …
Base Running: A Lost Art In Baseball,
2026
University of Mary Washington
Base Running: A Lost Art In Baseball, Ethan York
Departmental Honors & Graduate Capstone Projects
In an era of baseball dominated by home runs and launch angles, the subtle art of baserunning is often overlooked, despite its measurable impact on winning games. Baserunning Runs (BsR) addresses this gap by quantifying the number of runs a player contributes through performance on the basepaths, capturing value beyond traditional metrics like stolen bases. This study constructs multiple regression models that predict BsR for Major League Baseball (MLB) players based on baserunning-related statistics. The primary objective is to examine the association between BsR and key predictors, including stolen bases (SB), extra bases taken (EB), and sprint speed (SS), while …
The Item Response Warehouse: What It Is, How To Use It, And Targets For Potential Improvements,
2026
Stanford University
The Item Response Warehouse: What It Is, How To Use It, And Targets For Potential Improvements, Savira D. Nadela, Hansol Lee, Nishka Jain, Ayaan Gupta, Xingyi Zhang, Benjamin W. Domingue
Chinese/English Journal of Educational Measurement and Evaluation | 教育测量与评估双语期刊
The Item Response Warehouse (IRW) is a repository of harmonized item response datasets designed to support secondary analysis and methodological research in psychological and educational measurement. This paper serves as a practical guide for researchers interested in using the IRW. We describe the structure of IRW datasets and the quantitative and qualitative metadata available for dataset selection, and we demonstrate how researchers can navigate the IRW website to explore and compare available tables. We further show how the IRW R and Python packages can be used to filter datasets programmatically, download response-level data, and generate standardized citations for reproducible research …
Ownership Duration In The U.S. Business Jet Market,
2026
Fort Hays State University
Ownership Duration In The U.S. Business Jet Market, Yuchen Hu
SACAD: Scholarly Activities
This study analyzes ownership duration in the U.S. business jet market using FAA registry data as of February 16, 2026 (N=12,359). The analysis reveals a structured distribution with a mean of 5.86 years and a median of 5.00 years. Crucially, retention varies by acquisition status: new aircraft owners exhibit an average hold of 7.36 years, whereas pre-owned aircraft holders show a significantly higher turnover of 5.11 years, with most resales occurring within a 3–7-year window. These findings suggest that ownership behavior is driven by structured asset management and lifecycle planning, providing a predictive framework for identifying aircraft replacement and trade-in …
Efficacy Analysis In Clinical Trials: A Comprehensive Review Of Statistical And Machine Learning Approaches,
2026
Kennesaw State University
Efficacy Analysis In Clinical Trials: A Comprehensive Review Of Statistical And Machine Learning Approaches, Dhrubajyoti Ghosh, Samhita Pal
Faculty Articles
Efficacy testing is a cornerstone of clinical trials, ensuring that medical interventions achieve their intended therapeutic effects. Over the decades, a wide range of statistical methodologies have been developed to address the complexities of clinical trial data, including parametric, nonparametric, Bayesian, and machine learning approaches. Parametric methods, such as t-tests, ANOVA, and LMMs, have traditionally been the foundation of efficacy testing due to their efficiency under well-defined assumptions. Nonparametric techniques, including the Friedman test, Brunner-Munzel test, and modern extensions like nparLD, have emerged as robust alternatives, particularly for skewed, ordinal, or non-normal data. Bayesian methodologies have enabled the incorporation of …
Bayesball : A Comprehensive Framework For Predicting Ucl Injury,
2026
Belmont University
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
SPARK Symposium Presentations
Ulnar Collateral Ligament (UCL) reconstruction, commonly referred to as Tommy John Surgery, has seen a significant rise among Major League Baseball (MLB) pitchers, prompting growing interest in identifying the mechanical and performance-based factors that contribute to injury risk. While previous studies have examined these relationships using traditional frequentist approaches separately, this study combines multiple different model techniques to present a broad framework for finding significant predictors of UCL Surgery. These models include Lasso and Ridge Regression, Principal Component Regression (PCR) , Partial Least Squares Regression (PLS) , Random Forest, Multiple Linear Regression, and a Bayesian Statistical Model. Using these models, …
A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure,
2026
Rochester Institute of Technology
A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure, Ernest Fokoue
Articles
Ensemble learning is traditionally justified as a variance-reduction strategy, explaining its strong performance for unstable predictors such as decision trees. This explanation, however, does not account for ensembles constructed from intrinsically stable estimators-including smoothing splines, kernel ridge regression, Gaussian process regression, and other regularized reproducing kernel Hilbert space (RKHS) methods whose variance is already tightly controlled by regularization and spectral shrinkage. This paper develops a general weighting theory for ensemble learning that moves beyond classical variance-reduction arguments. We formalize ensembles as linear operators acting on a hypothesis space and endow the space of weighting sequences with geometric and spectral constraints. …
Learning Ordinal Geometry: Semantic–Aware Kernels For Ordered Categorical Data,
2026
Rochester Institute of Technology
Learning Ordinal Geometry: Semantic–Aware Kernels For Ordered Categorical Data, Ernest Fokoue
Articles
Ordinal data arise ubiquitously in survey research, psychology, medicine, economics, and recommender systems, yet kernel methods for such data typically rely on either nominal encodings or arbitrary numeric codings. The former discards order information; the lat- ter imposes a fictitious metric structure. This paper develops a principled framework for kernel design on ordinal scales and introduces a new class of Semantic–Aware Ordinal Ker- nels (SAOK) that simultaneously capture ordinal order and semantic proximity between categories. We begin by formalizing order–preserving embeddings of finite chains and characterizing a broad family of chain distances that are conditionally negative definite. Through Schoen- berg …
Selecting Without Replacement From A Population Of Bands Of Serially Connected Objects,
2026
Rochester Institute of Technology
Selecting Without Replacement From A Population Of Bands Of Serially Connected Objects, James E. Marengo, Dominick Banasik, Joseph Voelkel, David L. Farnsworth
Articles
The sampling procedure from a finite population of objects that are serially attached into bands is described and analyzed. One object is randomly selected and removed at a time, which results in that object’s band being broken into two bands or shortened by one object. The main result gives the probability of choosing an object that is part of a band of serially connected objects of any specified size at each stage of the selection process.
Empirical Comparisons Of Partial Dimension Reduction Algorithms In High-Dimensional Regression,
2026
California Polytechnic State University, San Luis Obispo
Empirical Comparisons Of Partial Dimension Reduction Algorithms In High-Dimensional Regression, Nathan Greenfield
Master's Theses
In high-dimensional regression problems, dimension reduction methods are often used to address the challenges of multicollinearity and estimation instability. Partial dimension reduction extends these ideas by applying dimension reduction to a subset of the predictors, while the remaining predictors are modeled without compression. This approach is particularly useful when it is important to retain variability and interpretability in certain predictors.
This thesis investigates the empirical performance of partial dimension reduction algorithms and introduces a novel algorithm, Iterative Partial Residual (IPR). Two algorithms are considered: a baseline algorithm, Marginal Residual (MR), and the proposed IPR method. Their predictive performance is evaluated …
Statistical Analysis Of Log Transformation Effectiveness In Air Traffic Movement Forecasting During Covid-19 In South Africa,
2026
University of the Free State
Statistical Analysis Of Log Transformation Effectiveness In Air Traffic Movement Forecasting During Covid-19 In South Africa, John Lehlaka Masekoameng
Journal of Aviation Technology and Engineering
This study evaluates the effectiveness of log transformation in enhancing multiple regression models used to forecast air traffic movements (ATMs) in South Africa during the COVID-19 pandemic. Using 60 monthly observations from October 2016 to September 2021, the analysis incorporates variables such as revenue, lockdown levels, COVID-19 metrics, exchange rates, gross domestic product, and population. Two models are compared: one using raw ATMs and another with log-transformed ATMs as the dependent variable.
While the untransformed model shows stronger explanatory power (R² = 0.904, adjusted R² = 0.891) compared to the log-transformed model (R² = 0.772, adjusted R² = 0.741), the …
Comparative Machine Learning Models For Disease Risk Prediction,
2026
Marshall University
Comparative Machine Learning Models For Disease Risk Prediction, Mercy Mawusi Agbley
Theses, Dissertations and Capstones
Accurate prediction of disease outcomes is crucial for improving clinical decision-making and enabling early intervention. This study compares the performance of various statistical and machine learning models for clinical risk prediction using two healthcare datasets: diabetic retinopathy and heart disease. The models assessed include Logistic Regression, LASSO, k-Nearest Neighbors (KNN), Support Vector Machines (SVM), Neural Networks, Random Forests, Gradient Boosting Machines (GBM), and a stacked ensemble model. Prior to modeling, datasets were split into train and test sets. Standardization was applied to numeric features whilst categorical features were one-hot encoded. These transformations were later applied to the test set. Principal …
Supplemental Bibliographic Details. From 2001 Mars Odyssey To Earth’S Climate Crisis: Integrating Gamma Spectroscopy, Martian Soil Simulants, And Plant Genomes For Agroecology, Anchored In Sri Lanka’S Mars-Context Serpentinites,
2026
Louisiana State University at Baton Rouge
Supplemental Bibliographic Details. From 2001 Mars Odyssey To Earth’S Climate Crisis: Integrating Gamma Spectroscopy, Martian Soil Simulants, And Plant Genomes For Agroecology, Anchored In Sri Lanka’S Mars-Context Serpentinites, Suniti Karunatillake, Maheshi Dassanayake, Carlos Gary Bicas
Planetary Science Lab
Bibliographic details follow to supplement hyperlinked citations in the multinational GANGOTRI-supporting project conceived by Karunatillake, Dassanayake, and Gary-Bicas
Interpretable Sample Uncertainty Measures For Ranked-Choice Election Polls,
2026
Claremont McKenna College
Interpretable Sample Uncertainty Measures For Ranked-Choice Election Polls, Jason Liang
CMC Senior Theses
Polling results from traditional single-choice plurality elections are readily interpretable. Simple frequentist population parameters are estimated, including each candidate’s total support and the size of the front runner's lead. If the point estimate for the size of the front runner's lead exceeds the margin of error of the lead, we can conclude that the poll shows a statistically significant front runner. However, the interpretability of these population statistics disappears when applied to ranked-choice voting elections. Because ballots rank multiple candidates and candidates are eliminated in rounds, simple population-wide parameters are not well-defined. In RCV elections, a candidate’s ability to win …
Flavor, Transverse Momentum, And Azimuthal Dependence Of Charged Pion Multiplicities In Semi-Inclusive Deep-Inelastic Scattering With 10.6 Gev Electrons,
2026
The College of William & Mary
Flavor, Transverse Momentum, And Azimuthal Dependence Of Charged Pion Multiplicities In Semi-Inclusive Deep-Inelastic Scattering With 10.6 Gev Electrons, P. Bosted, H. Bhatt, S. Jia, W. Armstrong, D. Dutta, R. Ent, D. Gaskell, E. Kinney, H. Mkrtchyan, S. Ali, R. Ambrose, D. Androic, C. Ayerbe Gayoso, A. Bandari, V. Berdnikov, D. Bhetuwal, D. Biswas, M. Boer, E. Brash, A. Camsonne, M. Cardonna, J. P. Chen, J. Chen, M. Chen, E. M. Christy, S. Covrig, S. Danagoulian, M. Diefenthaler, B. Duran, C. Elliot, H. Fenker, E. Fuchey, J. O. Hansen, F. Hauenstein, T. Horn, G. M. Huber, M. K. Jones, M. L. Kabir, A. Karki, B. Karki, S. J. D. Kay, C. Keppel, V. Kumar, N. Lashley-Colthirst, W. B. Li, D. Mack, S. Malace, P. Markowitz, M. Mccaughan, E. Mcclellan, D. Meekins, R. Michaels, A. Mkrtchyan, C. Morean, G. Niculescu, I. Niculescu, B. Pandey, S. Park, E. Pooser, B. Sawatzky, G. R. Smith, H. Szumila-Vance, A. S. Tadepalli, V. Tadevosyan, R. Trotta, H. Voskanyan, S. A. Wood, Z. Ye, C. Yero, X. Zheng, Hall C Sidis Collaboration
Physics Faculty Publications
Measurements of semi-inclusive deep-inelastic scattering multiplicities for 𝜋⁺ and 𝜋⁻ from proton and deuteron targets are reported on a grid of hadron kinematic variables 𝑧, 𝑃𝑇, and 𝜙* for leptonic kinematic variables in the range 0.3< 𝑥< 0.6 and 3< 𝑄²< 5GeV². Data were acquired in 2018 and 2019 at Jefferson Lab Hall C with a 10.6 GeV electron beam impinging on 10-cm-long liquid hydrogen and deuterium targets. Scattered electrons and charged pions were detected in the High Momentum Spectrometer and Super High Momentum Spectrometer, respectively. The multiplicities were fitted for each bin in (𝑥,𝑄²,𝑧,𝑃𝑡) to extract the 𝜙*—independent 𝑀₀ and the azimuthal modulations ⟨cos(𝜙*)⟩ and ⟨cos(2𝜙*)⟩. The 𝑃𝑡 dependence of the 𝑀₀ results was found to be remarkably consistent for the four cases studied: 𝑒𝑝→𝑒𝜋+𝑋, 𝑒𝑝→𝑒𝜋−𝑋, 𝑒𝑑→𝑒𝜋+𝑋, 𝑒𝑑→𝑒𝜋−𝑋 over the range 0GeV< 𝑃𝑡< 0.4GeV, as were the multiplicities evaluated near 𝜙*=180∘ over the extended range 0GeV< 𝑃𝑡< 0.7GeV. The Gaussian widths of the 𝑃𝑡 dependence exhibit a quadratic increase with 𝑧. The cos(𝜙*) modulations were found to be consistent with zero for 𝜋⁺, in agreement with previous world data, while the 𝜋− moments were, in many cases, significantly greater than zero. The cos(2𝜙*) modulations …
Modeling Housing Prices: Which Features Matter Most?,
2026
The University of Akron
Modeling Housing Prices: Which Features Matter Most?, Alex Ruvolo
Williams Honors College, Honors Research Projects
This paper attempts to find the biggest factors and traits that influence the cost of housing. This will include the lot size, type of street, utilities, neighborhood, year built, heating, electrical, yard size, number of different rooms, age, condition, and others. I will attempt to answer the question of whether the prices of houses have changed within the last 5 to 10 years, and obviously this is an easy question to answer. However, the bigger question beyond this is are the main factors affecting housing prices all important in explaining this relationship? Is one factor more important than the rest …
