Small Area Estimation Of Forest Biomass Via A Two-Stage Model For Continuous Zero-Inflated Data,
2025
Reed College
Small Area Estimation Of Forest Biomass Via A Two-Stage Model For Continuous Zero-Inflated Data, Grayson W. White, Josh K. Yamamoto, Dinan H. Elsyad, Julian F. Schmitt, Niels H. Korsgaard, Jie Hu, George C. Gaines Iii, Tracey S. Frescino, Kelly S. Mcconville
Faculty Journal Articles
Nationwide Forest Inventories (NFIs) collect data on and monitor the trends of forests across the globe. Users of NFI data are increasingly interested in monitoring forest attributes such as biomass at fine geographic and temporal scales, resulting in a need for assessment and development of small area estimation techniques in forest inventory. We implement a small area estimator and parametric bootstrap estimator that account for zero-inflation in biomass data via a two-stage model-based approach and compare the performance to a Horvitz–Thompson estimator, a post-stratified estimator, and to the unit- and area-level empirical best linear unbiased prediction (EBLUP) estimators. We conduct …
Majority Decision Using Top-Performing Neural Networks Models For Improved Credit Risk Prediction,
2025
Georgia Southern University
Majority Decision Using Top-Performing Neural Networks Models For Improved Credit Risk Prediction, Vincent Dey
College of Graduate Studies: Theses & Dissertations
Credit risk prediction remains both a challenging and high-interest problem due to the inherently unbalanced nature of financial datasets and the continuous drive for higher pre- dictive precision. In this work, I build upon previous advancements in credit risk modeling and introduce an ensemble-based Artificial Neural Network (ANN) architecture designed to enhance classification performance. By leveraging a selective ensemble of decision net- works, this approach not only improves prediction accuracy but also mitigates the chal- lenges posed by imbalanced data distributions. While the primary focus is on credit risk prediction, my analysis demonstrates that the proposed model can be effectively …
Further Results On Learning Quantum Measurement Classes: Quantum Pac Model For Povm Hypothesis Classes,
2025
University at Albany, State University of New York
Further Results On Learning Quantum Measurement Classes: Quantum Pac Model For Povm Hypothesis Classes, Arka Prabha Das
Electronic Theses & Dissertations (2024 - present)
This thesis investigates the problem of learning from quantum systems, where each example consists of a quantum state paired with a classical outcome. The task centers on choosing an effective measurement rule from a fixed set to enable accurate prediction of the classical outcome from the quantum state. A central focus lies in understanding whether joint measurement strategies that cannot be separated into local operations offer a real benefit in terms of the number of examples needed for successful learning. We examine conditions under which a non-separable measurement within a given hypothesis class achieves strictly better sample complexity bounds compared …
Theoretical Foundations And Applied Performance Of Periodicity-Aware Imputation: Variable Bandpass Block Bootstrap Methods For Incomplete Time Series,
2025
University at Albany, State University of New York
Theoretical Foundations And Applied Performance Of Periodicity-Aware Imputation: Variable Bandpass Block Bootstrap Methods For Incomplete Time Series, Asmaa Ahmad
Electronic Theses & Dissertations (2024 - present)
Time series data are prevalent across a wide range of disciplines, including health surveillance, public policy, and environmental monitoring. In the presence of underlying cyclical patterns, the integrity of time series analysis depends critically on the ability to detect, model, and impute structured missing data without compromising the temporal structure. This dissertation introduces and validates a novel imputation framework that integrates the Variable Bandpass Periodic Block Bootstrap (VBPBB) into multiple imputation procedures, improving the accuracy, robustness, and interpretability of time series models under high rates of missingness and noise. The overarching goal of this dissertation was to develop and evaluate …
Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam,
2025
Claremont Colleges
Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri
CMC Senior Theses
Over the past decades, the gaming industry has managed to evolve into a multi-billion-dollar enterprise. Gaming platforms such as Steam foster unprecedented amounts of engagement among players worldwide daily. In this thesis, we investigate the effect of incorporating sentiment-driven metrics, specifically YouTube view counts and positive reviews, into predictive models for game popularity. In addition, by comparing our linear regression sentiment-based approach to the Bayesian hierarchical folded normal model used by De Luisa et al. (2021), we can understand the many differences, strengths, and limitations of each methodology. In our thesis, we focus on three games. Each is of varying …
Forecasting Equity Betas Using Option-Implied Moments,
2025
Claremont McKenna College
Forecasting Equity Betas Using Option-Implied Moments, Ivan Kolesnikov
CMC Senior Theses
Traditional beta estimates are constructed from historical stock‑and‑market returns and therefore adjust only as fast as realized data accrue. This thesis investigates whether the forward‑looking information embedded in equity‑option prices can enhance beta forecasts. Using near‑end‑of‑day quotes for 236 S&P 500 firms between 2007 and 2024, I extract risk‑neutral variance and skewness, construct five alternative beta estimators (historical, option‑implied, and three hybrids), and evaluate them against realized betas over six‑, twelve‑, and twenty‑four‑month windows. Rolling‑OLS beta remains the most accurate benchmark at short horizons, yet option‑implied moments add economically and statistically significant value when systematic exposure is expected to change …
Analyzing Political Sentiment On Micro-Blogging Data: A Lexicon And Machine Learning Approach To The 2024 U.S. Presidential Election,
2025
Claremont McKenna College
Analyzing Political Sentiment On Micro-Blogging Data: A Lexicon And Machine Learning Approach To The 2024 U.S. Presidential Election, Ava Grey
CMC Senior Theses
This paper explores the trends in sentiment towards U.S. presidential candidates Kamala Harris and Donald Trump through micro-blogging social media text during the five months leading up to the election. Two datasets of varying sizes and origins were used to contextualize and validate analysis findings. The analyses include both a lexicon-based approach and a machine learning predictive method. Common sentiment analysis techniques like term frequency, term frequency inverse, various lexicons, and n-grams were utilized during the lexicon approach. During the modeling, a random forest was utilized in addition to the methods used during the lexicon approach. Results showed that overall …
In Search Of The Rational Voter In The 2020 Presidential Election: Understanding The Impact Of Voter Costs And Benefits On Turnout,
2025
Old Dominion University
In Search Of The Rational Voter In The 2020 Presidential Election: Understanding The Impact Of Voter Costs And Benefits On Turnout, Norou Diawara, Tiffany Henley, Samuel L. Brown, Md Iqbal Hossain
Mathematics & Statistics Faculty Publications
The ability to vote is one of the most valuable rights and privileges afforded by the Constitution of the United States to its citizens. For many, voting is not just a civic duty; it is also a choice. Voting is crucial to our democracy, and any changes to it may affect the efficiency of the democratic process. The bigger question is whether voters behave rationally by engaging in a cost-benefit calculus in deciding whether or not to vote. Using data science, this paper will examine the probability of voting and investigate its impact via cost and benefit among other variables …
The Presence Of Outer Giant Planets And Their Role In Inner Planet Formation With And Without Their Influence,
2025
Missouri State University
The Presence Of Outer Giant Planets And Their Role In Inner Planet Formation With And Without Their Influence, Mateo E. Guerra Toro
Graduate Theses/Dissertations
We performed dynamical simulations of the giant impact phase of planet formation to investigate the formation of inner terrestrial planets under the influence of 4 solar system-like outer giant planets. We developed a new code using the N-body simulation suite REBOUND and REBOUNDx (Rein et al. (2019) and Tamayo et al. (2019)) to simulate 2 stages of planetary formation: a residual gaseous protoplanetary disk phase and subsequent dynamical evolution after the disk photoevaporates. The initial conditions for the inner planetary embryos were taken by Morrison et al. (2020) based on a range of solid surface densities that produced Super-Earth terrestrial …
Applications Of Bayesian Functional Data Analysis,
2025
Northern Illinois University
Applications Of Bayesian Functional Data Analysis, Zhexuan Yang
Graduate Research Theses & Dissertations
Functional Data Analysis (FDA) is a statistical approach used to analyze data that vary across a domain, such as curves or functions. This dissertation investigates Bayesian Functional Data Analysis (BFDA) through three applications. First, we explore the use of BFDA in outcome-dependent follow-up (ODFL) studies. After conducting simulation studies, we apply our model to cardiotoxicity and kidney function data. Second, we extend BFDA to genetic data by modeling DNA methylation levels with a three-parameter skew-normal distribution and an alpha-skew generalized normal distribution. This study also introduces a novel Multistage Markov Chain Monte Carlo (MMCMC) method with the goal of identifying …
Covariance Matrix Forecasting Of Equity Portfolios,
2025
Northern Illinois University
Covariance Matrix Forecasting Of Equity Portfolios, Michael Nebor
Graduate Research Theses & Dissertations
This dissertation consists of two papers. The first paper introduces DCC-SVR, a hybrid Dynamic Conditional Correlation (DCC) and Support Vector Regression (SVR) method of forecasting the covariance matrix. This paper shows that DCC-SVR is able to outperform the traditional methods of DCC and rolling historical on multiple data sets. Performance is shown for both standard GARCH and GJR-GARCH methods. This paper also analyzes performance when dimensions are increased to 49 dimensions and when an application using equal weighted portfolio allocation is used.
The second paper introduces a covariance matrix forecasting method based on copula-GARCH simulated returns. The accuracy of this …
Efficient Algorithms For Nearest Correlation Matrix Computation With Missing Data,
2025
Northern Illinois University
Efficient Algorithms For Nearest Correlation Matrix Computation With Missing Data, Ibrahim Eniola Oyeyinka
Graduate Research Theses & Dissertations
This thesis investigates efficient algorithms for computing the Nearest Correlation Matrix (NCM) under incomplete financial data. Correlation matrices are vital in portfolio optimization and risk management, yet empirical estimates often violate symmetry, positive semidefiniteness, and unit diagonal conditions due to missing observations. Two projection-based methods are analyzed: the Modified Alternating Projections (MAP) and Anderson Acceleration (AA). Theoretical analysis using convex optimization and normal cone characterization supports numerical evaluation on synthetic and real-world stock-return matrices (550×550, 2020–2025). Missing data are modeled through Missing Completely at Random (MCAR) and Not Missing at Random (NMAR) mechanisms. The results show that AA converges faster …
A Modern Optimization Approach With Data-Driven Analytical Modeling For The Healthcare Business Segment (Hbs) From The S&P 500,
2025
Macon & Joan Brock Virginia Health Sciences at Old Dominion University
A Modern Optimization Approach With Data-Driven Analytical Modeling For The Healthcare Business Segment (Hbs) From The S&P 500, Aditya Chakraborty, Chris Tsokos
Epidemiology, Biostatistics, & Environmental Health Faculty Publications
Introduction: The S&P consists of eleven business segments, which are classified according to the type of industry. The current study focuses on developing a non-linear analytical model for the Healthcare Business Segment (HBS) of the S&P 500, as a function of different economic & financial indicators. Materials and Methods: The analytical model used six financial indicators together with four economic indicators to predict the weekly average closing price (WCP) of HBS stocks. Johnson’s SB transformation corrected skewness, while desirability-based optimization identified indicator values maximizing WCP. The model’s performance and generalizability were validated through repeated 10-fold cross-validation. Results: All attributable contributors …
Pediatric Cancer Incidence, Temporal Trends, And Mortality In The United States By Health Disparities Indicators, Seer (1973-2014),
2025
Macon & Joan Brock Virginia Health Sciences at Old Dominion University
Pediatric Cancer Incidence, Temporal Trends, And Mortality In The United States By Health Disparities Indicators, Seer (1973-2014), Prachi P. Chavan, Laurens Holmes Jr.
Epidemiology, Biostatistics, & Environmental Health Faculty Publications
Background: Pediatric cancer incidence has been increasing in the United States, despite improvement in pediatric cancer survival. This steady increase in incidence trends is not completely understood but maybe associated with social and environmental factors. In this study we aimed to assess the cumulative incidence, temporal trends, and mortality rates in pediatric cancer. Additionally, we examined sub-group variability in both incidence and mortality rates. Methods: Data from Surveillance, Epidemiology, and End Results (SEER) −18 from 1973–2014 were used for the purpose of analysis in this study. Age-adjusted incidence rates were used to assess temporal trends in cancer among children aged < 1–19 years. Univariable and multivariable binomial regression models were used to examine the association between race and cancer mortality while adjusting for potential confounders. Results: There were 92,594 cancer diagnoses during this period. White children comprised 74,758, (80.7%), black children 10,030, (10.8%), and other races 6648, (7.2%). Overall the age-adjusted cumulative incidence was slightly higher among white children (16.4%) than black children (12.4%) and other (13.0%). Children aged 15–19 years and those in metropolitan regions were more likely to be diagnosed with pediatric cancer. Relative to females, males were 16% more likely to die from the disease [adjusted Risk Ratio (aRR): 1.16, 95% Confidence Interval (CI): 1.09–1.22]. Additionally, compared to white children, black children had higher mortality rates [(aRR): 1.37, 99% CI: 1.23–1.52]. Conclusions: There is an increasing trend in pediatric cancer incidence; while white children have the highest incidence, black children and males indicated a survival disadvantage, indicative of racial and sex variability in overall pediatric cancer in the United States.
Estimating The Gender Wage Gap: A Comparative Analysis Of Different Estimators,
2025
Macalester College
Estimating The Gender Wage Gap: A Comparative Analysis Of Different Estimators, Xinran Zhang
Mathematics, Statistics, and Computer Science Honors Projects
The gender wage gap between males and females has been well studied by labor economists. We take a multi-prong approach to evaluate three estimators —a regression-imputation estimator, a weighting estimator, and a doubly robust estimator—in estimating the gender wage gap. Using the Panel Study of Income Dynamics, we conduct an empirical study of the estimators’ performances. In a simulation study, we evaluate the properties of estimators and study whether bootstrapping is an appropriate measure of the uncertainty of each estimator. The findings show that while the estimators provide different results, the doubly robust estimator provides reliable and consistent results under …
Action This Day: The Mathematics And Machinations That Bested The German Enigma,
2025
Dartmouth College
Action This Day: The Mathematics And Machinations That Bested The German Enigma, Jonah Weinbaum
Dartmouth College Master’s Theses
This thesis presents a comprehensive and chronological overview of cryptographic techniques designed to break Enigma, beginning in 1932 and culminating in the creation of the Turing-Welchman Bombe. We discuss the mathematical theory and electromechanical implements used to decode one of history's greatest ciphers.
Reexamining the Bombe through the lens of modern group theory, we critique Alan Turing's estimation of the number of "stops" that the Bombe produces for various plaintext-ciphertext pairing structures. To address its limitations, we introduce a new framework for estimating the number of stops by extending John Dixon's theorem concerning the probability that uniformly distributed elements of …
Food Insecurity Among Graduate Students At The University Of Alabama At Birmingham,
2025
University of Alabama at Birmingham
Food Insecurity Among Graduate Students At The University Of Alabama At Birmingham, Amy Elizabeth Callahan
All ETDs from UAB
College food insecurity (FI) has grown steadily as a field of research over the past fifteen years. Existing research primarily has demonstrated that FI is a persistent problem with various risk factors and myriad potential impacts on students, though graduate students are largely understudied in favor of undergraduate and aggregated student populations. Given the unique circumstances graduate students face, as well as expected demographic differences such as age and household characteristics, this is a serious gap in the literature. This dissertation aims to investigate graduate student food insecurity in the United States, and particularly at the University of Alabama at …
Bayesian Networks For Safety-Critical Systems,
2025
Technological University Dublin, Ireland
Bayesian Networks For Safety-Critical Systems, Joseph Mietkiewicz
Theses
This thesis addresses a operational challenge in modern industrial operations: the increasing complexity of systems and the consequent cognitive burden on operators. As industrial technologies advance, the human-computer interface has become the primary conduit for information flow, playing a pivotal role in operational decision-making. However, the proliferation of data often leads to information overload, potentially compromising rather than enhancing operator performance. This research explores an approach to this pressing issue through the application of Bayesian networks as decision support systems in safety- critical scenarios. Our study employs a multi-faceted approach, combining theoretical modeling with empirical testing. Through collaboration with industry …
Optimal Data Splitting Methods,
2025
Virginia Commonwealth University
Optimal Data Splitting Methods, Sujay Mudalgi
Theses and Dissertations
In predictive modeling, effective data splitting is crucial for creating statistically representative training and validation sets. The state-of-the-art data splitting methods are based on minimizing the energy distance between the split subsets. However, there are a number of limitations in the existing methods, which this dissertation aims to address. First, the existing methods were computationally inefficient. Thus, Chapter 2 proposes a method to scale up these approaches for big data. Here, we introduce scalable Twinning (s-Twinning), which significantly improves the execution speed of data splitting without sacrificing accuracy. Second, the existing methods did not consider the predictive relationship in the …
Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses,
2025
University of Thi Qar
Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi
Theses and Dissertations
Electronic Health Records (EHRs) constitute a comprehensive and high-dimensional repository of clinical data, encompassing a wide array of patient-level information such as diagnoses, procedures, medications, laboratory results, and unstructured clinical narratives. These data hold immense potential for advancing predictive modeling in healthcare, including tasks such as disease progression modeling, hospital readmission prediction, and length of stay (LoS) estimation. However, the intrinsic complexity of EHR data—manifested in its heterogeneity, sparsity, and temporal dynamics—poses significant analytical challenges that limit the generalizability and interpretability of conventional machine learning models. Recent methodological advancements in deep learning and graph-based learning, particularly Graph Neural Networks (GNNs), …
