Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Social and Behavioral Sciences (1395)
- Statistical Theory (1191)
- Statistical Models (374)
- Applied Mathematics (362)
- Mathematics (341)
-
- Statistical Methodology (274)
- Data Science (194)
- Computer Sciences (171)
- Biostatistics (155)
- Engineering (154)
- Medicine and Health Sciences (151)
- Probability (150)
- Multivariate Analysis (143)
- Life Sciences (139)
- Business (127)
- Longitudinal Data Analysis and Time Series (122)
- Categorical Data Analysis (115)
- Economics (113)
- Law (102)
- Other Statistics and Probability (93)
- Education (91)
- Environmental Sciences (86)
- Artificial Intelligence and Robotics (81)
- Design of Experiments and Sample Surveys (77)
- Econometrics (76)
- Psychology (65)
- Numerical Analysis and Scientific Computing (58)
- Institution
-
- Wayne State University (1097)
- Wright State University (138)
- Utah State University (103)
- Cornell University Law School (75)
- California Polytechnic State University, San Luis Obispo (59)
-
- Air Force Institute of Technology (55)
- Old Dominion University (55)
- University of Kentucky (54)
- Montclair State University (53)
- University of Arkansas, Fayetteville (51)
- Southern Methodist University (46)
- Western Kentucky University (46)
- University of Nebraska - Lincoln (43)
- Central Bank of Nigeria (41)
- City University of New York (CUNY) (39)
- Virginia Commonwealth University (39)
- Claremont Colleges (36)
- Illinois State University (35)
- Kennesaw State University (35)
- University of Richmond (34)
- Georgia Southern University (31)
- Louisiana Tech University (29)
- Stephen F. Austin State University (25)
- University of New Mexico (25)
- East Tennessee State University (24)
- University of Nevada, Las Vegas (24)
- Prairie View A&M University (22)
- The University of Akron (19)
- Technological University Dublin (17)
- Michigan Technological University (16)
- Keyword
-
- Statistics (122)
- Empirical legal studies (55)
- Simulation (49)
- Machine learning (42)
- Regression (40)
-
- Bias (35)
- Logistic regression (33)
- Monte Carlo simulation (30)
- Bootstrap (29)
- Power (29)
- Bayesian (28)
- Machine Learning (27)
- Reliability (27)
- Confidence interval (26)
- Western Kentucky University (26)
- Mean squared error (24)
- Estimation (22)
- Maximum likelihood estimation (22)
- Missing data (22)
- Monte Carlo (22)
- Nonparametric (21)
- Robustness (21)
- Sample size (21)
- Type I error (21)
- Effect size (20)
- Multicollinearity (20)
- Statistical analysis (20)
- Confidence intervals (19)
- Enrollment (19)
- Pure sciences (19)
- Publication Year
- Publication
-
- Journal of Modern Applied Statistical Methods (1093)
- Mathematics and Statistics Faculty Publications (137)
- Theses and Dissertations (91)
- Cornell Law Faculty Publications (75)
- Electronic Theses and Dissertations (58)
-
- Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works (49)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (46)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (42)
- CBN Journal of Applied Statistics (JAS) (40)
- Graduate Theses and Dissertations (38)
- Department of Math & Statistics Faculty Publications (34)
- Theses and Dissertations--Statistics (32)
- College of Graduate Studies: Theses & Dissertations (29)
- SMU Data Science Review (29)
- Master's Theses (27)
- Mathematics & Statistics Theses & Dissertations (26)
- WKU Administration Documents (26)
- Annual Symposium on Biomathematics and Ecology Education and Research (25)
- Symposium of Student Scholars (24)
- Applications and Applied Mathematics: An International Journal (AAM) (22)
- Articles (22)
- Statistics (21)
- Publications and Research (20)
- Williams Honors College, Honors Research Projects (19)
- Department of Statistics: Dissertations, Theses, and Student Research (17)
- CMC Senior Theses (16)
- Dissertations, Master's Theses and Master's Reports (16)
- Mathematics Senior Capstone Papers (15)
- Statistical Science Theses and Dissertations (15)
- Doctoral Dissertations (14)
- Publication Type
- File Type
Articles 181 - 210 of 2918
Full-Text Articles in Applied Statistics
A Monte Carlo Study Of The Efficacy Of The Mantel-Haenszel Method At Detecting Item Bias, Macy Daulton
A Monte Carlo Study Of The Efficacy Of The Mantel-Haenszel Method At Detecting Item Bias, Macy Daulton
Masters Theses & Specialist Projects
The Mantel-Haenszel (MH) method is a statistical method used for differential item functioning analysis. Previous differential item functioning (DIF) research of the MH method focused on Type I error rates and statistical power. Not yet studied in the literature is whether excessive numbers of biased items may cause the MH method to fail to identify biased items. The proposed study employed a Monte Carlo design to determine the extent of bias in the match items necessary for the MH method to fail to identify a biased item. Results showed that the threshold was greater than was hypothesized, requiring nearly all …
Gompertz Distribution On Time Scales, Wasiu Sule
Gompertz Distribution On Time Scales, Wasiu Sule
Theses, Dissertations and Capstones
We shall investigate Gompertz dynamic equations within the context of time scales calculus, by exploring the mathematical foundations and applications of the Gompertz model, which is commonly used to describe growth phenomena in various fields such as biology and economics. This research seeks to analyze the Gompertz cumulative distribution functions (CDF) and probability density functions (PDF) across different time scales, including the real numbers R and integer multiples hN. Probability techniques will be used to derive the CDF and PDF associated with the Gompertz dynamic equations, and we will examine how varying the time scale impacts the characteristics …
Predicting Superconducting Critical Temperature From Composition-Derived Features: A Transparent Linear And Regularized Regression Study, Md Ahiduzzaman
Predicting Superconducting Critical Temperature From Composition-Derived Features: A Transparent Linear And Regularized Regression Study, Md Ahiduzzaman
Data Science and Data Mining
We study prediction of superconducting critical temperature (Tc) from 81 composition-derived descriptors across 21,263 materials. To keep the analysis transparent and repro- ducible, we focus on linear models: Ordinary Least Squares (OLS), Ridge, Lasso, and Elastic Net (ENet). All models share a single evaluation protocol (5-fold cross-validation with standardized inputs) and are compared on RMSE, MAE, and R2. On this feature set, OLS attains the best cross-validated performance (RMSE = 17.6 K, MAE = 13.3 K , R2 = 0.735), with Lasso/ENet essentially tied next (RMSE ≈ 17.7 K , R2 ≈ 0.734); Ridge underperforms (RMSE = 18.9 K , …
Comparative Analysis Of Lasso, Ridge, And Elastic Net For Variable Selection In High-Dimensional Maize Data, Md Ahiduzzaman
Comparative Analysis Of Lasso, Ridge, And Elastic Net For Variable Selection In High-Dimensional Maize Data, Md Ahiduzzaman
Data Science and Data Mining
In high-dimensional genomic data analysis, traditional linear regression techniques often struggle due to the presence of a large number of predictor variables relative to observations. Penalized regression methods such as LASSO, Ridge, and Elastic Net have emerged as effective solutions by imposing regularization, which helps in managing multicollinearity and enhancing prediction accuracy. This study applies these techniques to the Maize dataset to model the time to male flowering, selecting relevant genetic markers as predictors. Our findings suggest that Elastic Net is particularly effective for high-dimensional data with correlated variables, achieving a balance between prediction accuracy and variable selection. The results …
Statistical Analysis Of Climate Trends And Variability In Tarrant County Using Annual, Decade, And Three-Decade Time Periods, Quinnton Debolt
Statistical Analysis Of Climate Trends And Variability In Tarrant County Using Annual, Decade, And Three-Decade Time Periods, Quinnton Debolt
Earth & Environmental Sciences Theses - Archive
This study is a statistical analysis of climate trends within the Dallas-Fort Worth metroplex according to several time-scale models to understand how the climate for the locality has changed and to provide a basis for projections of what future climate might look like. Trends in mean temperature and precipitation for North Central Texas generally correlate with corresponding global trends linked to natural variability and anthropogenic-induced climate change. Temperature rises in this region annually by 0.08°C and by 0.22°C for a three-decade average. Seasonal increases in three-decadal averages of temperature for North Central Texas relative to the average of the reference …
Predictive Inference For Ion Concentration With Machine Learning And Bayesian Methods, Alexandra B. Ulbing
Predictive Inference For Ion Concentration With Machine Learning And Bayesian Methods, Alexandra B. Ulbing
Theses and Dissertations
Ultraviolet--visible (UV--Vis) spectroscopy produces high-dimensional signals that are strongly collinear, shift with concentration, and exhibit heteroskedastic, non-Gaussian noise. These features make supervised regression from spectra to ionic concentrations statistically challenging and limit the reliability of methods that assume linear structure or homoscedastic errors.
This dissertation develops two complementary frameworks for prediction and uncertainty quantification in UV--Vis spectroscopic regression: (1) frequentist stacked ensembles combined with distribution-free conformal prediction, and (2) Bayesian hierarchical modeling and Bayesian stacking. Together, they provide a unified view of model-based and distribution-free uncertainty across nickel and nickel--cobalt datasets.
The frequentist component builds ensembles of Functional Data Analysis …
Analyzing Factors Influencing Employee Turnover In Tech Companies: A Predictive Modeling Approach, Shinjon Ghosh
Analyzing Factors Influencing Employee Turnover In Tech Companies: A Predictive Modeling Approach, Shinjon Ghosh
Theses and Dissertations
Employee turnover poses substantial challenges for technology firms, and understanding its key drivers through predictive modeling is essential for developing effective retention strategies. This study investigates factors influencing employee turnover in technology companies by implementing a predictive modeling approach on the IBM HR Analytics Employee Attrition dataset. The research aims were identifying key factors contributing to employee attrition, developing predictive models to forecast turnover risk, and analyzing interactions among significant predictors. By examining a range of features, the results highlight significant variables (Over Time, Monthly Income, Marital Status, etc.) of attrition and offer actionable insights for developing targeted employee retention …
Changes In Cancer Diagnosis And Survival In The United States During The Covid-19 Pandemic, Justin T. Burus
Changes In Cancer Diagnosis And Survival In The United States During The Covid-19 Pandemic, Justin T. Burus
Theses and Dissertations--Epidemiology and Biostatistics
The COVID-19 Pandemic led to global societal disruptions as political leaders and public health authorities attempted to control the spread of the newly discovered SARS- CoV-2 virus. While these measures were designed to lessen morbidity and mortality from a novel pathogen, their impact was also felt in many other, often unintended ways. The purpose of this dissertation is to use cancer surveillance research methods to examine the association between COVID-19 Pandemic-related disruptions and changes in the normal diagnosis and care of cancer in the United States.
The first two studies of this dissertation analyzed reductions in cancer diagnoses in the …
Mathematical Contributions To The Study Of Chemotaxis And Cell Signaling, Hajr Zam
Mathematical Contributions To The Study Of Chemotaxis And Cell Signaling, Hajr Zam
Graduate Theses, Dissertations, and Problem Reports (ETD)
This dissertation presents results from two mathematical projects concerned with the biology of cells. Chapter 1 provides biological background and places the two mathematical problems in the context of cell signaling. The larger project, with Prof. H. Hattori on a chemotaxis model is presented in Chapters 3 and 4. Work with Prof. \'{A}. Hal\'{a}sz on a chemical reaction network system with linear multimers and two types of labels is presented in Chapter 2. The chemotaxis system describes the one-dimensional dynamics of a species of cells with two chemical species, a chemo-attractant and chemo-repellent. The goal is to analyze the behavior …
Students’ Perceptions Of Self And Peers Predict Self-Reports Of Cheating, Amber M. Henslee, Luke Settles, Sara E. Johnson, Gayla R. Olbricht
Students’ Perceptions Of Self And Peers Predict Self-Reports Of Cheating, Amber M. Henslee, Luke Settles, Sara E. Johnson, Gayla R. Olbricht
Psychological Science Faculty Research & Creative Works
Academic dishonesty and how to address it are common concerns across higher education disciplines, but engineering students admit to higher rates of academic dishonesty than other students. However, first-year students may be particularly receptive to prevention efforts. Considering self-perception, social norming, and behavioral choice theories, we hypothesized that 1.) Students who perceived themself as ethical and more knowledgeable of the consequences for misconduct would be less likely to self-report cheating and 2.) Students who perceived cheating and plagiarism to be common would be more likely to self-report cheating. For this study, freshmen engineering students (N=703) reported their self-perception, perception of …
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
College of Graduate Studies: Theses & Dissertations
Intrusion Detection Systems (IDS) play a crucial role in computer network security by identifying malicious activities and potential cyberattacks. This thesis combines machine learning and cybersecurity by applying Reinforcement Learning (RL) in intrusion detection and response using the NSL-KDD dataset.
We designed and implemented a Q-learning framework where an agent learns to classify network traffic over time by interacting with the environment and receiving rewards based on detection accuracy. We also look at the importance of feature selection and classification techniques and how effective they are in improving model performance, reducing the complexity of computation, and producing more desirable results. …
Evaluation Of Practical Methods To Determine If A Karst Creek Is Gaining Or Losing: Case Study Of Leith Creek, Elizabeth Jones
Evaluation Of Practical Methods To Determine If A Karst Creek Is Gaining Or Losing: Case Study Of Leith Creek, Elizabeth Jones
Graduate Theses/Dissertations
Karst landscapes are abundant in Missouri, with features such as caves, springs, and sinkholes that form through the dissolution of limestone. Leith Creek is a small stream in Polk County, Missouri fed by two springs and the shallow unconfined Springfield Plateau aquifer, a highly karstified aquifer which is made up of limestone and minor interbedded shale-mudstone units. To determine if Leith Creek is gaining or losing, stream flow, water chemistry and temperature sensors were monitored. Stream flow results required multiple visits to take measurements while temperature sensors required two visits, one to install the dataloggers and another to remove the …
Theory And Applications Surrounding Markov Chains, Joseph J. Quisito Jr., Gallean Brown, Elijah Yoder
Theory And Applications Surrounding Markov Chains, Joseph J. Quisito Jr., Gallean Brown, Elijah Yoder
Capstone Showcase
This capstone project explores the Markov Chain – a mathematical model used to describe systems that transition between states based on probabilities. It begins by introducing the fundamental concepts, including transition matrices, state classifications, and stationary distributions. The paper then applies Markov Chain theory to real-world scenarios, such as simulating Snakes and Ladders games, predicting soccer match outcomes for Manchester United, and generating texts from movie lines. Finally, it discusses key findings, challenges, and potential areas for future research in the field.
Effects Of Chain Length, Saturation, And Bases On Saponification, Sarah Fenik
Effects Of Chain Length, Saturation, And Bases On Saponification, Sarah Fenik
Williams Honors College, Honors Research Projects
This project will analyze the effects of chain length and saturation of fatty acids on saponification processes, as well as the effects of the base used in the reaction. Stearic acid, lauric acid, and oleic acid will be used for the fatty acid comparisons, and sodium hydroxide and potassium hydroxide will be used for the base comparisons. Stearic acid is considered a long chain fatty acid, while lauric acid is considered a short chain fatty acid. Oleic acid is a monounsaturated fatty acid. Five soap products are be made: sodium stearate, sodium laurate, sodium oleate, potassium stearate, and potassium oleate. …
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Dissertations, Master's Theses and Master's Reports
Transcriptome-wide association studies (TWAS) have emerged as a powerful strategy to bridge genome-wide association studies (GWAS) with gene regulatory mechanisms by integrating genotypic data with gene expression data. While early TWAS methods typically rely on linear models and single-tissue expression references, recent advances underscore the need for flexible, multi-tissue approaches that can capture heterogeneous regulatory architectures and tissue-specific expression patterns. This dissertation introduces a three‑part research project that advances multi‑tissue transcriptome‑wide association studies (TWAS) along complementary axes of methodology, statistical power, and modelling flexibility.
In chapter One, TWAS‑CTL introduces a two‑stage cross‑tissue learner that trains any user‑chosen single‑tissue imputers (STLs) …
A Time Series Analysis Of The Macroeconomic Indicators, Mia Houston
A Time Series Analysis Of The Macroeconomic Indicators, Mia Houston
Honors Undergraduate Theses
Understanding inflation—particularly across regions and categories—is crucial for effective policymaking, strategic business decisions, and safeguarding vulnerable populations, as it highlights the diverse drivers and impacts of price changes within the economy. This has become increasingly crucial in recent years between the volatile inflation conditions introduced by the COVID-19 pandemic, energy price shocks, and renewed trade tensions and tariffs. This thesis analyzes 77 U.S. monthly inflation time series from 2003 to 2023 using two forecasting approaches: an elementwise Seasonal Autoregressive Integrated Moving Average (SARIMA) model and a Factor-Augmented Vector Autoregressive (FAVAR) model. The data obtained from the Bureau of Labor Statistics …
Role Of C4 Resources In Isotopic Variability In Diet Among Children From Kellis 2 Cemetery, Dakhleh Oasis, Egypt, Faith R. Hendrix
Role Of C4 Resources In Isotopic Variability In Diet Among Children From Kellis 2 Cemetery, Dakhleh Oasis, Egypt, Faith R. Hendrix
Honors Undergraduate Theses
Using stable carbon isotope analysis, this study investigates dietary diversity in children buried at the Kellis 2 Cemetery (c. AD 50–450) in Egypt's Dakhleh Oasis. From the analysis of δ¹³C isotope values in hair keratin and bone collagen, the study reconstructs short-term and long-term dietary signals in juvenile and adult subjects. The aim is to clarify the role of C₄ plants—particularly millet—in weaning and childhood diets in a Romano-Christian Egyptian village context. A total of 631 segmented hair and 54 bone collagen samples were analyzed from 127 juveniles and 97 adults. Juvenile individuals (i.e., under 15 years biological age) showed …
“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King
“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King
Pomona Senior Theses
The work of this thesis is twofold — first, qualitatively characterizing the confluence between the British eugenics and statistics movements in the late 19th and early 20th centuries, and second, quantitatively analyzing the effect of this foundation on pedagogical materials in the growing field of statistics between 1880 and 1970. Towards the first goal, the history of the method of least squares, state statistics, and positive and negative eugenics are outlined, followed by a close reading of the foundational texts authored by Francis Galton and Karl Pearson that introduced linear regression. Towards the latter goal, English-language statistics textbooks published between …
A Method For Empirically Assessing Small Area Estimators Via Bootstrap-Weighted K-Nearest-Neighbor Artificial Populations, With Applications To Forest Inventory, Grayson W. White, Jerzy Wieczorek, Zachariah W. Cody, Emily X. Tan, Jacqueline O. Chistolini, Kelly S. Mcconville, Tracey S. Frescino, Gretchen G. Moisen
A Method For Empirically Assessing Small Area Estimators Via Bootstrap-Weighted K-Nearest-Neighbor Artificial Populations, With Applications To Forest Inventory, Grayson W. White, Jerzy Wieczorek, Zachariah W. Cody, Emily X. Tan, Jacqueline O. Chistolini, Kelly S. Mcconville, Tracey S. Frescino, Gretchen G. Moisen
Faculty Journal Articles
National Forest Inventories monitor forest attributes across a variety of spatial and temporal scales in a given country. Increased interest in reporting and management at smaller scales has driven National Forest Inventories to investigate and adopt small area estimation (SAE) due to the promise of increased precision at these scales. However, comparing and evaluating SAE models for a given application is inherently difficult. Typically, many areas lack enough data to check unit-level modeling assumptions or to assess unit-level predictions empirically; and no ground truth is available for checking area-level estimates. Design-based simulation from artificial populations can help with each of …
Small Area Estimation Of Forest Biomass Via A Two-Stage Model For Continuous Zero-Inflated Data, Grayson W. White, Josh K. Yamamoto, Dinan H. Elsyad, Julian F. Schmitt, Niels H. Korsgaard, Jie Hu, George C. Gaines Iii, Tracey S. Frescino, Kelly S. Mcconville
Small Area Estimation Of Forest Biomass Via A Two-Stage Model For Continuous Zero-Inflated Data, Grayson W. White, Josh K. Yamamoto, Dinan H. Elsyad, Julian F. Schmitt, Niels H. Korsgaard, Jie Hu, George C. Gaines Iii, Tracey S. Frescino, Kelly S. Mcconville
Faculty Journal Articles
Nationwide Forest Inventories (NFIs) collect data on and monitor the trends of forests across the globe. Users of NFI data are increasingly interested in monitoring forest attributes such as biomass at fine geographic and temporal scales, resulting in a need for assessment and development of small area estimation techniques in forest inventory. We implement a small area estimator and parametric bootstrap estimator that account for zero-inflation in biomass data via a two-stage model-based approach and compare the performance to a Horvitz–Thompson estimator, a post-stratified estimator, and to the unit- and area-level empirical best linear unbiased prediction (EBLUP) estimators. We conduct …
Majority Decision Using Top-Performing Neural Networks Models For Improved Credit Risk Prediction, Vincent Dey
Majority Decision Using Top-Performing Neural Networks Models For Improved Credit Risk Prediction, Vincent Dey
College of Graduate Studies: Theses & Dissertations
Credit risk prediction remains both a challenging and high-interest problem due to the inherently unbalanced nature of financial datasets and the continuous drive for higher pre- dictive precision. In this work, I build upon previous advancements in credit risk modeling and introduce an ensemble-based Artificial Neural Network (ANN) architecture designed to enhance classification performance. By leveraging a selective ensemble of decision net- works, this approach not only improves prediction accuracy but also mitigates the chal- lenges posed by imbalanced data distributions. While the primary focus is on credit risk prediction, my analysis demonstrates that the proposed model can be effectively …
Theoretical Foundations And Applied Performance Of Periodicity-Aware Imputation: Variable Bandpass Block Bootstrap Methods For Incomplete Time Series, Asmaa Ahmad
Electronic Theses & Dissertations (2024 - present)
Time series data are prevalent across a wide range of disciplines, including health surveillance, public policy, and environmental monitoring. In the presence of underlying cyclical patterns, the integrity of time series analysis depends critically on the ability to detect, model, and impute structured missing data without compromising the temporal structure. This dissertation introduces and validates a novel imputation framework that integrates the Variable Bandpass Periodic Block Bootstrap (VBPBB) into multiple imputation procedures, improving the accuracy, robustness, and interpretability of time series models under high rates of missingness and noise. The overarching goal of this dissertation was to develop and evaluate …
Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri
Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri
CMC Senior Theses
Over the past decades, the gaming industry has managed to evolve into a multi-billion-dollar enterprise. Gaming platforms such as Steam foster unprecedented amounts of engagement among players worldwide daily. In this thesis, we investigate the effect of incorporating sentiment-driven metrics, specifically YouTube view counts and positive reviews, into predictive models for game popularity. In addition, by comparing our linear regression sentiment-based approach to the Bayesian hierarchical folded normal model used by De Luisa et al. (2021), we can understand the many differences, strengths, and limitations of each methodology. In our thesis, we focus on three games. Each is of varying …
Forecasting Equity Betas Using Option-Implied Moments, Ivan Kolesnikov
Forecasting Equity Betas Using Option-Implied Moments, Ivan Kolesnikov
CMC Senior Theses
Traditional beta estimates are constructed from historical stock‑and‑market returns and therefore adjust only as fast as realized data accrue. This thesis investigates whether the forward‑looking information embedded in equity‑option prices can enhance beta forecasts. Using near‑end‑of‑day quotes for 236 S&P 500 firms between 2007 and 2024, I extract risk‑neutral variance and skewness, construct five alternative beta estimators (historical, option‑implied, and three hybrids), and evaluate them against realized betas over six‑, twelve‑, and twenty‑four‑month windows. Rolling‑OLS beta remains the most accurate benchmark at short horizons, yet option‑implied moments add economically and statistically significant value when systematic exposure is expected to change …
Analyzing Political Sentiment On Micro-Blogging Data: A Lexicon And Machine Learning Approach To The 2024 U.S. Presidential Election, Ava Grey
CMC Senior Theses
This paper explores the trends in sentiment towards U.S. presidential candidates Kamala Harris and Donald Trump through micro-blogging social media text during the five months leading up to the election. Two datasets of varying sizes and origins were used to contextualize and validate analysis findings. The analyses include both a lexicon-based approach and a machine learning predictive method. Common sentiment analysis techniques like term frequency, term frequency inverse, various lexicons, and n-grams were utilized during the lexicon approach. During the modeling, a random forest was utilized in addition to the methods used during the lexicon approach. Results showed that overall …
Applications Of Bayesian Functional Data Analysis, Zhexuan Yang
Applications Of Bayesian Functional Data Analysis, Zhexuan Yang
Graduate Research Theses & Dissertations
Functional Data Analysis (FDA) is a statistical approach used to analyze data that vary across a domain, such as curves or functions. This dissertation investigates Bayesian Functional Data Analysis (BFDA) through three applications. First, we explore the use of BFDA in outcome-dependent follow-up (ODFL) studies. After conducting simulation studies, we apply our model to cardiotoxicity and kidney function data. Second, we extend BFDA to genetic data by modeling DNA methylation levels with a three-parameter skew-normal distribution and an alpha-skew generalized normal distribution. This study also introduces a novel Multistage Markov Chain Monte Carlo (MMCMC) method with the goal of identifying …
Covariance Matrix Forecasting Of Equity Portfolios, Michael Nebor
Covariance Matrix Forecasting Of Equity Portfolios, Michael Nebor
Graduate Research Theses & Dissertations
This dissertation consists of two papers. The first paper introduces DCC-SVR, a hybrid Dynamic Conditional Correlation (DCC) and Support Vector Regression (SVR) method of forecasting the covariance matrix. This paper shows that DCC-SVR is able to outperform the traditional methods of DCC and rolling historical on multiple data sets. Performance is shown for both standard GARCH and GJR-GARCH methods. This paper also analyzes performance when dimensions are increased to 49 dimensions and when an application using equal weighted portfolio allocation is used.
The second paper introduces a covariance matrix forecasting method based on copula-GARCH simulated returns. The accuracy of this …
A Modern Optimization Approach With Data-Driven Analytical Modeling For The Healthcare Business Segment (Hbs) From The S&P 500, Aditya Chakraborty, Chris Tsokos
A Modern Optimization Approach With Data-Driven Analytical Modeling For The Healthcare Business Segment (Hbs) From The S&P 500, Aditya Chakraborty, Chris Tsokos
Epidemiology, Biostatistics, & Environmental Health Faculty Publications
Introduction: The S&P consists of eleven business segments, which are classified according to the type of industry. The current study focuses on developing a non-linear analytical model for the Healthcare Business Segment (HBS) of the S&P 500, as a function of different economic & financial indicators. Materials and Methods: The analytical model used six financial indicators together with four economic indicators to predict the weekly average closing price (WCP) of HBS stocks. Johnson’s SB transformation corrected skewness, while desirability-based optimization identified indicator values maximizing WCP. The model’s performance and generalizability were validated through repeated 10-fold cross-validation. Results: All attributable contributors …
Optimal Data Splitting Methods, Sujay Mudalgi
Optimal Data Splitting Methods, Sujay Mudalgi
Theses and Dissertations
In predictive modeling, effective data splitting is crucial for creating statistically representative training and validation sets. The state-of-the-art data splitting methods are based on minimizing the energy distance between the split subsets. However, there are a number of limitations in the existing methods, which this dissertation aims to address. First, the existing methods were computationally inefficient. Thus, Chapter 2 proposes a method to scale up these approaches for big data. Here, we introduce scalable Twinning (s-Twinning), which significantly improves the execution speed of data splitting without sacrificing accuracy. Second, the existing methods did not consider the predictive relationship in the …
Climate Migration And Urban Survival: Evidence From Chittagong’S Slums, Mohammad Nur Nobi
Climate Migration And Urban Survival: Evidence From Chittagong’S Slums, Mohammad Nur Nobi
Graduate Research Theses & Dissertations
This study examines the impact of climate change on rural-to-urban migration in Chittagong, Bangladesh. Based on a primary survey of 400 respondents across 35 slums in the country's second-largest city, the analysis employs two estimation methods: a multinomial logit model to assess the influence of climate-related factors on migration decisions, and a logit model to evaluate the impact of migration on the living conditions of migrants. The results show that individuals involved in ‘agriculture and daily labor’ are most likely to migrate. Compared to the base category (‘Other Reasons’) for migration, the odds ratios for ‘floods’, ‘droughts’, and ‘better job …