Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Central Bank of Nigeria (21)
- Kennesaw State University (20)
- Southern Methodist University (16)
- University of Central Florida (8)
- California Polytechnic State University, San Luis Obispo (6)
-
- Chapman University (6)
- Georgia Southern University (6)
- Illinois State University (6)
- The University of Akron (5)
- Virginia Commonwealth University (5)
- City University of New York (CUNY) (4)
- Claremont Colleges (4)
- East Tennessee State University (4)
- Rochester Institute of Technology (4)
- University of Kentucky (4)
- Binghamton University (3)
- Louisiana State University (3)
- Murray State University (3)
- Purdue University (3)
- University of New Hampshire (3)
- West Virginia University (3)
- Bucknell University (2)
- Clemson University (2)
- Embry-Riddle Aeronautical University (2)
- Florida Institute of Technology (2)
- Michigan Technological University (2)
- Portland State University (2)
- South Dakota State University (2)
- Syracuse University (2)
- University at Albany, State University of New York (2)
- Keyword
-
- Machine learning (15)
- Statistics (15)
- Machine Learning (10)
- Data science (6)
- Regression (6)
-
- COVID-19 (5)
- Clustering (5)
- Forecasting (4)
- Time series (4)
- Classification (3)
- Logistic Regression (3)
- Logistic regression (3)
- NBA (3)
- Nigeria (3)
- Prediction (3)
- Analytics (2)
- Artificial Intelligence (2)
- Autocorrelation (2)
- Baseball (2)
- Basketball (2)
- Bayesian (2)
- Bayesian Analysis (2)
- Biostatistics (2)
- Cybersecurity (2)
- Data Science (2)
- Data analysis (2)
- Data generation (2)
- Economic Growth (2)
- Education (2)
- Electronic Health Records (2)
- Publication Year
- Publication
-
- CBN Journal of Applied Statistics (JAS) (20)
- Symposium of Student Scholars (19)
- SMU Data Science Review (11)
- Theses and Dissertations (9)
- Data Science and Data Mining (7)
-
- College of Graduate Studies: Theses & Dissertations (6)
- Electronic Theses and Dissertations (6)
- Master's Theses (6)
- Williams Honors College, Honors Research Projects (5)
- Annual Symposium on Biomathematics and Ecology Education and Research (4)
- Articles (4)
- Computational and Data Sciences (PhD) Dissertations (3)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (3)
- LSU Doctoral Dissertations (3)
- Northeast Journal of Complex Systems (NEJCS) (3)
- Statistical Science Theses and Dissertations (3)
- Theses and Dissertations--Statistics (3)
- Beyond: Undergraduate Research Journal (2)
- CMC Senior Theses (2)
- Dissertations, Master's Theses and Master's Reports (2)
- Electrical and Computer Engineering ETDs (2)
- Electronic Theses & Dissertations (2024 - present) (2)
- Faculty Journal Articles (2)
- Honors College Theses (2)
- Honors Theses and Capstones (2)
- Numeracy (2)
- Publications and Research (2)
- SDSU Data Science Symposium (2)
- Spora: A Journal of Biomathematics (2)
- Sport Management - All Scholarship (2)
- Publication Type
- File Type
Articles 31 - 60 of 193
Full-Text Articles in Data Science
Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens
Statistical Science Theses and Dissertations
Electronic Health Records (EHR) contain a wealth of structured and unstructured patient data that can be leveraged for computable phenotyping, the process of algorithmically identifying patient cohorts with specific diseases or conditions. Traditional rule-based phenotyping approaches, while interpretable, often struggle with scalability, portability across institutions, and effective use of unstructured clinical narratives. Recent advances in large language models (LLMs) present new opportunities for synthesizing complex free-text information into concise, clinically meaningful representations. However, integrating LLMs into phenotyping workflows requires careful design to maintain transparency, interpretability, and measurable uncertainty—features essential for clinical adoption and downstream applications such as decision support.
We …
Nba Player Types And Salaries: Assessing The Disparities In Pay, Nick Riccardi, Rodney J. Paul
Nba Player Types And Salaries: Assessing The Disparities In Pay, Nick Riccardi, Rodney J. Paul
Sport Management - All Scholarship
The purpose of this study was to identify player types that exist in the modern National Basketball Association (NBA), test whether player types are paid differently controlling for performance and other factors and construct successful rosters with cheaper payrolls.
We collected performance statistics and salary data for players and teams across five seasons (2018-19 to 2022-23). Cluster analysis is leveraged to group together player-seasons to identify the player types that exist in the NBA. Linear regression models are run to test for differences in pay by cluster membership while controlling for performance, age, and contractual details. Linear programming simulation models …
Experimental Design And Analysis For Decision Making: Methodology And Applications, Yezhuo Li
Experimental Design And Analysis For Decision Making: Methodology And Applications, Yezhuo Li
All Dissertations
This dissertation develops and applies advanced statistical and optimization frameworks to enhance decision-making under uncertainty, particularly in engineering and manufacturing contexts. First, we introduce an approach for the optimal design of controlled experiments that accounts for observational covariates, enabling more precise and personalized decisions. Second, we explore the application of constrained Bayesian optimization, using Gaussian process surrogate models, to optimize composite cure processes, significantly reducing computational effort while maintaining high predictive accuracy. Building on this foundation, we extend Bayesian optimization to bivariate Gaussian process models that capture correlations between objective and constraint functions, offering new insights into multidimensional decision landscapes. …
Online Prediction Of Streaming Data, Aleena Chanda
Online Prediction Of Streaming Data, Aleena Chanda
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
We present two new approaches for point prediction with streaming data based on a) the Count-Min sketch and b) Gaussian Process Priors with random bias. The methods are intended for the most general case where no true model can be usefully formulated for the data stream. In statistical contexts, this is often called the M open problem class. For the Count Min Sketch method we show that the predicted distribution function ^F converges to F under the assumption that the data consists of i.i.d samples from a fixed distribution function F. To implement the Gaussian Process Prior methods, we used …
Estimation Methods For Bayesian Exponential Random Graph Models Under The Horseshoe Prior., Pamela Linares
Estimation Methods For Bayesian Exponential Random Graph Models Under The Horseshoe Prior., Pamela Linares
Electronic Theses and Dissertations
Networks are powerful tools for modeling the complexity of social interactions, biological systems, and information spread. A leading statistical frameworks for analyzing network data are Exponential Random Graph Models (ERGMs), which provide a principled approach to capturing structural dependencies. However, ERGMs remain challenging to estimate, especially in sparse or high-dimensional settings where models suffer from degeneracy and unstable parameter inference. This paper proposes a penalized Bayesian approach to ERGMs that utilizes the horseshoe prior, a sparsity-inducing global-local shrinkage prior. This prior offers robust regularization while preserving important signals, improving estimation by shrinking irrelevant parameters and reducing the impact of extreme …
A Comparative Study Of Neural Networks And Xgboost Models For Flight Time Prediction, Ioannis Paraschos, Taryn E. Trimble, Eshna Bhargava, Jake Klingler, Benjamin R. Nicolai
A Comparative Study Of Neural Networks And Xgboost Models For Flight Time Prediction, Ioannis Paraschos, Taryn E. Trimble, Eshna Bhargava, Jake Klingler, Benjamin R. Nicolai
Beyond: Undergraduate Research Journal
Flight time prediction plays a crucial role in modern air travel, benefiting airlines and passengers alike. Accurate predictions enable airlines to optimize schedules, allocate resources effectively, and ensure passenger safety and satisfaction. In recent years, machine learning models, such as neural networks and XGBoost, have gained popularity for predicting flight times. This study aims to compare the performance of neural network and XGBoost models in predicting flight times, considering factors such as weather conditions, air traffic control, and aircraft performance. The results indicate that both models are effective, with XGBoost achieving slightly higher accuracy. However, neural networks offer advantages in …
Data Driven Analysis Of Samara Seed Kinematics And Dynamics, Shashwat Sparsh
Data Driven Analysis Of Samara Seed Kinematics And Dynamics, Shashwat Sparsh
Master's Theses
Samara Seeds are a class of fruit most famously belonging to the Acer species and are characterized by their single-bladed geometry and their auto-rotation response during descent. This steady-state auto-rotation response is the subject of aerodynamic analysis which aim to quantify the performance. The period prior to the beginning of steady-state auto-rotation is classified as the transition regime and has not been the subject of intense scrutiny.
This thesis employs a data-driven approach to analyzing the kinematic and dynamic response of these seeds during both the transition and auto-rotation stages of flight to quantify the performance with respect to the …
Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi
Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi
Open Educational Resources
Data analysis using standard statistical methods and relevant computer software. Emphasis on real-world data, interpretation, and misinterpretation of computer output.
This syllabus contains open source notebook about data analysis content.
Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson
Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson
Electronic Theses and Dissertations
The beef industry plays a vital role in global agriculture, with carcass quality and consumer preference being key determinants of market success. This thesis examines predictive modeling techniques for estimating the Total Score of beef carcasses, a composite measure representing yield and quality, primarily used by the Nebraska Cattlemen Association. Using data from the Nebraska Cattlemen’s Foundation Retail Value Steer Challenge (2000–2023), the study compares the performance of First Order Multiple Linear Regression (MLR) with three machine learning techniques: K-Nearest Neighbors (KNN), Random Forest, and Gradient Boosting Machine (GBM).
The analysis focuses on six key predictors: Hot Carcass Weight, Back …
Statistics - What Does My Data Say About Me?, Taylor Gadsden-Deterville
Statistics - What Does My Data Say About Me?, Taylor Gadsden-Deterville
Student Scholar Symposium Abstracts and Posters
For my Introduction to Statistics Class, I have been tasked with collecting unique, personal data to give insight into my daily routine. I decided to record nine different outcomes (two qualitative and seven quantitative). On February 6, 2025, I began with a blank Excel sheet, and so far, I have 57 full days of data collected. I will continue monitoring my findings for the remainder of the Spring 2025 Semester. Per my project instructions, I must include tables and graphs for my qualitative and quantitative outcomes. So far, I have collected daily quantitative data on my screen time (Instagram and …
Multimodal Benchmarking For Ncaa Basketball, Brendan Barnett
Multimodal Benchmarking For Ncaa Basketball, Brendan Barnett
Honors Scholar Theses
We present the first multimodal, multitask benchmark for NCAA basketball, synthesizing structured statistical features with large language model (LLM)-generated game summaries across 19,739 games spanning four NCAA Division I seasons (2021--2025). We evaluate three model families---XGBoost, deep neural networks, and Transformers---under tabular-only and early-fusion settings to measure the impact of LLM-derived textual embeddings. To assess practical utility, we simulate fixed-stake and Kelly criterion-based betting strategies using historical bookmaker odds, analyzing both profitability and downside risk via Monte Carlo simulation. Our results show that XGBoost with early-fusion achieves the highest return on investment and the lowest risk of loss. This work …
Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins
Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins
Honors College Theses
The financial crisis of the early 2000’s is a prime example of the severe consequences that mortgage default and borrower insolvency can have on economies at large. Mortgage default specifically is a prime case with the popularization of mortgage backed securities and the commonality of this loan structure. Multiple hypotheses and models have been formed to understand the reasons, causes, and consequences of mortgage default. This paper uses both machine learning and statistical classification models to inform an understanding of the variables most significant and impactful to the default outcome of mortgages. Consideration is given to both loan-level microeconomic variables …
Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun
Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun
SMU Data Science Review
Enhancing animal shelter operations through machine learning involves employing a variety of advanced techniques aimed at increasing efficiency, promoting animal welfare, and optimizing resource allocation. This paper explores predictive analytics for adoption rates using regression models to estimate the likelihood of adoption based on historical data, encompassing variables such as breed, health status, and previous adoption trends. Additionally, classification algorithms are utilized to categorize animals by adoption probability, facilitating better resources and marketing prioritization. Clustering algorithms are employed to group animals according to behavior patterns and/or physical health, enabling tailored medical care and enrichment activities that improve their mental and …
A Machine-Learning Tool-Supported Methodology For Nonprofit Donor Analysis, Corbin Weiss
A Machine-Learning Tool-Supported Methodology For Nonprofit Donor Analysis, Corbin Weiss
Campus Research Month
We developed a machine-learning tool-supported methodology for modeling the nonprofit donor relationship. This approach was demonstrated in the case of a US-based nonprofit. Conclusions were drawn from this example and tool-support provided for use by other nonprofits.
Linking Water Quality And Climate Change To Long-Term Trends In Species Abundance In Norwalk Harbor, Viktoria Savatorova, Aidan Kieft, Nicole C. Spiller, Kasey Burns
Linking Water Quality And Climate Change To Long-Term Trends In Species Abundance In Norwalk Harbor, Viktoria Savatorova, Aidan Kieft, Nicole C. Spiller, Kasey Burns
Spora: A Journal of Biomathematics
This study examines the effects of environmental changes on fish populations in Norwalk Harbor, focusing on winter flounder (Pseudopleuronectes americanus), cunner (Tautogolabrus adspersus), northern pipefish (Syngnathus fuscus), and naked goby (Gobiosoma bosci) as examples of species responding to climate-related shifts. We analyze how water temperature, salinity, and dissolved oxygen correlate with fish abundance. To assess statistically significant differences in catch per unit effort (CPUE) across harbor regions, we applied the Kruskal-Wallis test followed by Dunn's post-hoc test. Seasonal variations in CPUE were examined by comparing monthly catch data for each species. K-means …
Urban Heat Dynamics In Pune: The Influence Of Land Cover And Local Climate, Arpit Tiwari, Preethi Nanjundan, Ravi Ranjan Kumar, Ananya Karmakar, Satyaban Bishoyi Ratna
Urban Heat Dynamics In Pune: The Influence Of Land Cover And Local Climate, Arpit Tiwari, Preethi Nanjundan, Ravi Ranjan Kumar, Ananya Karmakar, Satyaban Bishoyi Ratna
Northeast Journal of Complex Systems (NEJCS)
Urban areas with high population density and extensive infrastructure development have been experiencing an increasing strain on the local heat budget, leading to a surge in heat-related illnesses and discomfort. This study examined the impact of climate and land use as heat islands in Pune, India, from 2012 to 2023 at six different locations representing varying degree of urbanization. Satellite land cover observations revealed that 55.17% of the total area was urbanized in the city itself, which was limited to 44.8% in 2012. This urbanization has significantly impacted the increasing tendency of maximum temperature (Tmax; 0.13℃ to 1.63℃ …
Analysis Of Systematic Trade-Offs Between Military And Healthcare Expenditure Alongside Gdp Growth Of Select Asian And Western Exporting Economies In The 21st Century, Rahul Balamurugan, Carlos Gershenson, Preethi Nanjundan, Hiroki Sayama
Analysis Of Systematic Trade-Offs Between Military And Healthcare Expenditure Alongside Gdp Growth Of Select Asian And Western Exporting Economies In The 21st Century, Rahul Balamurugan, Carlos Gershenson, Preethi Nanjundan, Hiroki Sayama
Northeast Journal of Complex Systems (NEJCS)
This study explores the complexity in the trade-offs between military expenditure, healthcare expenditure, and GDP growth across select Asian nations and major weapon-exporting countries, examining how nations allocate finite resources between national security and human well-being over the past two decades. Using a systems science approach, the research integrates Granger causality testing to analyze temporal and directional relationships among GDP growth, military expenditure, and healthcare expenditure, uncovering their dynamic interdependencies. The methodology includes trend and slope analysis, Granger causality testing, outlier detection, and clustering to identify heterogeneity in resource allocation strategies. Developed, weapon-exporting nations exhibit complementary trends, with strong causality …
Optimized Hiv/Aids Resource Allocation In Ohio: A Linear Programming Approach, Godfred Ahenkroa Kesse
Optimized Hiv/Aids Resource Allocation In Ohio: A Linear Programming Approach, Godfred Ahenkroa Kesse
Data Science and Data Mining
This study employs a linear and integer programming approach to optimize HIV resource allocation in Ohio, aiming to minimize new infections and enhance the impact of limited resources. With the advances in HIV prevention and treatment, Ohio faces challenges in addressing disparities in access to healthcare, particularly among high-risk populations. The proposed model integrates data on infection rates, transmission patterns, demographic factors, and cost-effectiveness to provide a decision-support framework for policymakers. Using epidemiological data and equity constraints, the model prioritizes high-risk regions and populations while ensuring fair resource distribution. Results indicate that increased funding allocations significantly enhance the potential to …
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Posters-at-the-Capitol
The grocery retail industry landscape has changed greatly in the wake of the pandemic. Specifically, delivery and pickup services have become more popular and customer buying habits have evolved. At the same time, improvements in data collection and analysis have allowed grocery marketing strategies to become highly individualized.
We worked with 84.51, an analytics firm, to identify customer segments for the Kroger Company based on data from 2023. Using clustering techniques, we organized customers into groups, or segments, based on similar characteristics. We identified and profiled four distinct groups of customers. Three segments were characterized by high frequency and spending …
Predicting Superconducting Critical Temperature From Composition-Derived Features: A Transparent Linear And Regularized Regression Study, Md Ahiduzzaman
Predicting Superconducting Critical Temperature From Composition-Derived Features: A Transparent Linear And Regularized Regression Study, Md Ahiduzzaman
Data Science and Data Mining
We study prediction of superconducting critical temperature (Tc) from 81 composition-derived descriptors across 21,263 materials. To keep the analysis transparent and repro- ducible, we focus on linear models: Ordinary Least Squares (OLS), Ridge, Lasso, and Elastic Net (ENet). All models share a single evaluation protocol (5-fold cross-validation with standardized inputs) and are compared on RMSE, MAE, and R2. On this feature set, OLS attains the best cross-validated performance (RMSE = 17.6 K, MAE = 13.3 K , R2 = 0.735), with Lasso/ENet essentially tied next (RMSE ≈ 17.7 K , R2 ≈ 0.734); Ridge underperforms (RMSE = 18.9 K , …
Comparative Analysis Of Lasso, Ridge, And Elastic Net For Variable Selection In High-Dimensional Maize Data, Md Ahiduzzaman
Comparative Analysis Of Lasso, Ridge, And Elastic Net For Variable Selection In High-Dimensional Maize Data, Md Ahiduzzaman
Data Science and Data Mining
In high-dimensional genomic data analysis, traditional linear regression techniques often struggle due to the presence of a large number of predictor variables relative to observations. Penalized regression methods such as LASSO, Ridge, and Elastic Net have emerged as effective solutions by imposing regularization, which helps in managing multicollinearity and enhancing prediction accuracy. This study applies these techniques to the Maize dataset to model the time to male flowering, selecting relevant genetic markers as predictors. Our findings suggest that Elastic Net is particularly effective for high-dimensional data with correlated variables, achieving a balance between prediction accuracy and variable selection. The results …
Analyzing Political Sentiment On Micro-Blogging Data: A Lexicon And Machine Learning Approach To The 2024 U.S. Presidential Election, Ava Grey
CMC Senior Theses
This paper explores the trends in sentiment towards U.S. presidential candidates Kamala Harris and Donald Trump through micro-blogging social media text during the five months leading up to the election. Two datasets of varying sizes and origins were used to contextualize and validate analysis findings. The analyses include both a lexicon-based approach and a machine learning predictive method. Common sentiment analysis techniques like term frequency, term frequency inverse, various lexicons, and n-grams were utilized during the lexicon approach. During the modeling, a random forest was utilized in addition to the methods used during the lexicon approach. Results showed that overall …
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
College of Graduate Studies: Theses & Dissertations
Intrusion Detection Systems (IDS) play a crucial role in computer network security by identifying malicious activities and potential cyberattacks. This thesis combines machine learning and cybersecurity by applying Reinforcement Learning (RL) in intrusion detection and response using the NSL-KDD dataset.
We designed and implemented a Q-learning framework where an agent learns to classify network traffic over time by interacting with the environment and receiving rewards based on detection accuracy. We also look at the importance of feature selection and classification techniques and how effective they are in improving model performance, reducing the complexity of computation, and producing more desirable results. …
A Method For Empirically Assessing Small Area Estimators Via Bootstrap-Weighted K-Nearest-Neighbor Artificial Populations, With Applications To Forest Inventory, Grayson W. White, Jerzy Wieczorek, Zachariah W. Cody, Emily X. Tan, Jacqueline O. Chistolini, Kelly S. Mcconville, Tracey S. Frescino, Gretchen G. Moisen
A Method For Empirically Assessing Small Area Estimators Via Bootstrap-Weighted K-Nearest-Neighbor Artificial Populations, With Applications To Forest Inventory, Grayson W. White, Jerzy Wieczorek, Zachariah W. Cody, Emily X. Tan, Jacqueline O. Chistolini, Kelly S. Mcconville, Tracey S. Frescino, Gretchen G. Moisen
Faculty Journal Articles
National Forest Inventories monitor forest attributes across a variety of spatial and temporal scales in a given country. Increased interest in reporting and management at smaller scales has driven National Forest Inventories to investigate and adopt small area estimation (SAE) due to the promise of increased precision at these scales. However, comparing and evaluating SAE models for a given application is inherently difficult. Typically, many areas lack enough data to check unit-level modeling assumptions or to assess unit-level predictions empirically; and no ground truth is available for checking area-level estimates. Design-based simulation from artificial populations can help with each of …
Small Area Estimation Of Forest Biomass Via A Two-Stage Model For Continuous Zero-Inflated Data, Grayson W. White, Josh K. Yamamoto, Dinan H. Elsyad, Julian F. Schmitt, Niels H. Korsgaard, Jie Hu, George C. Gaines Iii, Tracey S. Frescino, Kelly S. Mcconville
Small Area Estimation Of Forest Biomass Via A Two-Stage Model For Continuous Zero-Inflated Data, Grayson W. White, Josh K. Yamamoto, Dinan H. Elsyad, Julian F. Schmitt, Niels H. Korsgaard, Jie Hu, George C. Gaines Iii, Tracey S. Frescino, Kelly S. Mcconville
Faculty Journal Articles
Nationwide Forest Inventories (NFIs) collect data on and monitor the trends of forests across the globe. Users of NFI data are increasingly interested in monitoring forest attributes such as biomass at fine geographic and temporal scales, resulting in a need for assessment and development of small area estimation techniques in forest inventory. We implement a small area estimator and parametric bootstrap estimator that account for zero-inflation in biomass data via a two-stage model-based approach and compare the performance to a Horvitz–Thompson estimator, a post-stratified estimator, and to the unit- and area-level empirical best linear unbiased prediction (EBLUP) estimators. We conduct …
Majority Decision Using Top-Performing Neural Networks Models For Improved Credit Risk Prediction, Vincent Dey
Majority Decision Using Top-Performing Neural Networks Models For Improved Credit Risk Prediction, Vincent Dey
College of Graduate Studies: Theses & Dissertations
Credit risk prediction remains both a challenging and high-interest problem due to the inherently unbalanced nature of financial datasets and the continuous drive for higher pre- dictive precision. In this work, I build upon previous advancements in credit risk modeling and introduce an ensemble-based Artificial Neural Network (ANN) architecture designed to enhance classification performance. By leveraging a selective ensemble of decision net- works, this approach not only improves prediction accuracy but also mitigates the chal- lenges posed by imbalanced data distributions. While the primary focus is on credit risk prediction, my analysis demonstrates that the proposed model can be effectively …
Optimal Data Splitting Methods, Sujay Mudalgi
Optimal Data Splitting Methods, Sujay Mudalgi
Theses and Dissertations
In predictive modeling, effective data splitting is crucial for creating statistically representative training and validation sets. The state-of-the-art data splitting methods are based on minimizing the energy distance between the split subsets. However, there are a number of limitations in the existing methods, which this dissertation aims to address. First, the existing methods were computationally inefficient. Thus, Chapter 2 proposes a method to scale up these approaches for big data. Here, we introduce scalable Twinning (s-Twinning), which significantly improves the execution speed of data splitting without sacrificing accuracy. Second, the existing methods did not consider the predictive relationship in the …
Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem
Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem
Dissertations, Master's Theses and Master's Reports
Factor analysis is a powerful tool for modeling latent structures in high-dimensional data, traditional approaches assume a single global structure, limiting their ability to capture heterogeneity. The Mixture of Factor Analyzers (MFA) extends classical factor analysis by modeling data as a mixture of Gaussian-distributed local subspaces, effectively uncovering cluster-specific latent structures. However, MFA relies on Gaussian mixtures, making it sensitive to outliers and ill-suited for heavy-tailed data. The Mixture of $t$-Factor Analyzers (M$t$FA) addresses these limitations by incorporating multivariate $t$-distributions, improving robustness. Despite their advantages, both MFA and M$t$FA face significant computational challenges in high-dimensional settings, particularly due to costly …
Calculation And Statistical Analysis Of Wins Above Replacement, Joshua Taylor
Calculation And Statistical Analysis Of Wins Above Replacement, Joshua Taylor
Departmental Honors & Graduate Capstone Projects
The Wins Above Replacement (WAR) statistic in Major League Baseball is a prominent metric used to estimate player value by quantifying all aspects of play in terms of wins added to a baseball team. We will use R to calculate WAR for all players from 1871 to 2012 and use data from those years to construct multivariate predictive models to attempt to estimate WAR for players from 2013 to 2024. We find strong correlations between predicted and actual WAR values for most models, with the exception of the polynomial predictive model for non-qualified pitchers.
Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny
Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny
All Theses
High blood pressure, also known as hypertension, significantly increases the risk of heart disease and stroke, which are leading causes of death in the United States. While contributing to over 691,000 deaths in 2021 alone in the United States (U.S.), it also imposes immense economic burden on the healthcare system, costing approximately $131 billion annually. One way to address this issue is for increased self-care behaviors and medication adherence, both of which require sufficient health literacy. Despite the importance of health literacy, 90% of U.S. adults struggle with health-related subjects. Overcoming the issues associated with health literacy requires addressing the …