Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Statistical Methodology

Institution
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 30 of 109

Full-Text Articles in Data Science

Predicting Remaining Useful Life Using Multivariate Time-Series Data, Anayah Smith, Victoria Gaibor Aug 2026

Predicting Remaining Useful Life Using Multivariate Time-Series Data, Anayah Smith, Victoria Gaibor

Discovery Day - Daytona Beach

Accurate prediction of Remaining Useful Life (RUL) is critical for enabling predictive maintenance, improving system reliability, and reducing operational costs in degrading systems. This project addresses the problem of modeling and predicting RUL using multivariate time-series sensor data from the NASA CMAPSS turbofan engine dataset, with a focus on understanding how predictive performance changes across datasets of varying complexity. The objective is to develop a reproducible machine learning pipeline that captures degradation patterns and produces reliable time-to-failure predictions. The approach includes data preprocessing, exploratory data analysis, feature engineering, dimensionality reduction, and model evaluation. RUL values are computed and capped to …


Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski Jun 2026

Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski

Master's Theses

Humpback whale songs are notoriously complex. Identification of humpback whale song units requires bioacousticians to tediously listen, analyze, and annotate collected sound data. Even sparse data requires listening to the entirety of the collected acoustic data. In this study, three hours of audio containing over one-thousand humpback whale song units was collected in Monterey Bay, California.

Prior studies have seen success using convolutional neural networks by performing image classification on hundreds of hours worth of spectrograms. Our study uses traditional machine learning models, as they are less computationally demanding, and require less data.

We use time splitting and Mel-frequency cepstrum …


Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez May 2026

Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez

Statistical Science Theses and Dissertations

Impact evaluations of regional development programs often require estimating counterfactual outcomes for a small number of treated regions using survey-based areal data. In practice, evaluators typically rely on two-group quasi-experimental methods such as propensity score matching (PSM) and Difference-in-Differences (DiD). These approaches perform poorly when only a few regions receive treatment, and when the set of observed covariates is limited or only partially relevant. Moreover, they typically do not explicitly exploit the spatial and temporal dependence present in survey-based areal data such as in ACS (American Community Survey). This dissertation develops a family of Bayesian spatial predictive models for directly …


The Impatience Of Winning: An Analysis Of Time Discounting, Predictive Modeling, And The Nba Draft, Alec R. Plante May 2026

The Impatience Of Winning: An Analysis Of Time Discounting, Predictive Modeling, And The Nba Draft, Alec R. Plante

Business and Economics Honors Papers

This paper examines whether NBA draft decisions can be better explained by incorporating non-geometric time discounting into a model of general manager decision making. Using a dataset of 285 NBA draft prospects over a 12-year period, the impact of college statistics on Value Over Replacement Player (VORP) is determined, and these impact values are then used to create a “predicted” VORP for the first 4 seasons of each player’s career: a projection of what a general manager might think of a prospect’s future value given their college statistics. Following this, geometric and hyperbolic time discounting models are applied to estimate …


The Item Response Warehouse: What It Is, How To Use It, And Targets For Potential Improvements, Savira D. Nadela, Hansol Lee, Nishka Jain, Ayaan Gupta, Xingyi Zhang, Benjamin W. Domingue Apr 2026

The Item Response Warehouse: What It Is, How To Use It, And Targets For Potential Improvements, Savira D. Nadela, Hansol Lee, Nishka Jain, Ayaan Gupta, Xingyi Zhang, Benjamin W. Domingue

Chinese/English Journal of Educational Measurement and Evaluation | 教育测量与评估双语期刊

The Item Response Warehouse (IRW) is a repository of harmonized item response datasets designed to support secondary analysis and methodological research in psychological and educational measurement. This paper serves as a practical guide for researchers interested in using the IRW. We describe the structure of IRW datasets and the quantitative and qualitative metadata available for dataset selection, and we demonstrate how researchers can navigate the IRW website to explore and compare available tables. We further show how the IRW R and Python packages can be used to filter datasets programmatically, download response-level data, and generate standardized citations for reproducible research …


A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure, Ernest Fokoue Mar 2026

A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure, Ernest Fokoue

Articles

Ensemble learning is traditionally justified as a variance-reduction strategy, explaining its strong performance for unstable predictors such as decision trees. This explanation, however, does not account for ensembles constructed from intrinsically stable estimators-including smoothing splines, kernel ridge regression, Gaussian process regression, and other regularized reproducing kernel Hilbert space (RKHS) methods whose variance is already tightly controlled by regularization and spectral shrinkage. This paper develops a general weighting theory for ensemble learning that moves beyond classical variance-reduction arguments. We formalize ensembles as linear operators acting on a hypothesis space and endow the space of weighting sequences with geometric and spectral constraints. …


Learning Ordinal Geometry: Semantic–Aware Kernels For Ordered Categorical Data, Ernest Fokoue Mar 2026

Learning Ordinal Geometry: Semantic–Aware Kernels For Ordered Categorical Data, Ernest Fokoue

Articles

Ordinal data arise ubiquitously in survey research, psychology, medicine, economics, and recommender systems, yet kernel methods for such data typically rely on either nominal encodings or arbitrary numeric codings. The former discards order information; the lat- ter imposes a fictitious metric structure. This paper develops a principled framework for kernel design on ordinal scales and introduces a new class of Semantic–Aware Ordinal Ker- nels (SAOK) that simultaneously capture ordinal order and semantic proximity between categories. We begin by formalizing order–preserving embeddings of finite chains and characterizing a broad family of chain distances that are conditionally negative definite. Through Schoen- berg …


Statistical Analysis Of Log Transformation Effectiveness In Air Traffic Movement Forecasting During Covid-19 In South Africa, John Lehlaka Masekoameng Jan 2026

Statistical Analysis Of Log Transformation Effectiveness In Air Traffic Movement Forecasting During Covid-19 In South Africa, John Lehlaka Masekoameng

Journal of Aviation Technology and Engineering

This study evaluates the effectiveness of log transformation in enhancing multiple regression models used to forecast air traffic movements (ATMs) in South Africa during the COVID-19 pandemic. Using 60 monthly observations from October 2016 to September 2021, the analysis incorporates variables such as revenue, lockdown levels, COVID-19 metrics, exchange rates, gross domestic product, and population. Two models are compared: one using raw ATMs and another with log-transformed ATMs as the dependent variable.

While the untransformed model shows stronger explanatory power (R² = 0.904, adjusted R² = 0.891) compared to the log-transformed model (R² = 0.772, adjusted R² = 0.741), the …


Modeling Housing Prices: Which Features Matter Most?, Alex Ruvolo Jan 2026

Modeling Housing Prices: Which Features Matter Most?, Alex Ruvolo

Williams Honors College, Honors Research Projects

This paper attempts to find the biggest factors and traits that influence the cost of housing. This will include the lot size, type of street, utilities, neighborhood, year built, heating, electrical, yard size, number of different rooms, age, condition, and others. I will attempt to answer the question of whether the prices of houses have changed within the last 5 to 10 years, and obviously this is an easy question to answer. However, the bigger question beyond this is are the main factors affecting housing prices all important in explaining this relationship? Is one factor more important than the rest …


Hybrid Patchtst And Physics-Based Framework For Predicting Lithium-Ion Battery State Of Health, Pavan Ravuri Jan 2026

Hybrid Patchtst And Physics-Based Framework For Predicting Lithium-Ion Battery State Of Health, Pavan Ravuri

Honors Undergraduate Theses

Accurately forecasting the state of health of lithium-ion batteries is critical for improving performance, reliability and lifetime in energy storage applications. Battery capacity degrades into a nonlinear pattern over cycling due to electrochemical processes where neither purely data driven nor physics-based models can capture alone. This study looks at a hybrid framework combining that PatchTST patch-based transformer architecture with physics-based features derived from the solid electrolyte interphase and pseudo two-dimensional models. Physics inspired proxy features were computed from cycling data and concatenated with electrochemical measurements as added input channels before patch segmentation. There are five model configurations that were evaluated …


An Integrated Data-Driven Framework For Arctic Shipping: Analyzing Vessel Speed, Environmental And Ecological Factors Through Innovative Statistical Spatio-Temporal Methods, Inverse Optimization And Machine Learning, Mauli Pant Jan 2026

An Integrated Data-Driven Framework For Arctic Shipping: Analyzing Vessel Speed, Environmental And Ecological Factors Through Innovative Statistical Spatio-Temporal Methods, Inverse Optimization And Machine Learning, Mauli Pant

Theses and Dissertations

This dissertation develops an integrated data-driven framework to analyze vessel navigation and ecological risk in the United States Arctic from 2010 to 2019. As environmental change and maritime activity increase in the region, understanding how vessels respond to dynamic conditions and how those responses interact with marine ecosystems has become increasingly important. A central theme of this dissertation is the treatment of vessel speed as both an observed outcome and a decision variable reflecting trade- offs among operational, environmental, and ecological factors. The first chapter develops a predictive framework for vessel speed over ground (SOG) using Gaussian Process Boosting (GPBoost), …


Experimental Design And Analysis For Decision Making: Methodology And Applications, Yezhuo Li Aug 2025

Experimental Design And Analysis For Decision Making: Methodology And Applications, Yezhuo Li

All Dissertations

This dissertation develops and applies advanced statistical and optimization frameworks to enhance decision-making under uncertainty, particularly in engineering and manufacturing contexts. First, we introduce an approach for the optimal design of controlled experiments that accounts for observational covariates, enabling more precise and personalized decisions. Second, we explore the application of constrained Bayesian optimization, using Gaussian process surrogate models, to optimize composite cure processes, significantly reducing computational effort while maintaining high predictive accuracy. Building on this foundation, we extend Bayesian optimization to bivariate Gaussian process models that capture correlations between objective and constraint functions, offering new insights into multidimensional decision landscapes. …


Estimation Methods For Bayesian Exponential Random Graph Models Under The Horseshoe Prior., Pamela Linares Aug 2025

Estimation Methods For Bayesian Exponential Random Graph Models Under The Horseshoe Prior., Pamela Linares

Electronic Theses and Dissertations

Networks are powerful tools for modeling the complexity of social interactions, biological systems, and information spread. A leading statistical frameworks for analyzing network data are Exponential Random Graph Models (ERGMs), which provide a principled approach to capturing structural dependencies. However, ERGMs remain challenging to estimate, especially in sparse or high-dimensional settings where models suffer from degeneracy and unstable parameter inference. This paper proposes a penalized Bayesian approach to ERGMs that utilizes the horseshoe prior, a sparsity-inducing global-local shrinkage prior. This prior offers robust regularization while preserving important signals, improving estimation by shrinking irrelevant parameters and reducing the impact of extreme …


Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim Jun 2025

Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim

Master's Theses

The outcome of a search and rescue (SAR) operation is influenced by a complex, non-linear interplay among numerous factors, including geographic context, subject-specific characteristics, and environmental conditions. The high dimensionality and intricate dependencies among these variables pose significant challenges to traditional exploratory modeling approaches, limiting their ability to uncover meaningful patterns and relationships associated with mission success. This study introduces Rules Based Explanations for Generated neighborhoods Around Localized cases (REGAL), a novel adaptation of the Local Interpretable Model-agnostic Explanations (LIME) framework to explain deep multimodal neural networks and what key features it assesses to determine search and rescue success. REGAL …


Data Driven Analysis Of Samara Seed Kinematics And Dynamics, Shashwat Sparsh Jun 2025

Data Driven Analysis Of Samara Seed Kinematics And Dynamics, Shashwat Sparsh

Master's Theses

Samara Seeds are a class of fruit most famously belonging to the Acer species and are characterized by their single-bladed geometry and their auto-rotation response during descent. This steady-state auto-rotation response is the subject of aerodynamic analysis which aim to quantify the performance. The period prior to the beginning of steady-state auto-rotation is classified as the transition regime and has not been the subject of intense scrutiny.

This thesis employs a data-driven approach to analyzing the kinematic and dynamic response of these seeds during both the transition and auto-rotation stages of flight to quantify the performance with respect to the …


Statistical Study Of Solar Wind Conditions Prior To Substorm Onsets, Luke H. Francis Apr 2025

Statistical Study Of Solar Wind Conditions Prior To Substorm Onsets, Luke H. Francis

Doctoral Dissertations and Master's Theses

Due to complex, multi-region, coupled plasma systems, auroral substorm onsets have been historically difficult to predict. The northward turning of the interplanetary magnetic field was considered the primary candidate as an external triggering mechanism for substorm onsets. However, that was later shown to be coincidental in nature. This study is motivated by recent multi-spacecraft observations that show how several magnetosheath jets at the bow shock were heavily correlated to substorm onsets, indicated by a strongly radial IMF interval. In the past, studies have looked at small samples of substorms in order to make large-scale predictions. However in this study, a …


Predicting Superconducting Critical Temperature From Composition-Derived Features: A Transparent Linear And Regularized Regression Study, Md Ahiduzzaman Jan 2025

Predicting Superconducting Critical Temperature From Composition-Derived Features: A Transparent Linear And Regularized Regression Study, Md Ahiduzzaman

Data Science and Data Mining

We study prediction of superconducting critical temperature (Tc) from 81 composition-derived descriptors across 21,263 materials. To keep the analysis transparent and repro- ducible, we focus on linear models: Ordinary Least Squares (OLS), Ridge, Lasso, and Elastic Net (ENet). All models share a single evaluation protocol (5-fold cross-validation with standardized inputs) and are compared on RMSE, MAE, and R2. On this feature set, OLS attains the best cross-validated performance (RMSE = 17.6 K, MAE = 13.3 K , R2 = 0.735), with Lasso/ENet essentially tied next (RMSE ≈ 17.7 K , R2 ≈ 0.734); Ridge underperforms (RMSE = 18.9 K , …


Comparative Analysis Of Lasso, Ridge, And Elastic Net For Variable Selection In High-Dimensional Maize Data, Md Ahiduzzaman Jan 2025

Comparative Analysis Of Lasso, Ridge, And Elastic Net For Variable Selection In High-Dimensional Maize Data, Md Ahiduzzaman

Data Science and Data Mining

In high-dimensional genomic data analysis, traditional linear regression techniques often struggle due to the presence of a large number of predictor variables relative to observations. Penalized regression methods such as LASSO, Ridge, and Elastic Net have emerged as effective solutions by imposing regularization, which helps in managing multicollinearity and enhancing prediction accuracy. This study applies these techniques to the Maize dataset to model the time to male flowering, selecting relevant genetic markers as predictors. Our findings suggest that Elastic Net is particularly effective for high-dimensional data with correlated variables, achieving a balance between prediction accuracy and variable selection. The results …


Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh Jan 2025

Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh

Theses and Dissertations

Traditional models in psychiatric research often impose assumptions of causal homogeneity, treating population-level associations as reflective of uniform underlying mechanisms. This dissertation challenges that assumption by introducing statistical and machine learning frameworks designed to detect and model causal heterogeneity in the development of psychopathology. Central to this approach is the advancement of finite mixture structural equation modeling (FM-SEM) to identify latent subgroups characterized by distinct, and sometimes opposing, causal pathways.

The dissertation comprises three integrated empirical studies. The first introduces mixDoC, a finite mixture extension of the classical Direction of Causation (DoC) model applied to twin data, enabling the detection …


Modeling Non-Normal Distributions With Mixed Third-Order Polynomials Of Standard Normal And Logistic Variables, Mohan D. Pant, Aditya Chakraborty, Ismail El Moudden Jan 2025

Modeling Non-Normal Distributions With Mixed Third-Order Polynomials Of Standard Normal And Logistic Variables, Mohan D. Pant, Aditya Chakraborty, Ismail El Moudden

Epidemiology, Biostatistics, & Environmental Health Faculty Publications

Continuous data associated with many real-world events often exhibit non-normal characteristics, which contribute to the difficulty of accurately modeling such data with statistical procedures that rely on normality assumptions. Traditional statistical procedures often fail to accurately model non-normal distributions that are often observed in real-world data. This paper introduces a novel modeling approach using mixed third-order polynomials, which significantly enhances accuracy and flexibility in statistical modeling. The main objective of this study is divided into three parts: The first part is to introduce two new non-normal probability distributions by mixing standard normal and logistic variables using a piecewise function of …


Safeguard Cyberspace In Ransomware Era: Risk Analysis & Cyber Insurance, Li Huang Jan 2025

Safeguard Cyberspace In Ransomware Era: Risk Analysis & Cyber Insurance, Li Huang

Electronic Theses & Dissertations (2024 - present)

The increasing frequency and severity of ransomware attacks pose significant challenges for organizational cybersecurity. Fragmentation across disciplines in cyber defense has created practical gaps in the development of the necessary capabilities needed to address rapidly evolving cyber threats. This study explores the impact of ransomware attacks and the evolving role of cyber insurance as a proactive cybersecurity partner. Bridging the gap between actuarial science and cyber risk management, it proposes an interdisciplinary framework that quantifies the impact of ransomware and integrates cyber insurance into cybersecurity strategies.

The primary contribution of this study is methodology. We present a framework that remains …


Optimal Data Splitting Methods, Sujay Mudalgi Jan 2025

Optimal Data Splitting Methods, Sujay Mudalgi

Theses and Dissertations

In predictive modeling, effective data splitting is crucial for creating statistically representative training and validation sets. The state-of-the-art data splitting methods are based on minimizing the energy distance between the split subsets. However, there are a number of limitations in the existing methods, which this dissertation aims to address. First, the existing methods were computationally inefficient. Thus, Chapter 2 proposes a method to scale up these approaches for big data. Here, we introduce scalable Twinning (s-Twinning), which significantly improves the execution speed of data splitting without sacrificing accuracy. Second, the existing methods did not consider the predictive relationship in the …


Crime Modeling Using An Integrated Cnn–Lstm Architecture With Embedded Self-Excitation, Pawandeep Kaur Jan 2025

Crime Modeling Using An Integrated Cnn–Lstm Architecture With Embedded Self-Excitation, Pawandeep Kaur

Theses and Dissertations (Comprehensive)

It is often assumed that natural phenomena occur randomly over time. However, careful analysis reveals that these events typically form some series or sequences and exhibit distinctive temporal patterns. These patterns are not exclusive to nature. They also appear in human activities, often studied under the concept of bursty human dynamics. The statistical methods analyzing bursty human dynamics not only capture overall trends or seasonality but also explore how past events influence future ones. It makes the analysis more realistic and the results more closely aligned with reality. Bursty human dynamics can be studied at two levels: the individual level …


Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem Jan 2025

Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem

Dissertations, Master's Theses and Master's Reports

Factor analysis is a powerful tool for modeling latent structures in high-dimensional data, traditional approaches assume a single global structure, limiting their ability to capture heterogeneity. The Mixture of Factor Analyzers (MFA) extends classical factor analysis by modeling data as a mixture of Gaussian-distributed local subspaces, effectively uncovering cluster-specific latent structures. However, MFA relies on Gaussian mixtures, making it sensitive to outliers and ill-suited for heavy-tailed data. The Mixture of $t$-Factor Analyzers (M$t$FA) addresses these limitations by incorporating multivariate $t$-distributions, improving robustness. Despite their advantages, both MFA and M$t$FA face significant computational challenges in high-dimensional settings, particularly due to costly …


Forecasting Commercial Vehicle Miles Traveled (Vmt) In Urban California Areas, Steve Chung, Jaymin Kwon, Yushin Ahn Aug 2024

Forecasting Commercial Vehicle Miles Traveled (Vmt) In Urban California Areas, Steve Chung, Jaymin Kwon, Yushin Ahn

Mineta Transportation Institute

This study investigates commercial truck vehicle miles traveled (VMT) across six diverse California counties from 2000 to 2020. The counties—Imperial, Los Angeles, Riverside, San Bernardino, San Diego, and San Francisco—represent a broad spectrum of California’s demographics, economies, and landscapes. Using a rich dataset spanning demographics, economics, and pollution variables, we aim to understand the factors influencing commercial VMT. We first visually represent the geographic distribution of the counties, highlighting their unique characteristics. Linear regression models, particularly the least absolute shrinkage and selection operator (LASSO) and elastic net regressions are employed to identify key predictors of total commercial VMT. LASSO regression …


Extreme Value Statistics Analysis Of Process Defects In Additive Manufacturing Materials, Ayorinde E. Olatunde, Kristen Hernandez, Austin Ngo, Arafath Nihar, Thomas G. Ciardi, Rachel Yamamoto, Pawan K. Tripathi, Roger H. French, John J. Lewandowski, Anirban Mondal Jun 2024

Extreme Value Statistics Analysis Of Process Defects In Additive Manufacturing Materials, Ayorinde E. Olatunde, Kristen Hernandez, Austin Ngo, Arafath Nihar, Thomas G. Ciardi, Rachel Yamamoto, Pawan K. Tripathi, Roger H. French, John J. Lewandowski, Anirban Mondal

Faculty Scholarship

Fatigue and fracture studies focused on process defects that occur in Additive Manufacturing (AM) materials have shown that defect populations possess features which are better measured with extreme value statistics (EVS). In AM alloys, defect occurrences increase with material volume. This situation facilitates the need to model process defects in the path of fatigue crack growth with suitable statistical tools, such as EVS, which is more cost-effective when compared to destructive experiments. The application of EVS on defect space features helps determine the difference in defects present on fracture surfaces. As the fatigue quality of any material depends on its …


Stock Market Volatility In The United Kingdom: Simulating Post-Covid-19 Recovery, Bala A. Dahiru, Mohammed Shuaibu, Najibullah Hassanov Jun 2024

Stock Market Volatility In The United Kingdom: Simulating Post-Covid-19 Recovery, Bala A. Dahiru, Mohammed Shuaibu, Najibullah Hassanov

CBN Journal of Applied Statistics (JAS)

This paper investigates the time it would take for the FTSE-100 index to reach its post-COVID-19 peak. The paper utilises an exponential generalised autoregressive conditional heteroscedasticity (EGARCH) model that accounts for leverage effect and asymmetries. The preferred models amongst competing variants was the Autoregressive Moving Average (ARMA)-EGARCH(2,1) specification and was used to predict daily FTSE-100 data from 5th January 2000 to 21st June 2024. The empirical exercise showed that the COVID-19-induced financial crisis negatively affected the United Kingdom’s stock market performance. The results show that the FTSE100 index could reach its post-pandemic peak around 27th August, 2024 (two months after …


A Spatial Decision Support System For Rent Estimation Of Retail Spaces In Manhattan Using Geographically Weighted Regression And Spatial Regression, Andie M. Migden Miller May 2024

A Spatial Decision Support System For Rent Estimation Of Retail Spaces In Manhattan Using Geographically Weighted Regression And Spatial Regression, Andie M. Migden Miller

Theses and Dissertations

This report outlines an automated, three-phase Spatial Decision Support System that creates models to estimate rent of retail spaces across Manhattan. First, enrich data with predictors. Second, optimize spatially aware neighborhood-level models by combining GWR, spatial regression, and non-spatial regression. Finally, visualize results in an Esri-based WebApp.


A Novel Correction For The Multivariate Ljung-Box Test, Minhao Huang May 2024

A Novel Correction For The Multivariate Ljung-Box Test, Minhao Huang

Computational and Data Sciences (PhD) Dissertations

This research introduces an analytical improvement to the Multivariate Ljung-Box test that addresses significant deviations of the original test from the nominal Type I error rates under almost all scenarios. Prior attempts to mitigate this issue have been directed at modification of the test statistics or correction of the test distribution to achieve precise results in finite samples. In previous studies, focused on designing corrections to the univariate Ljung-Box, a method that specifically adjusts the test rejection region has been the most successful of attaining the best Type I error rates. We adopt the same approach for the more complex, …


Code For Care: Hypertension Prediction In Women Aged 18-39 Years, Kruti Sheth May 2024

Code For Care: Hypertension Prediction In Women Aged 18-39 Years, Kruti Sheth

Electronic Theses, Projects, and Dissertations

The longstanding prevalence of hypertension, often undiagnosed, poses significant risks of severe chronic and cardiovascular complications if left untreated. This study investigated the causes and underlying risks of hypertension in females aged between 18-39 years. The research questions were: (Q1.) What factors affect the occurrence of hypertension in females aged 18-39 years? (Q2.) What machine learning algorithms are suited for effectively predicting hypertension? (Q3.) How can SHAP values be leveraged to analyze the factors from model outputs? The findings are: (Q1.) Performing Feature selection using binary classification Logistic regression algorithm reveals an array of 30 most influential factors at an …