Open Access. Powered by Scholars. Published by Universities.®

Analysis Commons

Open Access. Powered by Scholars. Published by Universities.®

Applied Statistics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1 - 30 of 40

Full-Text Articles in Analysis

A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari May 2026

A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari

Theses and Dissertations

Student retention and degree completion remain central challenges for higher-education institutions, with significant implications for student success, institutional effectiveness, and public accountability. While advances in predictive analytics have enabled earlier identification of students at risk of withdrawal, many commonly used machine learning approaches suffer from limited interpretability, constraining their practical usefulness for advising, intervention, and policy decision making. This dissertation addresses the problem of predicting student persistence by developing and evaluating optimization based, interpretable classification models within the Logical Analysis of Data (LAD) framework. Building on existing LAD formulations, this research introduces two novel pattern generation models, the Best Term …


Multivariate Quantile Autoregression-Mixed Data Sampling (Mvqar-Midas) Modeling Of Cost Of Living And Supply Chain Dynamics In Canada., Patrick Gbolonyo Jan 2026

Multivariate Quantile Autoregression-Mixed Data Sampling (Mvqar-Midas) Modeling Of Cost Of Living And Supply Chain Dynamics In Canada., Patrick Gbolonyo

Theses and Dissertations (Comprehensive)

In recent years, the rising cost of living as a result of persistent inflationary pressures, disruptions in the global supply chains, and changes in the macroeconomic landscape has become a critical topic of discussion. To address this, we move beyond a mean-based framework and employ a quantile regression approach. This allows the persistence of each series and the transmis- sion of shocks between the Consumer Price Index (CPI) (the total CPI which is a percentage change over the past 12 months), the Interest Rate (IR)(the target for the overnight rate), the New Housing Price Index (NHPI), and high-frequency supply chain …


A Data-Driven Approach To Time Series Forecasting And Clustering Of U.S. Regional Drug Overdose Mortality, Koshali Hamy Muthunama Gonnage Jul 2025

A Data-Driven Approach To Time Series Forecasting And Clustering Of U.S. Regional Drug Overdose Mortality, Koshali Hamy Muthunama Gonnage

Mathematics & Statistics ETDs

The increasing rate of drug overdose deaths in the United States poses a critical public health challenge, particularly due to the surge in synthetic opioids and other high-risk substances. This study presents a data-driven framework that integrates time series forecasting and clustering techniques. Monthly mortality data for five key drug types: cocaine, fentanyl, heroin, methamphetamine, and oxycodone were analyzed using four time series forecasting models: ARIMA, ETS, TBATS, and NNAR. These models were evaluated using standard accuracy metrics RMSE, MAPE, and MAE to assess predictive performance. Signal decomposition approach based on Singular Value Decomposition and subspace modeling was employed to …


An In-Depth Analysis Of The Bellarmine University Student Athlete Experience, Mattingly E. Spalding Mar 2025

An In-Depth Analysis Of The Bellarmine University Student Athlete Experience, Mattingly E. Spalding

Undergraduate Theses

The goal of this thesis is to improve the student athlete experience at Bellarmine University through direct feedback from our current student athletes. By defining what parts of the student athlete experience they value most, Bellarmine can reflect on the support currently provided in these areas. Additionally, if it is concluded through the response that high valued areas are not being satisfied, Bellarmine can work to improve the support in these areas. Or, if there is abundant support in a low-valued area, Bellarmine could think to shift resources into a more needed concentration. Through the design sample, data was accumulated …


Analysis Of Systematic Trade-Offs Between Military And Healthcare Expenditure Alongside Gdp Growth Of Select Asian And Western Exporting Economies In The 21st Century, Rahul Balamurugan, Carlos Gershenson, Preethi Nanjundan, Hiroki Sayama Mar 2025

Analysis Of Systematic Trade-Offs Between Military And Healthcare Expenditure Alongside Gdp Growth Of Select Asian And Western Exporting Economies In The 21st Century, Rahul Balamurugan, Carlos Gershenson, Preethi Nanjundan, Hiroki Sayama

Northeast Journal of Complex Systems (NEJCS)

This study explores the complexity in the trade-offs between military expenditure, healthcare expenditure, and GDP growth across select Asian nations and major weapon-exporting countries, examining how nations allocate finite resources between national security and human well-being over the past two decades. Using a systems science approach, the research integrates Granger causality testing to analyze temporal and directional relationships among GDP growth, military expenditure, and healthcare expenditure, uncovering their dynamic interdependencies. The methodology includes trend and slope analysis, Granger causality testing, outlier detection, and clustering to identify heterogeneity in resource allocation strategies. Developed, weapon-exporting nations exhibit complementary trends, with strong causality …


Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari Jan 2025

Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari

Posters-at-the-Capitol

The grocery retail industry landscape has changed greatly in the wake of the pandemic. Specifically, delivery and pickup services have become more popular and customer buying habits have evolved. At the same time, improvements in data collection and analysis have allowed grocery marketing strategies to become highly individualized.

We worked with 84.51, an analytics firm, to identify customer segments for the Kroger Company based on data from 2023. Using clustering techniques, we organized customers into groups, or segments, based on similar characteristics. We identified and profiled four distinct groups of customers. Three segments were characterized by high frequency and spending …


Theoretical Foundations And Applied Performance Of Periodicity-Aware Imputation: Variable Bandpass Block Bootstrap Methods For Incomplete Time Series, Asmaa Ahmad Jan 2025

Theoretical Foundations And Applied Performance Of Periodicity-Aware Imputation: Variable Bandpass Block Bootstrap Methods For Incomplete Time Series, Asmaa Ahmad

Electronic Theses & Dissertations (2024 - present)

Time series data are prevalent across a wide range of disciplines, including health surveillance, public policy, and environmental monitoring. In the presence of underlying cyclical patterns, the integrity of time series analysis depends critically on the ability to detect, model, and impute structured missing data without compromising the temporal structure. This dissertation introduces and validates a novel imputation framework that integrates the Variable Bandpass Periodic Block Bootstrap (VBPBB) into multiple imputation procedures, improving the accuracy, robustness, and interpretability of time series models under high rates of missingness and noise. The overarching goal of this dissertation was to develop and evaluate …


Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe Jan 2024

Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe

Data Science and Data Mining

Cyberbullying refers to the act of bullying using electronic means and the internet. In recent years, this act has been identifed to be a major problem among young people and even adults. It can negatively impact one’s emotions and lead to adverse outcomes like depression, anxiety, harassment, and suicide, among others. This has led to the need to employ machine learning techniques to automatically detect cyberbullying and prevent them on various social media platforms. In this study, we want to analyze the combination of some Natural Language Processing (NLP) algorithms (such as Bag-of-Words and TFIDF) with some popular machine learning …


Multiscale Modelling Of Brain Networks And The Analysis Of Dynamic Processes In Neurodegenerative Disorders, Hina Shaheen Jan 2024

Multiscale Modelling Of Brain Networks And The Analysis Of Dynamic Processes In Neurodegenerative Disorders, Hina Shaheen

Theses and Dissertations (Comprehensive)

The complex nature of the human brain, with its intricate organic structure and multiscale spatio-temporal characteristics ranging from synapses to the entire brain, presents a major obstacle in brain modelling. Capturing this complexity poses a significant challenge for researchers. The complex interplay of coupled multiphysics and biochemical activities within this intricate system shapes the brain's capacity, functioning within a structure-function relationship that necessitates a specific mathematical framework. Advanced mathematical modelling approaches that incorporate the coupling of brain networks and the analysis of dynamic processes are essential for advancing therapeutic strategies aimed at treating neurodegenerative diseases (NDDs), which afflict millions of …


Movie Recommender System Using Matrix Factorization, Roland Fiagbe May 2023

Movie Recommender System Using Matrix Factorization, Roland Fiagbe

Data Science and Data Mining

Recommendation systems are a popular and beneficial field that can help people make informed decisions automatically. This technique assists users in selecting relevant information from an overwhelming amount of available data. When it comes to movie recommendations, two common methods are collaborative filtering, which compares similarities between users, and content-based filtering, which takes a user’s specific preferences into account. However, our study focuses on the collaborative filtering approach, specifically matrix factorization. Various similarity metrics are used to identify user similarities for recommendation purposes. Our project aims to predict movie ratings for unwatched movies using the MovieLens rating dataset. We developed …


Formula 101 Using 2022 Formula One Season Data To Understand The Race Results, Christopher Garcia, Oliver Lopez May 2023

Formula 101 Using 2022 Formula One Season Data To Understand The Race Results, Christopher Garcia, Oliver Lopez

Student Scholar Symposium Abstracts and Posters

The reason why I am interested in Formula One is that my friend showed me what Formula One was all about. It became interesting to see the action of the sport, including the battles the drivers have during the race and how fast they go through a corner. Also, when qualifying comes around, they push their car to the absolute limit to gain a few seconds off their opponents. The drivers only in the top 10 receive points from the winner getting 25 points, the last driver in the top 10 getting 1 point, and those below the top ten …


Employee Attrition: Analyzing Factors Influencing Job Satisfaction Of Ibm Data Scientists, Graham Nash Apr 2023

Employee Attrition: Analyzing Factors Influencing Job Satisfaction Of Ibm Data Scientists, Graham Nash

Symposium of Student Scholars

Employee attrition is a relevant issue that every business employer must consider when gauging the effectiveness of their employees. Whether or not an employee chooses to leave their job can come from a multitude of factors. As a result, employers need to develop methods in which they can measure attrition by calculating the several qualities of their employees. Factors like their age, years with the company, which department they work in, their level of education, their job role, and even their marital status are all considered by employers to assist in predicting employee attrition. This project will be analyzing a …


Session 5: Equipment Finance Credit Risk Modeling - A Case Study In Creative Model Development & Nimble Data Engineering, Edward Krueger, Landon Thompson, Josh Moore Feb 2022

Session 5: Equipment Finance Credit Risk Modeling - A Case Study In Creative Model Development & Nimble Data Engineering, Edward Krueger, Landon Thompson, Josh Moore

SDSU Data Science Symposium

This presentation will focus first on providing an overview of Channel and the Risk Analytics team that performed this case study. Given that context, we’ll then dive into our approach for building the modeling development data set, techniques and tools used to develop and implement the model into a production environment, and some of the challenges faced upon launch. Then, the presentation will pivot to the data engineering pipeline. During this portion, we will explore the application process and what happens to the data we collect. This will include how we extract & store the data along with how it …


Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu Jan 2022

Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu

Honors Theses and Capstones

COVID-19 caused state and nation-wide lockdowns, which altered human foot traffic, especially in restaurants. The seafood sector in particular suffered greatly as there was an increase in illegal fishing, it is made up of perishable goods, it is seasonal in some places, and imports and exports were slowed. Foot traffic data is useful for business owners to have to know how much to order, how many employees to schedule, etc. One issue is that the data is very expensive, hard to get, and not available until months after it is recorded. Our goal is to not only find covariates that …


Estimating The Statistics Of Operational Loss Through The Analyzation Of A Time Series, Maurice L. Brown Jan 2022

Estimating The Statistics Of Operational Loss Through The Analyzation Of A Time Series, Maurice L. Brown

Theses and Dissertations

In the world of finance, appropriately understanding risk is key to success or failure because it is a fundamental driver for institutional behavior. Here we focus on risk as it relates to the operations of financial institutions, namely operational risk. Quantifying operational risk begins with data in the form of a time series of realized losses, which can occur for a number of reasons, can vary over different time intervals, and can pose a challenge that is exacerbated by having to account for both frequency and severity of losses. We introduce a stochastic point process model for the frequency distribution …


Role Of Inhibition And Spiking Variability In Ortho- And Retronasal Olfactory Processing, Michelle F. Craft Jan 2022

Role Of Inhibition And Spiking Variability In Ortho- And Retronasal Olfactory Processing, Michelle F. Craft

Theses and Dissertations

Odor perception is the impetus for important animal behaviors, most pertinently for feeding, but also for mating and communication. There are two predominate modes of odor processing: odors pass through the front of nose (ortho) while inhaling and sniffing, or through the rear (retro) during exhalation and while eating and drinking. Despite the importance of olfaction for an animal’s well-being and specifically that ortho and retro naturally occur, it is unknown whether the modality (ortho versus retro) is transmitted to cortical brain regions, which could significantly instruct how odors are processed. Prior imaging studies show different …


An Alternative Class Of Ratio-Regression-Type Estimator Under Two-Phase Sampling Scheme, Muhammad Isah, Zakari Yahaya, Audu Ahmed Dec 2021

An Alternative Class Of Ratio-Regression-Type Estimator Under Two-Phase Sampling Scheme, Muhammad Isah, Zakari Yahaya, Audu Ahmed

CBN Journal of Applied Statistics (JAS)

In this study, a new exponential ratio-regression estimator is developed using an auxiliary variable for estimating the finite population mean under a two-phase sampling system. The Bias and Mean Square Error (MSE) of the proposed estimator are derived and compared with some of the estimators in extant literature. Thus, the conditions under which the proposed estimator is better than some existing estimators are provided. Empirically, using four real datasets and simulation study, the proposed estimator performs better than the classical ratio, classical regression, exponential ratio, and exponential regression cum ratio estimator when compared using the criteria of bias, mean square …


Applying Emotional Analysis For Automated Content Moderation, John Shelnutt May 2021

Applying Emotional Analysis For Automated Content Moderation, John Shelnutt

Computer Science and Computer Engineering Undergraduate Honors Theses

The purpose of this project is to explore the effectiveness of emotional analysis as a means to automatically moderate content or flag content for manual moderation in order to reduce the workload of human moderators in moderating toxic content online. In this context, toxic content is defined as content that features excessive negativity, rudeness, or malice. This often features offensive language or slurs. The work involved in this project included creating a simple website that imitates a social media or forum with a feed of user submitted text posts, implementing an emotional analysis algorithm from a word emotions dataset, designing …


On The Evolution Equation For Modelling The Covid-19 Pandemic, Jonathan Blackledge Jan 2021

On The Evolution Equation For Modelling The Covid-19 Pandemic, Jonathan Blackledge

Books/Book chapters

The paper introduces and discusses the evolution equation, and, based exclusively on this equation, considers random walk models for the time series available on the daily confirmed Covid-19 cases for different countries. It is shown that a conventional random walk model is not consistent with the current global pandemic time series data, which exhibits non-ergodic properties. A self-affine random walk field model is investigated, derived from the evolutionary equation for a specified memory function which provides the non-ergodic fields evident in the available Covid-19 data. This is based on using a spectral scaling relationship of the type 1/ωα where ω …


Neither “Post-War” Nor Post-Pregnancy Paranoia: How America’S War On Drugs Continues To Perpetuate Disparate Incarceration Outcomes For Pregnant, Substance-Involved Offenders, Becca S. Zimmerman Jan 2021

Neither “Post-War” Nor Post-Pregnancy Paranoia: How America’S War On Drugs Continues To Perpetuate Disparate Incarceration Outcomes For Pregnant, Substance-Involved Offenders, Becca S. Zimmerman

Pitzer Senior Theses

This thesis investigates the unique interactions between pregnancy, substance involvement, and race as they relate to the War on Drugs and the hyper-incarceration of women. Using ordinary least square regression analyses and data from the Bureau of Justice Statistics’ 2016 Survey of Prison Inmates, I examine if (and how) pregnancy status, drug use, race, and their interactions influence two length of incarceration outcomes: sentence length and amount of time spent in jail between arrest and imprisonment. The results collectively indicate that pregnancy decreases length of incarceration outcomes for those offenders who are not substance-involved but not evenhandedly -- benefitting white …


Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya May 2020

Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya

Electronic Theses and Dissertations

Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …


Evaluating An Ordinal Output Using Data Modeling, Algorithmic Modeling, And Numerical Analysis, Martin Keagan Wynne Brown Jan 2020

Evaluating An Ordinal Output Using Data Modeling, Algorithmic Modeling, And Numerical Analysis, Martin Keagan Wynne Brown

Murray State Theses and Dissertations

Data and algorithmic modeling are two different approaches used in predictive analytics. The models discussed from these two approaches include the proportional odds logit model (POLR), the vector generalized linear model (VGLM), the classification and regression tree model (CART), and the random forests model (RF). Patterns in the data were analyzed using trigonometric polynomial approximations and Fast Fourier Transforms. Predictive modeling is used frequently in statistics and data science to find the relationship between the explanatory (input) variables and a response (output) variable. Both approaches prove advantageous in different cases depending on the data set. In our case, the data …


Tropical Cyclone Hazards In Relation To Propagation Speed, Jiehao Huang Jan 2020

Tropical Cyclone Hazards In Relation To Propagation Speed, Jiehao Huang

Dissertations and Theses

As the population and infrastructure along the US East Coast increase, it becomes increasingly important to study the characteristics of tropical cyclones that can impact the coast. A recent study shows that the propagation speed of tropical cyclones has slowed over the past 60 years, which can lead to greater accumulation of precipitation and greater storm surge impacts. The study presented herein is meant to examine and analyze the relationships that exist between the propagation speed of tropical cyclones, their surface wind strength, displacement angles, and cyclone averaged winds. This analysis is focused on tropical cyclones spanning from 1950-2015 in …


Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater Jan 2019

Improving Vix Futures Forecasts Using Machine Learning Methods, James Hosker, Slobodan Djurdjevic, Hieu Nguyen, Robert Slater

SMU Data Science Review

The problem of forecasting market volatility is a difficult task for most fund managers. Volatility forecasts are used for risk management, alpha (risk) trading, and the reduction of trading friction. Improving the forecasts of future market volatility assists fund managers in adding or reducing risk in their portfolios as well as in increasing hedges to protect their portfolios in anticipation of a market sell-off event. Our analysis compares three existing financial models that forecast future market volatility using the Chicago Board Options Exchange Volatility Index (VIX) to six machine/deep learning supervised regression methods. This analysis determines which models provide best …


Yelp’S Review Filtering Algorithm, Yao Yao, Ivelin Angelov, Jack Rasmus-Vorrath, Mooyoung Lee, Daniel W. Engels Aug 2018

Yelp’S Review Filtering Algorithm, Yao Yao, Ivelin Angelov, Jack Rasmus-Vorrath, Mooyoung Lee, Daniel W. Engels

SMU Data Science Review

In this paper, we present an analysis of features influencing Yelp's proprietary review filtering algorithm. Classifying or misclassifying reviews as recommended or non-recommended affects average ratings, consumer decisions, and ultimately, business revenue. Our analysis involves systematically sampling and scraping Yelp restaurant reviews. Features are extracted from review metadata and engineered from metrics and scores generated using text classifiers and sentiment analysis. The coefficients of a multivariate logistic regression model were interpreted as quantifications of the relative importance of features in classifying reviews as recommended or non-recommended. The model classified review recommendations with an accuracy of 78%. We found that reviews …


Masked Instability: Within-Sector Financial Risk In The Presence Of Wealth Inequality, Youngna Choi Jun 2018

Masked Instability: Within-Sector Financial Risk In The Presence Of Wealth Inequality, Youngna Choi

Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works

We investigate masked financial instability caused by wealth inequality. When an economic sector is decomposed into two subsectors that possess a severe wealth inequality, the sector in entirety can look financially stable while the two subsectors possess extreme financially instabilities of opposite nature, one from excessive equity, the other from lack thereof. The unstable subsector can result in further financial distress and even trigger a financial crisis. The market instability indicator, an early warning system derived from dynamical systems applied to agent-based models, is used to analyze the subsectoral financial instabilities. Detailed mathematical analysis is provided to explain what financial …


Analysis Of 2016-17 Major League Soccer Season Data Using Poisson Regression With R, Ian D. Campbell May 2018

Analysis Of 2016-17 Major League Soccer Season Data Using Poisson Regression With R, Ian D. Campbell

Undergraduate Theses and Capstone Projects

To the outside observer, soccer is chaotic with no given pattern or scheme to follow, a random conglomeration of passes and shots that go on for 90 minutes. Yet, what if there was a pattern to the chaos, or a way to describe the events that occur in the game quantifiably. Sports statistics is a critical part of baseball and a variety of other of today’s sports, but we see very little statistics and data analysis done on soccer. Of this research, there has been looks into the effect of possession time on the outcome of a game, the difference …


Building A Better Risk Prevention Model, Steven Hornyak Mar 2018

Building A Better Risk Prevention Model, Steven Hornyak

National Youth Advocacy & Resilience Conference

This presentation chronicles the work of Houston County Schools in developing a risk prevention model built on more than ten years of longitudinal student data. In its second year of implementation, Houston At-Risk Profiles (HARP), has proven effective in identifying those students most in need of support and linking them to interventions and supports that lead to improved outcomes and significantly reduces the risk of failure.


Old English Character Recognition Using Neural Networks, Sattajit Sutradhar Jan 2018

Old English Character Recognition Using Neural Networks, Sattajit Sutradhar

College of Graduate Studies: Theses & Dissertations

Character recognition has been capturing the interest of researchers since the beginning of the twentieth century. While the Optical Character Recognition for printed material is very robust and widespread nowadays, the recognition of handwritten materials lags behind. In our digital era more and more historical, handwritten documents are digitized and made available to the general public. However, these digital copies of handwritten materials lack the automatic content recognition feature of their printed materials counterparts. We are proposing a practical, accurate, and computationally efficient method for Old English character recognition from manuscript images. Our method relies on a modern machine learning …


Statistical Analysis Of Momentum In Basketball, Mackenzi Stump Dec 2017

Statistical Analysis Of Momentum In Basketball, Mackenzi Stump

Honors Projects

The “hot hand” in sports has been debated for as long as sports have been around. The debate involves whether streaks and slumps in sports are true phenomena or just simply perceptions in the mind of the human viewer. This statistical analysis of momentum in basketball analyzes the distribution of time between scoring events for the BGSU Women’s Basketball team from 2011-2017. We discuss how the distribution of time between scoring events changes with normal game factors such as location of the game, game outcome, and several other factors. If scoring events during a game were always randomly distributed, or …