Open Access. Powered by Scholars. Published by Universities.®

Multivariate Analysis Commons

Open Access. Powered by Scholars. Published by Universities.®

Conference

Discipline
Institution
Keyword
Publication Year
Publication
File Type

Articles 1 - 17 of 17

Full-Text Articles in Multivariate Analysis

Predicting Remaining Useful Life Using Multivariate Time-Series Data, Anayah Smith, Victoria Gaibor Aug 2026

Predicting Remaining Useful Life Using Multivariate Time-Series Data, Anayah Smith, Victoria Gaibor

Discovery Day - Daytona Beach

Accurate prediction of Remaining Useful Life (RUL) is critical for enabling predictive maintenance, improving system reliability, and reducing operational costs in degrading systems. This project addresses the problem of modeling and predicting RUL using multivariate time-series sensor data from the NASA CMAPSS turbofan engine dataset, with a focus on understanding how predictive performance changes across datasets of varying complexity. The objective is to develop a reproducible machine learning pipeline that captures degradation patterns and produces reliable time-to-failure predictions. The approach includes data preprocessing, exploratory data analysis, feature engineering, dimensionality reduction, and model evaluation. RUL values are computed and capped to …


A Machine-Learning Tool-Supported Methodology For Nonprofit Donor Analysis, Corbin Weiss Apr 2025

A Machine-Learning Tool-Supported Methodology For Nonprofit Donor Analysis, Corbin Weiss

Campus Research Month

We developed a machine-learning tool-supported methodology for modeling the nonprofit donor relationship. This approach was demonstrated in the case of a US-based nonprofit. Conclusions were drawn from this example and tool-support provided for use by other nonprofits.


Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari Jan 2025

Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari

Posters-at-the-Capitol

The grocery retail industry landscape has changed greatly in the wake of the pandemic. Specifically, delivery and pickup services have become more popular and customer buying habits have evolved. At the same time, improvements in data collection and analysis have allowed grocery marketing strategies to become highly individualized.

We worked with 84.51, an analytics firm, to identify customer segments for the Kroger Company based on data from 2023. Using clustering techniques, we organized customers into groups, or segments, based on similar characteristics. We identified and profiled four distinct groups of customers. Three segments were characterized by high frequency and spending …


If You Can’T Beat Them Join Them: Empirical Assessment Into How Integrating Conventional Taxis On The Uber App Impacts Conventional Taxi Ridership, Shahmeer Mohsin Dec 2024

If You Can’T Beat Them Join Them: Empirical Assessment Into How Integrating Conventional Taxis On The Uber App Impacts Conventional Taxi Ridership, Shahmeer Mohsin

CBER Conference

Since the emergence of ride-hailing platforms like Uber, conventional taxi ridership has taken a severe hit. Taxi-hailing apps like Curb and Arro have allowed conventional taxis to jump on the platform economy bandwagon and offer a similar service to ride-hailing platforms. Despite the emergence of these taxi-hailing apps, strong lock-in effects and high switching costs of popular ride-hailing platforms (Uber, Lyft, etc.) restrict the ridership volumes of conventional taxis. Recently, the ride-hailing platform, Uber has started to add conventional taxis on its app under increasing pressure from Cities and conventional taxi associations. Such integrations have the potential of increasing conventional …


Principal Component Analysis With Application To Credit Card Data, Eleanor Cain, Semhar Michael, Gary Hatfield Feb 2024

Principal Component Analysis With Application To Credit Card Data, Eleanor Cain, Semhar Michael, Gary Hatfield

SDSU Data Science Symposium

Principal Component Analysis (PCA) is a type of dimension reduction technique used in data analysis to process the data before making a model. In general, dimension reduction allows analysts to make conclusions about large data sets by reducing the number of variables while retaining as much information as possible. Using the numerical variables from a data set, PCA aims to compute a smaller set of uncorrelated variables, called principal components, that account for a majority of the variability from the data. The purpose of this poster is to understand PCA as well as perform PCA on a large sample credit …


Session 6: Model-Based Clustering Analysis On The Spatial-Temporal And Intensity Patterns Of Tornadoes, Yana Melnykov, Yingying Zhang, Rong Zheng Feb 2024

Session 6: Model-Based Clustering Analysis On The Spatial-Temporal And Intensity Patterns Of Tornadoes, Yana Melnykov, Yingying Zhang, Rong Zheng

SDSU Data Science Symposium

Tornadoes are one of the nature’s most violent windstorms that can occur all over the world except Antarctica. Previous scientific efforts were spent on studying this nature hazard from facets such as: genesis, dynamics, detection, forecasting, warning, measuring, and assessing. While we want to model the tornado datasets by using modern sophisticated statistical and computational techniques. The goal of the paper is developing novel finite mixture models and performing clustering analysis on the spatial-temporal and intensity patterns of the tornadoes. To analyze the tornado dataset, we firstly try a Gaussian distribution with the mean vector and variance-covariance matrix represented as …


Expansionary Fiscal Contraction Hypothesis: An Evidence From Pakistan, Aisha Irum Nov 2023

Expansionary Fiscal Contraction Hypothesis: An Evidence From Pakistan, Aisha Irum

CBER Conference

The fiscal sector in Pakistan has been facing mule-layered challenges over several years. One of the reasons is the stubborn and unproductive nature of its public expenditure, and the other one is the lower tax revenues. This issue of hovering fiscal deficit is mostly dealt with the tools of fiscal contraction/austerity which can have a potential impact on the private sector of the economy. Thus, the question which has been addressed in this study is whether the Expansionary Fiscal Contraction (EFC) hypothesis holds in case of Pakistan. Fiscal contraction episodes have been identified using growth in the growth rates of …


Employee Attrition: Analyzing Factors Influencing Job Satisfaction Of Ibm Data Scientists, Graham Nash Apr 2023

Employee Attrition: Analyzing Factors Influencing Job Satisfaction Of Ibm Data Scientists, Graham Nash

Symposium of Student Scholars

Employee attrition is a relevant issue that every business employer must consider when gauging the effectiveness of their employees. Whether or not an employee chooses to leave their job can come from a multitude of factors. As a result, employers need to develop methods in which they can measure attrition by calculating the several qualities of their employees. Factors like their age, years with the company, which department they work in, their level of education, their job role, and even their marital status are all considered by employers to assist in predicting employee attrition. This project will be analyzing a …


Application Of Gaussian Mixture Models To Simulated Additive Manufacturing, Jason Hasse, Semhar Michael, Anamika Prasad Feb 2023

Application Of Gaussian Mixture Models To Simulated Additive Manufacturing, Jason Hasse, Semhar Michael, Anamika Prasad

SDSU Data Science Symposium

Additive manufacturing (AM) is the process of building components through an iterative process of adding material in specific designs. AM has a wide range of process parameters that influence the quality of the component. This work applies Gaussian mixture models to detect clusters of similar stress values within and across components manufactured with varying process parameters. Further, a mixture of regression models is considered to simultaneously find groups and also fit regression within each group. The results are compared with a previous naive approach.


R Shiny's Self-Organizing Map, Zury Betzab Marroquin, Joshua Walsh, Trenton Wesley Nov 2021

R Shiny's Self-Organizing Map, Zury Betzab Marroquin, Joshua Walsh, Trenton Wesley

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


Classification Of Coronary Artery Disease In Non-Diabetic Patients Using Artificial Neural Networks, Demond Handley Oct 2019

Classification Of Coronary Artery Disease In Non-Diabetic Patients Using Artificial Neural Networks, Demond Handley

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


Predicting Unplanned Medical Visits Among Patients With Diabetes Using Machine Learning, Arielle Selya, Eric L. Johnson Feb 2019

Predicting Unplanned Medical Visits Among Patients With Diabetes Using Machine Learning, Arielle Selya, Eric L. Johnson

SDSU Data Science Symposium

Diabetes poses a variety of medical complications to patients, resulting in a high rate of unplanned medical visits, which are costly to patients and healthcare providers alike. However, unplanned medical visits by their nature are very difficult to predict. The current project draws upon electronic health records (EMR’s) of adult patients with diabetes who received care at Sanford Health between 2014 and 2017. Various machine learning methods were used to predict which patients have had an unplanned medical visit based on a variety of EMR variables (age, BMI, blood pressure, # of prescriptions, # of diagnoses on problem list, A1C, …


Quantitative Electroencephalography For Detecting Concussions, Sara Krehbiel, Kathy Hoke, Joanna Wares Jun 2018

Quantitative Electroencephalography For Detecting Concussions, Sara Krehbiel, Kathy Hoke, Joanna Wares

Biology and Medicine Through Mathematics Conference

No abstract provided.


Building A Better Risk Prevention Model, Steven Hornyak Mar 2018

Building A Better Risk Prevention Model, Steven Hornyak

National Youth Advocacy & Resilience Conference

This presentation chronicles the work of Houston County Schools in developing a risk prevention model built on more than ten years of longitudinal student data. In its second year of implementation, Houston At-Risk Profiles (HARP), has proven effective in identifying those students most in need of support and linking them to interventions and supports that lead to improved outcomes and significantly reduces the risk of failure.


Statistically Analyzing Assembly Line Processing Times Through Incorporation Of Product Variation, Kyle Rehr, Matthew Farr Mar 2017

Statistically Analyzing Assembly Line Processing Times Through Incorporation Of Product Variation, Kyle Rehr, Matthew Farr

Scholars Week

Timing methods and performance metrics are important in the heavily industrialized world we live in. Industrial plants use metrics to measure quality of production, help make decisions, and drive the strategy of the organization. However, there are many factors to be considered when measuring performance based on a metric; of which we will be analyzing the importance of product variation. We will be analyzing assembly line timings, whilst controlling for product variance, to show the importance differences between products makes in one’s ability to predict performance. In addition, we will be analyzing the current “statistical” methods used by an industrial …


Model Selection For Gaussian Mixture Models For Uncertainty Qualification, Yiyi Chen, Guang Lin, Xuan Liu Aug 2015

Model Selection For Gaussian Mixture Models For Uncertainty Qualification, Yiyi Chen, Guang Lin, Xuan Liu

The Summer Undergraduate Research Fellowship (SURF) Symposium

Clustering is task of assigning the objects into different groups so that the objects are more similar to each other than in other groups. Gaussian Mixture model with Expectation Maximization method is the one of the most general ways to do clustering on large data set. However, this method needs the number of Gaussian mode as input(a cluster) so it could approximate the original data set. Developing a method to automatically determine the number of single distribution model will help to apply this method to more larger context. In the original algorithm, there is a variable represent the weight of …


Relationship Between Perceived And Actual Quality Of Data Checking, Hunter Speich, Sophia Karas, Dan Erosa, Kelly Grob, Kimberly A. Barchard Apr 2011

Relationship Between Perceived And Actual Quality Of Data Checking, Hunter Speich, Sophia Karas, Dan Erosa, Kelly Grob, Kimberly A. Barchard

Festival of Communities: UG Symposium (Posters)

Data quality is critical to reaching correct research conclusions. Researchers attempt to ensure that they have accurate data by checking the data after it has been entered. Previous research has demonstrated that some methods of data checking are better than others, but not all researchers use the best methods. Perhaps researchers continue to use less optimal data checking methods because they mistakenly believe that they are highly accurate. The purpose of this study was to examine the relationship between perceived data quality and actual data quality. A total of 29 participants completed this study. Participants checked that letters and numbers …