Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (9)
- Data Science (7)
- Statistical Models (7)
- Categorical Data Analysis (6)
- Statistical Methodology (5)
-
- Medicine and Health Sciences (4)
- Analysis (3)
- Biostatistics (3)
- Business (3)
- Business Analytics (3)
- Mathematics (3)
- Computer Sciences (2)
- Education (2)
- Longitudinal Data Analysis and Time Series (2)
- Other Statistics and Probability (2)
- Advertising and Promotion Management (1)
- Applied Mathematics (1)
- Artificial Intelligence and Robotics (1)
- Business Intelligence (1)
- Cardiovascular Diseases (1)
- Cognition and Perception (1)
- Curriculum and Instruction (1)
- Design of Experiments and Sample Surveys (1)
- Disability and Equity in Education (1)
- Diseases (1)
- Educational Administration and Supervision (1)
- Educational Assessment, Evaluation, and Research (1)
- Institution
- Keyword
-
- Analytics (1)
- Austerity (1)
- Conventional taxi integration (1)
- Data editing (1)
- Data integrity (1)
-
- Differencesin-differences analysis (1)
- EM algorithm (1)
- Fiscal contractions (1)
- Gaussian mixture models (1)
- Lock-in effects (1)
- Model selection (1)
- Multivariate Adaptive Regression Splines (1)
- Neuroscience (1)
- Penalized likelihood (1)
- Private Consumption (1)
- Private Credit (1)
- Private Investment (1)
- Public Expenditure (1)
- Quality control (1)
- Research (1)
- Ride-hailing platforms (1)
- Risk Prevention (1)
- Risk Profiles (1)
- Statistical Modeling (1)
- Taxi ridership trends (1)
- Urban mobility regulation (1)
- Publication Year
- Publication
- File Type
Articles 1 - 17 of 17
Full-Text Articles in Multivariate Analysis
Predicting Remaining Useful Life Using Multivariate Time-Series Data, Anayah Smith, Victoria Gaibor
Predicting Remaining Useful Life Using Multivariate Time-Series Data, Anayah Smith, Victoria Gaibor
Discovery Day - Daytona Beach
Accurate prediction of Remaining Useful Life (RUL) is critical for enabling predictive maintenance, improving system reliability, and reducing operational costs in degrading systems. This project addresses the problem of modeling and predicting RUL using multivariate time-series sensor data from the NASA CMAPSS turbofan engine dataset, with a focus on understanding how predictive performance changes across datasets of varying complexity. The objective is to develop a reproducible machine learning pipeline that captures degradation patterns and produces reliable time-to-failure predictions. The approach includes data preprocessing, exploratory data analysis, feature engineering, dimensionality reduction, and model evaluation. RUL values are computed and capped to …
A Machine-Learning Tool-Supported Methodology For Nonprofit Donor Analysis, Corbin Weiss
A Machine-Learning Tool-Supported Methodology For Nonprofit Donor Analysis, Corbin Weiss
Campus Research Month
We developed a machine-learning tool-supported methodology for modeling the nonprofit donor relationship. This approach was demonstrated in the case of a US-based nonprofit. Conclusions were drawn from this example and tool-support provided for use by other nonprofits.
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Posters-at-the-Capitol
The grocery retail industry landscape has changed greatly in the wake of the pandemic. Specifically, delivery and pickup services have become more popular and customer buying habits have evolved. At the same time, improvements in data collection and analysis have allowed grocery marketing strategies to become highly individualized.
We worked with 84.51, an analytics firm, to identify customer segments for the Kroger Company based on data from 2023. Using clustering techniques, we organized customers into groups, or segments, based on similar characteristics. We identified and profiled four distinct groups of customers. Three segments were characterized by high frequency and spending …
If You Can’T Beat Them Join Them: Empirical Assessment Into How Integrating Conventional Taxis On The Uber App Impacts Conventional Taxi Ridership, Shahmeer Mohsin
If You Can’T Beat Them Join Them: Empirical Assessment Into How Integrating Conventional Taxis On The Uber App Impacts Conventional Taxi Ridership, Shahmeer Mohsin
CBER Conference
Since the emergence of ride-hailing platforms like Uber, conventional taxi ridership has taken a severe hit. Taxi-hailing apps like Curb and Arro have allowed conventional taxis to jump on the platform economy bandwagon and offer a similar service to ride-hailing platforms. Despite the emergence of these taxi-hailing apps, strong lock-in effects and high switching costs of popular ride-hailing platforms (Uber, Lyft, etc.) restrict the ridership volumes of conventional taxis. Recently, the ride-hailing platform, Uber has started to add conventional taxis on its app under increasing pressure from Cities and conventional taxi associations. Such integrations have the potential of increasing conventional …
Principal Component Analysis With Application To Credit Card Data, Eleanor Cain, Semhar Michael, Gary Hatfield
Principal Component Analysis With Application To Credit Card Data, Eleanor Cain, Semhar Michael, Gary Hatfield
SDSU Data Science Symposium
Principal Component Analysis (PCA) is a type of dimension reduction technique used in data analysis to process the data before making a model. In general, dimension reduction allows analysts to make conclusions about large data sets by reducing the number of variables while retaining as much information as possible. Using the numerical variables from a data set, PCA aims to compute a smaller set of uncorrelated variables, called principal components, that account for a majority of the variability from the data. The purpose of this poster is to understand PCA as well as perform PCA on a large sample credit …
Session 6: Model-Based Clustering Analysis On The Spatial-Temporal And Intensity Patterns Of Tornadoes, Yana Melnykov, Yingying Zhang, Rong Zheng
Session 6: Model-Based Clustering Analysis On The Spatial-Temporal And Intensity Patterns Of Tornadoes, Yana Melnykov, Yingying Zhang, Rong Zheng
SDSU Data Science Symposium
Tornadoes are one of the nature’s most violent windstorms that can occur all over the world except Antarctica. Previous scientific efforts were spent on studying this nature hazard from facets such as: genesis, dynamics, detection, forecasting, warning, measuring, and assessing. While we want to model the tornado datasets by using modern sophisticated statistical and computational techniques. The goal of the paper is developing novel finite mixture models and performing clustering analysis on the spatial-temporal and intensity patterns of the tornadoes. To analyze the tornado dataset, we firstly try a Gaussian distribution with the mean vector and variance-covariance matrix represented as …
Expansionary Fiscal Contraction Hypothesis: An Evidence From Pakistan, Aisha Irum
Expansionary Fiscal Contraction Hypothesis: An Evidence From Pakistan, Aisha Irum
CBER Conference
The fiscal sector in Pakistan has been facing mule-layered challenges over several years. One of the reasons is the stubborn and unproductive nature of its public expenditure, and the other one is the lower tax revenues. This issue of hovering fiscal deficit is mostly dealt with the tools of fiscal contraction/austerity which can have a potential impact on the private sector of the economy. Thus, the question which has been addressed in this study is whether the Expansionary Fiscal Contraction (EFC) hypothesis holds in case of Pakistan. Fiscal contraction episodes have been identified using growth in the growth rates of …
Employee Attrition: Analyzing Factors Influencing Job Satisfaction Of Ibm Data Scientists, Graham Nash
Employee Attrition: Analyzing Factors Influencing Job Satisfaction Of Ibm Data Scientists, Graham Nash
Symposium of Student Scholars
Employee attrition is a relevant issue that every business employer must consider when gauging the effectiveness of their employees. Whether or not an employee chooses to leave their job can come from a multitude of factors. As a result, employers need to develop methods in which they can measure attrition by calculating the several qualities of their employees. Factors like their age, years with the company, which department they work in, their level of education, their job role, and even their marital status are all considered by employers to assist in predicting employee attrition. This project will be analyzing a …
Application Of Gaussian Mixture Models To Simulated Additive Manufacturing, Jason Hasse, Semhar Michael, Anamika Prasad
Application Of Gaussian Mixture Models To Simulated Additive Manufacturing, Jason Hasse, Semhar Michael, Anamika Prasad
SDSU Data Science Symposium
Additive manufacturing (AM) is the process of building components through an iterative process of adding material in specific designs. AM has a wide range of process parameters that influence the quality of the component. This work applies Gaussian mixture models to detect clusters of similar stress values within and across components manufactured with varying process parameters. Further, a mixture of regression models is considered to simultaneously find groups and also fit regression within each group. The results are compared with a previous naive approach.
R Shiny's Self-Organizing Map, Zury Betzab Marroquin, Joshua Walsh, Trenton Wesley
R Shiny's Self-Organizing Map, Zury Betzab Marroquin, Joshua Walsh, Trenton Wesley
Annual Symposium on Biomathematics and Ecology Education and Research
No abstract provided.
Classification Of Coronary Artery Disease In Non-Diabetic Patients Using Artificial Neural Networks, Demond Handley
Classification Of Coronary Artery Disease In Non-Diabetic Patients Using Artificial Neural Networks, Demond Handley
Annual Symposium on Biomathematics and Ecology Education and Research
No abstract provided.
Predicting Unplanned Medical Visits Among Patients With Diabetes Using Machine Learning, Arielle Selya, Eric L. Johnson
Predicting Unplanned Medical Visits Among Patients With Diabetes Using Machine Learning, Arielle Selya, Eric L. Johnson
SDSU Data Science Symposium
Diabetes poses a variety of medical complications to patients, resulting in a high rate of unplanned medical visits, which are costly to patients and healthcare providers alike. However, unplanned medical visits by their nature are very difficult to predict. The current project draws upon electronic health records (EMR’s) of adult patients with diabetes who received care at Sanford Health between 2014 and 2017. Various machine learning methods were used to predict which patients have had an unplanned medical visit based on a variety of EMR variables (age, BMI, blood pressure, # of prescriptions, # of diagnoses on problem list, A1C, …
Quantitative Electroencephalography For Detecting Concussions, Sara Krehbiel, Kathy Hoke, Joanna Wares
Quantitative Electroencephalography For Detecting Concussions, Sara Krehbiel, Kathy Hoke, Joanna Wares
Biology and Medicine Through Mathematics Conference
No abstract provided.
Building A Better Risk Prevention Model, Steven Hornyak
Building A Better Risk Prevention Model, Steven Hornyak
National Youth Advocacy & Resilience Conference
This presentation chronicles the work of Houston County Schools in developing a risk prevention model built on more than ten years of longitudinal student data. In its second year of implementation, Houston At-Risk Profiles (HARP), has proven effective in identifying those students most in need of support and linking them to interventions and supports that lead to improved outcomes and significantly reduces the risk of failure.
Statistically Analyzing Assembly Line Processing Times Through Incorporation Of Product Variation, Kyle Rehr, Matthew Farr
Statistically Analyzing Assembly Line Processing Times Through Incorporation Of Product Variation, Kyle Rehr, Matthew Farr
Scholars Week
Timing methods and performance metrics are important in the heavily industrialized world we live in. Industrial plants use metrics to measure quality of production, help make decisions, and drive the strategy of the organization. However, there are many factors to be considered when measuring performance based on a metric; of which we will be analyzing the importance of product variation. We will be analyzing assembly line timings, whilst controlling for product variance, to show the importance differences between products makes in one’s ability to predict performance. In addition, we will be analyzing the current “statistical” methods used by an industrial …
Model Selection For Gaussian Mixture Models For Uncertainty Qualification, Yiyi Chen, Guang Lin, Xuan Liu
Model Selection For Gaussian Mixture Models For Uncertainty Qualification, Yiyi Chen, Guang Lin, Xuan Liu
The Summer Undergraduate Research Fellowship (SURF) Symposium
Clustering is task of assigning the objects into different groups so that the objects are more similar to each other than in other groups. Gaussian Mixture model with Expectation Maximization method is the one of the most general ways to do clustering on large data set. However, this method needs the number of Gaussian mode as input(a cluster) so it could approximate the original data set. Developing a method to automatically determine the number of single distribution model will help to apply this method to more larger context. In the original algorithm, there is a variable represent the weight of …
Relationship Between Perceived And Actual Quality Of Data Checking, Hunter Speich, Sophia Karas, Dan Erosa, Kelly Grob, Kimberly A. Barchard
Relationship Between Perceived And Actual Quality Of Data Checking, Hunter Speich, Sophia Karas, Dan Erosa, Kelly Grob, Kimberly A. Barchard
Festival of Communities: UG Symposium (Posters)
Data quality is critical to reaching correct research conclusions. Researchers attempt to ensure that they have accurate data by checking the data after it has been entered. Previous research has demonstrated that some methods of data checking are better than others, but not all researchers use the best methods. Perhaps researchers continue to use less optimal data checking methods because they mistakenly believe that they are highly accurate. The purpose of this study was to examine the relationship between perceived data quality and actual data quality. A total of 29 participants completed this study. Participants checked that letters and numbers …