Modeling The Relationship Between Calories And Activity Metrics: A Regression Analysis With Variable Selection,
2025
University of Central Florida
Modeling The Relationship Between Calories And Activity Metrics: A Regression Analysis With Variable Selection, Felix Yeboah
Data Science and Data Mining
Physical activity monitors have become integral to daily routines, with wearable devices such as the Apple Watch and Fitbit offering continuous data on users’ physical activity. This study compares the measurement accuracy of these devices by examining how they record parameters relevant to fitness and health. Employing multiple linear regression, we modeled the relationship between calories expended and a set of explanatory variables, including heart rate, steps, distance, age, activity level, weight, and device type. Evaluation of all possible variable combinations identified heart rate, steps, distance, weight, and watch type as the most effective predictors of calorie expenditure. Although the …
Locational Data And The Public Interest,
2025
CUNY Hunter College
Locational Data And The Public Interest, William A. Herbert, Micahel Goodchild, Richard Appelbaum, Jeremy Crampton, Gary Langham, Krzysztof Janowicz, Mei-Po Kwan, Katina Michael, Lisa Schamess
Publications and Research
This article presents a paper developed by the AAG Organizing Committee on Locational Information and the Public Interest through a summit held in Santa Barbara, California in June 2022. The summit resulted in goals and ideas for addressing the issues that arise from the present environment for geodata, whereby public, private, and third-sector entities can tap into publicly available locational information with relatively little regulation on its access or use. The Committee articulates four goals: (1) develop a research agenda extending across disciplines, (2) outline educational resources and strategies to guide ethical practice, (3) devise a pathway to increase public …
Bayesian Machine Learning Approach For Corn Yield Prediction Using Satellite Imagery And Topographic Data,
2025
South Dakota State University
Bayesian Machine Learning Approach For Corn Yield Prediction Using Satellite Imagery And Topographic Data, Etornam Kwame Kunu, Hossein Moradi Rekabdarkolaee
SDSU Data Science Symposium
In an era of climate change and growing global food demand, accurate crop yield prediction is pivotal for leveraging advanced technologies to enhance crop management and sustainability. This study compares the prediction performance of several Bayesian Machine Learning method using high-resolution PlanetScope imagery and topographic data. In specific, the Bayesian Linear Regression, Bayesian Random Forest, Bayesian Splines, Bayesian Additive Regression Trees, and Bayesian Neural Network were developed to incorporate uncertainty quantification and achieve enhanced predictive accuracy. Our finding shows that the Bayesian Random Forest outperform the other model in term of crop yield prediction.
Classification Of Schizophrenia, Bipolar Disorder And Major Depressive Disorder With Comorbid Traits And Deep Learning Algorithms,
2025
The Texas Medical Center Library
Classification Of Schizophrenia, Bipolar Disorder And Major Depressive Disorder With Comorbid Traits And Deep Learning Algorithms, Xiangning Chen, Yimei Lu, Joan Manuel Cue, Mira V Han, Vishwajit L Nimgaonkar, Daniel R Weinberger, Shizhong Han, Zhongming Zhao, Jingchun Chen
Faculty, Staff and Student Publications
Many psychiatric disorders share genetic liabilities, but whether these shared liabilities can be utilized to classify and differentiate psychiatric disorders remains unclear. In this study, we use polygenic risk scores (PRSs) of 42 traits comorbid with schizophrenia (SCZ), bipolar disorder (BIP), and major depressive disorder (MDD) to evaluate their utilities. We found that combining target specific PRS with PRSs of comorbid traits can improve the classification of the target disorders. Importantly, without inclusion of PRSs from targeted disorders, we can still classify SCZ (accuracy 0.710 ± 0.008, AUC 0.789 ± 0.011), BIP (accuracy 0.782 ± 0.006, AUC 0.852 ± 0.004), …
Optimized Hiv/Aids Resource Allocation In Ohio: A Linear Programming Approach,
2025
University of Central Florida
Optimized Hiv/Aids Resource Allocation In Ohio: A Linear Programming Approach, Godfred Ahenkroa Kesse
Data Science and Data Mining
This study employs a linear and integer programming approach to optimize HIV resource allocation in Ohio, aiming to minimize new infections and enhance the impact of limited resources. With the advances in HIV prevention and treatment, Ohio faces challenges in addressing disparities in access to healthcare, particularly among high-risk populations. The proposed model integrates data on infection rates, transmission patterns, demographic factors, and cost-effectiveness to provide a decision-support framework for policymakers. Using epidemiological data and equity constraints, the model prioritizes high-risk regions and populations while ensuring fair resource distribution. Results indicate that increased funding allocations significantly enhance the potential to …
Assessing Vegetation Changes And Ecosystem Impacts Of Invasive Species In Florida's Water Conservation Area 3a (2004–2024) Using Landsat Data,
2025
Jacksonville State University
Assessing Vegetation Changes And Ecosystem Impacts Of Invasive Species In Florida's Water Conservation Area 3a (2004–2024) Using Landsat Data, Gideon Tandoh
Research, Papers & Creative Work
The spread of invasive species is a concrete problem, threatening biodiversity, ecosystem integrity, and the functional status of wetlands. As these species continue to advance, it is crucial to assess their spread as well as management success through the more sophisticated invasive monitoring methodologies. This research examines the change in vegetation communities and the spread of invasive plants in the Water Conservation Area 3A (WCA-3A) using Object-Based Image Analysis (OBIA) and Support Vector Machine (SVM) classifiers on time series Landsat images. Unlike traditional pixel-based methods, OBIA enhances classification accuracy by incorporating spatial, spectral, and contextual information, enabling precise differentiation between …
Testing Toxicity: Two Approaches To Toxicity Detection,
2025
CUNY Graduate Center
Testing Toxicity: Two Approaches To Toxicity Detection, Nelson E. Jarrin
Dissertations, Theses, and Capstone Projects
The current online environment is increasingly being polluted by polarization, echo chambers, and a lack of constructive dialogue between individuals with differing viewpoints. Social media algorithms, while intended to personalize user experiences and marketing areas, have been heavily criticized for inadvertently amplifying extremist content and contributing to the spread of toxic commentary. The current state of the internet hinders productive conversations, fuels misinformation, entices disinformation and contributes to online and offline societal division. This is fueled by algorithms that can create echo chambers and by the anonymity given by online platforms. While tools like the Perspective API exist to identify …
Quantified Lives: Data, Bias, And The Cost Of Categorization,
2025
CUNY Graduate Center
Quantified Lives: Data, Bias, And The Cost Of Categorization, Colin F. Geraghty
Dissertations, Theses, and Capstone Projects
This project critically examines how statistical methods and data visualization have historically been used to categorize, rank, and control human populations. By exploring the works of key figures like Francis Galton, Karl Pearson, and R.A. Fisher, it traces the evolution of categorization frameworks from their eugenic roots to modern applications in artificial intelligence and global metrics like the World Happiness Report. Through interactive visualizations and historical analysis, the project invites users to question the legitimacy of these frameworks and reflect on their impact on society. At its core, this work emphasizes the need to move beyond rigid classification systems, challenging …
Heartdj - Music Recommendation And Generation Through Biofeedback From Heart Rate Variability,
2025
Dartmouth College
Heartdj - Music Recommendation And Generation Through Biofeedback From Heart Rate Variability, Egemen Şahin
Dartmouth College Master’s Theses
This study investigates the integration of real-time physiological data with AI-generated music to enhance emotional well-being, stress regulation, and focus, using Heart Rate Variability (HRV) as a biomarker of autonomic function. Conducted in two phases—Stable Audio Open (SAO) and Suno (SUNO)—the research evaluates biofeedback-driven music interventions across varying daily music-listening habits.
In the SAO phase, short AI-generated instrumental tracks were compared with Spotify recommendations and guided meditation. Modest HRV improvements were observed in biofeedback conditions, but participants noted emotional limitations, citing short track lengths and abrupt transitions.
The SUNO phase addressed these limitations with longer, more complex AI-generated compositions combined …
In Memoriam - Nora Sabelli: Master Orchestrator Of Grant Programs And Mentor For Advancing The Interdisciplinary Learning Sciences Field,
2025
Pepperdine University
In Memoriam - Nora Sabelli: Master Orchestrator Of Grant Programs And Mentor For Advancing The Interdisciplinary Learning Sciences Field, Eric Hamilton, Jeremy Roschelle, Roy Pea, Barbara Means, Louis Gomez, Kim Gomez, Nancy Butler Songer
Education Division Scholarship
On Friday, September 6, 2024, the learning sciences field lost a giant in Dr. Nora Sabelli, 87 years old, a personal mentor to many researchers and an inspiration to so many learning scientists and STEM leaders. Nora’s first professional career was as a computational chemist, and later she became a passionate leader in research for improving STEM education. Nora’s time as a senior program officer at the National Science Foundation’s (NSF) Education and Human Resources (EHR) directorate was legendary; she was a force of nature who reshaped funding priorities for stronger science and a stronger connection of science to education …
Uncovering Acoustic Biomarkers To Classify Parkinson Disease Through Machine Learning,
2025
University of Central Florida
Uncovering Acoustic Biomarkers To Classify Parkinson Disease Through Machine Learning, Felix Yeboah
Data Science and Data Mining
The early detection of diseases profoundly influences treatment efficacy, and accurate classification methodologies are essential for effective disease identification. In this project, we examined fve different classifers—Logistic Regression, Gaussian Naive Bayes, K Nearest Neighbor (KNN), Extreme Gradient Boosting (XGBoost), and Support Vector Machines—and evaluated their performance in detecting Parkinson’s disease (PD) based on voice features. The study aims to identify the best classifier for detecting PD. XGBoost performed the best, with an accuracy of 91% on the full dataset. After variable selection, KNN had the best performance with an accuracy of 91%. These findings suggest that Machine learning algorithms(classifiers) can …
Evaluation Of Variable Selection Techniques On The Genetic Architecture Of Flowering Time In Maize,
2025
University of Central Florida
Evaluation Of Variable Selection Techniques On The Genetic Architecture Of Flowering Time In Maize, Felix Yeboah
Data Science and Data Mining
In this project, we investigate several variable selection procedures to give an overview of how well they perform on a genomic dataset using three different penalized regression approaches. Comparisons between different methods were performed. These methods include Ridge, lasso, and Elastic Net. We utilized 4494 observations with 7389 SNPs gene scores to predict time to male flowering (dtoa). We assessed the performance of these three models in terms of mean square error. Not surprisingly, Lasso and Elastic Net perform better than Ridge Regression. Overall, Elastic Net performed better in predicting the time of male flowering (dtoa).
Comparison Of Two Strategies Of Screening Experiments: Single-Shot Experiment Vs. Two-Stage Screening Experiment,
2025
Oklahoma State University - Main Campus
Comparison Of Two Strategies Of Screening Experiments: Single-Shot Experiment Vs. Two-Stage Screening Experiment, Kelvin Njuki, Emil Agbemade
Data Science and Data Mining
Experiments involving many factors are often complex, time-consuming, and expensive. Screening out the least important factors helps the experimenter(s) allocate the limited resources efciently to the most important factors. Supersaturated and orthogonal array designs are among the designs used to conduct screening experiments. Supersaturated designs (SSDs) are those where the number of runs (observations) is less than the number of factors, while orthogonal array (OA) designs are those where at least the columns are orthogonal to each other. In this study, we conduct a simulation study to compare two strategies of screening experiments. Strategy one is a single shot experiment …
Pull Or Play? The Interpretation Of A Novel, Ai-Powered On-Field Decision Support Tool.,
2025
Power of Patients
Pull Or Play? The Interpretation Of A Novel, Ai-Powered On-Field Decision Support Tool., Lynne Becker, Dafne Badilla, Osho Yonzon, Xiaoyu Wu, Roland Rocafort, Erik Viogan Phd, Devansh Manocha
Journal for Sports Neuroscience
The document titled "Pull or Play? The interpretation of a novel, AI-powered on-field decision support tool" explores the development and application of the Injury Impact Severity Score (IISS)™ for assessing traumatic brain injuries (TBIs), particularly in sports settings. It addresses the limitations of current assessment tools like the Glasgow Coma Scale (GCS) and proposes a more objective approach using patient-reported data and machine learning algorithms.
Key points include:
- Background: TBIs are a significant health concern with under-reported cases and a lack of effective research. Current assessment tools like the GCS have limitations in accuracy and speed, especially in dynamic …
Advanced Machine Learning Techniques For Cardiovascular Disease Risk Prediction,
2025
University of Central Florida
Advanced Machine Learning Techniques For Cardiovascular Disease Risk Prediction, Godfred Ahenkroa Kesse
Data Science and Data Mining
of mortality, necessitating advanced predictive models to aid early detection and prevention. This study explores the application of machine learning techniques, including Lo- gistic Regression, K-Nearest Neighbors (KNN), Random Forest, and XGBoost, to predict CVD risk using a dataset of 69,997 observations encompassing demographic, clinical, and lifestyle factors. Data preprocessing involved one-hot encoding of cat- egorical variables and scaling to ensure compatibility with all models. Model performance was evaluated using metrics such as accuracy, precision, recall, F1-score, and AUC-ROC. Among the models, XGBoost demonstrated the highest accuracy at 74%, leveraging its gradient-boosting framework to effectively handle feature interactions and imbalanced …
Predicting Blood Glucose Levels: A Linear Regression Approach For Non-Invasive Monitoring,
2025
University of Central Florida
Predicting Blood Glucose Levels: A Linear Regression Approach For Non-Invasive Monitoring, Godfred Ahenkroa Kesse
Data Science and Data Mining
Accurate monitoring of blood glucose levels is vital for the management of diabetes, a chronic condition affecting millions worldwide. This study explores a linear regression approach to estimate glucose levels non-invasively using a dataset enriched with demographic, physiological, and sensor-based variables. Following rigorous data preparation, including normalization and encoding, a Box-Cox transformation was applied to address violations of regression assumptions, stabilizing variance and improving model validity. Stepwise selection and hypothesis testing were employed to refne the model, retaining signifcant predictors such as AGE, GENDER, HEARTRATE, and DIABETIC, while excluding variables like NIR Reading and LAST EATEN for their minimal contribution. …
Handwritten Digit Recognition Using Naive Bayes And K-Nearest Neighbor Models,
2025
University of Central Florida
Handwritten Digit Recognition Using Naive Bayes And K-Nearest Neighbor Models, Godfred Ahenkroa Kesse
Data Science and Data Mining
This paper explores the performance of two fundamental classifcation algorithms. It uses Naive Bayes and K-Nearest Neighbors (KNN), framing it within the context of digit recognition of the MNIST dataset. The MNIST dataset has 70,00 grayscale images of handwritten digits, offering a standard for assessing classifcation models. This paper focuses on key performance metrics such as precision, accuracy, recall, and F1score to examine the effciency of each model. The results reveal that Naive Bayes has moderate accuracy and misclassifcations because of its notion of feature independence. The paper concludes that the KNN model performs better with the optimal k-value of …
Variable Selection Using Lasso Regression,
2025
University of Central Florida
Variable Selection Using Lasso Regression, Godfred Ahenkroa Kesse
Data Science and Data Mining
This study employs Lasso regression to analyze highdimensional genetic data for predicting flowering time in maize, specifically Days to Anthesis (DtoA). Lasso, or Least Absolute Shrinkage and Selection Operator, is a form of linear regression that introduces an L1 penalty to the model, encouraging sparsity by shrinking some coefficients to zero. This attribute makes Lasso ideal for feature selection in large datasets, as it highlights the most influential predictors while discarding irrelevant variables. Unlike Ridge regression, which applies an L2 penalty to minimize the squared magnitude of coefficients, Lasso’s L1 penalty induces sparsity, providing a clearer interpretation of the selected …
Classification And Evaluation Of Machine Learning Algorithms On The Mnist Dataset,
2025
University of Central Florida
Classification And Evaluation Of Machine Learning Algorithms On The Mnist Dataset, Felix Yeboah
Data Science and Data Mining
This paper discusses the use of machine learning algorithms in classifying the MNIST handwritten dataset. The MNIST dataset consists of 28x28 grayscale handwritten images with 10 classes from 0 to 9. The dataset was normalized by scaling the pixel values to a range between 0 and 1 by dividing each pixel value by 255. We compare and evaluate the K-nearest Neighbor and Naive Bayes algorithm based on performance metrics such as accuracy, error rate, f1-score, and precision. The K-nearest Neighbor algorithm achieved better performance in all the evaluation criteria.
Kroger Post-Pandemic Customer Segmentation,
2025
Northern Kentucky University
Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari
Posters-at-the-Capitol
The grocery retail industry landscape has changed greatly in the wake of the pandemic. Specifically, delivery and pickup services have become more popular and customer buying habits have evolved. At the same time, improvements in data collection and analysis have allowed grocery marketing strategies to become highly individualized.
We worked with 84.51, an analytics firm, to identify customer segments for the Kroger Company based on data from 2023. Using clustering techniques, we organized customers into groups, or segments, based on similar characteristics. We identified and profiled four distinct groups of customers. Three segments were characterized by high frequency and spending …
