Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 511 - 540 of 3231

Full-Text Articles in Data Science

Hyperparameter Tuning For Robust Autonomous Vehicle Vision, Nico D. De Ros Mar 2025

Hyperparameter Tuning For Robust Autonomous Vehicle Vision, Nico D. De Ros

Theses and Dissertations

Classification “flickering,” where the classification of an object changes inconsistently between consecutive video frames, remains a persistent issue in modern object classification algorithms. This problem undermines the reliability of autonomous vision systems and poses significant risks in high-stakes applications such as autonomous vehicles. This thesis explores the use of response surface methodology, a statistical design of experiments technique, to optimize hyperparameters across three object classification pipelines. The first pipeline combines YOLOv8 with SORT to establish a benchmark. The second integrates a Bayesian back-end, while the third employs an exponential smoothing back-end. Hyperparameter tuning was conducted using a two-step process: an …


Machine Learning Techniques To Predict Solar Particle Events And Radiation Of Aircrew, Haley Traub Mar 2025

Machine Learning Techniques To Predict Solar Particle Events And Radiation Of Aircrew, Haley Traub

Theses and Dissertations

Solar Particle Events (SPEs) are high-energy phenomena from the Sun that pose risks to technology, human health, and Air Force operations. Accurate prediction of SPEs exceeding 100 MeV is crucial for mitigating these risks. This thesis explores using Bayesian statistical models to predict such events, integrating prior knowledge from solar physics with the ability to update predictions based on new data. The research uses a dataset spanning three solar cycles (21–23) and incorporates attributes like flare fluence, peak flux, latitude, longitude, and class. Four Bayesian models (PyMC, Bnlearn, and two Dredge models) were compared to machine learning models. The Bayesian …


Geo-Spatial Mapping Of Sentiment Analysis With Transformer-Based Models, Dugan J. Turnbow Mar 2025

Geo-Spatial Mapping Of Sentiment Analysis With Transformer-Based Models, Dugan J. Turnbow

Theses and Dissertations

The public sentiment of events of interest, and their impacts, is vital for decision makers to allocate resources. This research develops a robust algorithm for aggregating sentiment analysis from social media and published articles, while contextualizing results through spatial and temporal mapping. The methodology employs two transformer-based language models for sentiment analysis and named entity recognition (NER). Sentiment scores are generated and augmented using explicit location data, such as latitude and longitude, and implicit location data derived through NER or location features. Results are mapped using a geo-tagged location dictionary, enabling visualization of sentiment trends at state and county levels …


Improving Zero Shot Learning By Linking Multi-Label Cnns With Llms, Michael A. Wegner Mar 2025

Improving Zero Shot Learning By Linking Multi-Label Cnns With Llms, Michael A. Wegner

Theses and Dissertations

Classifying previously unseen objects poses a significant challenge for traditional computer vision algorithms, which rely on extensive labeled training data. Zero-shot reasoning offers a way to overcome this limitation. This research explores a novel method for image recognition using the Animals with Attributes 2 (AWA2) dataset as a proof of concept. A multi-label ResNet50 model predicts core attributes like color, ear shape, or number of limbs. Those attributes then feed into ChatGPT which leverages its extensive knowledge base to classify the animal based on the provided attributes. This novel approach skips the need to train on every possible class. Instead, …


Comparative Evaluation Of Linear Regression, Cross Validation And Regularization Approaches In Multivariate Data Analysis, Ransford Owusu, Felix Yeboah, Francis Effah Boateng Feb 2025

Comparative Evaluation Of Linear Regression, Cross Validation And Regularization Approaches In Multivariate Data Analysis, Ransford Owusu, Felix Yeboah, Francis Effah Boateng

Data Science and Data Mining

This study evaluates linear regression and its enhanced variants incorporating cross-validation and regularization techniques for high-dimensional, multivariate datasets. We address challenges such as multicollinearity and overfitting. Methods including Ridge, LASSO, and Elastic Net are compared against ordinary least squares regression. Empirical analysis using an automobile dataset for fuel efficiency prediction shows that while OLS regression captures basic relationships, its limitations are mitigated through regularization and cross-validation, resulting in improved model interpretability. The findings provide a comprehensive framework for predictive modeling in complex data environments and offer insights into statistical methodology and practical applications in the automobile industry.


Toward Quantifying Interpolation Uncertainty In Set-Line Spacing Hydrographic Surveys, Elias Adediran, Christos Kastrisios, Kim Lowell, Glen Rice, Qi Zhang Feb 2025

Toward Quantifying Interpolation Uncertainty In Set-Line Spacing Hydrographic Surveys, Elias Adediran, Christos Kastrisios, Kim Lowell, Glen Rice, Qi Zhang

Faculty Publications

The oceans remain one of Earth’s last great unknowns, with about 74% still unmapped to modern standards. Consequently, interpolation is employed to create seamless digital bathymetric models (DBMs) from incomplete hydrographic datasets, but this introduces unquantified depth uncertainties. This study aims to estimate and characterize uncertainties arising from set-line spacing hydrographic surveys, which are important for nautical charting, navigational safety, and many other applications. By sampling at different line spacings four complete coverage testbeds that vary in slope and roughness, the study interpolates across entire testbed areas using Spline, Inverse Distance Weighting, and Linear interpolation. The resulting interpolation uncertainties are …


A Report On Health Care Access By The United States Citizens., Kelvin Njuki, Emil Agbemade Feb 2025

A Report On Health Care Access By The United States Citizens., Kelvin Njuki, Emil Agbemade

Data Science and Data Mining

Access to health care is a critical factor in ensuring public health. This study analyzes data from the National Health Interview Survey (NHIS) for the years 2015–2018 to examine the relationship between health care coverage, affordability, and costs among U.S. families. Re-sults indicate that families with at least one member covered by health insurance were more likely to afford medical care and incur lower health care costs. Despite a high proportion of families with health care coverage during this period, the number of insured family members declined over the years. These findings underscore the importance of health care coverage in …


Proxy Panels Enable Privacy-Aware Outsourcing Of Genotype Imputation, Degui Zhi, Xiaoqian Jiang, Arif Harmanci Feb 2025

Proxy Panels Enable Privacy-Aware Outsourcing Of Genotype Imputation, Degui Zhi, Xiaoqian Jiang, Arif Harmanci

Faculty, Staff and Student Publications

One of the major challenges in genomic data sharing is protecting participants' privacy in collaborative studies and in cases when genomic data are outsourced to perform analysis tasks, for example, genotype imputation services and federated collaborations genomic analysis. Although numerous cryptographic methods have been developed, these methods may not yet be practical for population-scale tasks in terms of computational requirements, rely on high-level expertise in security, and require each algorithm to be implemented from scratch. In this study, we focus on outsourcing of genotype imputation, a fundamental task that utilizes population-level reference panels, and develop protocols that rely on using …


Modeling The Relationship Between Calories And Activity Metrics: A Regression Analysis With Variable Selection, Felix Yeboah Feb 2025

Modeling The Relationship Between Calories And Activity Metrics: A Regression Analysis With Variable Selection, Felix Yeboah

Data Science and Data Mining

Physical activity monitors have become integral to daily routines, with wearable devices such as the Apple Watch and Fitbit offering continuous data on users’ physical activity. This study compares the measurement accuracy of these devices by examining how they record parameters relevant to fitness and health. Employing multiple linear regression, we modeled the relationship between calories expended and a set of explanatory variables, including heart rate, steps, distance, age, activity level, weight, and device type. Evaluation of all possible variable combinations identified heart rate, steps, distance, weight, and watch type as the most effective predictors of calorie expenditure. Although the …


Locational Data And The Public Interest, William A. Herbert, Micahel Goodchild, Richard Appelbaum, Jeremy Crampton, Gary Langham, Krzysztof Janowicz, Mei-Po Kwan, Katina Michael, Lisa Schamess Feb 2025

Locational Data And The Public Interest, William A. Herbert, Micahel Goodchild, Richard Appelbaum, Jeremy Crampton, Gary Langham, Krzysztof Janowicz, Mei-Po Kwan, Katina Michael, Lisa Schamess

Publications and Research

This article presents a paper developed by the AAG Organizing Committee on Locational Information and the Public Interest through a summit held in Santa Barbara, California in June 2022. The summit resulted in goals and ideas for addressing the issues that arise from the present environment for geodata, whereby public, private, and third-sector entities can tap into publicly available locational information with relatively little regulation on its access or use. The Committee articulates four goals: (1) develop a research agenda extending across disciplines, (2) outline educational resources and strategies to guide ethical practice, (3) devise a pathway to increase public …


Bayesian Machine Learning Approach For Corn Yield Prediction Using Satellite Imagery And Topographic Data, Etornam Kwame Kunu, Hossein Moradi Rekabdarkolaee Feb 2025

Bayesian Machine Learning Approach For Corn Yield Prediction Using Satellite Imagery And Topographic Data, Etornam Kwame Kunu, Hossein Moradi Rekabdarkolaee

SDSU Data Science Symposium

In an era of climate change and growing global food demand, accurate crop yield prediction is pivotal for leveraging advanced technologies to enhance crop management and sustainability. This study compares the prediction performance of several Bayesian Machine Learning method using high-resolution PlanetScope imagery and topographic data. In specific, the Bayesian Linear Regression, Bayesian Random Forest, Bayesian Splines, Bayesian Additive Regression Trees, and Bayesian Neural Network were developed to incorporate uncertainty quantification and achieve enhanced predictive accuracy. Our finding shows that the Bayesian Random Forest outperform the other model in term of crop yield prediction.


Classification Of Schizophrenia, Bipolar Disorder And Major Depressive Disorder With Comorbid Traits And Deep Learning Algorithms, Xiangning Chen, Yimei Lu, Joan Manuel Cue, Mira V Han, Vishwajit L Nimgaonkar, Daniel R Weinberger, Shizhong Han, Zhongming Zhao, Jingchun Chen Feb 2025

Classification Of Schizophrenia, Bipolar Disorder And Major Depressive Disorder With Comorbid Traits And Deep Learning Algorithms, Xiangning Chen, Yimei Lu, Joan Manuel Cue, Mira V Han, Vishwajit L Nimgaonkar, Daniel R Weinberger, Shizhong Han, Zhongming Zhao, Jingchun Chen

Faculty, Staff and Student Publications

Many psychiatric disorders share genetic liabilities, but whether these shared liabilities can be utilized to classify and differentiate psychiatric disorders remains unclear. In this study, we use polygenic risk scores (PRSs) of 42 traits comorbid with schizophrenia (SCZ), bipolar disorder (BIP), and major depressive disorder (MDD) to evaluate their utilities. We found that combining target specific PRS with PRSs of comorbid traits can improve the classification of the target disorders. Importantly, without inclusion of PRSs from targeted disorders, we can still classify SCZ (accuracy 0.710 ± 0.008, AUC 0.789 ± 0.011), BIP (accuracy 0.782 ± 0.006, AUC 0.852 ± 0.004), …


Optimized Hiv/Aids Resource Allocation In Ohio: A Linear Programming Approach, Godfred Ahenkroa Kesse Feb 2025

Optimized Hiv/Aids Resource Allocation In Ohio: A Linear Programming Approach, Godfred Ahenkroa Kesse

Data Science and Data Mining

This study employs a linear and integer programming approach to optimize HIV resource allocation in Ohio, aiming to minimize new infections and enhance the impact of limited resources. With the advances in HIV prevention and treatment, Ohio faces challenges in addressing disparities in access to healthcare, particularly among high-risk populations. The proposed model integrates data on infection rates, transmission patterns, demographic factors, and cost-effectiveness to provide a decision-support framework for policymakers. Using epidemiological data and equity constraints, the model prioritizes high-risk regions and populations while ensuring fair resource distribution. Results indicate that increased funding allocations significantly enhance the potential to …


Assessing Vegetation Changes And Ecosystem Impacts Of Invasive Species In Florida's Water Conservation Area 3a (2004–2024) Using Landsat Data, Gideon Tandoh Feb 2025

Assessing Vegetation Changes And Ecosystem Impacts Of Invasive Species In Florida's Water Conservation Area 3a (2004–2024) Using Landsat Data, Gideon Tandoh

Research, Papers & Creative Work

The spread of invasive species is a concrete problem, threatening biodiversity, ecosystem integrity, and the functional status of wetlands. As these species continue to advance, it is crucial to assess their spread as well as management success through the more sophisticated invasive monitoring methodologies. This research examines the change in vegetation communities and the spread of invasive plants in the Water Conservation Area 3A (WCA-3A) using Object-Based Image Analysis (OBIA) and Support Vector Machine (SVM) classifiers on time series Landsat images. Unlike traditional pixel-based methods, OBIA enhances classification accuracy by incorporating spatial, spectral, and contextual information, enabling precise differentiation between …


Testing Toxicity: Two Approaches To Toxicity Detection, Nelson E. Jarrin Feb 2025

Testing Toxicity: Two Approaches To Toxicity Detection, Nelson E. Jarrin

Dissertations, Theses, and Capstone Projects

The current online environment is increasingly being polluted by polarization, echo chambers, and a lack of constructive dialogue between individuals with differing viewpoints. Social media algorithms, while intended to personalize user experiences and marketing areas, have been heavily criticized for inadvertently amplifying extremist content and contributing to the spread of toxic commentary. The current state of the internet hinders productive conversations, fuels misinformation, entices disinformation and contributes to online and offline societal division. This is fueled by algorithms that can create echo chambers and by the anonymity given by online platforms. While tools like the Perspective API exist to identify …


Quantified Lives: Data, Bias, And The Cost Of Categorization, Colin F. Geraghty Feb 2025

Quantified Lives: Data, Bias, And The Cost Of Categorization, Colin F. Geraghty

Dissertations, Theses, and Capstone Projects

This project critically examines how statistical methods and data visualization have historically been used to categorize, rank, and control human populations. By exploring the works of key figures like Francis Galton, Karl Pearson, and R.A. Fisher, it traces the evolution of categorization frameworks from their eugenic roots to modern applications in artificial intelligence and global metrics like the World Happiness Report. Through interactive visualizations and historical analysis, the project invites users to question the legitimacy of these frameworks and reflect on their impact on society. At its core, this work emphasizes the need to move beyond rigid classification systems, challenging …


Heartdj - Music Recommendation And Generation Through Biofeedback From Heart Rate Variability, Egemen Şahin Jan 2025

Heartdj - Music Recommendation And Generation Through Biofeedback From Heart Rate Variability, Egemen Şahin

Dartmouth College Master’s Theses

This study investigates the integration of real-time physiological data with AI-generated music to enhance emotional well-being, stress regulation, and focus, using Heart Rate Variability (HRV) as a biomarker of autonomic function. Conducted in two phases—Stable Audio Open (SAO) and Suno (SUNO)—the research evaluates biofeedback-driven music interventions across varying daily music-listening habits.

In the SAO phase, short AI-generated instrumental tracks were compared with Spotify recommendations and guided meditation. Modest HRV improvements were observed in biofeedback conditions, but participants noted emotional limitations, citing short track lengths and abrupt transitions.

The SUNO phase addressed these limitations with longer, more complex AI-generated compositions combined …


In Memoriam - Nora Sabelli: Master Orchestrator Of Grant Programs And Mentor For Advancing The Interdisciplinary Learning Sciences Field, Eric Hamilton, Jeremy Roschelle, Roy Pea, Barbara Means, Louis Gomez, Kim Gomez, Nancy Butler Songer Jan 2025

In Memoriam - Nora Sabelli: Master Orchestrator Of Grant Programs And Mentor For Advancing The Interdisciplinary Learning Sciences Field, Eric Hamilton, Jeremy Roschelle, Roy Pea, Barbara Means, Louis Gomez, Kim Gomez, Nancy Butler Songer

Education Division Scholarship

On Friday, September 6, 2024, the learning sciences field lost a giant in Dr. Nora Sabelli, 87 years old, a personal mentor to many researchers and an inspiration to so many learning scientists and STEM leaders. Nora’s first professional career was as a computational chemist, and later she became a passionate leader in research for improving STEM education. Nora’s time as a senior program officer at the National Science Foundation’s (NSF) Education and Human Resources (EHR) directorate was legendary; she was a force of nature who reshaped funding priorities for stronger science and a stronger connection of science to education …


Uncovering Acoustic Biomarkers To Classify Parkinson Disease Through Machine Learning, Felix Yeboah Jan 2025

Uncovering Acoustic Biomarkers To Classify Parkinson Disease Through Machine Learning, Felix Yeboah

Data Science and Data Mining

The early detection of diseases profoundly influences treatment efficacy, and accurate classification methodologies are essential for effective disease identification. In this project, we examined fve different classifers—Logistic Regression, Gaussian Naive Bayes, K Nearest Neighbor (KNN), Extreme Gradient Boosting (XGBoost), and Support Vector Machines—and evaluated their performance in detecting Parkinson’s disease (PD) based on voice features. The study aims to identify the best classifier for detecting PD. XGBoost performed the best, with an accuracy of 91% on the full dataset. After variable selection, KNN had the best performance with an accuracy of 91%. These findings suggest that Machine learning algorithms(classifiers) can …


Evaluation Of Variable Selection Techniques On The Genetic Architecture Of Flowering Time In Maize, Felix Yeboah Jan 2025

Evaluation Of Variable Selection Techniques On The Genetic Architecture Of Flowering Time In Maize, Felix Yeboah

Data Science and Data Mining

In this project, we investigate several variable selection procedures to give an overview of how well they perform on a genomic dataset using three different penalized regression approaches. Comparisons between different methods were performed. These methods include Ridge, lasso, and Elastic Net. We utilized 4494 observations with 7389 SNPs gene scores to predict time to male flowering (dtoa). We assessed the performance of these three models in terms of mean square error. Not surprisingly, Lasso and Elastic Net perform better than Ridge Regression. Overall, Elastic Net performed better in predicting the time of male flowering (dtoa).


Comparison Of Two Strategies Of Screening Experiments: Single-Shot Experiment Vs. Two-Stage Screening Experiment, Kelvin Njuki, Emil Agbemade Jan 2025

Comparison Of Two Strategies Of Screening Experiments: Single-Shot Experiment Vs. Two-Stage Screening Experiment, Kelvin Njuki, Emil Agbemade

Data Science and Data Mining

Experiments involving many factors are often complex, time-consuming, and expensive. Screening out the least important factors helps the experimenter(s) allocate the limited resources efciently to the most important factors. Supersaturated and orthogonal array designs are among the designs used to conduct screening experiments. Supersaturated designs (SSDs) are those where the number of runs (observations) is less than the number of factors, while orthogonal array (OA) designs are those where at least the columns are orthogonal to each other. In this study, we conduct a simulation study to compare two strategies of screening experiments. Strategy one is a single shot experiment …


Pull Or Play? The Interpretation Of A Novel, Ai-Powered On-Field Decision Support Tool., Lynne Becker, Dafne Badilla, Osho Yonzon, Xiaoyu Wu, Roland Rocafort, Erik Viogan Phd, Devansh Manocha Jan 2025

Pull Or Play? The Interpretation Of A Novel, Ai-Powered On-Field Decision Support Tool., Lynne Becker, Dafne Badilla, Osho Yonzon, Xiaoyu Wu, Roland Rocafort, Erik Viogan Phd, Devansh Manocha

Journal for Sports Neuroscience

The document titled "Pull or Play? The interpretation of a novel, AI-powered on-field decision support tool" explores the development and application of the Injury Impact Severity Score (IISS)™ for assessing traumatic brain injuries (TBIs), particularly in sports settings. It addresses the limitations of current assessment tools like the Glasgow Coma Scale (GCS) and proposes a more objective approach using patient-reported data and machine learning algorithms.

Key points include:

  • Background: TBIs are a significant health concern with under-reported cases and a lack of effective research. Current assessment tools like the GCS have limitations in accuracy and speed, especially in dynamic …


Advanced Machine Learning Techniques For Cardiovascular Disease Risk Prediction, Godfred Ahenkroa Kesse Jan 2025

Advanced Machine Learning Techniques For Cardiovascular Disease Risk Prediction, Godfred Ahenkroa Kesse

Data Science and Data Mining

of mortality, necessitating advanced predictive models to aid early detection and prevention. This study explores the application of machine learning techniques, including Lo- gistic Regression, K-Nearest Neighbors (KNN), Random Forest, and XGBoost, to predict CVD risk using a dataset of 69,997 observations encompassing demographic, clinical, and lifestyle factors. Data preprocessing involved one-hot encoding of cat- egorical variables and scaling to ensure compatibility with all models. Model performance was evaluated using metrics such as accuracy, precision, recall, F1-score, and AUC-ROC. Among the models, XGBoost demonstrated the highest accuracy at 74%, leveraging its gradient-boosting framework to effectively handle feature interactions and imbalanced …


Predicting Blood Glucose Levels: A Linear Regression Approach For Non-Invasive Monitoring, Godfred Ahenkroa Kesse Jan 2025

Predicting Blood Glucose Levels: A Linear Regression Approach For Non-Invasive Monitoring, Godfred Ahenkroa Kesse

Data Science and Data Mining

Accurate monitoring of blood glucose levels is vital for the management of diabetes, a chronic condition affecting millions worldwide. This study explores a linear regression approach to estimate glucose levels non-invasively using a dataset enriched with demographic, physiological, and sensor-based variables. Following rigorous data preparation, including normalization and encoding, a Box-Cox transformation was applied to address violations of regression assumptions, stabilizing variance and improving model validity. Stepwise selection and hypothesis testing were employed to refne the model, retaining signifcant predictors such as AGE, GENDER, HEARTRATE, and DIABETIC, while excluding variables like NIR Reading and LAST EATEN for their minimal contribution. …


Handwritten Digit Recognition Using Naive Bayes And K-Nearest Neighbor Models, Godfred Ahenkroa Kesse Jan 2025

Handwritten Digit Recognition Using Naive Bayes And K-Nearest Neighbor Models, Godfred Ahenkroa Kesse

Data Science and Data Mining

This paper explores the performance of two fundamental classifcation algorithms. It uses Naive Bayes and K-Nearest Neighbors (KNN), framing it within the context of digit recognition of the MNIST dataset. The MNIST dataset has 70,00 grayscale images of handwritten digits, offering a standard for assessing classifcation models. This paper focuses on key performance metrics such as precision, accuracy, recall, and F1score to examine the effciency of each model. The results reveal that Naive Bayes has moderate accuracy and misclassifcations because of its notion of feature independence. The paper concludes that the KNN model performs better with the optimal k-value of …


Variable Selection Using Lasso Regression, Godfred Ahenkroa Kesse Jan 2025

Variable Selection Using Lasso Regression, Godfred Ahenkroa Kesse

Data Science and Data Mining

This study employs Lasso regression to analyze highdimensional genetic data for predicting flowering time in maize, specifically Days to Anthesis (DtoA). Lasso, or Least Absolute Shrinkage and Selection Operator, is a form of linear regression that introduces an L1 penalty to the model, encouraging sparsity by shrinking some coefficients to zero. This attribute makes Lasso ideal for feature selection in large datasets, as it highlights the most influential predictors while discarding irrelevant variables. Unlike Ridge regression, which applies an L2 penalty to minimize the squared magnitude of coefficients, Lasso’s L1 penalty induces sparsity, providing a clearer interpretation of the selected …


Classification And Evaluation Of Machine Learning Algorithms On The Mnist Dataset, Felix Yeboah Jan 2025

Classification And Evaluation Of Machine Learning Algorithms On The Mnist Dataset, Felix Yeboah

Data Science and Data Mining

This paper discusses the use of machine learning algorithms in classifying the MNIST handwritten dataset. The MNIST dataset consists of 28x28 grayscale handwritten images with 10 classes from 0 to 9. The dataset was normalized by scaling the pixel values to a range between 0 and 1 by dividing each pixel value by 255. We compare and evaluate the K-nearest Neighbor and Naive Bayes algorithm based on performance metrics such as accuracy, error rate, f1-score, and precision. The K-nearest Neighbor algorithm achieved better performance in all the evaluation criteria.


Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari Jan 2025

Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari

Posters-at-the-Capitol

The grocery retail industry landscape has changed greatly in the wake of the pandemic. Specifically, delivery and pickup services have become more popular and customer buying habits have evolved. At the same time, improvements in data collection and analysis have allowed grocery marketing strategies to become highly individualized.

We worked with 84.51, an analytics firm, to identify customer segments for the Kroger Company based on data from 2023. Using clustering techniques, we organized customers into groups, or segments, based on similar characteristics. We identified and profiled four distinct groups of customers. Three segments were characterized by high frequency and spending …


Scproatlas: An Atlas Of Multiplexed Single-Cell Spatial Proteomics Imaging In Human Tissues, Tiangang Wang, Xuanmin Chen, Yujuan Han, Jiahao Yi, Xi Liu, Pora Kim, Liyu Huang, Kexin Huang, Xiaobo Zhou Jan 2025

Scproatlas: An Atlas Of Multiplexed Single-Cell Spatial Proteomics Imaging In Human Tissues, Tiangang Wang, Xuanmin Chen, Yujuan Han, Jiahao Yi, Xi Liu, Pora Kim, Liyu Huang, Kexin Huang, Xiaobo Zhou

Faculty, Staff and Student Publications

Spatial proteomics can visualize and quantify protein expression profiles within tissues at single-cell resolution. Although spatial proteomics can only detect a limited number of proteins compared to spatial transcriptomics, it provides comprehensive spatial information with single-cell resolution. By studying the spatial distribution of cells, we can clearly obtain the spatial context within tissues at multiple scales. Spatial context includes the spatial composition of cell types, the distribution of functional structures, and the spatial communication between functional regions, all of which are crucial for the patterns of cellular distribution. Here, we constructed a comprehensive spatial proteomics functional annotation knowledgebase, scProAtlas (https://relab.xidian.edu.cn/scProAtlas/#/), …


Aspdb: An Integrative Knowledgebase Of Human Protein Isoforms From Experimental And Ai-Predicted Structures, Yuntao Yang, Himansu Kumar, Yuhan Xie, Zhao Li, Rongbin Li, Wenbo Chen, Chiamaka S Diala, Meer A Ali, Yi Xu, Albon Wu, Sayed-Rzgar Hosseini, Erfei Bi, Hongyu Zhao, Pora Kim, W Jim Zheng Jan 2025

Aspdb: An Integrative Knowledgebase Of Human Protein Isoforms From Experimental And Ai-Predicted Structures, Yuntao Yang, Himansu Kumar, Yuhan Xie, Zhao Li, Rongbin Li, Wenbo Chen, Chiamaka S Diala, Meer A Ali, Yi Xu, Albon Wu, Sayed-Rzgar Hosseini, Erfei Bi, Hongyu Zhao, Pora Kim, W Jim Zheng

Faculty, Staff and Student Publications

Alternative splicing is a crucial cellular process in eukaryotes, enabling the generation of multiple protein isoforms with diverse functions from a single gene. To better understand the impact of alternative splicing on protein structures, protein-protein interaction and human diseases, we developed ASpdb (https://biodataai.uth.edu/ASpdb/), a comprehensive database integrating experimentally determined structures and AlphaFold 2-predicted models for human protein isoforms. ASpdb includes over 3400 canonical isoforms, each represented by both experimentally resolved and predicted structures, and >7200 alternative isoforms with AlphaFold 2 predictions. In addition to detailed splicing events, 3D structures, sequence variations and functional annotations, ASpdb uniquely offers comparative analyses and …