Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Machine learning

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 31 - 60 of 241

Full-Text Articles in Data Science

Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun Apr 2025

Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun

SMU Data Science Review

Enhancing animal shelter operations through machine learning involves employing a variety of advanced techniques aimed at increasing efficiency, promoting animal welfare, and optimizing resource allocation. This paper explores predictive analytics for adoption rates using regression models to estimate the likelihood of adoption based on historical data, encompassing variables such as breed, health status, and previous adoption trends. Additionally, classification algorithms are utilized to categorize animals by adoption probability, facilitating better resources and marketing prioritization. Clustering algorithms are employed to group animals according to behavior patterns and/or physical health, enabling tailored medical care and enrichment activities that improve their mental and …


Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul Apr 2025

Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul

School of Public Health Faculty Publications

Diabetes is a growing global health concern, affecting millions and leading to severe complications if not properly managed. The primary challenge in diabetes management is maintaining blood glucose levels (BGLs) within a safe range to prevent complications such as renal failure, cardiovascular disease, and neuropathy. Traditional methods, such as finger-prick testing, often result in low patient adherence due to discomfort, invasiveness, and inconvenience. Consequently, there is an increasing need for non-invasive techniques that provide accurate BGL measurements. Photoplethysmography (PPG), a photosensitive method that detects blood volume variations, has shown promise for non-invasive glucose monitoring. Deep neural networks (DNNs) applied to …


Leveraging Large Language Models For Knowledge-Free Weak Supervision In Clinical Natural Language Processing, Enshuo Hsu, Kirk Roberts Mar 2025

Leveraging Large Language Models For Knowledge-Free Weak Supervision In Clinical Natural Language Processing, Enshuo Hsu, Kirk Roberts

Faculty, Staff and Student Publications

The performance of deep learning-based natural language processing systems is based on large amounts of labeled training data which, in the clinical domain, are not easily available or affordable. Weak supervision and in-context learning offer partial solutions to this issue, particularly using large language models (LLMs), but their performance still trails traditional supervised methods with moderate amounts of gold-standard data. In particular, inferencing with LLMs is computationally heavy. We propose an approach leveraging fine-tuning LLMs and weak supervision with virtually no domain knowledge that still achieves consistently dominant performance. Using a prompt-based approach, the LLM is used to generate weakly-labeled …


Machine Learning With Flight Data Recorder Data For Flight Fuel Consumption Predictions, Adam C. Levandowski Mar 2025

Machine Learning With Flight Data Recorder Data For Flight Fuel Consumption Predictions, Adam C. Levandowski

Theses and Dissertations

This study applies advanced Machine Learning (ML) to Flight Data Recorder (FDR) data for fuel consumption predictions. It explores feature engineering, model selection, and Hyper-Parameter Optimization (HPO) across all flight phases. Baseline models like Ordinary Least Squares (OLS) regression, Multi- Layer Perceptrons (MLPs), and decision trees are compared to Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs) with Gated Recurrent Unit (GRU) layers, and XGBoost. Results analyze segmentation strategies, tailored features, and model performance. A counterfactual analysis compares ML models to operational fuel predictions, demonstrating their deployment potential. Findings establish a foundation for future ML-driven advancements in aviation fuel optimization.


Toward Quantifying Interpolation Uncertainty In Set-Line Spacing Hydrographic Surveys, Elias Adediran, Christos Kastrisios, Kim Lowell, Glen Rice, Qi Zhang Feb 2025

Toward Quantifying Interpolation Uncertainty In Set-Line Spacing Hydrographic Surveys, Elias Adediran, Christos Kastrisios, Kim Lowell, Glen Rice, Qi Zhang

Faculty Publications

The oceans remain one of Earth’s last great unknowns, with about 74% still unmapped to modern standards. Consequently, interpolation is employed to create seamless digital bathymetric models (DBMs) from incomplete hydrographic datasets, but this introduces unquantified depth uncertainties. This study aims to estimate and characterize uncertainties arising from set-line spacing hydrographic surveys, which are important for nautical charting, navigational safety, and many other applications. By sampling at different line spacings four complete coverage testbeds that vary in slope and roughness, the study interpolates across entire testbed areas using Spline, Inverse Distance Weighting, and Linear interpolation. The resulting interpolation uncertainties are …


Classification Of Variable Stars Using Convolutional Neural Network, Abhina Premachandran Bindu Jan 2025

Classification Of Variable Stars Using Convolutional Neural Network, Abhina Premachandran Bindu

Dissertations and Theses

This research focuses on developing Convolutional Neural Networks (CNNs), for the process of classifying and identifying variable stars through the analysis of unprocessed light curves from Transiting Exoplanet Survey Satellite (TESS). As astronomical data is becoming increasingly complex, and as advanced missions deploy sophisticated instruments for data collection, both the quality and quantity of the data are improving at a rapid pace. This has created an urgent need to automate the analysis process using efficient and effective methods, such as those based on machine learning. While previous research has explored machine learning approaches, there has been limited focus on implementing …


Comparative Analysis Of Machine Learning Models For Glioblastoma Survival., Muna Awel Jan 2025

Comparative Analysis Of Machine Learning Models For Glioblastoma Survival., Muna Awel

All Graduate Theses, Dissertations, and Other Capstone Projects

Glioblastoma multiforme (GBM) remains one of the most lethal brain tumors, necessitating improved survival prediction models that integrate clinical and molecular data. This study develops a comprehensive machine learning pipeline leveraging TCGA-derived multi-omics datasets to predict binary survival outcomes. The framework integrates four classifiers Logistic Regression, Random Forest, XGBoost, and Support Vector Machine (SVM) and includes rigorous preprocessing with MCAR testing, KNN imputation, feature scaling, and hyperparameter optimization via GridSearchCV. SMOTE was applied to mitigate class imbalance and enhance model robustness for minority survival classes. Comparative performance analyses revealed Random Forest and XGBoost as top performers, achieving the highest recall …


Deep Learning-Based Ensemble Two-Step Classification Of Medical Images Using Cnn Architectures And Ensemble Methods, Noreliz Alorico Jan 2025

Deep Learning-Based Ensemble Two-Step Classification Of Medical Images Using Cnn Architectures And Ensemble Methods, Noreliz Alorico

Master's Theses or Doctor of Nursing Practice

Breast cancer remains one of the most common cancers amongst women globally. Early detection is crucial for improving survival rates. While mammography is widely used and an effective imaging technique, it can sometimes yield false positive or false negatives. Mammogram interpretation is highly operator-dependent, introducing variability and the potential for diagnostic errors. Additionally, mammographic images have limitations, such as low contrast in breast tissue and overlapping structures that can obscure lesions or mimic abnormalities. These limitations can lead to unnecessary biopsies or delayed diagnosis. These challenges highlight the needs for advanced and data driven diagnostic tools to support and enhance …


Tour Demand Forecasting In Ireland: Development And Evaluation Of Classical, Deep Learning, And Hybrid Models, Ruben Elias Charleston Montfort Jan 2025

Tour Demand Forecasting In Ireland: Development And Evaluation Of Classical, Deep Learning, And Hybrid Models, Ruben Elias Charleston Montfort

ICT

Tourism plays a significant role in global economies by supporting employment, infrastructure, and national development. As international travel continues to grow, accurate tourism demand forecasting has become increasingly important for effective planning and decision-making. In Ireland, tourism is a key economic sector attracting millions of visitors annually. For tour operators such as Irish Day Tours, reliable demand forecasting is essential for optimizing logistics, resource allocation, marketing strategies, and customer satisfaction. Advances in machine learning and deep learning techniques offer new opportunities to improve forecasting accuracy and support data-driven decision-making within the tourism industry.


European Air Pollution And The Proposed Timelines Of Implementing The World Health Organization 2021 Air Quality Guidelines Ca3., Lukia Hartin Jan 2025

European Air Pollution And The Proposed Timelines Of Implementing The World Health Organization 2021 Air Quality Guidelines Ca3., Lukia Hartin

ICT

This research examines Ireland’s air pollution trends and evaluates whether current reductions in PM2.5, PM10, and NO₂ are sufficient to meet the WHO 2021 Air Quality Guidelines by 2040. Using four years of EPA-validated pollutant data (2020–2023), alongside Building Energy Rating (BER) and national transport datasets, the study applies the CRISP-DM methodology to guide analysis, preprocessing, modelling, and evaluation. Extensive data cleaning and alignment were required due to inconsistent station coverage, varying formats, and missing values. Forecasting models—including Random Forest, Gradient Boosting, SVR, and linear regression—were assessed using MSE, R², and trend significance to project pollutant levels across different Irish …


Customer Service Support. Utilizing Machine Learning To Classify, Prioritize And Summarize Issues., Hoai Nhan Nguyen Jan 2025

Customer Service Support. Utilizing Machine Learning To Classify, Prioritize And Summarize Issues., Hoai Nhan Nguyen

ICT

This capstone project investigates the application of machine learning and natural language processing (NLP) to enhance customer support operations through automated ticket classification, prioritization, and summarization. Using the multilingual Customer Support Emails dataset from Kaggle, the project follows the CRISP-DM methodology, performing extensive data cleaning, preprocessing, feature engineering, and class balancing. Five machine learning models—Decision Tree, KNN, LinearSVC, Naive Bayes, and Random Forest—were evaluated using hyperparameter tuning, cross-validation, confusion matrix analysis, and learning curves. LinearSVC demonstrated the strongest performance for both queue and priority classification, achieving accuracies of 89.8% and 81.2% respectively, with consistent generalization across folds. For summarization, extractive …


Predicting Early Customer Inactivity In The Banking Sector Using Machine Learning: A Churn Prevention, Ivana Mc Fadden Jan 2025

Predicting Early Customer Inactivity In The Banking Sector Using Machine Learning: A Churn Prevention, Ivana Mc Fadden

ICT

Customer churn, when customers stop using a company’s services, is a challenge for the banking sector (Singh et al., 2023). High churn rates often signal poor customer experiences, resulting in revenue losses and increased costs to obtain new clients. Goyal and Srivastava (2015) stress that fostering loyalty through exceptional service and understanding customer needs is important for long-term retention.

This project aims to predict early customer inactivity, an indicator of churn, by using machine learning. Early identification of at-risk customers will allow banks to apply targeted interventions, reduce acquisition costs, and improve customer satisfaction (Singh et al., 2023). By analysing …


The Actuarial Applications Of Machine Learning And Big Data In The Life Assurance Industry: Managing Customer Retention And Customer Outcomes By The Application Of Data Science., Brian Cunningham Jan 2025

The Actuarial Applications Of Machine Learning And Big Data In The Life Assurance Industry: Managing Customer Retention And Customer Outcomes By The Application Of Data Science., Brian Cunningham

ICT

Lapses are an issue in the insurance industry in general. They affect a company’s profitability, cash flows and solvency. High levels of lapses can cause reputational damage that could provoke a cycle of even more lapses. It is therefore incumbent on a company to do its utmost to retain the business it has written for the term it was written for.

If a company could predict which of its policies were about to lapse, it could proactively attempt to prevent them by contacting the policyholder and engaging in a discussion to ascertain the likelihood of their choosing to leave. In …


Premier League Results Predictions, Laura Consuegra Jan 2025

Premier League Results Predictions, Laura Consuegra

ICT

This project applies machine learning techniques to predict outcomes in the English Premier League, one of the most prestigious and widely followed football competitions worldwide. By analysing historical and real-time match data, including performance metrics such as goals scored, shots on target, and cards received, predictive models are developed to forecast match results with higher accuracy. The study evaluates the effectiveness of these models and explores the influence of key statistical features on team performance. The findings aim to provide strategic insights for fans, bookmakers, and coaching staff, supporting performance evaluation, tactical decision-making, and a deeper understanding of the factors …


Predicting Repeat Purchases In E-Commerce Using Interpretable Machine Learning, Zahid Bhatti Jan 2025

Predicting Repeat Purchases In E-Commerce Using Interpretable Machine Learning, Zahid Bhatti

ICT

This project investigates the prediction of repeat purchase behaviour in e-commerce using machine learning, with a focus on balancing predictive accuracy and interpretability. Large volumes of transactional and behavioural data are analysed to identify customer-level features that drive loyalty and repeat purchases. Various supervised learning models, including Random Forests and Logistic Regression, are evaluated for predictive performance, while SHAP (SHapley Additive Explanations) is employed to provide both global and local interpretability. The study aims to generate actionable insights for customer relationship management and marketing strategy, demonstrating how advanced predictive models can support informed business decisions without sacrificing transparency.


Predicting Customer Churn And Enhancing Retention Strategies Through Machine Learning., Swan Saung Lwin Jan 2025

Predicting Customer Churn And Enhancing Retention Strategies Through Machine Learning., Swan Saung Lwin

ICT

Customer Churn is a critical challenge faced by businesses across industries, especially in the digital market. Many companies struggle to predict customer churn accurately and have difficulties in carrying out effective retention strategies. Key challenges include ineffective traditional methods, lack of insights into impact of different services, generalized retention strategies, and the need to have cost-effective retention strategies. This project aims to predict customer churn using machine learning and identify the impact of key services offered by a telecommunication company.


Fake News Detection Using Machine Learning Models., Anne Higgins Jan 2025

Fake News Detection Using Machine Learning Models., Anne Higgins

ICT

This project investigates the use of machine learning to detect fake news, addressing the societal and political challenges posed by the widespread dissemination of false and misleading information. Using automated classification techniques, the project analyses news content to predict the likelihood of intentional deception. The methodology follows the CRISP-DM framework, encompassing data preparation, model development, and evaluation. By leveraging machine learning, the study aims to support organisations, governments, and digital platforms in mitigating misinformation, while also considering ethical, interpretability, and strategic implications. The findings provide actionable insights for enhancing content moderation and reducing the influence of disinformation in digital media …


Data Analysis For Maintenance Reliability., Romulo Menezes Santos Jan 2025

Data Analysis For Maintenance Reliability., Romulo Menezes Santos

ICT

This project explores the application of predictive maintenance (PdM) in modern manufacturing, highlighting its strategic role in reducing downtime, optimising resources, and improving operational efficiency. Traditional reactive and preventive maintenance approaches are insufficient for Industry 4.0 environments, where machine failures can cause significant financial and operational losses. By leveraging sensor data, historical performance records, and machine learning techniques, PdM enables early detection of potential equipment failures, allowing maintenance teams to act proactively. The approach not only enhances operational reliability and safety but also supports strategic decision-making, cost control, and competitive advantage in globalised manufacturing contexts.


Credit Card Fraud Detection., Sonia Ndonga Jan 2025

Credit Card Fraud Detection., Sonia Ndonga

ICT

This capstone project applies machine learning to detect credit card fraud, addressing a critical financial threat to banks and payment providers. Using an anonymised dataset of 284,807 transactions, which is highly imbalanced with only 0.172% fraudulent cases, three models—Logistic Regression, Random Forest, and Gradient Boosting—were developed and evaluated. The pipeline incorporates data preprocessing, feature engineering, hyperparameter tuning, cross-validation, and interpretability analysis using SHAP values, SHAPASH, and permutation importance. Random Forest achieved the highest performance with an ROC AUC of 0.97 and Average Precision of 0.66. The study also considers fairness, threshold optimisation, and practical deployment strategies, providing a robust automated …


Using Machine Learning To Predict Credit Card Fraud, Pedro Henrique Das Chagas Morais Jan 2025

Using Machine Learning To Predict Credit Card Fraud, Pedro Henrique Das Chagas Morais

ICT

Credit card fraud poses a significant challenge to financial institutions, leading to substantial financial losses and declining customer trust. This project develops and evaluates machine learning models to detect fraudulent credit card transactions using a large, realistic synthetic dataset. Following data preprocessing, exploratory analysis, and class-balancing using SMOTE, four supervised models—Logistic Regression, Decision Tree, Random Forest, and XGBoost—were trained and compared. Performance was assessed using metrics suited to imbalanced classification, including AUC, Recall, Precision, F1-score, and Average Precision. Results show that XGBoost, particularly after hyperparameter optimisation, delivered the strongest performance (AUC 0.99, Recall 0.83, AP 0.70), outperforming other models and …


Credit Card Default Prediction Using Machine Learning., Yassine Zohair Jan 2025

Credit Card Default Prediction Using Machine Learning., Yassine Zohair

ICT

This project investigates the use of machine learning to predict credit card payment defaults, aiming to help banks and credit card companies mitigate financial losses. Using historical customer data, including demographics, income, education, and previous payment behaviour, three machine learning algorithms were implemented to forecast the likelihood of default. Techniques such as cross-validation, hyperparameter tuning, and SHAPASH were applied to improve model performance and interpretability. Accurate prediction of potential defaulters enables financial institutions to take proactive measures, such as adjusting credit limits or providing targeted financial guidance, thereby enhancing risk management and customer retention.


Exploring The Determinants Of Life Expectancy At Birth: Predicting And Forecasting Global Health Trends Using Statistical, Machine Learning, And Deep Learning Models., Emma Rath Jan 2025

Exploring The Determinants Of Life Expectancy At Birth: Predicting And Forecasting Global Health Trends Using Statistical, Machine Learning, And Deep Learning Models., Emma Rath

ICT

Accurate life expectancy forecasting is essential for health policy planning, yet research comparing statistical, machine learning, and deep learning approaches under real-world constraints remains limited. This study evaluates ARIMA/ARIMAX, tree-based, and neural network models using Irish and global datasets, considering small samples, missing data, and COVID-19 shocks. ARIMAX with lagged socioeconomic variables outperformed LSTM and other ML/DL methods. Income-based stratification improved predictive accuracy and interpretability, with SHAP analysis highlighting GDP per capita for developed countries and school enrolment and trade indicators for developing contexts. Results provide practical guidance for policymakers and establish limits for model complexity under constrained health data.


Evaluating Aspect-Based Sentiment Analysis In Healthcare Drug Reviews Across Machine Learning, Deep Neural Networks, And Transformer Models, Eun Soo Park Jan 2025

Evaluating Aspect-Based Sentiment Analysis In Healthcare Drug Reviews Across Machine Learning, Deep Neural Networks, And Transformer Models, Eun Soo Park

All Graduate Theses, Dissertations, and Other Capstone Projects

Sentiment analysis has become a critical area of research in Natural Language Processing (NLP), enabling insights from unstructured text. Within this field, Aspect-Based Sentiment Analysis (ABSA) plays a practical role in domains such as healthcare, where patients drug reviews often contain diverse opinions across multiple aspects, including overall comments, perceived benefits, and side effects. However, aspect-level classification remains challenging due to class imbalance, subtle sentiment expression, and the limitations of traditional models. This research investigates the performance of three modeling paradigms: traditional machine learning (SVM, SVC, and XGBoost), deep learning (CNN-BiLSTM), and transformer-based approaches (DistilBERT sentence-pair classification). Using the UCI …


A Bayesian Deep Segmentation Framework For Glioblastoma Tumor Segmentation Using Follow-Up Mris, Tanjida Kabir, Kang-Lin Hsieh, Luis Nunez, Yu-Chun Hsu, Juan C Rodriguez Quintero, Octavio Arevalo, Kangyi Zhao, Jay-Jiguang Zhu, Roy F Riascos, Mahboubeh Madadi, Xiaoqian Jiang, Shayan Shams Jan 2025

A Bayesian Deep Segmentation Framework For Glioblastoma Tumor Segmentation Using Follow-Up Mris, Tanjida Kabir, Kang-Lin Hsieh, Luis Nunez, Yu-Chun Hsu, Juan C Rodriguez Quintero, Octavio Arevalo, Kangyi Zhao, Jay-Jiguang Zhu, Roy F Riascos, Mahboubeh Madadi, Xiaoqian Jiang, Shayan Shams

Faculty, Staff and Student Publications

Background: Glioblastoma (GBM) is the most common malignant brain tumor with an abysmal prognosis. Since complete tumor cell removal is impossible due to the infiltrative nature of GBM, accurate measurement is paramount for GBM assessment. Preoperative magnetic resonance images (MRIs) are crucial for initial diagnosis and surgical planning, while follow-up MRIs are vital for evaluating treatment response. The structural changes in the brain caused by surgical and therapeutic measures create significant differences between preoperative and follow-up MRIs. In clinical research, advanced deep learning models trained on preoperative MRIs are often applied to assess follow-up scans, but their effectiveness in this …


Fares On Fairness: Using A Total Error Framework To Examine The Role Of Measurement And Representation In Training Data On Model Fairness And Bias, Patrick Oliver Schenk, Christoph Kern, Trent D. Buskirk Jan 2025

Fares On Fairness: Using A Total Error Framework To Examine The Role Of Measurement And Representation In Training Data On Model Fairness And Bias, Patrick Oliver Schenk, Christoph Kern, Trent D. Buskirk

Data Science Faculty Publications

Data-driven decisions, often based on predictions from machine learning (ML) models are becoming ubiquitous. For these decisions to be just, the underlying ML models must be fair, i.e., work equally well for all parts of the population such as groups defined by gender or age. What are the logical next steps if, however, a trained model is accurate but not fair? How can we guide the whole data pipeline such that we avoid training unfair models based on inadequate data, recognizing possible sources of unfairness early on? How can the concepts of data-based sources of unfairness that exist in the …


Training Set Augmentation And Harmonization Enables Radiomic Models To Detect Early Onset Of Lung Cancer, Claire Huchthausen, Menglin Shi, Gabriel L.A. Sousa, James Larner, Einsley Janowski, Jonathan Colen, Krishni Wijesooriya Jan 2025

Training Set Augmentation And Harmonization Enables Radiomic Models To Detect Early Onset Of Lung Cancer, Claire Huchthausen, Menglin Shi, Gabriel L.A. Sousa, James Larner, Einsley Janowski, Jonathan Colen, Krishni Wijesooriya

Data Science Faculty Publications

Radiomics-based machine learning models have the potential to detect lung cancer at inception from CT scans and transform patient outcomes. Low malignancy rates in early-development pulmonary nodules (PNs) and variable image acquisition hinder development of clinically applicable radiomics-based early detection models. To address these challenges, we augmented training using later-development PNs and harmonized for acquisition effects. We first trained machine learning models to predict PN malignancy using radiomic features from scans of early-development benign and malignant PNs (n = 187) harmonized using ComBat. Observing near-chance performance, we augmented training with later-development benign and malignant PNs (n = 225). We evaluated …


T3-Ciders: Train-The-Trainer And Community Building To Increase Cyberinfrastructure Adoption In Cybersecurity Research And Education, Wirawan Purwanto, Mohan Yang, Peng Jiang, Shanan Chappell Moots, Masha Sosonkina, Hongyi Wu Jan 2025

T3-Ciders: Train-The-Trainer And Community Building To Increase Cyberinfrastructure Adoption In Cybersecurity Research And Education, Wirawan Purwanto, Mohan Yang, Peng Jiang, Shanan Chappell Moots, Masha Sosonkina, Hongyi Wu

University Administration Publications

T³-CIDERS is a train-the-trainer program to increase the adoption of advanced cyberinfrastructure (CI) and data skills into the fabric of research and education in cybersecurity and cyber-related disciplines. T³-CIDERS trains faculty, researchers, and students as “future trainers” (FTs) with hands-on technical and instructional skills to enable more people to effectively leverage CI in cybersecurity. The program includes a series of technical pre-training modules, a weeklong summer institute, ongoing learning engagements conducted over an academic year; it culminates with the FTs conducting locally tailored CI-infused training events at their respective home institutions. Ultimately, T³-CIDERS aims to build a “CI+cybersecurity” community of …


T3-Ciders: Fostering A Community Of Practice In Ci-And Data Enabled Cybersecurity Research Through A Train-The-Trainer Program, Wirawan Purwanto, Mohan Yang, Peng Jiang, Masha Sosonkina Jan 2025

T3-Ciders: Fostering A Community Of Practice In Ci-And Data Enabled Cybersecurity Research Through A Train-The-Trainer Program, Wirawan Purwanto, Mohan Yang, Peng Jiang, Masha Sosonkina

Electrical & Computer Engineering Faculty Publications

We present a training program named T³-CIDERS, the Train- The-Trainer approach to fostering cyberinfrastructure (CI)- and Data-Enabled Research in CyberSecurity. T³-CIDERS is a train-the-trainer program for advanced cyberinfrastructure (CI) skills that is designed to be synergistic with research, teaching, and learning activities in cybersecurity and cyber-related disciplines. The participants, termed 'future trainers' (FTs), are trained in effective instructional design and CI hands-on materials from DeapSECURE, developed in a previous CyberTraining program. T³-CIDERS aims to enhance cybersecurity research and education through broader adoption of advanced CI techniques such as artificial intelligence, big data, parallel programming, and platforms like high-performance computing (HPC) …


High-Fidelity Soh Prediction In Lithium-Ion Batteries Using Hybrid Ml Networks, Shafiyee Islam, Gon Namkoong Jan 2025

High-Fidelity Soh Prediction In Lithium-Ion Batteries Using Hybrid Ml Networks, Shafiyee Islam, Gon Namkoong

Electrical & Computer Engineering Faculty Publications

Accurate and efficient prediction of lithium-ion battery state of health (SOH) is critical for ensuring reliability in electric vehicles, grid storage, and aerospace systems. Traditional SOH estimation methods often struggle with nonlinear degradation behaviors and lack sensitivity to subtle electrochemical signals, limiting their real-world deployment. To address these challenges, this study examines hybrid deep learning models that integrate differential capacity (dQ/dV) analysis to enhance predictive accuracy. Four hybrid architectures - hybrid CNN-LSTM multihead, CNN extractor for LSTM, DNN-LSTM, and DNN Bi-LSTM - were developed and evaluated using the NASA randomized battery usage dataset, offering a realistic benchmark under diverse operational …


Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi Jan 2025

Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi

Theses and Dissertations

Electronic Health Records (EHRs) constitute a comprehensive and high-dimensional repository of clinical data, encompassing a wide array of patient-level information such as diagnoses, procedures, medications, laboratory results, and unstructured clinical narratives. These data hold immense potential for advancing predictive modeling in healthcare, including tasks such as disease progression modeling, hospital readmission prediction, and length of stay (LoS) estimation. However, the intrinsic complexity of EHR data—manifested in its heterogeneity, sparsity, and temporal dynamics—poses significant analytical challenges that limit the generalizability and interpretability of conventional machine learning models. Recent methodological advancements in deep learning and graph-based learning, particularly Graph Neural Networks (GNNs), …