Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (130)
- Artificial Intelligence and Robotics (75)
- Medicine and Health Sciences (46)
- Engineering (41)
- Life Sciences (38)
-
- Statistics and Probability (36)
- Bioinformatics (22)
- Social and Behavioral Sciences (22)
- Biomedical Informatics (17)
- Applied Statistics (15)
- Electrical and Computer Engineering (14)
- Applied Mathematics (13)
- Databases and Information Systems (13)
- Medical Sciences (13)
- Other Computer Sciences (13)
- Medical Specialties (12)
- Business (11)
- Numerical Analysis and Scientific Computing (11)
- Theory and Algorithms (11)
- Environmental Sciences (9)
- Mathematics (9)
- Statistical Models (9)
- Computer Engineering (8)
- Education (8)
- Biostatistics (7)
- Earth Sciences (7)
- Information Security (7)
- Operations Research, Systems Engineering and Industrial Engineering (7)
- Institution
-
- Old Dominion University (18)
- The Texas Medical Center Library (18)
- CCT College Dublin (14)
- Southern Methodist University (12)
- Air Force Institute of Technology (10)
-
- City University of New York (CUNY) (10)
- Virginia Commonwealth University (10)
- New Jersey Institute of Technology (9)
- Clemson University (7)
- West Virginia University (7)
- Chapman University (6)
- Minnesota State University, Mankato (6)
- Dartmouth College (5)
- Embry-Riddle Aeronautical University (5)
- University of Kentucky (5)
- University of Louisville (5)
- California Polytechnic State University, San Luis Obispo (4)
- Kennesaw State University (4)
- Technological University Dublin (4)
- University of Nebraska - Lincoln (4)
- University of New Mexico (4)
- Louisiana State University (3)
- Marshall University (3)
- Purdue University (3)
- Singapore Management University (3)
- University of Arkansas, Fayetteville (3)
- University of Montana (3)
- University of Texas at Arlington (3)
- Western Michigan University (3)
- DePaul University (2)
- Publication Year
- Publication
-
- Theses and Dissertations (19)
- Faculty, Staff and Student Publications (15)
- ICT (14)
- Dissertations (11)
- SMU Data Science Review (11)
-
- Electronic Theses and Dissertations (7)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (7)
- All Dissertations (6)
- All Graduate Theses, Dissertations, and Other Capstone Projects (6)
- Dissertations, Theses, and Capstone Projects (6)
- Faculty Publications (6)
- Electrical & Computer Engineering Faculty Publications (5)
- Computer Science Faculty Publications (4)
- Master's Theses (4)
- Computational and Data Sciences (PhD) Dissertations (3)
- Computer Science Senior Theses (3)
- Dissertations and Theses (3)
- Dissertations and Theses (Open Access) (3)
- Graduate Student Theses, Dissertations, & Professional Papers (3)
- LSU Doctoral Dissertations (3)
- Theses, Dissertations and Capstones (3)
- Articles (2)
- College of Computing and Digital Media Dissertations (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Computer Science Faculty Scholarship (2)
- Conference papers (2)
- Data Science Faculty Publications (2)
- Data Science Undergraduate Honors Theses (2)
- Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023– (2)
- Dissertations, Master's Theses and Master's Reports (2)
- Publication Type
- File Type
Articles 31 - 60 of 241
Full-Text Articles in Data Science
Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun
Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun
SMU Data Science Review
Enhancing animal shelter operations through machine learning involves employing a variety of advanced techniques aimed at increasing efficiency, promoting animal welfare, and optimizing resource allocation. This paper explores predictive analytics for adoption rates using regression models to estimate the likelihood of adoption based on historical data, encompassing variables such as breed, health status, and previous adoption trends. Additionally, classification algorithms are utilized to categorize animals by adoption probability, facilitating better resources and marketing prioritization. Clustering algorithms are employed to group animals according to behavior patterns and/or physical health, enabling tailored medical care and enrichment activities that improve their mental and …
Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul
Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul
School of Public Health Faculty Publications
Diabetes is a growing global health concern, affecting millions and leading to severe complications if not properly managed. The primary challenge in diabetes management is maintaining blood glucose levels (BGLs) within a safe range to prevent complications such as renal failure, cardiovascular disease, and neuropathy. Traditional methods, such as finger-prick testing, often result in low patient adherence due to discomfort, invasiveness, and inconvenience. Consequently, there is an increasing need for non-invasive techniques that provide accurate BGL measurements. Photoplethysmography (PPG), a photosensitive method that detects blood volume variations, has shown promise for non-invasive glucose monitoring. Deep neural networks (DNNs) applied to …
Leveraging Large Language Models For Knowledge-Free Weak Supervision In Clinical Natural Language Processing, Enshuo Hsu, Kirk Roberts
Leveraging Large Language Models For Knowledge-Free Weak Supervision In Clinical Natural Language Processing, Enshuo Hsu, Kirk Roberts
Faculty, Staff and Student Publications
The performance of deep learning-based natural language processing systems is based on large amounts of labeled training data which, in the clinical domain, are not easily available or affordable. Weak supervision and in-context learning offer partial solutions to this issue, particularly using large language models (LLMs), but their performance still trails traditional supervised methods with moderate amounts of gold-standard data. In particular, inferencing with LLMs is computationally heavy. We propose an approach leveraging fine-tuning LLMs and weak supervision with virtually no domain knowledge that still achieves consistently dominant performance. Using a prompt-based approach, the LLM is used to generate weakly-labeled …
Machine Learning With Flight Data Recorder Data For Flight Fuel Consumption Predictions, Adam C. Levandowski
Machine Learning With Flight Data Recorder Data For Flight Fuel Consumption Predictions, Adam C. Levandowski
Theses and Dissertations
This study applies advanced Machine Learning (ML) to Flight Data Recorder (FDR) data for fuel consumption predictions. It explores feature engineering, model selection, and Hyper-Parameter Optimization (HPO) across all flight phases. Baseline models like Ordinary Least Squares (OLS) regression, Multi- Layer Perceptrons (MLPs), and decision trees are compared to Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs) with Gated Recurrent Unit (GRU) layers, and XGBoost. Results analyze segmentation strategies, tailored features, and model performance. A counterfactual analysis compares ML models to operational fuel predictions, demonstrating their deployment potential. Findings establish a foundation for future ML-driven advancements in aviation fuel optimization.
Toward Quantifying Interpolation Uncertainty In Set-Line Spacing Hydrographic Surveys, Elias Adediran, Christos Kastrisios, Kim Lowell, Glen Rice, Qi Zhang
Toward Quantifying Interpolation Uncertainty In Set-Line Spacing Hydrographic Surveys, Elias Adediran, Christos Kastrisios, Kim Lowell, Glen Rice, Qi Zhang
Faculty Publications
The oceans remain one of Earth’s last great unknowns, with about 74% still unmapped to modern standards. Consequently, interpolation is employed to create seamless digital bathymetric models (DBMs) from incomplete hydrographic datasets, but this introduces unquantified depth uncertainties. This study aims to estimate and characterize uncertainties arising from set-line spacing hydrographic surveys, which are important for nautical charting, navigational safety, and many other applications. By sampling at different line spacings four complete coverage testbeds that vary in slope and roughness, the study interpolates across entire testbed areas using Spline, Inverse Distance Weighting, and Linear interpolation. The resulting interpolation uncertainties are …
Classification Of Variable Stars Using Convolutional Neural Network, Abhina Premachandran Bindu
Classification Of Variable Stars Using Convolutional Neural Network, Abhina Premachandran Bindu
Dissertations and Theses
This research focuses on developing Convolutional Neural Networks (CNNs), for the process of classifying and identifying variable stars through the analysis of unprocessed light curves from Transiting Exoplanet Survey Satellite (TESS). As astronomical data is becoming increasingly complex, and as advanced missions deploy sophisticated instruments for data collection, both the quality and quantity of the data are improving at a rapid pace. This has created an urgent need to automate the analysis process using efficient and effective methods, such as those based on machine learning. While previous research has explored machine learning approaches, there has been limited focus on implementing …
Comparative Analysis Of Machine Learning Models For Glioblastoma Survival., Muna Awel
Comparative Analysis Of Machine Learning Models For Glioblastoma Survival., Muna Awel
All Graduate Theses, Dissertations, and Other Capstone Projects
Glioblastoma multiforme (GBM) remains one of the most lethal brain tumors, necessitating improved survival prediction models that integrate clinical and molecular data. This study develops a comprehensive machine learning pipeline leveraging TCGA-derived multi-omics datasets to predict binary survival outcomes. The framework integrates four classifiers Logistic Regression, Random Forest, XGBoost, and Support Vector Machine (SVM) and includes rigorous preprocessing with MCAR testing, KNN imputation, feature scaling, and hyperparameter optimization via GridSearchCV. SMOTE was applied to mitigate class imbalance and enhance model robustness for minority survival classes. Comparative performance analyses revealed Random Forest and XGBoost as top performers, achieving the highest recall …
Deep Learning-Based Ensemble Two-Step Classification Of Medical Images Using Cnn Architectures And Ensemble Methods, Noreliz Alorico
Deep Learning-Based Ensemble Two-Step Classification Of Medical Images Using Cnn Architectures And Ensemble Methods, Noreliz Alorico
Master's Theses or Doctor of Nursing Practice
Breast cancer remains one of the most common cancers amongst women globally. Early detection is crucial for improving survival rates. While mammography is widely used and an effective imaging technique, it can sometimes yield false positive or false negatives. Mammogram interpretation is highly operator-dependent, introducing variability and the potential for diagnostic errors. Additionally, mammographic images have limitations, such as low contrast in breast tissue and overlapping structures that can obscure lesions or mimic abnormalities. These limitations can lead to unnecessary biopsies or delayed diagnosis. These challenges highlight the needs for advanced and data driven diagnostic tools to support and enhance …
Tour Demand Forecasting In Ireland: Development And Evaluation Of Classical, Deep Learning, And Hybrid Models, Ruben Elias Charleston Montfort
Tour Demand Forecasting In Ireland: Development And Evaluation Of Classical, Deep Learning, And Hybrid Models, Ruben Elias Charleston Montfort
ICT
Tourism plays a significant role in global economies by supporting employment, infrastructure, and national development. As international travel continues to grow, accurate tourism demand forecasting has become increasingly important for effective planning and decision-making. In Ireland, tourism is a key economic sector attracting millions of visitors annually. For tour operators such as Irish Day Tours, reliable demand forecasting is essential for optimizing logistics, resource allocation, marketing strategies, and customer satisfaction. Advances in machine learning and deep learning techniques offer new opportunities to improve forecasting accuracy and support data-driven decision-making within the tourism industry.
European Air Pollution And The Proposed Timelines Of Implementing The World Health Organization 2021 Air Quality Guidelines Ca3., Lukia Hartin
ICT
This research examines Ireland’s air pollution trends and evaluates whether current reductions in PM2.5, PM10, and NO₂ are sufficient to meet the WHO 2021 Air Quality Guidelines by 2040. Using four years of EPA-validated pollutant data (2020–2023), alongside Building Energy Rating (BER) and national transport datasets, the study applies the CRISP-DM methodology to guide analysis, preprocessing, modelling, and evaluation. Extensive data cleaning and alignment were required due to inconsistent station coverage, varying formats, and missing values. Forecasting models—including Random Forest, Gradient Boosting, SVR, and linear regression—were assessed using MSE, R², and trend significance to project pollutant levels across different Irish …
Customer Service Support. Utilizing Machine Learning To Classify, Prioritize And Summarize Issues., Hoai Nhan Nguyen
Customer Service Support. Utilizing Machine Learning To Classify, Prioritize And Summarize Issues., Hoai Nhan Nguyen
ICT
This capstone project investigates the application of machine learning and natural language processing (NLP) to enhance customer support operations through automated ticket classification, prioritization, and summarization. Using the multilingual Customer Support Emails dataset from Kaggle, the project follows the CRISP-DM methodology, performing extensive data cleaning, preprocessing, feature engineering, and class balancing. Five machine learning models—Decision Tree, KNN, LinearSVC, Naive Bayes, and Random Forest—were evaluated using hyperparameter tuning, cross-validation, confusion matrix analysis, and learning curves. LinearSVC demonstrated the strongest performance for both queue and priority classification, achieving accuracies of 89.8% and 81.2% respectively, with consistent generalization across folds. For summarization, extractive …
Predicting Early Customer Inactivity In The Banking Sector Using Machine Learning: A Churn Prevention, Ivana Mc Fadden
Predicting Early Customer Inactivity In The Banking Sector Using Machine Learning: A Churn Prevention, Ivana Mc Fadden
ICT
Customer churn, when customers stop using a company’s services, is a challenge for the banking sector (Singh et al., 2023). High churn rates often signal poor customer experiences, resulting in revenue losses and increased costs to obtain new clients. Goyal and Srivastava (2015) stress that fostering loyalty through exceptional service and understanding customer needs is important for long-term retention.
This project aims to predict early customer inactivity, an indicator of churn, by using machine learning. Early identification of at-risk customers will allow banks to apply targeted interventions, reduce acquisition costs, and improve customer satisfaction (Singh et al., 2023). By analysing …
The Actuarial Applications Of Machine Learning And Big Data In The Life Assurance Industry: Managing Customer Retention And Customer Outcomes By The Application Of Data Science., Brian Cunningham
ICT
Lapses are an issue in the insurance industry in general. They affect a company’s profitability, cash flows and solvency. High levels of lapses can cause reputational damage that could provoke a cycle of even more lapses. It is therefore incumbent on a company to do its utmost to retain the business it has written for the term it was written for.
If a company could predict which of its policies were about to lapse, it could proactively attempt to prevent them by contacting the policyholder and engaging in a discussion to ascertain the likelihood of their choosing to leave. In …
Premier League Results Predictions, Laura Consuegra
Premier League Results Predictions, Laura Consuegra
ICT
This project applies machine learning techniques to predict outcomes in the English Premier League, one of the most prestigious and widely followed football competitions worldwide. By analysing historical and real-time match data, including performance metrics such as goals scored, shots on target, and cards received, predictive models are developed to forecast match results with higher accuracy. The study evaluates the effectiveness of these models and explores the influence of key statistical features on team performance. The findings aim to provide strategic insights for fans, bookmakers, and coaching staff, supporting performance evaluation, tactical decision-making, and a deeper understanding of the factors …
Predicting Repeat Purchases In E-Commerce Using Interpretable Machine Learning, Zahid Bhatti
Predicting Repeat Purchases In E-Commerce Using Interpretable Machine Learning, Zahid Bhatti
ICT
This project investigates the prediction of repeat purchase behaviour in e-commerce using machine learning, with a focus on balancing predictive accuracy and interpretability. Large volumes of transactional and behavioural data are analysed to identify customer-level features that drive loyalty and repeat purchases. Various supervised learning models, including Random Forests and Logistic Regression, are evaluated for predictive performance, while SHAP (SHapley Additive Explanations) is employed to provide both global and local interpretability. The study aims to generate actionable insights for customer relationship management and marketing strategy, demonstrating how advanced predictive models can support informed business decisions without sacrificing transparency.
Predicting Customer Churn And Enhancing Retention Strategies Through Machine Learning., Swan Saung Lwin
Predicting Customer Churn And Enhancing Retention Strategies Through Machine Learning., Swan Saung Lwin
ICT
Customer Churn is a critical challenge faced by businesses across industries, especially in the digital market. Many companies struggle to predict customer churn accurately and have difficulties in carrying out effective retention strategies. Key challenges include ineffective traditional methods, lack of insights into impact of different services, generalized retention strategies, and the need to have cost-effective retention strategies. This project aims to predict customer churn using machine learning and identify the impact of key services offered by a telecommunication company.
Fake News Detection Using Machine Learning Models., Anne Higgins
Fake News Detection Using Machine Learning Models., Anne Higgins
ICT
This project investigates the use of machine learning to detect fake news, addressing the societal and political challenges posed by the widespread dissemination of false and misleading information. Using automated classification techniques, the project analyses news content to predict the likelihood of intentional deception. The methodology follows the CRISP-DM framework, encompassing data preparation, model development, and evaluation. By leveraging machine learning, the study aims to support organisations, governments, and digital platforms in mitigating misinformation, while also considering ethical, interpretability, and strategic implications. The findings provide actionable insights for enhancing content moderation and reducing the influence of disinformation in digital media …
Data Analysis For Maintenance Reliability., Romulo Menezes Santos
Data Analysis For Maintenance Reliability., Romulo Menezes Santos
ICT
This project explores the application of predictive maintenance (PdM) in modern manufacturing, highlighting its strategic role in reducing downtime, optimising resources, and improving operational efficiency. Traditional reactive and preventive maintenance approaches are insufficient for Industry 4.0 environments, where machine failures can cause significant financial and operational losses. By leveraging sensor data, historical performance records, and machine learning techniques, PdM enables early detection of potential equipment failures, allowing maintenance teams to act proactively. The approach not only enhances operational reliability and safety but also supports strategic decision-making, cost control, and competitive advantage in globalised manufacturing contexts.
Credit Card Fraud Detection., Sonia Ndonga
Credit Card Fraud Detection., Sonia Ndonga
ICT
This capstone project applies machine learning to detect credit card fraud, addressing a critical financial threat to banks and payment providers. Using an anonymised dataset of 284,807 transactions, which is highly imbalanced with only 0.172% fraudulent cases, three models—Logistic Regression, Random Forest, and Gradient Boosting—were developed and evaluated. The pipeline incorporates data preprocessing, feature engineering, hyperparameter tuning, cross-validation, and interpretability analysis using SHAP values, SHAPASH, and permutation importance. Random Forest achieved the highest performance with an ROC AUC of 0.97 and Average Precision of 0.66. The study also considers fairness, threshold optimisation, and practical deployment strategies, providing a robust automated …
Using Machine Learning To Predict Credit Card Fraud, Pedro Henrique Das Chagas Morais
Using Machine Learning To Predict Credit Card Fraud, Pedro Henrique Das Chagas Morais
ICT
Credit card fraud poses a significant challenge to financial institutions, leading to substantial financial losses and declining customer trust. This project develops and evaluates machine learning models to detect fraudulent credit card transactions using a large, realistic synthetic dataset. Following data preprocessing, exploratory analysis, and class-balancing using SMOTE, four supervised models—Logistic Regression, Decision Tree, Random Forest, and XGBoost—were trained and compared. Performance was assessed using metrics suited to imbalanced classification, including AUC, Recall, Precision, F1-score, and Average Precision. Results show that XGBoost, particularly after hyperparameter optimisation, delivered the strongest performance (AUC 0.99, Recall 0.83, AP 0.70), outperforming other models and …
Credit Card Default Prediction Using Machine Learning., Yassine Zohair
Credit Card Default Prediction Using Machine Learning., Yassine Zohair
ICT
This project investigates the use of machine learning to predict credit card payment defaults, aiming to help banks and credit card companies mitigate financial losses. Using historical customer data, including demographics, income, education, and previous payment behaviour, three machine learning algorithms were implemented to forecast the likelihood of default. Techniques such as cross-validation, hyperparameter tuning, and SHAPASH were applied to improve model performance and interpretability. Accurate prediction of potential defaulters enables financial institutions to take proactive measures, such as adjusting credit limits or providing targeted financial guidance, thereby enhancing risk management and customer retention.
Exploring The Determinants Of Life Expectancy At Birth: Predicting And Forecasting Global Health Trends Using Statistical, Machine Learning, And Deep Learning Models., Emma Rath
ICT
Accurate life expectancy forecasting is essential for health policy planning, yet research comparing statistical, machine learning, and deep learning approaches under real-world constraints remains limited. This study evaluates ARIMA/ARIMAX, tree-based, and neural network models using Irish and global datasets, considering small samples, missing data, and COVID-19 shocks. ARIMAX with lagged socioeconomic variables outperformed LSTM and other ML/DL methods. Income-based stratification improved predictive accuracy and interpretability, with SHAP analysis highlighting GDP per capita for developed countries and school enrolment and trade indicators for developing contexts. Results provide practical guidance for policymakers and establish limits for model complexity under constrained health data.
Evaluating Aspect-Based Sentiment Analysis In Healthcare Drug Reviews Across Machine Learning, Deep Neural Networks, And Transformer Models, Eun Soo Park
All Graduate Theses, Dissertations, and Other Capstone Projects
Sentiment analysis has become a critical area of research in Natural Language Processing (NLP), enabling insights from unstructured text. Within this field, Aspect-Based Sentiment Analysis (ABSA) plays a practical role in domains such as healthcare, where patients drug reviews often contain diverse opinions across multiple aspects, including overall comments, perceived benefits, and side effects. However, aspect-level classification remains challenging due to class imbalance, subtle sentiment expression, and the limitations of traditional models. This research investigates the performance of three modeling paradigms: traditional machine learning (SVM, SVC, and XGBoost), deep learning (CNN-BiLSTM), and transformer-based approaches (DistilBERT sentence-pair classification). Using the UCI …
A Bayesian Deep Segmentation Framework For Glioblastoma Tumor Segmentation Using Follow-Up Mris, Tanjida Kabir, Kang-Lin Hsieh, Luis Nunez, Yu-Chun Hsu, Juan C Rodriguez Quintero, Octavio Arevalo, Kangyi Zhao, Jay-Jiguang Zhu, Roy F Riascos, Mahboubeh Madadi, Xiaoqian Jiang, Shayan Shams
A Bayesian Deep Segmentation Framework For Glioblastoma Tumor Segmentation Using Follow-Up Mris, Tanjida Kabir, Kang-Lin Hsieh, Luis Nunez, Yu-Chun Hsu, Juan C Rodriguez Quintero, Octavio Arevalo, Kangyi Zhao, Jay-Jiguang Zhu, Roy F Riascos, Mahboubeh Madadi, Xiaoqian Jiang, Shayan Shams
Faculty, Staff and Student Publications
Background: Glioblastoma (GBM) is the most common malignant brain tumor with an abysmal prognosis. Since complete tumor cell removal is impossible due to the infiltrative nature of GBM, accurate measurement is paramount for GBM assessment. Preoperative magnetic resonance images (MRIs) are crucial for initial diagnosis and surgical planning, while follow-up MRIs are vital for evaluating treatment response. The structural changes in the brain caused by surgical and therapeutic measures create significant differences between preoperative and follow-up MRIs. In clinical research, advanced deep learning models trained on preoperative MRIs are often applied to assess follow-up scans, but their effectiveness in this …
Fares On Fairness: Using A Total Error Framework To Examine The Role Of Measurement And Representation In Training Data On Model Fairness And Bias, Patrick Oliver Schenk, Christoph Kern, Trent D. Buskirk
Fares On Fairness: Using A Total Error Framework To Examine The Role Of Measurement And Representation In Training Data On Model Fairness And Bias, Patrick Oliver Schenk, Christoph Kern, Trent D. Buskirk
Data Science Faculty Publications
Data-driven decisions, often based on predictions from machine learning (ML) models are becoming ubiquitous. For these decisions to be just, the underlying ML models must be fair, i.e., work equally well for all parts of the population such as groups defined by gender or age. What are the logical next steps if, however, a trained model is accurate but not fair? How can we guide the whole data pipeline such that we avoid training unfair models based on inadequate data, recognizing possible sources of unfairness early on? How can the concepts of data-based sources of unfairness that exist in the …
Training Set Augmentation And Harmonization Enables Radiomic Models To Detect Early Onset Of Lung Cancer, Claire Huchthausen, Menglin Shi, Gabriel L.A. Sousa, James Larner, Einsley Janowski, Jonathan Colen, Krishni Wijesooriya
Training Set Augmentation And Harmonization Enables Radiomic Models To Detect Early Onset Of Lung Cancer, Claire Huchthausen, Menglin Shi, Gabriel L.A. Sousa, James Larner, Einsley Janowski, Jonathan Colen, Krishni Wijesooriya
Data Science Faculty Publications
Radiomics-based machine learning models have the potential to detect lung cancer at inception from CT scans and transform patient outcomes. Low malignancy rates in early-development pulmonary nodules (PNs) and variable image acquisition hinder development of clinically applicable radiomics-based early detection models. To address these challenges, we augmented training using later-development PNs and harmonized for acquisition effects. We first trained machine learning models to predict PN malignancy using radiomic features from scans of early-development benign and malignant PNs (n = 187) harmonized using ComBat. Observing near-chance performance, we augmented training with later-development benign and malignant PNs (n = 225). We evaluated …
T3-Ciders: Train-The-Trainer And Community Building To Increase Cyberinfrastructure Adoption In Cybersecurity Research And Education, Wirawan Purwanto, Mohan Yang, Peng Jiang, Shanan Chappell Moots, Masha Sosonkina, Hongyi Wu
T3-Ciders: Train-The-Trainer And Community Building To Increase Cyberinfrastructure Adoption In Cybersecurity Research And Education, Wirawan Purwanto, Mohan Yang, Peng Jiang, Shanan Chappell Moots, Masha Sosonkina, Hongyi Wu
University Administration Publications
T³-CIDERS is a train-the-trainer program to increase the adoption of advanced cyberinfrastructure (CI) and data skills into the fabric of research and education in cybersecurity and cyber-related disciplines. T³-CIDERS trains faculty, researchers, and students as “future trainers” (FTs) with hands-on technical and instructional skills to enable more people to effectively leverage CI in cybersecurity. The program includes a series of technical pre-training modules, a weeklong summer institute, ongoing learning engagements conducted over an academic year; it culminates with the FTs conducting locally tailored CI-infused training events at their respective home institutions. Ultimately, T³-CIDERS aims to build a “CI+cybersecurity” community of …
T3-Ciders: Fostering A Community Of Practice In Ci-And Data Enabled Cybersecurity Research Through A Train-The-Trainer Program, Wirawan Purwanto, Mohan Yang, Peng Jiang, Masha Sosonkina
T3-Ciders: Fostering A Community Of Practice In Ci-And Data Enabled Cybersecurity Research Through A Train-The-Trainer Program, Wirawan Purwanto, Mohan Yang, Peng Jiang, Masha Sosonkina
Electrical & Computer Engineering Faculty Publications
We present a training program named T³-CIDERS, the Train- The-Trainer approach to fostering cyberinfrastructure (CI)- and Data-Enabled Research in CyberSecurity. T³-CIDERS is a train-the-trainer program for advanced cyberinfrastructure (CI) skills that is designed to be synergistic with research, teaching, and learning activities in cybersecurity and cyber-related disciplines. The participants, termed 'future trainers' (FTs), are trained in effective instructional design and CI hands-on materials from DeapSECURE, developed in a previous CyberTraining program. T³-CIDERS aims to enhance cybersecurity research and education through broader adoption of advanced CI techniques such as artificial intelligence, big data, parallel programming, and platforms like high-performance computing (HPC) …
High-Fidelity Soh Prediction In Lithium-Ion Batteries Using Hybrid Ml Networks, Shafiyee Islam, Gon Namkoong
High-Fidelity Soh Prediction In Lithium-Ion Batteries Using Hybrid Ml Networks, Shafiyee Islam, Gon Namkoong
Electrical & Computer Engineering Faculty Publications
Accurate and efficient prediction of lithium-ion battery state of health (SOH) is critical for ensuring reliability in electric vehicles, grid storage, and aerospace systems. Traditional SOH estimation methods often struggle with nonlinear degradation behaviors and lack sensitivity to subtle electrochemical signals, limiting their real-world deployment. To address these challenges, this study examines hybrid deep learning models that integrate differential capacity (dQ/dV) analysis to enhance predictive accuracy. Four hybrid architectures - hybrid CNN-LSTM multihead, CNN extractor for LSTM, DNN-LSTM, and DNN Bi-LSTM - were developed and evaluated using the NASA randomized battery usage dataset, offering a realistic benchmark under diverse operational …
Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi
Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi
Theses and Dissertations
Electronic Health Records (EHRs) constitute a comprehensive and high-dimensional repository of clinical data, encompassing a wide array of patient-level information such as diagnoses, procedures, medications, laboratory results, and unstructured clinical narratives. These data hold immense potential for advancing predictive modeling in healthcare, including tasks such as disease progression modeling, hospital readmission prediction, and length of stay (LoS) estimation. However, the intrinsic complexity of EHR data—manifested in its heterogeneity, sparsity, and temporal dynamics—poses significant analytical challenges that limit the generalizability and interpretability of conventional machine learning models. Recent methodological advancements in deep learning and graph-based learning, particularly Graph Neural Networks (GNNs), …