Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

2025

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 481 - 504 of 504

Full-Text Articles in Data Science

High-Fidelity Soh Prediction In Lithium-Ion Batteries Using Hybrid Ml Networks, Shafiyee Islam, Gon Namkoong Jan 2025

High-Fidelity Soh Prediction In Lithium-Ion Batteries Using Hybrid Ml Networks, Shafiyee Islam, Gon Namkoong

Electrical & Computer Engineering Faculty Publications

Accurate and efficient prediction of lithium-ion battery state of health (SOH) is critical for ensuring reliability in electric vehicles, grid storage, and aerospace systems. Traditional SOH estimation methods often struggle with nonlinear degradation behaviors and lack sensitivity to subtle electrochemical signals, limiting their real-world deployment. To address these challenges, this study examines hybrid deep learning models that integrate differential capacity (dQ/dV) analysis to enhance predictive accuracy. Four hybrid architectures - hybrid CNN-LSTM multihead, CNN extractor for LSTM, DNN-LSTM, and DNN Bi-LSTM - were developed and evaluated using the NASA randomized battery usage dataset, offering a realistic benchmark under diverse operational …


Out Of Order And Causally Correct: Ready-Event Discovery Through Data-Dependence Analysis, Erik John Jensen, James Leathrum Jr., Christopher Lynch, Katherine Smith, Ross Gore Jan 2025

Out Of Order And Causally Correct: Ready-Event Discovery Through Data-Dependence Analysis, Erik John Jensen, James Leathrum Jr., Christopher Lynch, Katherine Smith, Ross Gore

Electrical & Computer Engineering Faculty Publications

Data-dependence analysis can identify causally-unordered events in a pending event set. The execution of these events is independent from all other scheduled events, making them ready for execution. These events can be executed out of order or in parallel. This approach may find and utilize more parallelism than spatial-decomposition parallelization methods, which are limited by the number of subdomains and by synchronization methods. This work provides formal definitions that use data-dependence analysis to find causally-unordered events and uses these definitions to measure parallelism in several discrete-event simulation models. A variant of the event-graph formalism is proposed, which assists with identifying …


Enhancing Channel Data Savings And Information Transfer Efficiency In Ultrasound Imaging, Sai Konda, Hicham Chaoui Jan 2025

Enhancing Channel Data Savings And Information Transfer Efficiency In Ultrasound Imaging, Sai Konda, Hicham Chaoui

Electrical & Computer Engineering Faculty Publications

Ultrasound is a popular imaging technique mainly due to its non-invasive nature. And so, it is being used in a variety of applications. Due to plane wave imaging technique in ultrasound, frame rate of ultrasound imaging has the potential for being very high. Due to which, many channel data frames are being generated within a few seconds. As a result, tasks such as storing data frames and transferring them from front end ultrasonic system to processing computers are presenting significant challenges. Our current research work minimized these issues. We proposed and implemented: (a) Data encoding technique - We combined every …


Computational Bridges: Enhancing Natural Language Processing Of Swahili., Joyce Murungi Jan 2025

Computational Bridges: Enhancing Natural Language Processing Of Swahili., Joyce Murungi

Harrisburg University Other Works

Swahili remains significantly underrepresented in natural language processing (NLP) despite being one of the most widely spoken languages in Africa. Computational Bridges: Enhancing Natural Language Processing of Swahili addresses this gap through computational linguistics, corpus creation, and large-scale analysis of Swahili syntax and lexical structure. Central to this study is GUMZO, a novel corpus developed from spontaneous conversational data collected from YouTube videos, television panel discussions, political speeches, religious discourse, and unscripted broadcasts. Unlike many existing datasets that rely on formal or translated text, GUMZO captures authentic language use and provides a stronger foundation for NLP research involving low-resource languages. …


A Bounded Custom Gpt Structures Operational Knowledge For Small-Industry Decision Support, Ikhwan Arief, Alizar Hasan, Nilda Tri Putri, Hafiz Rahmann Jan 2025

A Bounded Custom Gpt Structures Operational Knowledge For Small-Industry Decision Support, Ikhwan Arief, Alizar Hasan, Nilda Tri Putri, Hafiz Rahmann

Knowledge Engineering and Data Science

Small manufacturing and craft-based firms increasingly use Generative Artificial Intelligence (GenAI) through public chat interfaces, low-cost tools, and informal experimentation. However, these firms often make operational decisions with incomplete records, tacit owner knowledge, fragmented spreadsheets, and limited managerial capacity. Under such conditions, open-ended chatbots may generate fluent but unsafe recommendations by overlooking missing information, contradictory evidence, feasibility constraints, and implementation constraints. This study presents Asisten Cerdas Industri Kecil as a bounded Custom GPT artifact for operational diagnosis, priority selection, and short-horizon action planning in small industries. Using a Design Science Research approach, the study develops a documented artifact corpus comprising …


Optimal Data Splitting Methods, Sujay Mudalgi Jan 2025

Optimal Data Splitting Methods, Sujay Mudalgi

Theses and Dissertations

In predictive modeling, effective data splitting is crucial for creating statistically representative training and validation sets. The state-of-the-art data splitting methods are based on minimizing the energy distance between the split subsets. However, there are a number of limitations in the existing methods, which this dissertation aims to address. First, the existing methods were computationally inefficient. Thus, Chapter 2 proposes a method to scale up these approaches for big data. Here, we introduce scalable Twinning (s-Twinning), which significantly improves the execution speed of data splitting without sacrificing accuracy. Second, the existing methods did not consider the predictive relationship in the …


Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi Jan 2025

Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi

Theses and Dissertations

Electronic Health Records (EHRs) constitute a comprehensive and high-dimensional repository of clinical data, encompassing a wide array of patient-level information such as diagnoses, procedures, medications, laboratory results, and unstructured clinical narratives. These data hold immense potential for advancing predictive modeling in healthcare, including tasks such as disease progression modeling, hospital readmission prediction, and length of stay (LoS) estimation. However, the intrinsic complexity of EHR data—manifested in its heterogeneity, sparsity, and temporal dynamics—poses significant analytical challenges that limit the generalizability and interpretability of conventional machine learning models. Recent methodological advancements in deep learning and graph-based learning, particularly Graph Neural Networks (GNNs), …


Learning From Non-Stationary Data Streams, Gabriel Jonas Aguiar Jan 2025

Learning From Non-Stationary Data Streams, Gabriel Jonas Aguiar

Theses and Dissertations

The rapid growth of data from sources such as mobile applications, sensors, and network monitoring has increased the need for machine learning algorithms capable of handling non-stationary data streams. However, learning from such streams presents significant challenges due to their evolving nature and the presence of concept drift. One of the most complex issues is learning from imbalanced data streams, where shifting data distributions, combined with feature space drifts, complicate continuous adaptation. These challenges become even more pronounced in multi-class scenarios, which are common in real-world applications. Detecting concept drift in such contexts is particularly demanding, as it requires tracking …


Multi-Lingual And Cross-Domain Frontiers In Machine-Generated Content Detection, Gurunameh Singh Chhatwal Jan 2025

Multi-Lingual And Cross-Domain Frontiers In Machine-Generated Content Detection, Gurunameh Singh Chhatwal

Theses and Dissertations (Comprehensive)

The rapid advancement of generative artificial intelligence, particularly Large Language Models (LLMs) such as GPT-4 and their multilingual capabilities, has significantly blurred the distinction between human-authored and machine-generated content. This technological evolution introduces critical challenges concerning the detection and attribution of textual authenticity and authorship, exacerbating societal issues like misinformation proliferation and compromising academic and professional integrity. Traditional detection methodologies, predominantly monolingual and heuristic-based, have demonstrated inadequate generalizability and efficacy against the sophisticated, multilingual capabilities of contemporary generative models.

This thesis addresses two major problems arising from these advancements. Firstly, it introduces novel multilingual detection methodologies explicitly designed to differentiate …


Crime Modeling Using An Integrated Cnn–Lstm Architecture With Embedded Self-Excitation, Pawandeep Kaur Jan 2025

Crime Modeling Using An Integrated Cnn–Lstm Architecture With Embedded Self-Excitation, Pawandeep Kaur

Theses and Dissertations (Comprehensive)

It is often assumed that natural phenomena occur randomly over time. However, careful analysis reveals that these events typically form some series or sequences and exhibit distinctive temporal patterns. These patterns are not exclusive to nature. They also appear in human activities, often studied under the concept of bursty human dynamics. The statistical methods analyzing bursty human dynamics not only capture overall trends or seasonality but also explore how past events influence future ones. It makes the analysis more realistic and the results more closely aligned with reality. Bursty human dynamics can be studied at two levels: the individual level …


Advancing Multivariate Time Series Similarity Assessment: An Integrated Computational Approach, Franck B.N. Tonle, Henri E.Z. Tonnang, Milliam M.Z. Ndadji, Maurice Tchoupe Tchendji, Armand Nzeukou, Kennedy Senagi, Saliou Niassy Jan 2025

Advancing Multivariate Time Series Similarity Assessment: An Integrated Computational Approach, Franck B.N. Tonle, Henri E.Z. Tonnang, Milliam M.Z. Ndadji, Maurice Tchoupe Tchendji, Armand Nzeukou, Kennedy Senagi, Saliou Niassy

All Peer-Reviewed Publications

Data mining, particularly multivariate time series data analysis, is crucial in extracting insights from complex systems and supporting informed decision-making across diverse domains. However, assessing the similarity of multivariate time series data presents several challenges, including dealing with large datasets, addressing temporal misalignments, and necessitating efficient and comprehensive analytical frameworks. A novel integrated computational approach, Multivariate Time series Alignment and Similarity Assessment (MTASA) is proposed to address these challenges. MTASA is built upon a hybrid methodology designed to optimise time series alignment, complemented by a multiprocessing engine that enhances the utilisation of computational resources. This integrated approach comprises four key …


Privshap: A Finer-Granularity Network Linearization Method For Private Inference, Xiangrui Xu, Zhenzhen Wang, Rui Ning, Chunsheng Xiu, Hongyi Wu Jan 2025

Privshap: A Finer-Granularity Network Linearization Method For Private Inference, Xiangrui Xu, Zhenzhen Wang, Rui Ning, Chunsheng Xiu, Hongyi Wu

Computer Science Faculty Publications

Private inference applies cryptographic techniques like homomorphic encryption, garble circuit and secret sharing to keep both sides privacy in a client-server setting during inference. It is often hindered by the high communication overheads, especially at non-linear activation layers such as ReLU. Hence ReLU pruning has been widely recognized as an efficient way to accelerate private inference. Existing approaches to ReLU pruning typically rely on coarse hypothesis, which assume an inverse correlation between the importance of ReLU and linear layers or shallow activation layers have less importance for universal models, to assign the budgets according to the layer while preserving the …


Heterogeneous Clustering Of Multiomics Data For Breast Cancer Subgroup Classification And Detection, Joseph Pateras, Musaddiq Lodi, Pratip Rana, Preetam Ghosh Jan 2025

Heterogeneous Clustering Of Multiomics Data For Breast Cancer Subgroup Classification And Detection, Joseph Pateras, Musaddiq Lodi, Pratip Rana, Preetam Ghosh

Computer Science Faculty Publications

The rapid growth of diverse -omics datasets has made multiomics data integration crucial in cancer research. This study adapts the expectation–maximization routine for the joint latent variable modeling of multiomics patient profiles. By combining this approach with traditional biological feature selection methods, this study optimizes latent distribution, enabling efficient patient clustering from well-studied cancer types with reduced computational expense. The proposed optimization subroutines enhance survival analysis and improve runtime performance. This article presents a framework for distinguishing cancer subtypes and identifying potential biomarkers for breast cancer. Key insights into individual subtype expression and function were obtained through differentially expressed gene …


A Data-Driven Sliding-Window Pairwise Comparative Approach For The Estimation Of Transmission Fitness Of Sars-Cov-2 Variants And The Construction Of The Evolution Fitness Landscape, Md Jubair Pantho, Richard Annan, Landen Alexander Bauder, Sophia Huang, Letu Qingge, Hong Qin Jan 2025

A Data-Driven Sliding-Window Pairwise Comparative Approach For The Estimation Of Transmission Fitness Of Sars-Cov-2 Variants And The Construction Of The Evolution Fitness Landscape, Md Jubair Pantho, Richard Annan, Landen Alexander Bauder, Sophia Huang, Letu Qingge, Hong Qin

Computer Science Faculty Publications

Estimating the transmission fitness of SARS-CoV-2 variants and understanding their evolutionary fitness trends are important for epidemiological forecasting. Existing methods are often constrained by their parametric natures and do not satisfactorily align with the observations during COVID-19. Here, we introduce a sliding-window data-driven pairwise comparison method, the differential population growth rate (DPGR) that uses viral strains as internal controls to mitigate sampling biases. DPGR is applicable in time windows in which the logarithmic ratio of two variant subpopulations is approximately linear. We apply DPGR to genomic surveillance data and focus on variants of concern (VOCs) in multiple countries and regions. …


Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun Jan 2025

Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun

Computer Science Faculty Publications

Triple-negative breast cancer (TNBC) requires detailed cellular mapping given its aggressive nature, immense tumor heterogeneity and genetic diversity. We integrated 156,794 cells from six scRNA-seq datasets—including tumors, metastases, and cell lines—to build a TNBC scRNA cell atlas, focusing on batch effect mitigation while maintaining biological and molecular details. Preprocessing f ilters noise, normalizes data, and leverages PCA for integration readiness. We utilized scANVI, a semi-supervised tool, to align datasets, preserving TNBC’s complex tumor heterogeneity via marker annotations [1]. UMAPs demonstrate biological clustering in integrated data, contrasted with datasetdriven unintegrated patterns. Assessments verifying effective batch correction. This method aligns with NASA’s …


Ai For Nuclear Physics: The Exclaim Project, S. Liuti, D. Adams, M. Boër, G. W. Chern, M. Cuic, M. Engelhardt, G. R. Goldstein, B. Kriesten, Y. Li, H. W. Lin, M. Sievert, D. Sivers Jan 2025

Ai For Nuclear Physics: The Exclaim Project, S. Liuti, D. Adams, M. Boër, G. W. Chern, M. Cuic, M. Engelhardt, G. R. Goldstein, B. Kriesten, Y. Li, H. W. Lin, M. Sievert, D. Sivers

Computer Science Faculty Publications

An overview of the recent activity of the newly funded EXCLusives with AI and Machine learning (EXCLAIM) collaboration is presented. The main goal of the collaboration is to develop a framework to implement AI and machine learning techniques in problems emerging from the phenomenology of high energy exclusive scattering processes from nucleons and nuclei, maximizing the information that can be extracted from various sets of experimental data, while implementing theoretical constraints from lattice QCD. A specific perspective embraced by EXCLAIM is to use the methods of theoretical physics to understand the working of ML, beyond its standardized applications to physics …


Neural Topic Modeling Via Contextual And Graph Information Fusion, Jiyuan Liu, Jiaxing Yan, Chunjiang Zhu, Xingyu Liu, Qing Li, Yanghui Rao Jan 2025

Neural Topic Modeling Via Contextual And Graph Information Fusion, Jiyuan Liu, Jiaxing Yan, Chunjiang Zhu, Xingyu Liu, Qing Li, Yanghui Rao

Computer Science Faculty Publications

Topic modeling is a powerful unsupervised tool for knowledge discovery. However, existing work struggles with generating limited-quality topics that are uninformative and incoherent, which hindering interpretable insights from managing textual data. In this paper, we improve the original variational autoencoder framework by incorporating contextual and graph information to address the above issues. First, the encoder utilizes topic fusion techniques to combine contextual and bag-of-words information well, and meanwhile exploits the constraints of topic alignment and topic sharpening to generate informative topics. Second, we develop a simple word co-occurrence graph information fusion strategy that efficiently increases topic coherence. On three benchmark …


Decode The Workload: Training Deep Learning Models For Efficient Compute Cluster Representation, Ahmed Hossam Mohammed, Mark Jones, Diana Mcspadden, Malachi Schram, Bryan Hess, Kishansingh Rajput Jan 2025

Decode The Workload: Training Deep Learning Models For Efficient Compute Cluster Representation, Ahmed Hossam Mohammed, Mark Jones, Diana Mcspadden, Malachi Schram, Bryan Hess, Kishansingh Rajput

Computer Science Faculty Publications

In this study, we address the mounting challenge of monitoring high throughput computing clusters running computationally intensive jobs, which increasingly strains system administrators. We develop autoencoders that analyze traces of Linux kernel CPU metrics to capture salient system features by producing robust compressed embeddings for various downstream tasks. In addition, we employ graph neural networks to incorporate contextual information from surrounding CPUs and assess their performance. We also demonstrate the enhanced job differentiation achieved by increasing the sampling rate of these traces. Our models are evaluated based on their ability to generate meaningful latent representations, detect anomalies, and distinguish between …


S²Il: Structurally Stable Incremental Learning, S. Balasubramanian, P. Yedu Krishna, Talasu Sai Sriram, M. Sai Subramaniam, Manepalli Pranav Phanindra Sai, Ravi Mukkamala Jan 2025

S²Il: Structurally Stable Incremental Learning, S. Balasubramanian, P. Yedu Krishna, Talasu Sai Sriram, M. Sai Subramaniam, Manepalli Pranav Phanindra Sai, Ravi Mukkamala

Computer Science Faculty Publications

Feature Distillation (FD) strategies are proven to be effective in mitigating Catastrophic Forgetting (CF) seen in Class Incremental Learning (CIL). However, current FD approaches enforce strict alignment of feature magnitudes and directions across incremental steps, limiting the model’s ability to adapt to new knowledge. In this paper, we propose Structurally Stable Incremental Learning (S²IL), a FD method for CIL that mitigates forgetting by focusing on preserving the overall spatial patterns of features which promote flexible (plasticity) yet stable representations that preserve old knowledge (stability). We also demonstrate that our proposed method S²IL achieves strong incremental accuracy and outperforms other FD …


Effective Pii Extraction From Llms Through Augmented Few-Shot Learning, Shuai Cheng, Shu Meng, Haitao Xu, Haoran Zhang, Shuai Hao, Chuan Yue, Wenrui Ma, Meng Han, Fang Zhang, Zhao Li Jan 2025

Effective Pii Extraction From Llms Through Augmented Few-Shot Learning, Shuai Cheng, Shu Meng, Haitao Xu, Haoran Zhang, Shuai Hao, Chuan Yue, Wenrui Ma, Meng Han, Fang Zhang, Zhao Li

Computer Science Faculty Publications

Large Language Models (LLMs) exhibit strong natural language processing capabilities but also pose significant privacy risks, particularly regarding the leakage of Personally Identifiable Information (PII) embedded in their training data. Existing PII extraction methods suffer from the limitations of low success rates or impracticality for large-scale PII extraction. In this study, we propose a novel PII extraction approach based on enhanced few-shot learning techniques, which achieves efficient and cost-effective PII retrieval without relying on fine-tuning or jailbreaking. We evaluated our approach on both open-source and closed-source LLMs. The experimental results demonstrate that, for non-targeted PII extraction, the attack success rate …


Energy-Based Deep Incomplete Multi-View Clustering, Ziyu Wang, Yiming Du, Rui Ning, Lusi Li Jan 2025

Energy-Based Deep Incomplete Multi-View Clustering, Ziyu Wang, Yiming Du, Rui Ning, Lusi Li

Computer Science Faculty Publications

Incomplete multi-view clustering (IMVC) deals with real-world scenarios where certain views are partially missing, posing significant challenges to effective clustering. Most existing IMVC approaches face a trade-off: imputation-free methods suffer from information bias and imbalance, while full-imputation methods risk introducing and propagating noise. To overcome these limitations, we propose Energy-Based Deep Incomplete Multi-View Clustering (Energy-DIMC), a novel selective-imputation framework that leverages energy-based models (EBMs) to guide reliable imputations and robust clustering. EBMs assess data compatibility by assigning lower energy to more coherent structures, effectively modeling complex inter-view and inter-sample dependencies. Inspired by EBMs, Energy-DIMC integrates four key components: 1) a …


Icu-Length Of Stay Prediction On Electronic Health Records Using Graph Neural Networks And Homogeneous Similarity Graphs, Ahmad F. Al Musawi, Pratip Rana, Sibtanu Raha, Joshua Braunstein, William C. Sleeman Iv, Rishabh Kapoor, Preetam Ghosh Jan 2025

Icu-Length Of Stay Prediction On Electronic Health Records Using Graph Neural Networks And Homogeneous Similarity Graphs, Ahmad F. Al Musawi, Pratip Rana, Sibtanu Raha, Joshua Braunstein, William C. Sleeman Iv, Rishabh Kapoor, Preetam Ghosh

Computer Science Faculty Publications

Predicting the length of stay (LoS) is important for hospital administration, as it helps allocate proper resources, such as bed management and hospital staffing. Patients' Electronic Health Records (EHRs) contain highly relevant data for LoS prediction; however, their integration and effective use in predictive modeling for accurately estimating LoS remain challenging. To address this, we propose a homogeneous Graph Neural Network (GNN)-based framework for predicting LoS. This method employs a comprehensive data fusion strategy based on the hospital Visit-based Similarity Graph (VSG), which integrates diverse multi-modal clinical features into a coherent, homogeneous graph representation. Next, this VSG is fed into …


Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem Jan 2025

Hybrid Mixtures Of Factor Analyzers For High Dimensional Data, Kazeem Abiodun Kareem

Dissertations, Master's Theses and Master's Reports

Factor analysis is a powerful tool for modeling latent structures in high-dimensional data, traditional approaches assume a single global structure, limiting their ability to capture heterogeneity. The Mixture of Factor Analyzers (MFA) extends classical factor analysis by modeling data as a mixture of Gaussian-distributed local subspaces, effectively uncovering cluster-specific latent structures. However, MFA relies on Gaussian mixtures, making it sensitive to outliers and ill-suited for heavy-tailed data. The Mixture of $t$-Factor Analyzers (M$t$FA) addresses these limitations by incorporating multivariate $t$-distributions, improving robustness. Despite their advantages, both MFA and M$t$FA face significant computational challenges in high-dimensional settings, particularly due to costly …


Generalizing Medical Image Segmentation Task With Efficient Deep Learning Models, Abel A. Reyes-Angulo Jan 2025

Generalizing Medical Image Segmentation Task With Efficient Deep Learning Models, Abel A. Reyes-Angulo

Dissertations, Master's Theses and Master's Reports

Medical Image Segmentation is a critical task in the field of medical imaging, playing a crucial role in diagnostics, treatment planning, and disease monitoring. The emergence of Deep Learning (DL) has ushered in a new era in Artificial Intelligence (AI), propelling remarkable advancements in key domains like language translation, object recognition, and recommendation systems. This evolution has been accompanied by continuous enhancements in computational efficiency and improvements in predictive accuracy. The introduction of sophisticated algorithms, such as convolutional neural networks (CNNs) and transformers, exemplifies these advancements. DL algorithms have demonstrated exceptional efficacy in medical image segmentation tasks, showcasing the potential …