Computational Bridges: Enhancing Natural Language Processing Of Swahili.,
2025
Harrisburg University of Science and Technology
Computational Bridges: Enhancing Natural Language Processing Of Swahili., Joyce Murungi
Harrisburg University Other Works
Swahili remains significantly underrepresented in natural language processing (NLP) despite being one of the most widely spoken languages in Africa. Computational Bridges: Enhancing Natural Language Processing of Swahili addresses this gap through computational linguistics, corpus creation, and large-scale analysis of Swahili syntax and lexical structure. Central to this study is GUMZO, a novel corpus developed from spontaneous conversational data collected from YouTube videos, television panel discussions, political speeches, religious discourse, and unscripted broadcasts. Unlike many existing datasets that rely on formal or translated text, GUMZO captures authentic language use and provides a stronger foundation for NLP research involving low-resource languages. …
A Bounded Custom Gpt Structures Operational Knowledge For Small-Industry Decision Support,
2025
Industrial Engineering Department, Faculty of Engineering, Universitas Andalas, Padang, West Sumatra, Indonesia
A Bounded Custom Gpt Structures Operational Knowledge For Small-Industry Decision Support, Ikhwan Arief, Alizar Hasan, Nilda Tri Putri, Hafiz Rahmann
Knowledge Engineering and Data Science
Small manufacturing and craft-based firms increasingly use Generative Artificial Intelligence (GenAI) through public chat interfaces, low-cost tools, and informal experimentation. However, these firms often make operational decisions with incomplete records, tacit owner knowledge, fragmented spreadsheets, and limited managerial capacity. Under such conditions, open-ended chatbots may generate fluent but unsafe recommendations by overlooking missing information, contradictory evidence, feasibility constraints, and implementation constraints. This study presents Asisten Cerdas Industri Kecil as a bounded Custom GPT artifact for operational diagnosis, priority selection, and short-horizon action planning in small industries. Using a Design Science Research approach, the study develops a documented artifact corpus comprising …
Optimal Data Splitting Methods,
2025
Virginia Commonwealth University
Optimal Data Splitting Methods, Sujay Mudalgi
Theses and Dissertations
In predictive modeling, effective data splitting is crucial for creating statistically representative training and validation sets. The state-of-the-art data splitting methods are based on minimizing the energy distance between the split subsets. However, there are a number of limitations in the existing methods, which this dissertation aims to address. First, the existing methods were computationally inefficient. Thus, Chapter 2 proposes a method to scale up these approaches for big data. Here, we introduce scalable Twinning (s-Twinning), which significantly improves the execution speed of data splitting without sacrificing accuracy. Second, the existing methods did not consider the predictive relationship in the …
Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses,
2025
University of Thi Qar
Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi
Theses and Dissertations
Electronic Health Records (EHRs) constitute a comprehensive and high-dimensional repository of clinical data, encompassing a wide array of patient-level information such as diagnoses, procedures, medications, laboratory results, and unstructured clinical narratives. These data hold immense potential for advancing predictive modeling in healthcare, including tasks such as disease progression modeling, hospital readmission prediction, and length of stay (LoS) estimation. However, the intrinsic complexity of EHR data—manifested in its heterogeneity, sparsity, and temporal dynamics—poses significant analytical challenges that limit the generalizability and interpretability of conventional machine learning models. Recent methodological advancements in deep learning and graph-based learning, particularly Graph Neural Networks (GNNs), …
Learning From Non-Stationary Data Streams,
2025
Virginia Commonwealth University
Learning From Non-Stationary Data Streams, Gabriel Jonas Aguiar
Theses and Dissertations
The rapid growth of data from sources such as mobile applications, sensors, and network monitoring has increased the need for machine learning algorithms capable of handling non-stationary data streams. However, learning from such streams presents significant challenges due to their evolving nature and the presence of concept drift. One of the most complex issues is learning from imbalanced data streams, where shifting data distributions, combined with feature space drifts, complicate continuous adaptation. These challenges become even more pronounced in multi-class scenarios, which are common in real-world applications. Detecting concept drift in such contexts is particularly demanding, as it requires tracking …
Multi-Lingual And Cross-Domain Frontiers In Machine-Generated Content Detection,
2025
Wilfrid Laurier University
Multi-Lingual And Cross-Domain Frontiers In Machine-Generated Content Detection, Gurunameh Singh Chhatwal
Theses and Dissertations (Comprehensive)
The rapid advancement of generative artificial intelligence, particularly Large Language Models (LLMs) such as GPT-4 and their multilingual capabilities, has significantly blurred the distinction between human-authored and machine-generated content. This technological evolution introduces critical challenges concerning the detection and attribution of textual authenticity and authorship, exacerbating societal issues like misinformation proliferation and compromising academic and professional integrity. Traditional detection methodologies, predominantly monolingual and heuristic-based, have demonstrated inadequate generalizability and efficacy against the sophisticated, multilingual capabilities of contemporary generative models.
This thesis addresses two major problems arising from these advancements. Firstly, it introduces novel multilingual detection methodologies explicitly designed to differentiate …
Crime Modeling Using An Integrated Cnn–Lstm Architecture With Embedded Self-Excitation,
2025
Wilfrid Laurier University
Crime Modeling Using An Integrated Cnn–Lstm Architecture With Embedded Self-Excitation, Pawandeep Kaur
Theses and Dissertations (Comprehensive)
It is often assumed that natural phenomena occur randomly over time. However, careful analysis reveals that these events typically form some series or sequences and exhibit distinctive temporal patterns. These patterns are not exclusive to nature. They also appear in human activities, often studied under the concept of bursty human dynamics. The statistical methods analyzing bursty human dynamics not only capture overall trends or seasonality but also explore how past events influence future ones. It makes the analysis more realistic and the results more closely aligned with reality. Bursty human dynamics can be studied at two levels: the individual level …
Advancing Multivariate Time Series Similarity Assessment: An Integrated Computational Approach,
2025
International Centre of Insect Physiology and Ecology Nairobi
Advancing Multivariate Time Series Similarity Assessment: An Integrated Computational Approach, Franck B.N. Tonle, Henri E.Z. Tonnang, Milliam M.Z. Ndadji, Maurice Tchoupe Tchendji, Armand Nzeukou, Kennedy Senagi, Saliou Niassy
All Peer-Reviewed Publications
Data mining, particularly multivariate time series data analysis, is crucial in extracting insights from complex systems and supporting informed decision-making across diverse domains. However, assessing the similarity of multivariate time series data presents several challenges, including dealing with large datasets, addressing temporal misalignments, and necessitating efficient and comprehensive analytical frameworks. A novel integrated computational approach, Multivariate Time series Alignment and Similarity Assessment (MTASA) is proposed to address these challenges. MTASA is built upon a hybrid methodology designed to optimise time series alignment, complemented by a multiprocessing engine that enhances the utilisation of computational resources. This integrated approach comprises four key …
Privshap: A Finer-Granularity Network Linearization Method For Private Inference,
2025
Old Dominion University
Privshap: A Finer-Granularity Network Linearization Method For Private Inference, Xiangrui Xu, Zhenzhen Wang, Rui Ning, Chunsheng Xiu, Hongyi Wu
Computer Science Faculty Publications
Private inference applies cryptographic techniques like homomorphic encryption, garble circuit and secret sharing to keep both sides privacy in a client-server setting during inference. It is often hindered by the high communication overheads, especially at non-linear activation layers such as ReLU. Hence ReLU pruning has been widely recognized as an efficient way to accelerate private inference. Existing approaches to ReLU pruning typically rely on coarse hypothesis, which assume an inverse correlation between the importance of ReLU and linear layers or shallow activation layers have less importance for universal models, to assign the budgets according to the layer while preserving the …
Heterogeneous Clustering Of Multiomics Data For Breast Cancer Subgroup Classification And Detection,
2025
Virginia Commonwealth University
Heterogeneous Clustering Of Multiomics Data For Breast Cancer Subgroup Classification And Detection, Joseph Pateras, Musaddiq Lodi, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
The rapid growth of diverse -omics datasets has made multiomics data integration crucial in cancer research. This study adapts the expectation–maximization routine for the joint latent variable modeling of multiomics patient profiles. By combining this approach with traditional biological feature selection methods, this study optimizes latent distribution, enabling efficient patient clustering from well-studied cancer types with reduced computational expense. The proposed optimization subroutines enhance survival analysis and improve runtime performance. This article presents a framework for distinguishing cancer subtypes and identifying potential biomarkers for breast cancer. Key insights into individual subtype expression and function were obtained through differentially expressed gene …
A Data-Driven Sliding-Window Pairwise Comparative Approach For The Estimation Of Transmission Fitness Of Sars-Cov-2 Variants And The Construction Of The Evolution Fitness Landscape,
2025
University of Tennessee at Chattanooga
A Data-Driven Sliding-Window Pairwise Comparative Approach For The Estimation Of Transmission Fitness Of Sars-Cov-2 Variants And The Construction Of The Evolution Fitness Landscape, Md Jubair Pantho, Richard Annan, Landen Alexander Bauder, Sophia Huang, Letu Qingge, Hong Qin
Computer Science Faculty Publications
Estimating the transmission fitness of SARS-CoV-2 variants and understanding their evolutionary fitness trends are important for epidemiological forecasting. Existing methods are often constrained by their parametric natures and do not satisfactorily align with the observations during COVID-19. Here, we introduce a sliding-window data-driven pairwise comparison method, the differential population growth rate (DPGR) that uses viral strains as internal controls to mitigate sampling biases. DPGR is applicable in time windows in which the logarithmic ratio of two variant subpopulations is approximately linear. We apply DPGR to genomic surveillance data and focus on variants of concern (VOCs) in multiple countries and regions. …
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas,
2025
Old Dominion University
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Computer Science Faculty Publications
Triple-negative breast cancer (TNBC) requires detailed cellular mapping given its aggressive nature, immense tumor heterogeneity and genetic diversity. We integrated 156,794 cells from six scRNA-seq datasets—including tumors, metastases, and cell lines—to build a TNBC scRNA cell atlas, focusing on batch effect mitigation while maintaining biological and molecular details. Preprocessing f ilters noise, normalizes data, and leverages PCA for integration readiness. We utilized scANVI, a semi-supervised tool, to align datasets, preserving TNBC’s complex tumor heterogeneity via marker annotations [1]. UMAPs demonstrate biological clustering in integrated data, contrasted with datasetdriven unintegrated patterns. Assessments verifying effective batch correction. This method aligns with NASA’s …
Ai For Nuclear Physics: The Exclaim Project,
2025
University of Virginia
Ai For Nuclear Physics: The Exclaim Project, S. Liuti, D. Adams, M. Boër, G. W. Chern, M. Cuic, M. Engelhardt, G. R. Goldstein, B. Kriesten, Y. Li, H. W. Lin, M. Sievert, D. Sivers
Computer Science Faculty Publications
An overview of the recent activity of the newly funded EXCLusives with AI and Machine learning (EXCLAIM) collaboration is presented. The main goal of the collaboration is to develop a framework to implement AI and machine learning techniques in problems emerging from the phenomenology of high energy exclusive scattering processes from nucleons and nuclei, maximizing the information that can be extracted from various sets of experimental data, while implementing theoretical constraints from lattice QCD. A specific perspective embraced by EXCLAIM is to use the methods of theoretical physics to understand the working of ML, beyond its standardized applications to physics …
Neural Topic Modeling Via Contextual And Graph Information Fusion,
2025
Sun Yat-sen University
Neural Topic Modeling Via Contextual And Graph Information Fusion, Jiyuan Liu, Jiaxing Yan, Chunjiang Zhu, Xingyu Liu, Qing Li, Yanghui Rao
Computer Science Faculty Publications
Topic modeling is a powerful unsupervised tool for knowledge discovery. However, existing work struggles with generating limited-quality topics that are uninformative and incoherent, which hindering interpretable insights from managing textual data. In this paper, we improve the original variational autoencoder framework by incorporating contextual and graph information to address the above issues. First, the encoder utilizes topic fusion techniques to combine contextual and bag-of-words information well, and meanwhile exploits the constraints of topic alignment and topic sharpening to generate informative topics. Second, we develop a simple word co-occurrence graph information fusion strategy that efficiently increases topic coherence. On three benchmark …
Decode The Workload: Training Deep Learning Models For Efficient Compute Cluster Representation,
2025
Thomas Jefferson National Accelerator Facility
Decode The Workload: Training Deep Learning Models For Efficient Compute Cluster Representation, Ahmed Hossam Mohammed, Mark Jones, Diana Mcspadden, Malachi Schram, Bryan Hess, Kishansingh Rajput
Computer Science Faculty Publications
In this study, we address the mounting challenge of monitoring high throughput computing clusters running computationally intensive jobs, which increasingly strains system administrators. We develop autoencoders that analyze traces of Linux kernel CPU metrics to capture salient system features by producing robust compressed embeddings for various downstream tasks. In addition, we employ graph neural networks to incorporate contextual information from surrounding CPUs and assess their performance. We also demonstrate the enhanced job differentiation achieved by increasing the sampling rate of these traces. Our models are evaluated based on their ability to generate meaningful latent representations, detect anomalies, and distinguish between …
S²Il: Structurally Stable Incremental Learning,
2025
Sri Sathya Sai Institute of Higher Learning
S²Il: Structurally Stable Incremental Learning, S. Balasubramanian, P. Yedu Krishna, Talasu Sai Sriram, M. Sai Subramaniam, Manepalli Pranav Phanindra Sai, Ravi Mukkamala
Computer Science Faculty Publications
Feature Distillation (FD) strategies are proven to be effective in mitigating Catastrophic Forgetting (CF) seen in Class Incremental Learning (CIL). However, current FD approaches enforce strict alignment of feature magnitudes and directions across incremental steps, limiting the model’s ability to adapt to new knowledge. In this paper, we propose Structurally Stable Incremental Learning (S²IL), a FD method for CIL that mitigates forgetting by focusing on preserving the overall spatial patterns of features which promote flexible (plasticity) yet stable representations that preserve old knowledge (stability). We also demonstrate that our proposed method S²IL achieves strong incremental accuracy and outperforms other FD …
Effective Pii Extraction From Llms Through Augmented Few-Shot Learning,
2025
Zhejiang University
Effective Pii Extraction From Llms Through Augmented Few-Shot Learning, Shuai Cheng, Shu Meng, Haitao Xu, Haoran Zhang, Shuai Hao, Chuan Yue, Wenrui Ma, Meng Han, Fang Zhang, Zhao Li
Computer Science Faculty Publications
Large Language Models (LLMs) exhibit strong natural language processing capabilities but also pose significant privacy risks, particularly regarding the leakage of Personally Identifiable Information (PII) embedded in their training data. Existing PII extraction methods suffer from the limitations of low success rates or impracticality for large-scale PII extraction. In this study, we propose a novel PII extraction approach based on enhanced few-shot learning techniques, which achieves efficient and cost-effective PII retrieval without relying on fine-tuning or jailbreaking. We evaluated our approach on both open-source and closed-source LLMs. The experimental results demonstrate that, for non-targeted PII extraction, the attack success rate …
Energy-Based Deep Incomplete Multi-View Clustering,
2025
Old Dominion University
Energy-Based Deep Incomplete Multi-View Clustering, Ziyu Wang, Yiming Du, Rui Ning, Lusi Li
Computer Science Faculty Publications
Incomplete multi-view clustering (IMVC) deals with real-world scenarios where certain views are partially missing, posing significant challenges to effective clustering. Most existing IMVC approaches face a trade-off: imputation-free methods suffer from information bias and imbalance, while full-imputation methods risk introducing and propagating noise. To overcome these limitations, we propose Energy-Based Deep Incomplete Multi-View Clustering (Energy-DIMC), a novel selective-imputation framework that leverages energy-based models (EBMs) to guide reliable imputations and robust clustering. EBMs assess data compatibility by assigning lower energy to more coherent structures, effectively modeling complex inter-view and inter-sample dependencies. Inspired by EBMs, Energy-DIMC integrates four key components: 1) a …
Icu-Length Of Stay Prediction On Electronic Health Records Using Graph Neural Networks And Homogeneous Similarity Graphs,
2025
Virginia Commonwealth University
Icu-Length Of Stay Prediction On Electronic Health Records Using Graph Neural Networks And Homogeneous Similarity Graphs, Ahmad F. Al Musawi, Pratip Rana, Sibtanu Raha, Joshua Braunstein, William C. Sleeman Iv, Rishabh Kapoor, Preetam Ghosh
Computer Science Faculty Publications
Predicting the length of stay (LoS) is important for hospital administration, as it helps allocate proper resources, such as bed management and hospital staffing. Patients' Electronic Health Records (EHRs) contain highly relevant data for LoS prediction; however, their integration and effective use in predictive modeling for accurately estimating LoS remain challenging. To address this, we propose a homogeneous Graph Neural Network (GNN)-based framework for predicting LoS. This method employs a comprehensive data fusion strategy based on the hospital Visit-based Similarity Graph (VSG), which integrates diverse multi-modal clinical features into a coherent, homogeneous graph representation. Next, this VSG is fed into …
Precision Haptics For Rehabilitation: Quantifying Directional Bias And Force Threshold Effects On Motor Learning,
2025
University of North Florida
Precision Haptics For Rehabilitation: Quantifying Directional Bias And Force Threshold Effects On Motor Learning, Conor J. Nolan
UNF Graduate Theses and Dissertations
Hand rehabilitation represents a critical challenge in modern physical therapy, with significant implications for patients' quality of life and functional independence. Despite technological advancements across healthcare, traditional hand assessment methods often rely on subjective measures or lack task-specific biomechanical assessment capabilities. This thesis investigates the effects of haptic force feedback, handedness, and rotation direction on circle-tracing task completion using a within-subjects factorial design with 20 university participants examining varying resistance levels (0.0N, 0.5N, 1.2N) across different movement configurations using a 3D Systems Touch X haptic device with 0.023mm precision. Results demonstrate that moderate haptic force (0.5N) significantly enhanced spatial accuracy …
