Machine Learning Classification Of Prostate Cancer Genomic Sequences Using K-Mer And Sequence-Derived Features,
2026
Elizabeth City State University
Machine Learning Classification Of Prostate Cancer Genomic Sequences Using K-Mer And Sequence-Derived Features, Kuldeep Rawat, Hirendra Nath Banerjee, Jamie Noble, Saa Naudia Deloatch, Satyendra Banerjee, Sachin Shetty, Soumya Banerjee
VMASC Publications
Prostate cancer disproportionately impacts African American men, who experience significantly higher mortality rates and earlier disease onset than other populations. Current diagnostic approaches, including prostate-specific antigen testing and biopsy, lack sufficient specificity and sensitivity, underscoring the need for accurate, molecular-level classification tools. This paper presents a machine learning framework for binary classification of genomic DNA sequences as cancerous or healthy. A dataset of 1684 FASTA-formatted sequences obtained from the National Library of Medicine - GenBank was analyzed, with 1662 sequences retained after quality control filtering. Feature engineering yielded 67 attributes, including GC content, Shannon entropy, sequence length, and trinucleotide k-mer …
Beyond Full Fine-Tuning: The New Playbook For Adapting Deep Neural Networks,
2026
University of Central Florida
Beyond Full Fine-Tuning: The New Playbook For Adapting Deep Neural Networks, Cristian S. Mcgee
Honors Undergraduate Theses
Fine-tuning is the process of teaching and specializing a pre-trained neural network on a downstream task. Fine-tuning is a rapidly growing topic in artificial intelligence domains; however, many fine-tuning endeavors are highly specialized without a coherent framework connecting them. This work presents a unified perspective on fine-tuning methods and performance metrics. Our perspective organizes the methods in terms of how they are applied to fine-tuning. This framework showcases methods that (i) update effective subspaces of the pre-trained model, (ii) change the adaptation optimization procedure, and (iii) alter the representations of the embedded input. Additionally, we present unconventional metrics such as …
An Empirical Framework For Evaluating Semantic Preservation Using Hugging Face,
2026
CUNY Graduate Center
An Empirical Framework For Evaluating Semantic Preservation Using Hugging Face, Nan Jia, Anita Raja, Raffi Khatchadourian
Publications and Research
As machine learning (ML) becomes an integral part of high-autonomy systems, it is critical to ensure the trustworthiness of learning-enabled software systems (LESS). Yet, the nondeterministic and run-time-defined semantics of ML complicate traditional software refactoring. We define semantic preservation in LESS as the property that optimizations of intelligent components do not alter the system's overall functional behavior. This paper introduces an empirical framework to evaluate semantic preservation in LESS by mining model evolution data from HuggingFace. We extract commit histories, $\textit{Model Cards}$, and performance metrics from a large number of models. To establish baselines, we conducted case studies in three …
Clinical Subtypes Of Co-Morbid Insomnia And Obstructive Sleep Apnea (Comisa): Results Of A Cluster Analysis,
2026
Sichuan University
Clinical Subtypes Of Co-Morbid Insomnia And Obstructive Sleep Apnea (Comisa): Results Of A Cluster Analysis, Yuan Shi, Xujun Feng, Fengyi Hao, Yuru Nie, Yihui Zhang, Zhaohua Chen, Siqi Guan, Larry D. Sanford, Michael V. Vitiello, Xiangdong Tang
Department of Pathology & Anatomy Faculty Publications
Background
Variations in the bidirectional relationship between obstructive sleep apnea (OSA) and insomnia in co-morbid insomnia and OSA (COMISA) may form distinct subtypes of COMISA, which have not been previously characterized. This study aims to identify and characterize subtypes of COMISA.
Methods
From a community-recruited COMISA cohort 256 individuals who met diagnosis for COMISA were used to identify subtypes using a two-step clustering methodology. Demographics and multidimension clinical characteristics were collected and compared among obtained subtypes. Logistic models were used to evaluate whether these subtypes were associated with cardiometabolic and mental disorders. A clinical cohort of 1816 COMISA patients was …
Enlem: Ensemble Learning-Based Model To Detect Phishing Websites,
2026
Bangladesh University of Business and Technology
Enlem: Ensemble Learning-Based Model To Detect Phishing Websites, Most Nilufa Yeasmin, Md Abu Rumman Refat, Bikash Chandra Singh, Zulfikar Alom, Zeyar Aung, Mohammad Azim
School of Cybersecurity Faculty Publications
Phishing involves manipulating individuals into revealing private data, e.g., user IDs, bank details, and passwords. The observed surge in fraud is related to increased deception, impersonation, and advanced online attacks. Thus, effective phishing detection methods are required to mitigate escalating global phishing threats. Existing methods (e.g., heuristics-based, signature-based, and visual similarity-based methods) attempt to detect phishing sites, and machine learning (ML) and deep learning (DL) methods are effective in the cybersecurity context in terms of learning from data, offering insights, and forecasting. However, independent ML algorithms are limited when handling complex data, and DL techniques surpass traditional ML methods in …
Game-Based Learning For Asynchronous Ai Literacy Course: Approach To Improve Students' Cognitive, Behavioural, Affective, And Ethical Learning Of Ai,
2026
Old Dominion University
Game-Based Learning For Asynchronous Ai Literacy Course: Approach To Improve Students' Cognitive, Behavioural, Affective, And Ethical Learning Of Ai, Jinhee Kim, Guang Yang, Wing Sha Chan, Xi Lin, Yukyeong Song
STEMPS Faculty Publications
Educators in higher education face persistent challenges in scaling AI literacy across disciplines and helping novice learners understand abstract AI concepts. Although research on game-based learning (GBL) reports mixed outcomes, few studies have examined its large-scale use in mandatory, asynchronous AI literacy courses for diverse undergraduate populations. Addressing this gap, this study investigates a scalable GBL-based AI literacy course delivered to 4898 first-year undergraduates across disciplines. Using a mixed-methods design with 311 valid pre- and post-survey responses and 20 interviews, the study evaluates students' cognitive, behavioural, affective, and ethical learning of AI. Quantitative results show significant improvements in overall AI …
A New Parallel-In-Time Direct Inverse Method For Nonlinear Differential Equations,
2026
Old Dominion University
A New Parallel-In-Time Direct Inverse Method For Nonlinear Differential Equations, Nail K. Yamaleev, Subhash Paudel
Mathematics & Statistics Faculty Publications
We propose a new method for parallelization of the first-order backward difference discretization (BDF1) of the first-order time derivative in nonlinear partial differential equations, such as conservation law equations. The time derivative term is discretized by using the method of lines based on the implicit BDF1 scheme, while the inviscid and viscous terms are approximated by conventional 2nd-order central discretizations of the 1st- and 2nd-order derivatives in each spatial direction. The global system of nonlinear discrete equations in the space-time domain is solved by the Newton method for all time levels simultaneously. For the BDF1 discretization, this all-at-once system at …
The Role Of Education In Reducing Social Inequality: A Systems-Level Analysis Of Socio-Technical Infrastructures And Policy Governance,
2026
Old Dominion University
The Role Of Education In Reducing Social Inequality: A Systems-Level Analysis Of Socio-Technical Infrastructures And Policy Governance, Aisling O'Shea, Batzorig Dashnyam, Ximena Quintanilla
Women's & Gender Studies Faculty Publications
Social inequality remains one of the most persistent challenges to global systemic stability, threatening the robustness of democratic institutions and economic sustainability. Education has long been theorized as the primary mechanism for social mobility and the mitigation of disparate life outcomes; however, its role within modern socio-technical infrastructures is increasingly complex and often contradictory. This paper provides a comprehensive systems-level analysis of the relationship between educational architecture and social stratification. By examining the structural trade-offs inherent in contemporary pedagogical deployment, the research evaluates how institutional governance, digital infrastructure, and policy mandates either facilitate or hinder the reduction of inequality. The …
Qubit Lattice Algorithm Simulations Of The Scattering Of A Bounded Two Dimensional Electromagnetic Pulse From The Infinite Planar Dielectric Interface,
2026
Rogers State University
Qubit Lattice Algorithm Simulations Of The Scattering Of A Bounded Two Dimensional Electromagnetic Pulse From The Infinite Planar Dielectric Interface, Min Soe, George Vahala, Linda Vahala, Efstratios Koukoutsis, Abhay K. Ram, Kyriakos Hizanidis
Electrical & Computer Engineering Faculty Publications
Qubit lattice algorithm (QLA) simulations are performed for a two-dimensional spatially bounded pulse propagating onto a plane interface between two dielectric slabs. QLA is an initial value scheme that consists of a sequence of unitary collision and streaming operators, with appropriate potential operators, that recover Maxwell equations in inhomogeneous dielectric media to the second order in the lattice discreteness. For the case of total internal reflection, there is transient energy transfer into the second medium due to the evanescent fields as the Poynting unit vector of the pulse is rotated from its incident to reflected direction. Because of the finite …
Reconstructing Lost Voices,
2026
The University of Akron
Reconstructing Lost Voices, Lana Tamim
Williams Honors College, Honors Research Projects
This project uses digital text mining tools (OCR, NLP, sentiment analysis, and topic modeling) to analyze 19th–20th-century newspaper archives, focusing on how marginalized groups (women, immigrants, or labor workers) were historically portrayed. Many historical newspapers were dominated by elite voices, so this project aims to recover silenced or misrepresented perspectives by identifying hidden patterns in language, frequency of coverage, sentiment, and shifts in public perception over time. Using machine learning and visualization tools, the project will create interactive maps and timelines showing how representation evolved across regions.
Machine Learning For Economists,
2026
Portland State University
Machine Learning For Economists, John Luke Gallup
Economics Faculty Publications and Presentations
Explication of machine learning algorithms and their usefulness for economic research. The prediction algorithms of Random Forest, Gradient Boost Machines, Neural Networks and Support Vector Machines are built from simple steps applied at large scale to generate surprisingly precise nonlinear estimates. Although useful for processing and interpreting new forms of data, their application to economics research is limited because they do not provide readily interpretable evidence of the causes of outcomes.
To Print A Remake: An Analysis Of Hollywood Remakes And Their Cultural Value,
2026
University of Central Florida
To Print A Remake: An Analysis Of Hollywood Remakes And Their Cultural Value, Connor F. Seaton
Honors Undergraduate Theses
In recent years, audiences and movie critics have expressed concern that Hollywood’s growing reliance on remakes, sequels, franchises, and similar adaptations has led to a broader worry that originality is fading from modern cinema and that the industry is instead focused on using adaptations to maximize profits. Although adaptation is often seen as a commercially driven framework for reproducing existing intellectual property in a new media format, this thesis argues that it should be recognized as an autonomous cultural category with its own artistic, historical, and social significance and merit. By analyzing adaptation scholarship and reviewing its complex historical development, …
A General Algorithm For Assortment Optimization Under Random Utility Choice Models,
2026
Singapore Management University
A General Algorithm For Assortment Optimization Under Random Utility Choice Models, Tien Mai, Andrea Lodi
Research Collection School Of Computing and Information Systems
This work concerns the assortment optimization problem that refers to selecting a subset of items that maximizes the expected revenue in the presence of the substitution behavior of consumers specified by a random utility choice model. The key challenge lies in the computational difficulty of finding the best subset solution, which often requires exhaustive search. The literature on constrained assortment optimization lacks a practically efficient method that is general to deal with different types of customer choice models (e.g., the multinomial logit, mixed logit or general multivariate extreme value models). In this work, we propose a new approach that allows …
Algorithmic Trading In Idiosyncratic-Payoff Markets: A Multi-Agent System For On-Chain Prediction Contracts,
2026
Claremont McKenna College
Algorithmic Trading In Idiosyncratic-Payoff Markets: A Multi-Agent System For On-Chain Prediction Contracts, Saif Aldeen A.K. Agha
CMC Senior Theses
This thesis documents the design, deployment, and forward-test evaluation of an evolutionary multi-agent algorithmic trading system on Polymarket, the largest decentralized prediction market. The system pairs a locally-hosted 72-billion-parameter language model with a gradient-boosted statistical filter and an evolutionary selection mechanism that maintains a population of approximately 500 autonomous trading agents. Each agent generates a probability estimate for an event, compares it to the prevailing market price, and trades the resulting disagreement.
The central empirical exercise estimates a panel regression of trade-level profit on the absolute disagreement between the agent's probability estimate and the market price, controlling for agent identity, …
Utilizing Machine Learning Techniques For Computer-Aided Covid-19 Screening Based On Clinical Data,
2026
The University of Texas at El Paso
Utilizing Machine Learning Techniques For Computer-Aided Covid-19 Screening Based On Clinical Data, Honglun Xu, Andrews T. Anum, Michael Pokojovy, Sreenath Chalil Madathil, Yuxin Wen, Md Fashiar Rahman, Tzu-Liang (Bill) Tseng, Scott Moen, Eric Walser
Mathematics & Statistics Faculty Publications
The COVID-19 pandemic has highlighted the importance of rapid clinical decision-making to facilitate the efficient usage of healthcare resources. Over the past decade, machine learning (ML) has caused a tectonic shift in healthcare, empowering data-driven prediction and decision-making. Recent research demonstrates how ML was used to respond to the COVID-19 pandemic. This paper puts forth new computer-aided COVID-19 disease screening techniques using six classes of ML algorithms (including penalized logistic regression, random forest, artificial neural networks, and support vector machines) and evaluates their performance when applied to a real-world clinical dataset containing patients’ demographic information and vital indices (such as …
A Systematic Review And Characterization Of Privacy Noncompliance In Real-World Applications,
2026
University of Central Florida
A Systematic Review And Characterization Of Privacy Noncompliance In Real-World Applications, Alexander E. Charkiewicz
Graduate Studies Theses and Dissertations 2026
Software applications increasingly rely on user data to provide their functionality, but improper handling of such data can lead to serious privacy noncompliance with applicable regulations and policies. A prominent example is the Facebook–Cambridge Analytica scandal, in which a third-party application collected the personal data of approximately 87 million Facebook users without users' consent. Despite growing attention to privacy compliance, two key challenges hinder the systematic understanding and analysis of privacy noncompliance. First, unlike security vulnerabilities, which have been systematically categorized through taxonomies such as the Common Weakness Enumeration (CWE), privacy noncompliance lacks a technical taxonomy describing how it manifests …
Towards Sample-Efficient Deep Reinforcement Learning,
2026
Virginia Commonwealth University
Towards Sample-Efficient Deep Reinforcement Learning, Guang Yang
Theses and Dissertations
Deep reinforcement learning (DRL), combining reinforcement learning and high-performance function approximations such as deep neural networks (DNN), is a powerful approach to solving complex sequential decision-making problems. However, due to the complex solution space of the sequential decision-making problems and the inefficient design of the DRL algorithms, DRL algorithms usually require a prohibitively large number of data samples to train effective strategies. Consequently, it is difficult to apply these DRL algorithms to complex real-world problems that require high costs to collect a large volume of data samples. This dissertation proposes new mechanisms to address this sample inefficiency issue, realizing sample-efficient …
Mg-Spair: Multi-Grade Sparse-Guided Implicit Representation For Training-Data-Free Image Restoration,
2026
Syracuse University
Mg-Spair: Multi-Grade Sparse-Guided Implicit Representation For Training-Data-Free Image Restoration, Jianmin Liao, Lei Huang, Ronglong Fang, Ashley Prater-Bennette, Lixin Shen, Yuesheng Xu
Mathematics & Statistics Faculty Publications
MG-SpaIR is a training-data-free framework for restoring a clean image from a single observation corrupted by a mixture of blur, downsampling, noise, and missing pixels. Building on implicit neural representations (INRs), we introduce a multi-grade residual hierarchy that progressively refines the reconstruction from low to high spatial frequencies across grades, improving representational fidelity and mitigating spectral limitations. To stabilize reconstruction optimization and suppress INR-induced artifacts, we further propose an explicit sparse proximal regularization (e.g., ℓ0 type) applied directly in the high-resolution image domain, which discourages spurious high-frequency patterns while preserving sharp structures. The resulting optimization is solved efficiently via a …
Ai-Driven Penetration Testing For Arm Systems: A Comprehensive Framework With Experimental Validation,
2026
Georgia Southern University
Ai-Driven Penetration Testing For Arm Systems: A Comprehensive Framework With Experimental Validation, Matthew Ragsdale
College of Graduate Studies: Theses & Dissertations
The convergence of artificial intelligence and cybersecurity presents new opportunities for automated penetration testing capable of discovering, prioritizing, and remediating vulnerabilities at machine speed. However, deployment on resource-constrained ARM platforms remains unexplored despite ARM’s dominance in mobile, IoT, and edge computing with over 280 billion chips deployed globally. This thesis presents systematic experimental evaluation of AI-driven penetration testing across four paradigms—traditional machine learning, deep learning, large language models, and reinforcement learning—on three ARM platform tiers: Raspberry Pi 5 (8GB, Cortex-A76), Radxa ROCK 5B Plus (16GB LPDDR5 with NPU), and NVIDIA Jetson Nano (4GB with Maxwell GPU). The experimental framework generates …
Optimizing Maintenance Routes For Highway Infrastructure Using Leader-Follower Autonomous Vehicles,
2026
Old Dominion University
Optimizing Maintenance Routes For Highway Infrastructure Using Leader-Follower Autonomous Vehicles, Qing Tang, Chenxi Chen, Xianbiao Hu, Yuxin Ding, Tianjia Yang
Civil & Environmental Engineering Faculty Publications
The Autonomous Truck Mounted Attenuator (ATMA), a leader–follower style connected and automated vehicle system, enhances safety during transportation infrastructure maintenance in work zones. However, the significantly lower speed of ATMA, compared to regular vehicles, causes moving bottlenecks that reduce roadway capacity and prolong queuing, leading to further delays. Different ATMA routes lead to varying patterns of time-dependent capacity drop, affecting the user equilibrium traffic assignment and resulting in differing system costs. This study aims to optimize ATMA routing within a network to minimize the system cost associated with its slow-moving operation. To this end, a queuing-based traffic assignment approach is …
