Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (4407)
- TÜBİTAK (1233)
- Purdue University (1089)
- Old Dominion University (945)
- Missouri University of Science and Technology (916)
-
- University of Nebraska - Lincoln (870)
- Air Force Institute of Technology (846)
- Wright State University (817)
- San Jose State University (616)
- Dartmouth College (581)
- Edith Cowan University (567)
- Brigham Young University (506)
- California Polytechnic State University, San Luis Obispo (448)
- City University of New York (CUNY) (444)
- University of Texas at El Paso (387)
- Technological University Dublin (367)
- Portland State University (354)
- New Jersey Institute of Technology (342)
- University of Central Florida (333)
- Washington University in St. Louis (313)
- Embry-Riddle Aeronautical University (311)
- China Simulation Federation (308)
- Nova Southeastern University (295)
- Kennesaw State University (280)
- University of Arkansas, Fayetteville (266)
- Utah State University (259)
- University of South Florida (252)
- University of Texas at Arlington (252)
- Zayed University (252)
- Syracuse University (245)
- Keyword
-
- Machine learning (1031)
- Deep learning (593)
- Artificial intelligence (509)
- Machine Learning (497)
- Security (364)
-
- Computer Science (342)
- Deep Learning (292)
- Artificial Intelligence (261)
- Classification (260)
- Cybersecurity (243)
- Computer vision (229)
- Department of Computer Science and Engineering (211)
- Privacy (211)
- Neural networks (205)
- Algorithms (203)
- Data mining (193)
- Natural language processing (186)
- Computer science (181)
- Applied sciences (171)
- College for Professional Studies (165)
- Software engineering (163)
- School of Computer & Information Science (150)
- Simulation (146)
- AI (142)
- Optimization (138)
- Blockchain (133)
- Image processing (133)
- Natural Language Processing (133)
- Clustering (128)
- Education (123)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (4205)
- Theses and Dissertations (1686)
- Turkish Journal of Electrical Engineering and Computer Sciences (1233)
- Department of Computer Science Technical Reports (906)
- Master's Projects (565)
-
- Computer Science Faculty Publications (435)
- Computer Science Technical Reports (404)
- Computer Science Faculty Research & Creative Works (385)
- Electronic Theses and Dissertations (385)
- Dissertations (362)
- The R Journal (353)
- Faculty Publications (351)
- Journal of System Simulation (308)
- CCAC Theses and Dissertations (277)
- Browse all Theses and Dissertations (262)
- Theses (253)
- All Works (252)
- USF Tampa Graduate Theses and Dissertations (230)
- Computer Science Faculty Publications and Presentations (229)
- Chulalongkorn University Theses and Dissertations (Chula ETD) (217)
- Masters Theses (215)
- Computer Science & Engineering Syllabi (214)
- Departmental Technical Reports (CS) (212)
- Walden Dissertations and Doctoral Studies (210)
- Master's Theses (209)
- All Computer Science and Engineering Research (205)
- Kno.e.sis Publications (196)
- Computer Science and Software Engineering (180)
- Open Access Theses & Dissertations (171)
- Regis University Student Publications (comprehensive collection) (166)
- Publication Type
Articles 121 - 150 of 27587
Full-Text Articles in Entire DC Network
Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Although Large Language Models (LLMs) demonstrate significant capabilities, their reliance on parametric knowledge often leads to inaccuracies. Retrieval Augmented Generation (RAG) mitigates this by incorporating external knowledge, but these methods may introduce irrelevant retrieved documents, leading to inaccurate responses. While the integration methods filter out incorrect answers from multiple responses, but lack external knowledge like RAG methods, and their high costs require balancing overhead with performance gains. To address these issues, we propose an Efficient Test-Time Retrieval-Augmented Generation Framework named ET2RAG to improve the performance of LLMs while maintaining efficiency. Specifically, ET2RAG is a training-free method, that first retrieves the …
Lessons From The Club Homeschool Capstone: Testing, Data Discipline, And The Computer Science Curriculum, Shane Brown
Lessons From The Club Homeschool Capstone: Testing, Data Discipline, And The Computer Science Curriculum, Shane Brown
University Honors Theses
This thesis looks at the CLUB Homeschool Capstone project to argue that Portland State University's Computer Science curriculum should introduce testing and data quality discipline earlier and more intentionally than it does now. As team lead of a seven-person team, I coordinated sprint planning, communicated with the sponsor, and developed custom Discourse plugins that enhanced an existing forum platform instead of creating a separate application database, as requested by the sponsor. The project's requirements document called for a formal testing plan, but our team lacked the practical experience to implement one. This gap became evident through my internships as a …
Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen
Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen
Dissertations, Theses, and Projects
The increasing adoption of the Internet of Medical Things (IoMT) has improved healthcare delivery through connected medical devices while simultaneously expanding the cybersecurity risks facing healthcare organizations. Although machine learning based intrusion detection systems have demonstrated high detection accuracy, their ability to respond reliably to previously unseen cyberattacks remains uncertain. This study investigated how a Neural Network model and a Logistic Regression model classified novel cyberattacks within the IoMT environment. The Neural Network and Logistic Regression models were both trained and tested using a subset of the CICIoMT2024 benchmark dataset. The Neural Network achieved 99.82% test accuracy and a 0.94 …
Evaluating Generative Ai-Based User Interfaces Using An Integrated Neutrosophic Multi-Criteria Decision-Making Framework, Nada Mohamed, Alshaimaa A. Tantawy
Evaluating Generative Ai-Based User Interfaces Using An Integrated Neutrosophic Multi-Criteria Decision-Making Framework, Nada Mohamed, Alshaimaa A. Tantawy
Neutrosophic Systems with Applications
User Interface (UI) design can be seen as an essential aspect of human-computer interaction (HCI) and makes communication easier between people and technology. In today's digital economy, interface quality has become one of the most important business concerns, since it has a direct impact on customer satisfaction and retention while affecting revenue. Although creating user-centered and accessible interfaces is crucial, doing so is a difficult and time-consuming process, which leads to burnout for many usability professionals. Although conventional artificial intelligence (AI) was utilized for design assessment and automation, the arrival of generative AI technology has created new possibilities for automated …
Where Do Ai Coding Agents Fail? An Empirical Study Of Failed Agentic Pull Requests In Github, Ramtin Ehsani, Sakshi Pathak, Shriya Rawal, Abdullah Al Mujahid, Mia Mohammad Imran, Preetha Chatterjee
Where Do Ai Coding Agents Fail? An Empirical Study Of Failed Agentic Pull Requests In Github, Ramtin Ehsani, Sakshi Pathak, Shriya Rawal, Abdullah Al Mujahid, Mia Mohammad Imran, Preetha Chatterjee
Computer Science Faculty Research & Creative Works
AI coding agents are now submitting pull requests (PRs) to software projects, acting not just as assistants but as autonomous contributors. As these agentic contributions are rapidly increasing across real repositories, little is known about how they behave in practice and why many of them fail to be merged. In this paper, we conduct a large-scale study of 33k agent-authored PRs made by five coding agents across GitHub. (RQ1) We first quantitatively characterize merged and not-merged PRs along four broad dimensions: 1) merge outcomes across task types, 2) code changes, 3) CI build results, and 4) review dynamics. We observe …
Attention-Based Ensemble Deep Learning Model For Arabic And English Fake News Classification, Ameer Alhaq Alshamery
Attention-Based Ensemble Deep Learning Model For Arabic And English Fake News Classification, Ameer Alhaq Alshamery
Journal of Intelligent Informatics, Networking, and Cybersecurity
It is difficult to classify articles as fake news since one article may consist of true facts with only some statements being fake. Moreover, classification becomes complicated for the Arabic language owing to its morphology and several ways of spelling, as well as the lack of well-classified and marked data sets. This paper presents an Ensemble Deep Learning Model (EDLM) used for Arabic and English fake news classification. The EDLM consists of CNN, Bi-LSTM with attention, and Bi-GRU with attention networks. Each of them produces one probability of the article, which is then summed up to a final probability via …
A Mathematical Decision-Making Framework For Athlete Development In A Collegiate Taekwondo Community: Prioritizing Coaching Interventions Using Statistical Analysis And The Analytic Hierarchy Process, King Harold A. Recto, Hazel Jade L. Antonio, Jhyrald Anthony P. Dalida
A Mathematical Decision-Making Framework For Athlete Development In A Collegiate Taekwondo Community: Prioritizing Coaching Interventions Using Statistical Analysis And The Analytic Hierarchy Process, King Harold A. Recto, Hazel Jade L. Antonio, Jhyrald Anthony P. Dalida
Electronics, Computer, and Communications Engineering Faculty Publications
Athlete development within collegiate sports communities requires informed decisions regarding the prioritization of coaching interventions and allocation of developmental resources. However, such decisions are frequently guided by experience and intuition, limiting opportunities for systematic and evidence-based decision-making. This study develops a mathematical decision-making framework for athlete development by integrating statistical analysis and the Analytic Hierarchy Process (AHP) within a collegiate taekwondo community. Data were collected from 25 collegiate taekwondo athletes who satisfied established eligibility criteria, including participation in University Athletic Association of the Philippines (UAAP) competitions during the previous three seasons. Athletes evaluated coaching practices across five dimensions: Training and …
A Data-Driven Framework For Mitigating Breast Cancer Overdiagnosis: From Estimation To Risk-Adjusted Computer-Aided Diagnosis, William M. Brown Jr.
A Data-Driven Framework For Mitigating Breast Cancer Overdiagnosis: From Estimation To Risk-Adjusted Computer-Aided Diagnosis, William M. Brown Jr.
LSU Doctoral Dissertations
In Computer-Aided Diagnosis (CAD) of cancer, standard cost metrics (false-positives and false-negatives) fundamentally fail to account for overdiagnosis. Overdiagnosis is a critical scenario where a disease is correctly detected (true-positive) but is biologically indolent and would never have caused the patient harm or symptoms. While widely recognized in the medical community as a major healthcare crisis driving stressful and invasive overtreatment, overdiagnosis remains severely under-researched within computer science and engineering. This dissertation addresses this interdisciplinary gap by defining the three key computational challenges of overdiagnosis: (i) accurate estimation, (ii) harm quantification, and (iii) algorithmic mitigation. To overcome the estimation challenge, …
Information Theory Analysis Of Water Vapor Stable Isotopes From The Sail Campaign, Matthew John Rybecky
Information Theory Analysis Of Water Vapor Stable Isotopes From The Sail Campaign, Matthew John Rybecky
Earth and Planetary Sciences ETDs
Understanding the processes that control water vapor isotopic composition in mountain environ- ments is essential for interpreting isotope records and predicting water resource responses to cli- mate change. This thesis applies information theory to continuous, high-resolution water vapor stable isotope measurements from the Surface Atmosphere Integrated Field Laboratory (SAIL) campaign in the East River watershed of Colorado’s Upper Gunnison Basin, spanning the winter- to-spring transition of 2022–2023. The analysis employs Shannon entropy, mutual information, transfer entropy, and joint transfer en- tropy (JTE) to quantify how environmental variables, including surface meteorology, radiation, tur- bulent fluxes, and ERA5 reanalysis products, transfer information …
Ai-Powered Resume Screening, Sang Suh, Numery Zaber
Ai-Powered Resume Screening, Sang Suh, Numery Zaber
Faculty Publications
Traditional resume screening is manual, slow, and susceptible to bias, and it struggles to keep pace with today’s application volumes. This paper presents a dual-engine, AI-powered resume screening system designed for transparency and reproducibility. The primary (classical) pipeline encodes resumes and job descriptions using Sentence-BERT (SBERT), computes a resume–job match score via cosine similarity, classifies candidates into 25 job categories using XGBoost, and provides model interpretability through SHAP. In parallel, a prompted large language model (LLM) baseline (GPT-4o/4o-mini) outputs a match score and predicted category for comparative analysis. A Streamlit-based interface integrates both engines to support recruiter workflows and human-in-the-loop …
Stop Blaming My Users: Illumination Of The Technocentric Mythos Bias, Ervin H. Frenzel, Richard Lightcap
Stop Blaming My Users: Illumination Of The Technocentric Mythos Bias, Ervin H. Frenzel, Richard Lightcap
Journal of Cybersecurity Education, Research and Practice
Abstract -This conceptual essay addresses the need for systemic and systematic transdisciplinary analytical techniques within cybersecurity and technical security. This conceptual essay is contingent upon recognition that cybersecurity is not simply technical in nature, it does not need an adversary, and more importantly it is based upon systems engineering and systems thinking. The essay contributes a socio-technical attribution chain and field-specific ontology/taxonomy which distinguish user-triggered events from root causes, latent conditions, technical debt, validation failures, governance failures, and attribution bias before assigning responsibility to end users. It systematically defines an ontology inclusive of developer technical debt, organizational debt arising from …
Escaping The Cyberstorm: A Gamified Social Engineering Training Program, Noah Mcclanahan, Fadi Abu-Amara, Ali Khattab, Travis Jett, Andre Jackson
Escaping The Cyberstorm: A Gamified Social Engineering Training Program, Noah Mcclanahan, Fadi Abu-Amara, Ali Khattab, Travis Jett, Andre Jackson
Journal of Cybersecurity Education, Research and Practice
In this research work, we explored the effectiveness of gamification in improving cybersecurity awareness and training users on targeted social engineering attacks. Traditional cybersecurity training focuses on lectures and videos. These training methods may not actively engage employees, which reduces their knowledge retention and ability to recognize social engineering attacks. This lack of involvement is a concern, as social engineering continues to be one of the most prevalent attack methods faced by end-users. A gamified training program, Escaping the Cyberstorm, was developed using the Godot game engine to address key challenges in spreading cybersecurity awareness. The game includes real-life …
Match Made In Ml: Developing Compatibility Relationships In Evidential Reasoning Approaches With Machine Learning, Ella Jolie Thomas
Match Made In Ml: Developing Compatibility Relationships In Evidential Reasoning Approaches With Machine Learning, Ella Jolie Thomas
Master's Theses
The presented expectation maximization informed evidential reasoning model extends the ability of the evidential reasoning calculus to support decision making by integrating an adaptive model learning capability. Compatibility relationships in Evidential Reasoning models are traditionally built by human domain experts. This process is labor-intensive, especially for large and complex models. Additionally, when new data becomes available, compatibility relationships must be reconstructed. Using machine learning and the expectation maximization algorithm, it is demonstrated that compatibility relationships can be constructed that learn relationships between domain knowledge that is used to make decisions. Using drug development as a domain of application, a traditional …
Accessibility Fairness Practices In Ai Applications For People With Disabilities, Megan Gross
Accessibility Fairness Practices In Ai Applications For People With Disabilities, Megan Gross
Master's Theses
Advancements in artificial intelligence (AI) improve the lives of people every day with tools like the auto-captioning of videos, improved screen-reader capabilities, and advanced mobility control through speech. However, are all groups of people benefiting from AI or are some being overlooked and left out? Although AI tools made for people with disabilities (PWDs) have improved their lives, AI for the general population generally ignores the experiences of PWDs, making them unable to interact with and benefit from technology. This research evaluates ChatGPT and Gemini in Gmail for usability fairness and analyzes how current regulations and development processes fail to …
Can Machines Testify? Llms And The Boundaries Of Testimonial Epistemology, Michael J. Cummins
Can Machines Testify? Llms And The Boundaries Of Testimonial Epistemology, Michael J. Cummins
Philosophy Summer Fellows
As Large Language Models and AI chatbots become increasingly prevalent, pressing questions are raised about whether beliefs formed through LLM interactions carry the same epistemic weight as beliefs formed through human testimony. How we answer this question depends on whether LLMs can function as testifiers, a role which is typically assumed to require a human or human-like agent. This assumption has gone largely unexamined, yet its consequences are significant: if LLM outputs cannot constitute testimony, then the justificatory tools of testimonial epistemology are unavailable to any beliefs formed through LLM interaction. This paper challenges that assumption. It first argues that …
Entity Labels Are Not Entity Signals: A Framework For Observable Relevance In Document Re-Ranking, Utshab Kumar Ghosh, Shubham Chatterjee
Entity Labels Are Not Entity Signals: A Framework For Observable Relevance In Document Re-Ranking, Utshab Kumar Ghosh, Shubham Chatterjee
Computer Science Faculty Research & Creative Works
Entity-aware document retrieval uses query-associated entities as ranking signals, assuming that semantically relevant entities are also useful retrieval signals. We show this assumption is insufficient - and explain why. Unlike terms, which are ground-truth observations, entity links are hypotheses produced by an imperfect linker: an entity can be topically central yet provide no discriminative signal if the linker fires indiscriminately across relevant and non-relevant documents. We formalize this as a distinction between Conceptual Entity Relevance (CER) - whether an entity is topically related to a query - and Observable Entity Relevance (OER) - whether its observed presence in a collection …
Stylometric And Formal Patterns In The Scholarly Impact Of Scientific Literature, Joshua Ange, Eric Godat, Rajani Sudan
Stylometric And Formal Patterns In The Scholarly Impact Of Scientific Literature, Joshua Ange, Eric Godat, Rajani Sudan
SMU Journal of Undergraduate Research
Scientific communication is typically tied to promoting public engagement and interest in science, increasing scientific literacy, and playing an essential role in policymaking. The success of public communication of scientific findings is largely associated with secondary characteristics of research (e.g. the style of writing and presentation), rather than the primary content or research quality. But it is unclear to what extent the success of scientific literature intended for working scientists is influenced by those same secondary characteristics. Does the writing style of scientific articles impact their success in academic spheres? In this study, we explore the stylometric and formal characteristics …
Image Fusion Based On Deep Learning With Different Locations And Sizes Of Objects, Baneen Al-Kalabi, Tawfiq Al-Assadi
Image Fusion Based On Deep Learning With Different Locations And Sizes Of Objects, Baneen Al-Kalabi, Tawfiq Al-Assadi
Journal of Intelligent Informatics, Networking, and Cybersecurity
Multi-focus image fusion combines partially focused images into a single all-in-focus composite. Existing object-based methods assume precise spatial and scale alignment across source images, an assumption that frequently fails in Misaligned Multi-Focus Dataset scenarios due to camera displacement and focal length variation. This paper proposes a novel training-free, object-aware fusion framework to address this limitation through a five-stage pipeline: YOLOv8x detection, SAM2-L segmentation, LoFTR correspondence matching, a novel Scale-Aware Area Resize Algorithm, and GLCM-guided MSB/LSB bit-level fusion. The framework was evaluated on the EDMF benchmark (20 image pairs, synthetically modified to simulate Misaligned Multi-Focus Dataset shifts) and a Misaligned Multi-Focus …
A Blockchain-Integrated Federated Learning Model And Autoencoder-Based Feature Reduction For Improving Iot Intrusion Detection, Tahseen A. Wotaifi
A Blockchain-Integrated Federated Learning Model And Autoencoder-Based Feature Reduction For Improving Iot Intrusion Detection, Tahseen A. Wotaifi
Journal of Intelligent Informatics, Networking, and Cybersecurity
The rapid growth of Internet of Things (IoT) environments has brought forth a wealth of security challenges in detecting network intrusions in diverse and resource-restricted systems. Privacy, scalability, and single point of failure issues plague traditional centralized intrusion detection solutions. To address these challenges, the study proposes a secure and adaptive intrusion detection model using Federated Learning (FL) and Blockchain, augmented with autoencoder-based feature reduction. The ToN-IoT dataset is pre-processed, and then an unsupervised autoencoder is used to build informative low-dimensional feature representations. The processed data is deployed to various clients to mimic a real federated situation. Every client will …
Tapping Into The Ocean’S Hidden Energy: Feasibility Of A 5 Mw Otec Installation In North Bali, Widodo Setiyo Pranowo, Yani Permanawati, Gisela Malya Asoka Anindita, Agung Kurniawan, Albertus Sulaiman, Johar Setiyadi, Safri Burhanuddin, Ivonne Milichristi Radjawane, Hansan Park, Endro Sigit Kurniawan
Tapping Into The Ocean’S Hidden Energy: Feasibility Of A 5 Mw Otec Installation In North Bali, Widodo Setiyo Pranowo, Yani Permanawati, Gisela Malya Asoka Anindita, Agung Kurniawan, Albertus Sulaiman, Johar Setiyadi, Safri Burhanuddin, Ivonne Milichristi Radjawane, Hansan Park, Endro Sigit Kurniawan
Karbala International Journal of Modern Science
Indonesia is an archipelagic country surrounded by water and faces energy challenges due to the low use of renewable energy. Among the potential renewable options, marine energy is a particularly suitable resource given the country’s geographical nature. Based on this discussion, seawater temperature can be used as an alternative ocean thermal en-ergy known as ocean thermal energy conversion (OTEC) by using the difference in sea surface and deep-sea water temperatures. Therefore, this study aimed to examine OTEC installations in North Bali waters using closed-cycle OTEC calculations for a 5 MW system. Following this objective, we examined the water conditions by …
Ai-Powered Knowledge Engines As Research Infrastructure For Systematic Knowledge Discovery, Gary Welz
Ai-Powered Knowledge Engines As Research Infrastructure For Systematic Knowledge Discovery, Gary Welz
Publications and Research
This paper proposes knowledge engines as a framework for understanding how intelligent systems — both human and artificial — systematically discover, integrate, and generate knowledge. We argue that history’s greatest scientific minds functioned as knowledge engines, processing information through iterative cycles of ingestion, analysis, synthesis, and communication, guided by curiosity and willingness to challenge established beliefs.
We propose a taxonomy of nine integrated capabilities — ingestion, digestion, analysis, calculation, comparison, connection, association, analogy, and multimodal communication — that any serious knowledge engine must combine systematically. The argument is deliberately integrative: achieving ambitious research goals requires orchestrating all nine capabilities within …
Can An Ai System Be Creative? A Critical Perspective From Art And Engineering, Ivan Magrin-Chagnolleau
Can An Ai System Be Creative? A Critical Perspective From Art And Engineering, Ivan Magrin-Chagnolleau
Presidential Fellows Articles and Research
This paper examines the question of whether artificial intelligence (AI) systems can be creative, approached from the dual perspective of a researcher trained in electrical engineering, pattern recognition, machine learning, and neural networks, who has also spent most of his life engaged in the arts as actor, stage and film director, writer, composer, and visual artist, and in philosophy. Drawing on Margaret Boden’s foundational framework — both her three properties of creativity (novelty, surprise, and value) and her three types of creative processes (combinatorial, exploratory, and transformational) — the paper argues that AI systems are structurally incapable of creativity in …
One-For-All Community Search On Unseen Graphs, Mo Li, Zhaosong Zhao, Linlin Ding, Renata Borovica-Gajic, Zhongming Yao, Jianxin Li
One-For-All Community Search On Unseen Graphs, Mo Li, Zhaosong Zhao, Linlin Ding, Renata Borovica-Gajic, Zhongming Yao, Jianxin Li
Research outputs 2022 to 2026
Community search is a fundamental graph-based retrieval problem that aims to identify a query-dependent subgraph whose nodes exhibit strong internal connectivity. While recent learning-based methods improve retrieval effectiveness via graph representation learning, they follow a ''one-use-one-train'' paradigm that requires retraining or fine-tuning for each target graph, leading to high data dependency, high training costs, and limited generalization. To handle this, we propose OFA-CS, a ''one-for-all'' community search framework trained once on source datasets and directly deployed to arbitrary unseen graphs without retraining or fine-tuning, while preserving strong performance. Specifically, we introduce a Spectral-Aware Feature Alignment module to unify feature dimensionality …
Reproduction Beyond Benchmarks: Constbert And Colbert-V2 Across Backends And Query Distributions, Utshab Kumar Ghosh, Ashish David, Shubham Chatterjee
Reproduction Beyond Benchmarks: Constbert And Colbert-V2 Across Backends And Query Distributions, Utshab Kumar Ghosh, Ashish David, Shubham Chatterjee
Computer Science Faculty Research & Creative Works
Reproducibility must validate architectural robustness, not just numerical accuracy. We evaluate ColBERT-v2 and ConstBERT across five dimensions, finding that while ConstBERT reproduces within 0.05% MRR@10 on MS-MARCO, both models show a drop of 86-97% on long, narrative queries (TREC ToT 2025). Ablations prove this failure is architectural: performance plateaus at 20 words because the MaxSim operator's uniform token weighting cannot distinguish signal from filler noise. Furthermore, undocumented backend parameters create an 8-point gap due to ConstBERT's sparse centroid coverage, and fine-tuning with 3x more data actually degrades performance by up to 29%. We conclude that architectural constraints in multi-vector retrieval …
Depro: Understanding The Role Of Llms In Debugging Competitive Programming Code, Nabiha Parvez, Md Tanvin Sarkar Pallab, Mia Mohammad Imran, Tarannum Shaila Zaman
Depro: Understanding The Role Of Llms In Debugging Competitive Programming Code, Nabiha Parvez, Md Tanvin Sarkar Pallab, Mia Mohammad Imran, Tarannum Shaila Zaman
Computer Science Faculty Research & Creative Works
Debugging consumes a substantial portion of the software development lifecycle, yet researchers do not yet understand well the effectiveness of Large Language Models (LLMs) in this task. Competitive programming offers a rich benchmark for such evaluation, given its diverse problem domains and strict efficiency requirements. We present an empirical study of LLM-based debugging on competitive programming problems and introduce DePro, a test-case-driven approach that assists programmers by correcting existing code rather than generating new solutions. DePro combines brute-force reference generation, stress testing, and iterative LLM-guided refinement to efficiently identify and resolve errors. Experiments on 13 faulty user submissions from Codeforces …
Llm-Enabled Open-Source Systems In The Wild: An Empirical Study Of Vulnerabilities In Github Security Advisories, Fariha Tanjim Shifat, Hariswar Baburaj, Ce Zhou, Jaydeb Sarker, Mia Mohammad Imran
Llm-Enabled Open-Source Systems In The Wild: An Empirical Study Of Vulnerabilities In Github Security Advisories, Fariha Tanjim Shifat, Hariswar Baburaj, Ce Zhou, Jaydeb Sarker, Mia Mohammad Imran
Computer Science Faculty Research & Creative Works
Large language models (LLMs) are increasingly embedded in open-source software (OSS) ecosystems, creating complex interactions among natural language prompts, probabilistic model outputs, and execution-capable components. However, it remains unclear whether traditional vulnerability disclosure frameworks adequately capture these model-mediated risks. To investigate this, we analyze 295 GitHub Security Advisories published between January 2025 and January 2026 that reference LLM-related components, and we manually annotate a sample of 100 advisories using the OWASP Top 10 for LLM Applications 2025.We find no evidence of new implementation-level weakness classes specific to LLM systems. Most advisories map to established CWEs, particularly injection and deserialization weaknesses. …
Improving Cancer Diagnosis And Patient Outcomes With Deep Learning Models, Mariana Arriz-Jorquiera
Improving Cancer Diagnosis And Patient Outcomes With Deep Learning Models, Mariana Arriz-Jorquiera
USF Tampa Graduate Theses and Dissertations
Cancer care depends on timely and reliable decisions, from detection and diagnosis to treatment planning and patient monitoring. These decisions are often made under uncertainty because medical images and healthcare data may be noisy, incomplete, or difficult to interpret. In breast cancer imaging, ultrasound is widely used because it is safe, accessible, and complementary to other imaging modalities. However, variations in image quality, acquisition conditions, and noise can obscure lesion boundaries and texture, affecting human interpretation and artificial intelligence reliability. This dissertation develops deep learning, image-analysis, and optimization methods to improve healthcare decisions under imperfect information. Its primary focus is …
From Automation To Adjudication: Evaluating The Role Of Artificial Intelligence In Dispute Settlement, Karem Sayed Aboelazm, Muayad Ahmad Obeidat, Raghda Raafat, Nada Zuhair Alfil, Fady Tawakol
From Automation To Adjudication: Evaluating The Role Of Artificial Intelligence In Dispute Settlement, Karem Sayed Aboelazm, Muayad Ahmad Obeidat, Raghda Raafat, Nada Zuhair Alfil, Fady Tawakol
All Works
This paper explores the evolving transition from automation to adjudication by examining the role of artificial intelligence (AI) in dispute settlement processes. It assesses how AI can enhance procedural efficiency, support judicial reasoning, and improve access to justice. Adopting a qualitative and interpretive approach, the study analyzes academic scholarship, policy frameworks, and comparative international practices to understand the integration of AI within judicial and quasi-judicial settings (Abedi et al., 2025). The findings suggest that while AI significantly improves administrative processes and provides valuable decision-support tools, it also raises critical concerns regarding algorithmic bias, lack of transparency, and the risk of …
"Todo: Fix The Mess Gemini Created": Towards Understanding Genai-Induced Self-Admitted Technical Debt, Abdullah Al Mujahid, Mia Mohammad Imran
"Todo: Fix The Mess Gemini Created": Towards Understanding Genai-Induced Self-Admitted Technical Debt, Abdullah Al Mujahid, Mia Mohammad Imran
Computer Science Faculty Research & Creative Works
As large language models (LLMs) such as ChatGPT, Copilot, Claude, and Gemini become integrated into software development workflows, developers increasingly leave traces of AI involvement in their code comments. Among these, some comments explicitly acknowledge both the use of generative AI and the presence of technical shortcomings. Analyzing 6,540 LLM-referencing code comments from public Python and JavaScript-based GitHub repositories (November 2022-July 2025), we identified 81 that also self-admit technical debt (SATD). Developers most often describe postponed testing, incomplete adaptation, and limited understanding of AI-generated code, suggesting that AI assistance affects both when and why technical debt emerges. We term GenAI-Induced …
Assessing Flaws In Captcha Security Through Progress In Ai, Jaydon Stanislowski
Assessing Flaws In Captcha Security Through Progress In Ai, Jaydon Stanislowski
Scholarly Horizons: University of Minnesota, Morris Undergraduate Journal
Protecting the internet from the threat of malicious bot activity is an important problem as AI tools become more powerful and commonplace over time. To that end, security measures are employed across websites in the form of CAPTCHAs, short challenges designed to identify and block fake web traffic. Yet, they become less effective over time as AI becomes more powerful, and thus more capable of solving them. This paper examines recent research on the threat to CAPTCHA security posed by current AI models and how this security can be reinforced over time, focusing primarily on Google’s reCAPTCHA v3.