Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Engineering (5391)
- Computer Engineering (4370)
- Operations Research, Systems Engineering and Industrial Engineering (4244)
- Numerical Analysis and Scientific Computing (4163)
- Systems Science (3895)
-
- Social and Behavioral Sciences (982)
- Databases and Information Systems (629)
- Medicine and Health Sciences (605)
- Data Science (527)
- Theory and Algorithms (485)
- Business (447)
- Graphics and Human Computer Interfaces (414)
- Electrical and Computer Engineering (402)
- Education (382)
- Software Engineering (373)
- Arts and Humanities (337)
- Public Affairs, Public Policy and Public Administration (297)
- Information Security (293)
- Other Computer Sciences (274)
- Life Sciences (270)
- Law (241)
- Statistics and Probability (203)
- Medical Specialties (179)
- Library and Information Science (165)
- Psychology (157)
- Robotics (157)
- Programming Languages and Compilers (154)
- Institution
-
- China Simulation Federation (3880)
- Singapore Management University (1897)
- Old Dominion University (641)
- San Jose State University (277)
- MBZUAI (233)
-
- City University of New York (CUNY) (184)
- Technological University Dublin (157)
- Air Force Institute of Technology (137)
- Chapman University (125)
- California Polytechnic State University, San Luis Obispo (116)
- Chinese Academy of Sciences (113)
- University of Arkansas, Fayetteville (102)
- Lindenwood University (97)
- Edith Cowan University (92)
- Embry-Riddle Aeronautical University (92)
- University of Nebraska - Lincoln (78)
- University of Kentucky (76)
- University of South Florida (71)
- University of Nevada, Las Vegas (63)
- Dartmouth College (62)
- Clemson University (60)
- University of Denver (59)
- University of Michigan Law School (57)
- Utah State University (57)
- The Texas Medical Center Library (54)
- Thomas Jefferson University (54)
- New Jersey Institute of Technology (53)
- University of Malaya (50)
- Purdue University (48)
- Missouri University of Science and Technology (47)
- Keyword
-
- Artificial intelligence (778)
- Machine learning (685)
- Deep learning (435)
- Artificial Intelligence (359)
- Machine Learning (359)
-
- AI (239)
- Deep Learning (201)
- Simulation (160)
- Computer vision (157)
- Reinforcement learning (140)
- Generative AI (134)
- Neural networks (128)
- Large language models (109)
- Natural language processing (108)
- Robotics (97)
- Natural Language Processing (90)
- ChatGPT (89)
- Path planning (89)
- Optimization (82)
- Large Language Models (77)
- Computer Vision (76)
- Classification (71)
- Neural network (67)
- Neural Networks (65)
- Virtual reality (64)
- Reinforcement Learning (63)
- Computer Science (59)
- Cybersecurity (59)
- Genetic algorithm (58)
- Algorithms (57)
- Publication Year
- Publication
-
- Journal of System Simulation (3880)
- Research Collection School Of Computing and Information Systems (1664)
- Master's Projects (248)
- Theses and Dissertations (183)
- Computer Science Faculty Publications (125)
-
- Bulletin of Chinese Academy of Sciences (Chinese Version) (113)
- Faculty Scholarship (108)
- Publications and Research (99)
- Computer Vision Faculty Publications (98)
- Master's Theses (96)
- Conference papers (92)
- Electrical & Computer Engineering Faculty Publications (90)
- Machine Learning Faculty Publications (86)
- Electronic Theses and Dissertations (85)
- Faculty Publications (77)
- Dissertations (70)
- Research outputs 2022 to 2026 (64)
- USF Tampa Graduate Theses and Dissertations (59)
- Dissertations and Theses Collection (Open Access) (57)
- Articles (54)
- Dissertations, Theses, and Capstone Projects (53)
- Theses and Dissertations--Computer Science (48)
- Natural Language Processing Faculty Publications (46)
- Teaching and Generative AI: Pedagogical Possibilities and Productive Tensions (46)
- Graduate Theses and Dissertations (44)
- Open Access Theses & Dissertations (42)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (40)
- Theses (40)
- Electrical & Computer Engineering Theses & Dissertations (39)
- Publications (39)
- Publication Type
- File Type
Articles 541 - 570 of 11180
Full-Text Articles in Artificial Intelligence and Robotics
Rethinking News Classification Through A Multi-Dimensional Framework, Luana De Jesus Ferreira
Rethinking News Classification Through A Multi-Dimensional Framework, Luana De Jesus Ferreira
Honors Theses
This thesis proposes a multi-dimensional framework for news classification that evaluates articles across three independent dimensions: headline accuracy, language neutrality, and content reliability. These dimensions produce both a continuous reliability score and a five-tier interpretive scale, while additionally classifying articles by genre and topic. To operationalize this framework, a structured annotation protocol was developed and applied to a dataset of 373 news articles drawn from 79 outlets spanning a wide range of contemporary media ecosystem. A binary Logistic Regression classifier trained on the ISOT Fake News Dataset was then evaluated against this dataset to examine how a model trained on …
Deep Learning-Based Automated Pneumonia Detection From Chest X-Rays: A Comparative Study Of Custom Cnn And Transfer Learning Architectures, Ahmed Sajim
Honors Theses
Pneumonia is a leading global cause of mortality, claiming approximately 2.5 million lives an-nually and placing exceptional diagnostic pressure on radiologists in resource-limited settings. Manual interpretation of chest X-ray (CXR) images is time-consuming, subject to inter-observer variability, and limited by radiologist availability. This thesis presents a systematic investiga-tion into deep learning-based automated pneumonia detection comparing five convolutional neural network (CNN) architectures: a custom-designed 2D CNN and four pretrained transfer learning models—ResNet, DenseNet, MobileNet, and VGG19.
A targeted data augmentation pipeline addresses the severe class imbalance in the Kag-gle Chest X-Ray Pneumonia dataset, expanding the Normal class from 1,583 to 9,495 …
Sentra, David Aguilar
Sentra, David Aguilar
Posters - 2026
In the fast-paced space of event organization, fostering continuous collaboration among participants is essential. However, organizers often lose valuable time monitoring multiple, disconnected systems once an event is underway. Enter Sentra: an all-in-one Discord bot tailored specifically for weekend events like hackathons. Sentra bridges the gap between participants and organizers by consolidating seamless team matchmaking, robust support ticketing, and automated AI moderation into a single, unified interface.
Synapse, Nicolas Diaz, Alexander Murphy, Sonia Cerrillo, Naomi Ramirez, Jesse Kemmer
Synapse, Nicolas Diaz, Alexander Murphy, Sonia Cerrillo, Naomi Ramirez, Jesse Kemmer
Posters - 2026
People tend to accumulate a great deal of notes throughout their lives with no coherent way to organize them. Even with the built-in notes app, the notes eventually accumulate until it becomes borderline impossible to find what is needed. Our proposed solution is Synapse, an LLM powered notes app with a tagging system that allows notes to be sorted by topic. The LLM will be able to read the user's notes and recommend tags
Ai Dependence And Its Impact On Human Decision-Making Quality And Supply Chain Efficiency, Jesus Salazar, Leonardo Fabbri, Axel Villegas
Ai Dependence And Its Impact On Human Decision-Making Quality And Supply Chain Efficiency, Jesus Salazar, Leonardo Fabbri, Axel Villegas
Posters - 2026
- Artificial Intelligence (AI) is transforming supply chain management by enabling:
- Data-driven decision-making
- Improved forecasting accuracy
- Enhanced operational efficiency (Choudhary et al., 2023; Ivanov & Dolgui, 2021)
- AI applications such as predictive analytics support:
- Inventory optimization, Logistics planning
- Procurement decisions in real time
- However, increasing reliance on AI introduces risks:
- Automation bias (over-trusting AI outputs)
- Reduced human critical thinking
- Overdependence on algorithmic recommendations (Raisch & Krakowski, 2021)
- This study examines the dual impact of AI dependence on:
- Decision-making quality
- Supply chain efficiency
- Objective:
- Identify whether AI improves performance or reduces human effectiveness
- Determine the optimal balance between AI support and human …
From Attention To Reasoning: Beyond Accuracy In Multimodal Ai, Wayner Barrios
From Attention To Reasoning: Beyond Accuracy In Multimodal Ai, Wayner Barrios
Dartmouth College Ph.D Dissertations
Multimodal large language models have achieved impressive performance on vision-language benchmarks by integrating visual encoders with large language models. Yet a critical gap persists between benchmark accuracy and genuine multimodal understanding: current evaluation frameworks assess performance by final answers alone, rewarding confident predictions while leaving systematic reasoning failures undetected.
This thesis addresses this gap through a unified framework that progresses from understanding to reasoning, using video as the most comprehensive multimodal testbed. Video inherently combines vision, audio, and language with temporal dynamics and massive token redundancy; techniques developed for video's comprehensive challenges transfer naturally to simpler multimodal tasks.
On understanding …
A Real Account Of Deep Fakes, Benjamin L.W Sobel
A Real Account Of Deep Fakes, Benjamin L.W Sobel
Michigan Law Review
Laws regulating pornographic deepfakes are written to prohibit “digital forgeries,” “false” images, or media “indistinguishable” from “authentic” recordings. Yet the typical anti-deepfake law covers materials that aren’t forgeries, aren’t false, and that reasonable observers can easily distinguish from authentic recordings. Though drafted as if they regulate statements of fact, anti-deepfake laws actually target certain outrageous depictions per se—and rightly so, because pornographic deepfakes cause harm irrespective of their truth or falsity. However, the inapposite language of facts results in statutes with crucial ambiguities. Moreover, because anti-deepfake laws ban outrageous depictions irrespective of the factual assertions they make, they differ fundamentally …
Enhancing Financial Audit Operations Through Ai Anomaly Detection, Nadya Cousin, Rebekah Garza, Talisa Gomez
Enhancing Financial Audit Operations Through Ai Anomaly Detection, Nadya Cousin, Rebekah Garza, Talisa Gomez
Posters - 2026
❖ Financial auditing plays a critical role in ensuring accuracy, regulatory compliance, and fraud detection in financial reporting
❖ Traditional audit approaches rely heavily on sampling and manual review processes, limiting their ability to scale with increasing data complexity
❖ The rapid growth of high-volume, high-velocity financial data (big data) has exposed significant limitations in traditional auditing, including:
- Incomplete data coverage
- Delayed anomaly detection
- Increased risk of material misstatements
❖ These limitations create a need for scalable, automated, and data-driven audit solutions
❖ Artificial Intelligence (AI), particularly anomaly detection models, enables:
- Full-population testing
- Real-time pattern recognition
- Proactive risk identification
The Psychology Behind Ai-Generated Phishing And Social Engineering Attacks, A’Shya Reynolds
The Psychology Behind Ai-Generated Phishing And Social Engineering Attacks, A’Shya Reynolds
School of Cybersecurity Master's Level Projects and Papers
Cybercrime has evolved significantly with the integration of artificial intelligence (AI), transforming traditional phishing and social engineering attacks into highly sophisticated and personalized threats. While early phishing attempts relied on generic messaging and low success rates, modern AI-driven attacks leverage advanced data analytics, natural language processing, and behavioral prediction to manipulate victims more effectively.
This research examines how cybercriminals utilize AI to enhance psychological manipulation techniques in phishing and social engineering attacks, increasing victim susceptibility. Drawing from interdisciplinary literature in cybersecurity and psychology, this study explores key psychological mechanisms, including cognitive biases, emotional triggers, and decision-making processes that influence victim …
Deep Learning Based Approaches For Low Cost Defense Detection, Adele J. Noel-Rickert
Deep Learning Based Approaches For Low Cost Defense Detection, Adele J. Noel-Rickert
All NMU Master's Theses
Pulmonary fibrosis is a progressive interstitial lung disease characterized by the accumulation of fibrotic tissue within the lungs, leading to impaired respiratory function and reduced quality of life. Early detection is important for disease management; however, accurate diagnosis often relies on high-resolution computed tomography (CT), which may not be accessible in all clinical settings. Chest radiography provides a lower-cost and widely available imaging modality, but interpretation of chest X-rays for fibrotic disease can be challenging due to subtle radiographic patterns and overlapping anatomical structures. This thesis investigates the use of multimodal deep learning techniques to assist in pul- monary fibrosis …
Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy
Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy
Publications and Research
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools—training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures—encodes a set of assumptions that systematically disadvantages speakers of underrepresented languages before a single model is trained. This paper examines those assumptions through the lens of Bengali, one of the world’s most widely spoken languages with roughly 285 million speakers (Ethnologue, 2025; International Communication and Leadership School, 2026), and the structural barriers that emerge when attempting to build AI-assisted educational tools for Bengali-speaking learners in low-connectivity …
Generative Ai In Enterprises: Optimizing Applications With Large Language Models, Donghao Huang
Generative Ai In Enterprises: Optimizing Applications With Large Language Models, Donghao Huang
Dissertations and Theses Collection (Open Access)
This dissertation investigates how to deploy Large Language Models (LLMs) effectively in enterprise settings, where accuracy, reliability, cost, privacy, and operational constraints often matter more than benchmark performance alone. Drawing on seventeen peer-reviewed publications (eleven published and six accepted for publication), the work develops and validates optimization strategies across three connected themes: retrieval-augmented generation (RAG), agentic AI for workflow automation, and deployment guidelines for real-world enterprise environments.
First, we study RAG optimization through systematic evaluation of open and proprietary models, highlighting conditions under which efficient open-weight models can match or exceed proprietary alternatives. To address a pervasive failure mode in …
Low-Complexity Structured Neural Networks And Their Usage In Image And Signal Processing, Adam Kuzmicki
Low-Complexity Structured Neural Networks And Their Usage In Image And Signal Processing, Adam Kuzmicki
Doctoral Dissertations and Master's Theses
Conventional neural networks face significant challenges due to high computational costs, large parameter counts, and reliance on backpropagation, which restricts their application in resource-constrained and real-time settings. To address these challenges, this thesis proposes three structured neural network (NN) architectures grounded in the theories of sparse and self-contained factorizations of transforms, with applications to image compression, reconstruction, classification, encryption, and also adaptive wideband multi-beam beamforming. The first neural network architecture, named DCTrix-Net, replaces conventional spatial con- volution with highly sparse factorization of the discrete Cosine transform (DCT) complemented by Toeplitz-structured weight initialization, achieving at least 97% FLOP reduction over CNNs, …
Ai Interpretability In Healthcare Communication, Ananya Jeyappragash
Ai Interpretability In Healthcare Communication, Ananya Jeyappragash
Dartmouth College Master’s Theses
Artificial intelligence has increasingly been adopted in healthcare, largely for specialized tasks and under significant human oversight. The use of large black-box systems raises important concerns about transparency in high-stakes environments such as clinical decision-making. Clinical communication is fundamentally human-centered, and failures in judgment can have serious consequences for patient care. Overestimating the reasoning abilities of large language models may lead to undue trust in fabricated or “hallucinated” outputs, while rejecting AI-assisted tools altogether may preserve inefficient workflows and contribute to missed or delayed diagnoses. These concerns reflect a broader tradeoff between accuracy and interpretability: although more complex models may …
Generative Artificial Intelligence With A Human Touch: Building Hana, Conrad Johnson
Generative Artificial Intelligence With A Human Touch: Building Hana, Conrad Johnson
Faculty Scholarship
This Essay examines how generative artificial intelligence (GenAI) can be integrated into legal education and public interest law practice in a way that meaningfully enhances — rather than diminishes — human judgment, professional responsibility, and access to justice. Drawing on the experience of Columbia Law School’s Lawyering in the Digital Age Clinic, the Essay situates GenAI within an experiential pedagogy that emphasizes competence, ethical awareness, and collaborative problem-solving. It argues that law students and lawyers must move beyond a passive or uncritical use of GenAI tools; toward a deeper understanding of how these systems operate, the risks they pose, and …
Match-A-Fit, Adan Diaz De Leon, Juan Marco Saca Dada, Brianna Mendoza, Arsalan Kataneh, Theophile Nsabimana, Pedro Jacobo
Match-A-Fit, Adan Diaz De Leon, Juan Marco Saca Dada, Brianna Mendoza, Arsalan Kataneh, Theophile Nsabimana, Pedro Jacobo
Presentations - 2026
Welcome to Match-a-Fit! Match-a-Fit is an iOS application that allows the user to create a digital closet by uploading images of their clothing items. With AI, the program can generate outfits based on the digital closet, the time, and the occasion. Match-a-Fit’s purpose is designed to help users who struggle to get ready, run out of time, or can’t decide on an outfit, by easily generating outfit options based on the occasion.
Comparative Analysis Of Mlp And Cnn Models For Cardiac Arrhythmia Classification, Veltman Okey-Ejowhor, Vahid Emamian
Comparative Analysis Of Mlp And Cnn Models For Cardiac Arrhythmia Classification, Veltman Okey-Ejowhor, Vahid Emamian
Posters - 2026
Electrocardiogram (ECG) is a record of the electric activity of the heart over time. ECG analysis plays a pivotal role in diagnosing critical heart conditions. Significant developments have been made in the realm of deep learning and applied artificial intelligence. These deep learning models have been utilized heavily because of their ability to analyze deep morphological features of each signal. The model architecture used in this study is a convolutional neural network (CNN) combined with a multi-layered perceptron (MLP). The MLP acts as an input filter that classifies normal heartbeat signals from abnormal. The CNN is the second filter in …
Using Ai-Based Predictive Scheduling To Improve Patient Flow And Reduce Wait Times In Healthcare Clinics, Oscar Martinez
Using Ai-Based Predictive Scheduling To Improve Patient Flow And Reduce Wait Times In Healthcare Clinics, Oscar Martinez
Posters - 2026
- Healthcare systems face increasing challenges in patient access and wait times
- Average wait times for specialist care continue to rise, creating:
- Delays in treatment
- Reduced patient satisfaction
- Increased system inefficiencies (Sanford, 2025)
- A major contributor is operational bottlenecks, defined as:
- Points of congestion that slow or disrupt service flow
- Hospitals typically operate under process layouts, which:
- Handle diverse patient needs
- Reduce specialization efficiency
- Contributing factors to bottlenecks:
- Physician shortages and burnout
- Administrative burden
- Inefficient scheduling systems (Moura & Pinho, 2025)
- AI offers potential solutions through:
- Predictive scheduling
- Automation of administrative processes
- Data-driven optimization of patient flow
Optimizing Retail Grocery Inventory Using Ai And Large Language Models: Evidence On Forecast Accuracy, Waste Reduction, And Cost Efficiency, Robert Miller, Stephen Garcia, Brandon Ermis
Optimizing Retail Grocery Inventory Using Ai And Large Language Models: Evidence On Forecast Accuracy, Waste Reduction, And Cost Efficiency, Robert Miller, Stephen Garcia, Brandon Ermis
Posters - 2026
Aim: To evaluate how AI and LLMs improve forecasting accuracy, reduce waste, and enhance inventory decision-making
Beyond Technology Acceptance: Ai Adoption In The United States And Korean Accounting Firms, Madeline Ortega, Brenda Vazquez Saavedra
Beyond Technology Acceptance: Ai Adoption In The United States And Korean Accounting Firms, Madeline Ortega, Brenda Vazquez Saavedra
Posters - 2026
A comparison of AI adoption between the United States and South Korea
Think First, Chatgpt Later: Guiding Human-Ai Collaboration For Learning Gains In Independent Human Creativity, Sarah Shi Hui Wong, Sophia Xuefei Qiu
Think First, Chatgpt Later: Guiding Human-Ai Collaboration For Learning Gains In Independent Human Creativity, Sarah Shi Hui Wong, Sophia Xuefei Qiu
Research Collection School of Social Sciences
Generative artificial intelligence (AI) tools such as ChatGPT can boost creative performance, but do these boosts translate into learning gains? This study examined whether the benefits of ChatGPT for creativity persist even when its assistance is removed, and how people can effectively use ChatGPT to enhance their learning and independent creativity. University students (N = 196) solved a creative product improvement task either independently (human-only group) or using ChatGPT freely (general-AI group) or using ChatGPT in a guided way (regulated-AI group). Specifically, the regulated-AI group used a novel “think first, ChatGPT later” approach—they first generated their own ideas, then collaborated …
Pushing High-Performance Private Inference Towards Resource-Constrained Edge Clients, Xiangrui Xu
Pushing High-Performance Private Inference Towards Resource-Constrained Edge Clients, Xiangrui Xu
Computer Science Theses & Dissertations
The widespread adoption of Machine Learning as a Service (MLaaS) has enabled resource constrained edge clients, such as mobile and IoT devices, to leverage powerful deep learning mod els hosted on the cloud. However, this paradigm introduces critical privacy challenges regarding the client’s sensitive input data and the server’s proprietary model parameters. While cryptographic techniques like Homomorphic Encryption (HE) and Multi-Party Computation (MPC) enable Private Inference (PI), existing frameworks impose prohibitive computational and communication overheads that render them impractical for edge deployment. This dissertation introduces three novel frameworks—SPOT, LUTless, and PrivShap—to systematically address the efficiency bottlenecks of PI in edge …
Knowledge Distillation From A Large Vision-Language Model To Compact Students For Architectural Floor Plan Understanding, Kiran Silwal
Knowledge Distillation From A Large Vision-Language Model To Compact Students For Architectural Floor Plan Understanding, Kiran Silwal
Honors Theses
In this research, the use of a large vision-language model to train smaller, deployable models for architectural floor plan question answering is investigated. Reading a floor plan today requires either a human expert or a paid query to a proprietary model, and neither option is practical for real-estate platforms that must process thousands of units at scale. To address this problem, a knowledge distillation approach is employed in which a large teacher model (GPT-4.1-mini) generates labeled question-answer pairs from floor plan images, and smaller student models learn from those labels. The teacher produced 37,027 labeled pairs from 12,343 floor plan …
Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng
Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng
Research Collection Lee Kong Chian School Of Business
Existing research on AI cultural biases predominantly focuses on Western models, overlooking critical gaps in non-Western models. We conduct a comparative analysis of AI models – ChatGPT (U.S. developed) and ErnieBot (China developed) – from different cultures to investigate how corresponding cultural biases manifest in their outputs. Additionally, we examine how cultural alignment between human users and AI models impacts their collaborative creative performance and the underlying psychological mechanisms. In Study 1, multi-choice prompt with zero-shot technique was used to evaluate cultural biases in four widely used AI models – ChatGPT-3.5/4, ErnieBot-3.5/4 – comparing their responses to established cultural psychometric …
Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong
Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong
Research Collection School Of Computing and Information Systems
Conventionally, human intuition defines vision as a modality of passive optical sensing, relying on ambient light to perceive the environment. However, active optical sensing, which involves emitting and receiving signals, offers unique advantages by capturing both radiometric and geometric properties of the environment, independent of external illumination conditions. This work focuses on advancing active optical sensing using Light Detection and Ranging (LiDAR), which captures intensity data, enabling the estimation of surface reflectance that remains invariant under varying illumination. Such properties are crucial for robotic perception tasks, including detection, recognition, segmentation, and Simultaneous Localization and Mapping (SLAM). A key challenge with …
Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo
Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo
Research Collection School Of Computing and Information Systems
Penetration testing is essential for identifying vulnerabilities in web applications before real adversaries can exploit them. Recent work has explored automating this process with Large Language Model (LLM)-powered agents, but existing approaches either rely on a single generic agent that struggles in complex scenarios or narrowly specialized agents that cannot adapt to diverse vulnerability types. We therefore introduce PenForge, a framework that dynamically constructs expert agents during testing rather than relying on those prepared beforehand. By integrating automated reconnaissance of potential attack surfaces with agents instantiated on the fly for context-aware exploitation, PenForge achieves a 30.0% exploit success rate (12/40) …
Semat: Semantic Enhanced Natural Image Interactive Matting, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Qianru Sun, Yang Tang, Bo Li, Pan Zhou
Semat: Semantic Enhanced Natural Image Interactive Matting, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Qianru Sun, Yang Tang, Bo Li, Pan Zhou
Research Collection School Of Computing and Information Systems
Recent approaches attempt to adapt powerful interactive segmentation models, such as SAM, to interactive matting and fine-tune the models based on synthetic matting datasets. However, models trained on synthetic data fail to generalize to complex and occlusion scenes. We address this challenge by proposing a new matting dataset based on the COCO dataset, namely COCO-Matting. It selects real-world complex images from COCO and converts semantic segmentation masks to matting labels. The built COCO-Matting comprises an extensive collection of 36,980 human instance-level alpha mattes in complex natural scenarios. Furthermore, existing SAM-based matting methods extract intermediate features and masks from a frozen …
Stacked From One: Multi-Scale Self-Injection For Context Window Extension, Wei Han, Pan Zhou, Shuicheng Yan
Stacked From One: Multi-Scale Self-Injection For Context Window Extension, Wei Han, Pan Zhou, Shuicheng Yan
Research Collection School Of Computing and Information Systems
The limited context window of contemporary large language models (LLMs) remains a primary bottleneck for their broader application across diverse domains. Although continual pre-training on long-context data offers a straightforward solution, it incurs prohibitive data acquisition and computational costs. To address this challenge, we propose SHAREDLLM, a novel framework based on multi-grained context compression and query-aware information acquisition. SHAREDLLM comprises two stacked short-context LLMs: a lower model serving as a compressor and an upper model acting as a decoder. The lower model compresses long inputs into compact, multi-grained representations, which are then forwarded to the upper model for context-aware processing. …
Distributional Vision-Language Alignment By Cauchy-Schwarz Divergence, Wenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu, Jiayi Shen, Jan-Jakob Sonke, Stratis Gavves
Distributional Vision-Language Alignment By Cauchy-Schwarz Divergence, Wenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu, Jiayi Shen, Jan-Jakob Sonke, Stratis Gavves
Research Collection School Of Computing and Information Systems
Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairwise samples across modalities while overlooking distributional differences. In addition, InfoNCE has inherent conflict in terms of alignment and uniformity in multimodality, leading to suboptimal alignment with modality gaps. To overcome the limitations, we propose CS-Aligner, a novel framework that performs distributional vision-language alignment by integrating Cauchy-Schwarz (CS) divergence with mutual information. CS-Aligner captures both the global distribution information of each modality and the pairwise semantic relationships. We find that the CS divergence …
Bridging Draft Policy Misalignment: Group Tree Optimization For Speculative Decoding, Shijing Hu, Jingyang Li, Zhihui Lu, Pan Zhou
Bridging Draft Policy Misalignment: Group Tree Optimization For Speculative Decoding, Shijing Hu, Jingyang Li, Zhihui Lu, Pan Zhou
Research Collection School Of Computing and Information Systems
Speculative decoding accelerates large language model (LLM) inference by letting a lightweight draft model propose multiple tokens that the target model verifies in parallel. Yet existing training objectives optimize only a single greedy draft path, while decoding follows a tree policy that re-ranks and verifies multiple branches. This draft policy misalignment limits achievable speedups. We introduce Group Tree Optimization (GTO), which aligns training with the decoding-time tree policy through two components: (i) Draft Tree Reward, a sampling-free objective equal to the expected acceptance length of the draft tree under the target model, directly measuring decoding performance; (ii) Group-based Draft Policy …