Open Access. Powered by Scholars. Published by Universities.®

Artificial Intelligence and Robotics Commons™

Open Access. Powered by Scholars. Published by Universities.®

11,188 Full-Text Articles 24,563 Authors 5,758,021 Downloads 274 Institutions

All Articles in Artificial Intelligence and Robotics

Faceted Search

11,188 full-text articles. Page 29 of 542.

Think First, Chatgpt Later: Guiding Human-Ai Collaboration For Learning Gains In Independent Human Creativity, Sarah Shi Hui WONG, Sophia Xuefei QIU 2026 Singapore Management University

Think First, Chatgpt Later: Guiding Human-Ai Collaboration For Learning Gains In Independent Human Creativity, Sarah Shi Hui Wong, Sophia Xuefei Qiu

Research Collection School of Social Sciences

Generative artificial intelligence (AI) tools such as ChatGPT can boost creative performance, but do these boosts translate into learning gains? This study examined whether the benefits of ChatGPT for creativity persist even when its assistance is removed, and how people can effectively use ChatGPT to enhance their learning and independent creativity. University students (N = 196) solved a creative product improvement task either independently (human-only group) or using ChatGPT freely (general-AI group) or using ChatGPT in a guided way (regulated-AI group). Specifically, the regulated-AI group used a novel “think first, ChatGPT later” approach—they first generated their own ideas, then collaborated …


Learning Techniques In Prediction Of Functional Epigenomic Events, Mohammad Shiri 2026 Old Dominion University

Learning Techniques In Prediction Of Functional Epigenomic Events, Mohammad Shiri

Computer Science Theses & Dissertations

Accurately predicting functional epigenomic events from DNA sequences is critical to understanding gene regulation and the functional impact of non-coding variants. Despite considerable progress, critical challenges hamper the effectiveness and efficiency of existing deep learning approaches. These challenges include negative transfer in multi-task learning (MTL), suboptimal network architectures, and pervasive label noise, particularly the positive-unlabeled problem arising from data sparsity in single-cell assays. This dissertation presents a cohesive framework of novel learning techniques to effectively address these challenges. First, a highly scalable task grouping framework is presented to mitigate negative transfer in deep MTL. This method clusters tasks based on …


Behavioral, System, And Informational Cyberattacks: A Human-In-The-Loop Driving Simulator Experiment, Samuel Petkac 2026 Old Dominion University

Behavioral, System, And Informational Cyberattacks: A Human-In-The-Loop Driving Simulator Experiment, Samuel Petkac

Psychology Theses & Dissertations

Advanced technologies such as sensors and AI/ML algorithms have enabled increasing levels of automated driving system that detects, responds, and even predicts changes in a driving environment supported by wireless connectivity to nearby vehicles and infrastructure. Such connected and automated vehicles (CAVs) can be particularly vulnerable to cyberattacks targeting not only infotainment systems but also firmware and other applications, critically compromising driver safety. As we anticipate a “mixed” traffic where vehicles with various levels of automated technologies share the road for the foreseeable future, it is urgent to systematically examine types of possible cyberattacks and control human behaviors in such …


Personality Predictors Of Cybersecurity Vulnerability: Insights From Self-Reports And Stimulated Threat Scenarios, Saroja Roy Grandhi 2026 Old Dominion University

Personality Predictors Of Cybersecurity Vulnerability: Insights From Self-Reports And Stimulated Threat Scenarios, Saroja Roy Grandhi

Psychology Theses & Dissertations

In this cyber dependent and enabled era, understanding the role of human factors in digital security is essential. This study investigates the relationship between Big-Five personality traits and cybersecurity behaviors by examining both self-reported and stimulated behaviors in security threat scenarios. Participants completed validated questionnaires to report their personality traits, cybersecurity practices and engage in task-based stimulations to capture behaviors such as phishing detection, password creation, and response to security alerts. The study tested whether higher conscientiousness, openness, and agreeableness would be associated with stronger cybersecurity practices and smaller discrepancies between self-reported and observed behaviors. And, whether greater extraversion and …


Pushing High-Performance Private Inference Towards Resource-Constrained Edge Clients, Xiangrui Xu 2026 Old Dominion University

Pushing High-Performance Private Inference Towards Resource-Constrained Edge Clients, Xiangrui Xu

Computer Science Theses & Dissertations

The widespread adoption of Machine Learning as a Service (MLaaS) has enabled resource constrained edge clients, such as mobile and IoT devices, to leverage powerful deep learning mod els hosted on the cloud. However, this paradigm introduces critical privacy challenges regarding the client’s sensitive input data and the server’s proprietary model parameters. While cryptographic techniques like Homomorphic Encryption (HE) and Multi-Party Computation (MPC) enable Private Inference (PI), existing frameworks impose prohibitive computational and communication overheads that render them impractical for edge deployment. This dissertation introduces three novel frameworks—SPOT, LUTless, and PrivShap—to systematically address the efficiency bottlenecks of PI in edge …


Super Lidar Intensity For Robotic Perception, Wei GAO, Jie ZHANG, Mingle ZHAO, Zhiyuan ZHANG, Shu KONG, Maani GHAFFARI, Dezhen SONG, Chengzhong XU, Hui KONG 2026 Singapore Management University

Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong

Research Collection School Of Computing and Information Systems

Conventionally, human intuition defines vision as a modality of passive optical sensing, relying on ambient light to perceive the environment. However, active optical sensing, which involves emitting and receiving signals, offers unique advantages by capturing both radiometric and geometric properties of the environment, independent of external illumination conditions. This work focuses on advancing active optical sensing using Light Detection and Ranging (LiDAR), which captures intensity data, enabling the estimation of surface reflectance that remains invariant under varying illumination. Such properties are crucial for robotic perception tasks, including detection, recognition, segmentation, and Simultaneous Localization and Mapping (SLAM). A key challenge with …


Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui HUANG, Jieke SHI, Junkai CHEN, Ting ZHANG, Yikun LI, Chengran YANG, Eng Lieh OUH, Lwin Khin SHAR, David LO 2026 Singapore Management University

Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo

Research Collection School Of Computing and Information Systems

Penetration testing is essential for identifying vulnerabilities in web applications before real adversaries can exploit them. Recent work has explored automating this process with Large Language Model (LLM)-powered agents, but existing approaches either rely on a single generic agent that struggles in complex scenarios or narrowly specialized agents that cannot adapt to diverse vulnerability types. We therefore introduce PenForge, a framework that dynamically constructs expert agents during testing rather than relying on those prepared beforehand. By integrating automated reconnaissance of potential attack surfaces with agents instantiated on the fly for context-aware exploitation, PenForge achieves a 30.0% exploit success rate (12/40) …


Semat: Semantic Enhanced Natural Image Interactive Matting, Ruihao XIA, Yu LIANG, Peng-Tao JIANG, Hao ZHANG, Qianru SUN, Yang TANG, Bo LI, Pan ZHOU 2026 Singapore Management University

Semat: Semantic Enhanced Natural Image Interactive Matting, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Qianru Sun, Yang Tang, Bo Li, Pan Zhou

Research Collection School Of Computing and Information Systems

Recent approaches attempt to adapt powerful interactive segmentation models, such as SAM, to interactive matting and fine-tune the models based on synthetic matting datasets. However, models trained on synthetic data fail to generalize to complex and occlusion scenes. We address this challenge by proposing a new matting dataset based on the COCO dataset, namely COCO-Matting. It selects real-world complex images from COCO and converts semantic segmentation masks to matting labels. The built COCO-Matting comprises an extensive collection of 36,980 human instance-level alpha mattes in complex natural scenarios. Furthermore, existing SAM-based matting methods extract intermediate features and masks from a frozen …


Stacked From One: Multi-Scale Self-Injection For Context Window Extension, Wei HAN, Pan ZHOU, Shuicheng YAN 2026 Singapore Management University

Stacked From One: Multi-Scale Self-Injection For Context Window Extension, Wei Han, Pan Zhou, Shuicheng Yan

Research Collection School Of Computing and Information Systems

The limited context window of contemporary large language models (LLMs) remains a primary bottleneck for their broader application across diverse domains. Although continual pre-training on long-context data offers a straightforward solution, it incurs prohibitive data acquisition and computational costs. To address this challenge, we propose SHAREDLLM, a novel framework based on multi-grained context compression and query-aware information acquisition. SHAREDLLM comprises two stacked short-context LLMs: a lower model serving as a compressor and an upper model acting as a decoder. The lower model compresses long inputs into compact, multi-grained representations, which are then forwarded to the upper model for context-aware processing. …


Distributional Vision-Language Alignment By Cauchy-Schwarz Divergence, Wenzhe YIN, Zehao XIAO, Pan ZHOU, Shujian YU, Jiayi SHEN, Jan-Jakob SONKE, Stratis GAVVES 2026 Singapore Management University

Distributional Vision-Language Alignment By Cauchy-Schwarz Divergence, Wenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu, Jiayi Shen, Jan-Jakob Sonke, Stratis Gavves

Research Collection School Of Computing and Information Systems

Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairwise samples across modalities while overlooking distributional differences. In addition, InfoNCE has inherent conflict in terms of alignment and uniformity in multimodality, leading to suboptimal alignment with modality gaps. To overcome the limitations, we propose CS-Aligner, a novel framework that performs distributional vision-language alignment by integrating Cauchy-Schwarz (CS) divergence with mutual information. CS-Aligner captures both the global distribution information of each modality and the pairwise semantic relationships. We find that the CS divergence …


Bridging Draft Policy Misalignment: Group Tree Optimization For Speculative Decoding, Shijing HU, Jingyang LI, Zhihui LU, Pan ZHOU 2026 Singapore Management University

Bridging Draft Policy Misalignment: Group Tree Optimization For Speculative Decoding, Shijing Hu, Jingyang Li, Zhihui Lu, Pan Zhou

Research Collection School Of Computing and Information Systems

Speculative decoding accelerates large language model (LLM) inference by letting a lightweight draft model propose multiple tokens that the target model verifies in parallel. Yet existing training objectives optimize only a single greedy draft path, while decoding follows a tree policy that re-ranks and verifies multiple branches. This draft policy misalignment limits achievable speedups. We introduce Group Tree Optimization (GTO), which aligns training with the decoding-time tree policy through two components: (i) Draft Tree Reward, a sampling-free objective equal to the expected acceptance length of the draft tree under the target model, directly measuring decoding performance; (ii) Group-based Draft Policy …


Dragging With Geometry: From Pixels To Geometry-Guided Image Editing, Xinyu PU, Hongsong WANG, Jie GUI, Pan ZHOU 2026 Singapore Management University

Dragging With Geometry: From Pixels To Geometry-Guided Image Editing, Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou

Research Collection School Of Computing and Information Systems

Interactive point-based image editing serves as a controllable editor, enabling precise and flexible manipulation of image content. However, most drag-based methods operate primarily on the 2D pixel plane with limited use of 3D cues. As a result, they often produce imprecise and inconsistent edits, particularly in geometry-intensive scenarios such as rotations and perspective transformations. To address these limitations, we propose a novel geometry-guided drag-based image editing method—GeoDrag, which addresses three key challenges: 1) incorporating 3D geometric cues into pixel-level editing, 2) mitigating discontinuities caused by geometry-only guidance, and 3) resolving conflicts arising from multi-point dragging. Built upon a unified displacement …


Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong ZOU, Ruihao XIA, Hongsong WANG, Pan ZHOU 2026 Singapore Management University

Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou

Research Collection School Of Computing and Information Systems

While text-to-3D generation has attracted growing interest, existing methods often struggle to produce 3D assets that align well with human preferences. Current preference alignment techniques for 3D content typically rely on hardly-collected preference-paired multi-view 2D images to train 2D reward models, when then guide 3D generation — leading to geometric artifacts, such as the Janus face problem and geometric incompleteness, due to their inherent 2D bias. To address these limitations, we construct 3D-MeshPref, the first large-scale unpaired 3D preference dataset, featuring diverse 3D meshes annotated by a large language model and refined by human evaluators. We then develop RewardCS, the …


From Spatial To Actions: Grounding Vision-Language-Action Model In Spatial Foundation Priors, Zhengshen ZHANG, Hao LI, Yalun DAI, Zhengbang ZHU, Lei ZHOU, Chenchen LIU, Dong WANG, Francis E. H. TAY, Sijin CHEN, Ziwei LIU, Yuxiao LIU, Xinghang LI, Pan ZHOU 2026 Singapore Management University

From Spatial To Actions: Grounding Vision-Language-Action Model In Spatial Foundation Priors, Zhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu, Lei Zhou, Chenchen Liu, Dong Wang, Francis E. H. Tay, Sijin Chen, Ziwei Liu, Yuxiao Liu, Xinghang Li, Pan Zhou

Research Collection School Of Computing and Information Systems

Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require specialized sensors and transfer poorly across modalities, or inject weak cues that lack geometry and degrade vision-language alignment. In this work, we introduce FALCON (From Spatial to Action), a novel paradigm that injects rich 3D spatial tokens into the action head. FALCON leverages spatial foundation models to deliver strong geometric priors from RGB alone, and includes an Embodied Spatial Model that can optionally fuse depth, or pose …


Where Did It Go Wrong? Attributing Undesirable Llm Behaviors Via Representation Gradient Tracing, Zhe LI, Wei ZHAO, Yige LI, Jun SUN 2026 Singapore Management University

Where Did It Go Wrong? Attributing Undesirable Llm Behaviors Via Representation Gradient Tracing, Zhe Li, Wei Zhao, Yige Li, Jun Sun

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their deployment is frequently undermined by undesirable behaviors such as generating harmful content, factual inaccuracies, and societal biases. Diagnosing the root causes of these failures poses a critical challenge for AI safety. Existing attribution methods, particularly those based on parameter gradients, often fall short due to prohibitive noisy signals and computational complexity. In this work, we introduce a novel and efficient framework that diagnoses a range of undesirable LLM behaviors by analyzing representation and its gradients, which operates directly in the model's activation space to provide a semantically meaningful signal linking …


Propaganda Ai: An Analysis Of Semantic Divergence In Large Language Models, Nay Myat MIN, Long H. PHAM, Yige LI, Jun SUN 2026 Singapore Management University

Propaganda Ai: An Analysis Of Semantic Divergence In Large Language Models, Nay Myat Min, Long H. Pham, Yige Li, Jun Sun

Research Collection School Of Computing and Information Systems

Large language models (LLMs) can exhibit concept-conditioned semantic divergence: common high-level cues (e.g., ideologies, public figures) elicit unusually uniform, stance-like responses that evade token-trigger audits. This behavior falls in a blind spot of current safety evaluations, yet carries major societal stakes, as such concept cues can steer content exposure at scale. We formalize this phenomenon and present RAVEN (Response Anomaly Vigilance), a black-box audit that flags cases where a model is simultaneously highly certain and atypical among peers by coupling semantic entropy over paraphrastic samples with cross-model disagreement. In a controlled LoRA fine-tuning study, we implant a concept-conditioned stance using …


Reducing Class-Wise Performance Disparity Via Margin Regularization, Beier ZHU, Kesen ZHAO, Jiequan CUI, Qianru SUN, Yuan ZHOU, Xun YANG, Hanwang ZHANG 2026 Singapore Management University

Reducing Class-Wise Performance Disparity Via Margin Regularization, Beier Zhu, Kesen Zhao, Jiequan Cui, Qianru Sun, Yuan Zhou, Xun Yang, Hanwang Zhang

Research Collection School Of Computing and Information Systems

Deep neural networks often exhibit substantial disparities in class-wise accuracy, even when trained on class-balanced data—posing concerns for reliable deployment. While prior efforts have explored empirical remedies, a theoretical understanding of such performance disparities in classification remains limited. In this work, we present Margin Regularization for performance disparity Reduction (MR2 ), a theoretically principled regularization for classification by dynamically adjusting margins in both the logit and representation spaces. Our analysis establishes a margin-based, class-sensitive generalization bound that reveals how per-class feature variability contributes to error, motivating the use of larger margins for “hard” classes. Guided by this insight, MR2 optimizes …


Real-Time Motion-Controllable Autoregressive Video Diffusion, Kesen ZHAO, Jiaxin SHI, Beier ZHU, Junbao ZHOU, Xiaolong SHEN, Yuan ZHOU, Qianru SUN, Hanwang ZHANG 2026 Singapore Management University

Real-Time Motion-Controllable Autoregressive Video Diffusion, Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou, Xiaolong Shen, Yuan Zhou, Qianru Sun, Hanwang Zhang

Research Collection School Of Computing and Information Systems

Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to simple control signals or text-to-video generation, and often suffer from quality degradation and motion artifacts in few-step generation. To address these challenges, we propose AR-Drag, the first RL-enhanced few-step AR video diffusion model for real-time image-to-video generation with diverse motion control. We first fine-tune a base I2V model to support basic motion control, then further improve it via reinforcement learning with a trajectory-based reward model. Our design preserves the …


Llmqua: Practical Backdoor Injection On Large Language Model Quantization, Xiangxiang CHEN, Peixin ZHANG, Jun SUN, Jin Song DONG, Wenhai WANG, Jingyi WANG 2026 Singapore Management University

Llmqua: Practical Backdoor Injection On Large Language Model Quantization, Xiangxiang Chen, Peixin Zhang, Jun Sun, Jin Song Dong, Wenhai Wang, Jingyi Wang

Research Collection School Of Computing and Information Systems

Quantization is widely used to enable local deployment of large language models (LLMs) on resource-constrained devices. Recent work (e.g., QuRA) shows quantization can be exploited via rounding manipulation to implant backdoors. However, such an attack has been evaluated only on small models and does not directly apply to LLMs due to three key constraints: (1) limited poisoning data from small, task-agnostic calibration sets; (2) layer-wise quantization restricting adversarial access to global representations; and (3) lack of gradient access in quantization pipelines, blocking gradient-based attacks.We propose LLMQuA, a practical quantization-phase backdoor attack tailored to the LLM setting. LLMQuA (i) injects backdoors …


Be Responsible In Your Answers! Monitoring Out-Of-Domain Behaviors In Domain-Specific Llms, Boquan LI, Chenzhe LOU, Zhe REN, Peixin ZHANG, Zirui FU, Jun SUN, Yaowen ZHENG 2026 Singapore Management University

Be Responsible In Your Answers! Monitoring Out-Of-Domain Behaviors In Domain-Specific Llms, Boquan Li, Chenzhe Lou, Zhe Ren, Peixin Zhang, Zirui Fu, Jun Sun, Yaowen Zheng

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have accelerated the rapid development of chatbot web applications in various domains, such as coding, biomedicine and psychology. Compared to general LLMs like ChatGPT, domain-specific LLMs require a greater sense of responsibility. For instance, if a programming LLM casually answers medical or psychological questions, it not only misleads the public but also poses legal risks. This highlights new demands for monitoring and preventing such irresponsible behaviors. Existing efforts attempt to monitor LLMs from multiple aspects, such as lying, jailbreaks, and toxic content, while overlooking out-of-domain behaviors. In this work, we propose an innovative LLM domain monitoring …


Digital Commons powered by bepress