Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (360)
- Engineering (296)
- Operations Research, Systems Engineering and Industrial Engineering (256)
- Business (177)
- Graphics and Human Computer Interfaces (176)
-
- Social and Behavioral Sciences (173)
- Software Engineering (137)
- Numerical Analysis and Scientific Computing (112)
- Theory and Algorithms (103)
- Public Affairs, Public Policy and Public Administration (75)
- Transportation (63)
- Programming Languages and Compilers (62)
- Information Security (43)
- Medicine and Health Sciences (43)
- Technology and Innovation (38)
- OS and Networks (37)
- Education (36)
- Law (36)
- Asian Studies (35)
- Computer Engineering (35)
- International and Area Studies (35)
- Health Information Technology (30)
- Science and Technology Law (22)
- Psychology (21)
- Library and Information Science (20)
- Finance and Financial Management (18)
- Higher Education (18)
- Keyword
-
- Artificial intelligence (98)
- Machine learning (55)
- Reinforcement learning (42)
- Deep learning (38)
- Artificial Intelligence (30)
-
- Large Language Models (30)
- Generative AI (29)
- ChatGPT (23)
- Large Language Model (23)
- Large language models (23)
- Singapore (22)
- Computer vision (19)
- Large language model (18)
- Optimization (18)
- Reinforcement Learning (18)
- Scheduling (18)
- Anomaly detection (17)
- Natural language processing (17)
- Deep reinforcement learning (16)
- Deep Learning (15)
- LLMs (15)
- Machine Learning (15)
- Vehicle routing problem (15)
- AI (14)
- Neural networks (13)
- Uncertainty (13)
- Software engineering (12)
- Graph neural networks (11)
- Metaverse (10)
- Multi-agent systems (10)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (1664)
- Dissertations and Theses Collection (Open Access) (57)
- Research Collection Lee Kong Chian School Of Business (33)
- Research Collection Yong Pung How School Of Law (31)
- Research Collection School of Social Sciences (22)
-
- Asian Management Insights (16)
- FORCE 2026 (14)
- Perspectives@SMU (11)
- Research Collection College of Integrative Studies (10)
- Research Collection Library (8)
- PhD Student’s Publications Collection (6)
- MITB Thought Leadership Series (4)
- 2024 AI for Research Week (3)
- CCX Research (3)
- LARC Research Publications (2)
- Research Collection School Of Accountancy (2)
- CASTLe: Collection of Articles on Scholarship for Teaching and Learning (1)
- Centre for AI & Data Governance (2019-2025) (1)
- Centre for Computational Law (2022-2025) (1)
- ROSA Journal Articles and Publications (1)
- Research Collection Office of Research (1)
- Research Collection School Of Economics (1)
- Research@SMU Infographics (1)
- Research@SMU: Connecting the Dots (1)
- SMU Press Releases and News (1)
- Sim Kee Boon Institute for Financial Economics (1)
- Student Publications (1)
- Publication Type
- File Type
Articles 91 - 120 of 1897
Full-Text Articles in Artificial Intelligence and Robotics
“Alexa, Do Not Say That In Front Of My Boss!” A Cross-Cultural Comparison Of User And Ai Preferences For Privacy-Aware Smart Speaker Interactions Across Contexts, Lynne Warin, Emily Aurelia, Anthony Tang, Emily Aurelia, Delphine Reinhardt
“Alexa, Do Not Say That In Front Of My Boss!” A Cross-Cultural Comparison Of User And Ai Preferences For Privacy-Aware Smart Speaker Interactions Across Contexts, Lynne Warin, Emily Aurelia, Anthony Tang, Emily Aurelia, Delphine Reinhardt
Research Collection School Of Computing and Information Systems
Due to their limited ability to reason about the social context in which they are used, smart speakers pose significant privacy risks by responding in ways that may violate people's implicit social boundaries. We conducted a cross-cultural vignette study (N = 944) in Germany and Singapore to investigate how situational factors—specifically social context (bystander relationships and closeness), physical context (location), and interaction context (topic and deceptive intent)—regulate user preferences for smart speaker responses. Our results demonstrate that these factors are superior predictors of response preferences than dispositional user traits (i.e., intrinsic personal traits). We identify two distinct social dynamics: a …
Cfalr: Collaborative Filtering-Augmented Large Language Model For Personalized Fashion Outfit Recommendation, Yujuan Ding, Junrong Liao, Yunshan Ma, Yi Bin, Wenqi Fan, Tat-Seng Chua, Qing Li
Cfalr: Collaborative Filtering-Augmented Large Language Model For Personalized Fashion Outfit Recommendation, Yujuan Ding, Junrong Liao, Yunshan Ma, Yi Bin, Wenqi Fan, Tat-Seng Chua, Qing Li
Research Collection School Of Computing and Information Systems
Personalized outfit recommendation poses a significant challenge in e-commerce and social media platforms, requiring systems that balance user preferences with aesthetic compatibility. Collaborative filtering (CF) provides a traditional solution for this, but it struggles with data-sparse scenarios and complex user-item-outfit relationships. Meanwhile, existing template-based approaches are constrained by rigid pre-designed structures. To bridge these research gaps, we introduce CFALR (Collaborative Filtering-Augmented Large Language Model for Recommendation), a novel framework that synergizes collaborative filtering with large language models for personalized outfit recommendation. Specifically, CFALR describes user-outfit interactions in natural language and leverages LLMs to capture fashion semantics while employing CF-enhanced embeddings …
How Do Machine Learning Models Change?, Joel Castaño, Rafael Cabañas, Antonio Salmerón, David Lo, Silverio Martínez-Fernández
How Do Machine Learning Models Change?, Joel Castaño, Rafael Cabañas, Antonio Salmerón, David Lo, Silverio Martínez-Fernández
Research Collection School Of Computing and Information Systems
The proliferation of Machine Learning (ML) models and their open source implementations has transformed AI research and applications. Platforms like Hugging Face (HF) enable this evolving ecosystem, yet a large-scale longitudinal study of how these models change is lacking. This study addresses this gap by analyzing over 680,000 commits from 100,000 models and 2,251 releases from 202 of these models on HF using repository mining and longitudinal methods. We apply an extended ML change taxonomy to classify commits and use Bayesian networks to model temporal patterns in commit and release activities. Our findings show that commit activities align with established …
Sequential Robustness In Adversarial Reinforcement Learning, Roman Lok-Ming Belaire
Sequential Robustness In Adversarial Reinforcement Learning, Roman Lok-Ming Belaire
Dissertations and Theses Collection (Open Access)
My goal is to build autonomous systems that expand the reach of human capability in challenging domains such as undersea and space exploration, disaster response, and large-scale infrastructure. In everyday settings, these systems will increasingly appear in safety-critical applications such as autonomous driving, robotics, and industrial manufacturing. A central requirement for these systems is the ability to operate reliably under uncertainty, particularly when the environment behaves in unanticipated ways.
The robust handling of unforeseen environment dynamics is therefore a technical cornerstone of autonomous decision-making; Adversarial attacks provide a useful and principled lens through which to study this problem. Adversarial \textit{robustness}, …
Perspectives On Interpretability For Neural Text Representations, Jia Peng Lim
Perspectives On Interpretability For Neural Text Representations, Jia Peng Lim
Dissertations and Theses Collection (Open Access)
In this dissertation, we investigate interpretability in the three elements of learning neural text representations: inputs, passed into models, to produce probabilistic outputs. We emphasise perspectives as we present alternative novel methods to mine and organise meaning in this work.
Models. We initiate our investigation by examining Neural Topic Models (NTM), proposing an alternate angle of interpreting its word-topic distribution, producing better topic representations for interpretation. Our method maps the problem of finding these better interpretations to classical NP-hard graph problems, enabling examination of topic distributions in a composite manner. Next, we apply our previous findings to extract interpretations from …
Benchmarking Gaslighting Attacks Against Speech Large Language Models, Jinyang Wu, Bin Zhu, Xiandong Zou, Qiquan Zhang
Benchmarking Gaslighting Attacks Against Speech Large Language Models, Jinyang Wu, Bin Zhu, Xiandong Zou, Qiquan Zhang
PhD Student’s Publications Collection
As Speech Large Language Models (Speech LLMs) become increasingly integrated into voice-based applications, ensuring their robustness against manipulative or adversarial input becomes critical. Although prior work has studied adversarial attacks in text-based LLMs and vision-language models, the unique cognitive and perceptual challenges of speech-based interaction remain underexplored. In contrast, speech presents inherent ambiguity, continuity, and perceptual diversity, which make adversarial attacks more difficult to detect. In this paper, we introduce gaslighting attacks, strategically crafted prompts designed to mislead, override, or distort model reasoning as a means to evaluate the vulnerability of Speech LLMs. Specifically, we construct five manipulation strategies: Anger, …
Teacher-Student Diffusion Model For Text-Driven 3d Hand Motion Generation, Ching Lam Cheng, Bin Zhu, Shengfeng He
Teacher-Student Diffusion Model For Text-Driven 3d Hand Motion Generation, Ching Lam Cheng, Bin Zhu, Shengfeng He
PhD Student’s Publications Collection
Generating realistic 3D hand motion from natural language is vital for VR, robotics, and human-computer interaction. Existing methods either focus on full-body motion, overlooking detailed hand gestures, or require explicit 3D object meshes, limiting generality. We propose TSHaMo, a model-agnostic teacher-student diffusion framework for text-driven hand motion generation. The student model learns to synthesize motions from text alone, while the teacher leverages auxiliary signals (e.g., MANO parameters) to provide structured guidance during training. A co-training strategy enables the student to benefit from the teacher’s intermediate predictions while remaining text-only at inference. Evaluated using two diffusion backbones on GRAB and H2O, …
Understanding Critical Thinking In Generative Artificial Intelligence Use: Development, Validation, And Correlates Of The Critical Thinking In Ai Use Scale, Gabriel R. Lau, Wei Yan Low, Louis Tay, Ysabel Thereze Ang Guevarra, Dragon Gašević, Andree Hartanto
Understanding Critical Thinking In Generative Artificial Intelligence Use: Development, Validation, And Correlates Of The Critical Thinking In Ai Use Scale, Gabriel R. Lau, Wei Yan Low, Louis Tay, Ysabel Thereze Ang Guevarra, Dragon Gašević, Andree Hartanto
Research Collection School of Social Sciences
Generative AI tools are increasingly embedded in everyday work and learning, yet their fluency, opacity, and propensity to hallucinate mean that users must critically evaluate AI outputs rather than accept them at face value. The present research conceptualises critical thinking in AI use as a dispositional tendency to verify the source and content of AI-generated information, to understand how models work and where they fail, and to reflect on the broader implications of relying on AI. Across six studies ( N = 1341), we developed and validated the 13-item critical thinking in AI use scale and mapped its nomological network. …
Digital Grief Technology To Support Bereavement: A Systematic Review Of Potential Benefits And Risks, Xun Ci Soh, Adalia Yin Hui Goh, Paye Shin Koh, Andree Hartanto
Digital Grief Technology To Support Bereavement: A Systematic Review Of Potential Benefits And Risks, Xun Ci Soh, Adalia Yin Hui Goh, Paye Shin Koh, Andree Hartanto
Research Collection School of Social Sciences
Grief is a universal and inevitable experience. However, the way we support the bereaved is changing, especially in the digital era. This systematic review examines the potential benefits and risks associated with various digital grief technologies, including online grief support groups, generative AI chatbots, online memorials, online therapy interventions, virtual reality, and digitally reproduced visuals or audio of the deceased. A systematic search was conducted in seven databases, and 30 articles were included in the final review. Findings indicate that digital grief technologies offer several benefits, such as reductions in grief and depressive symptoms, enhanced social support, greater accessibility, and …
Mease: Multi-Agent Episodic Action Sequence Explanation, Phyo Wai Khaing, Minghong Geng, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan
Mease: Multi-Agent Episodic Action Sequence Explanation, Phyo Wai Khaing, Minghong Geng, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan
Research Collection School Of Computing and Information Systems
Multi-agent reinforcement learning (MARL) achieves remarkable performance in complex coordination tasks, yet interpreting the emergent behaviors of trained agents remains a fundamental challenge. Most current explainability methods focus on individual agent decisions, overlooking the critical interplay of joint strategiesand temporal coordination patterns that define successful multi-agent policies. We present MEASE (Multi-agent Episodic Action Sequence Explanation), a novel explainable MARL (XMARL) framework that explains trained MARL policies as human-interpretable emergent cooperative joint behaviors. MEASE employs a cognition-inspired episodic memory model to learn spatio-temporal multi-agent interaction patterns, coupled with abstraction algorithms that identify significant cooperative agent behaviors. We evaluate MEASE on diverse …
Synthesis And Evaluation Of Long-Term History-Aware Medical Dialogue, Hebin Hu, Renke Dai, Ah-Hwee Tan, Yilin Kang
Synthesis And Evaluation Of Long-Term History-Aware Medical Dialogue, Hebin Hu, Renke Dai, Ah-Hwee Tan, Yilin Kang
Research Collection School Of Computing and Information Systems
An effective healthcare agent must be able to recall and reason over a patient’s longitudinal medical history. However, the absence of datasets with realistic long-term dialogue timelines limits systematic evaluation. Real clinical text is constrained by privacy and ethics, while existing benchmarks focus on isolated interactions, failing to capture cross-session reasoning. We introduce a framework for synthesizing high-quality, long-term medical dialogues with LLMs. Our approach entails a knowledge-guided decomposition into three stages: constructing synthetic patient profiles with diverse disease and complication trajectories, generating multiturn dialogues per encounter, and integrating them into a coherent longitudinal history dataset, MediLongChat. We establish three …
Market Reactions To Deceptive Language In Fake News: Implications From Language Expectancy Theory And Transfer Learning, Ka Chung Ng, Ping Fan Ke, Ping Fan, Mike So, Tam, Kar Yan
Market Reactions To Deceptive Language In Fake News: Implications From Language Expectancy Theory And Transfer Learning, Ka Chung Ng, Ping Fan Ke, Ping Fan, Mike So, Tam, Kar Yan
Research Collection School Of Computing and Information Systems
The advent of generative artificial intelligence (AI) has heightened the proliferation of fake news. A key challenge is the limited real-world data to investigate the societal impact of fake news produced by generative AI. In this paper, we examine stock market reactions to financial news articles that exhibit stylometric similarity to human-crafted and AI-crafted fake financial news. Grounded in language expectancy theory, we employ a style-based transfer learning model, pre-trained to recognizing deceptive language employed in various types of fake news intricacies. We then apply this model to a comprehensive dataset of financial news, assigning a “veracity style score” to …
Fighting Against Recruitment Scams: Theory-Driven Supervised Learning And Empirical Analysis For Digital Fraudulent Recruitment Posting Behavior, Tom (Tianteng) Wang, David (Jingjun) Xu, Keng Siau, Zhongju (John) Zhang
Fighting Against Recruitment Scams: Theory-Driven Supervised Learning And Empirical Analysis For Digital Fraudulent Recruitment Posting Behavior, Tom (Tianteng) Wang, David (Jingjun) Xu, Keng Siau, Zhongju (John) Zhang
Research Collection School Of Computing and Information Systems
The number of recruitment postings on digital recruitment hiring platforms has increased since the COVID-19 pandemic. However, the weak surveillance and operations of these platforms, combined with the fact that most job seekers have relatively low vigilance and a strong desire for recruitment offers, enable scammers to easily deceive job seekers for their money and confidential information. In this work, we combine prevailing text mining techniques (i.e., ChatGPT with prompting engineering and supervised machine learning) with interpersonal deception theory (IDT) from social science to design an interpretable IT system to predict fraudulent recruitment postings on digital recruitment-hiring platforms. We compare …
Enhancing Action And Ingredient Modeling For Semantically Grounded Recipe Generation, Guoshan Liu, Bin Zhu, Yian Li, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang
Enhancing Action And Ingredient Modeling For Semantically Grounded Recipe Generation, Guoshan Liu, Bin Zhu, Yian Li, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
Recent advances in Multimodal Large Language Models (MLMMs) have enabled recipe generation from food images, yet outputs often contain semantically incorrect actions or ingredients despite high lexical scores (e.g., BLEU, ROUGE). To address this gap, we propose a semantically grounded framework that predicts and validates actions and ingredients as internal context for instruction generation. Our two-stage pipeline combines supervised fine-tuning (SFT) with reinforcement fine-tuning (RFT): SFT builds foundational accuracy using an Action-Reasoning dataset and ingredient corpus, while RFT employs frequency-aware rewards to improve long-tail action prediction and ingredient generalization. A Semantic Confidence Scoring and Rectification (SCSR) module further filters and …
Selective Concolic Testing, Guofeng Zhang, Zhenbang Chen, Ziqi Shuai, Jun Sun, Weijiang Hong, Yufeng Zhang, Ji Wang, Yang Liu
Selective Concolic Testing, Guofeng Zhang, Zhenbang Chen, Ziqi Shuai, Jun Sun, Weijiang Hong, Yufeng Zhang, Ji Wang, Yang Liu
Research Collection School Of Computing and Information Systems
The principled combination of symbolic execution and random testing lacks a formal foundation, especially in deciding which inputs to symbolize. We propose selective concolic testing, a cost-aware framework that formulates this choice as an optimized policy problem of a MDP (Markov Decision Process). We model program exploration over a finite control-flow graph, where MDP states represent covered statements, actions partition path constraints into symbolic and random fragments, rewards reflect coverage gain, and costs account for SMT solving effort and sampling inefficiency. Our framework yields the first formal characterization of selective symbolization as policy synthesis in a probabilistic system. We prove …
Detecting Doubt In Reflective Learning: A Learning Analytics Study With Large And Small Language Models, Eng Lieh Ouh, Kar Way Tan, Siaw Ling Lo, Yuhao Zhang
Detecting Doubt In Reflective Learning: A Learning Analytics Study With Large And Small Language Models, Eng Lieh Ouh, Kar Way Tan, Siaw Ling Lo, Yuhao Zhang
Research Collection School Of Computing and Information Systems
Reflective learning enhances understanding, especially when instructors promptly address difficulties raised in student reflections. Automated doubt detection can reduce time for instructors, yet existing classification approaches take substantial time for manual annotation and model training. This paper investigates whether large and small language models (LLMs, SLMs) can automate doubt detection without time-consuming training. Using a dataset of anonymized student reflections, we evaluate zeroshot, few-shot prompting, and multi-step reasoning against prior supervised classification baselines. We show that LLMs (GPT-4o, Claude-4, Gemini-2.5) surpass earlier F1 scores without prompting, while prompting further improves their performance. However, using proprietary LLMs can raise cost and …
Gencode: A Generic Data Augmentation Framework For Boosting Deep Learning-Based Code Understanding, Zeming Dong, Qiang Hu, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Yves Le Traon, Jianjun Zhao
Gencode: A Generic Data Augmentation Framework For Boosting Deep Learning-Based Code Understanding, Zeming Dong, Qiang Hu, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Yves Le Traon, Jianjun Zhao
Research Collection School Of Computing and Information Systems
Pre-trained code models lead the era of code intelligence, with multiple models designed with impressive performance. However, one important problem, data augmentation for code data that automatically helps developers prepare training data lacks study in this field. In this paper, we introduce a generic data augmentation framework, GenCode, to enhance the training of code understanding models. Simply speaking, GenCode follows a generation-and-selection paradigm to prepare useful training code data. Specifically, it employs code augmentation techniques to generate new code candidates first and then identifies important ones as the training data by influence scores. To evaluate the effectiveness of GenCode, we …
Architecture-Agnostic Test-Time Adaptation Via Backprop-Free Embedding Alignment, Xiao Ma, Young D. Kwon, Pan Zhou, Dong Ma
Architecture-Agnostic Test-Time Adaptation Via Backprop-Free Embedding Alignment, Xiao Ma, Young D. Kwon, Pan Zhou, Dong Ma
PhD Student’s Publications Collection
Test-Time Adaptation (TTA) adapts a deployed model during online inference to mitigate the impact of domain shift. While achieving strong accuracy, most existing methods rely on backpropagation, which is memory and computation intensive, making them unsuitable for resource-constrained devices. Recent attempts to reduce this overhead often suffer from high latency or are tied to specific architectures such as ViT-only or CNN-only. In this work, we revisit domain shift from an embedding perspective. Our analysis reveals that domain shift induces three distinct structural changes in the embedding space: translation (mean shift), scaling (variance shift), and rotation (covariance shift). Based on this …
Scalable Multi-Task Low-Rank Model Adaptation, Zichen Tian, Antoine Ledent, Qianru Sun
Scalable Multi-Task Low-Rank Model Adaptation, Zichen Tian, Antoine Ledent, Qianru Sun
PhD Student’s Publications Collection
Scaling multi-task low-rank adaptation (LoRA) to a large number of tasks induces catastrophic performance degradation, such as an accuracy drop from 88.2% to 2.0% on DOTA when scaling from 5 to 15 tasks. This failure is due to parameter and representation misalignment. We find that existing solutions, like regularization and dynamic routing, fail at scale because they are constrained by a fundamental trade-off: strengthening regularization to reduce inter-task conflict inadvertently suppresses the essential feature discrimination required for effective routing. In this work, we identify two root causes for this trade-off. First, uniform regularization disrupts inter-task knowledge sharing: shared underlying knowledge …
Generative Ai In Enterprises: Optimizing Applications With Large Language Models, Donghao Huang
Generative Ai In Enterprises: Optimizing Applications With Large Language Models, Donghao Huang
Dissertations and Theses Collection (Open Access)
This dissertation investigates how to deploy Large Language Models (LLMs) effectively in enterprise settings, where accuracy, reliability, cost, privacy, and operational constraints often matter more than benchmark performance alone. Drawing on seventeen peer-reviewed publications (eleven published and six accepted for publication), the work develops and validates optimization strategies across three connected themes: retrieval-augmented generation (RAG), agentic AI for workflow automation, and deployment guidelines for real-world enterprise environments.
First, we study RAG optimization through systematic evaluation of open and proprietary models, highlighting conditions under which efficient open-weight models can match or exceed proprietary alternatives. To address a pervasive failure mode in …
Think First, Chatgpt Later: Guiding Human-Ai Collaboration For Learning Gains In Independent Human Creativity, Sarah Shi Hui Wong, Sophia Xuefei Qiu
Think First, Chatgpt Later: Guiding Human-Ai Collaboration For Learning Gains In Independent Human Creativity, Sarah Shi Hui Wong, Sophia Xuefei Qiu
Research Collection School of Social Sciences
Generative artificial intelligence (AI) tools such as ChatGPT can boost creative performance, but do these boosts translate into learning gains? This study examined whether the benefits of ChatGPT for creativity persist even when its assistance is removed, and how people can effectively use ChatGPT to enhance their learning and independent creativity. University students (N = 196) solved a creative product improvement task either independently (human-only group) or using ChatGPT freely (general-AI group) or using ChatGPT in a guided way (regulated-AI group). Specifically, the regulated-AI group used a novel “think first, ChatGPT later” approach—they first generated their own ideas, then collaborated …
Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng
Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng
Research Collection Lee Kong Chian School Of Business
Existing research on AI cultural biases predominantly focuses on Western models, overlooking critical gaps in non-Western models. We conduct a comparative analysis of AI models – ChatGPT (U.S. developed) and ErnieBot (China developed) – from different cultures to investigate how corresponding cultural biases manifest in their outputs. Additionally, we examine how cultural alignment between human users and AI models impacts their collaborative creative performance and the underlying psychological mechanisms. In Study 1, multi-choice prompt with zero-shot technique was used to evaluate cultural biases in four widely used AI models – ChatGPT-3.5/4, ErnieBot-3.5/4 – comparing their responses to established cultural psychometric …
Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong
Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong
Research Collection School Of Computing and Information Systems
Conventionally, human intuition defines vision as a modality of passive optical sensing, relying on ambient light to perceive the environment. However, active optical sensing, which involves emitting and receiving signals, offers unique advantages by capturing both radiometric and geometric properties of the environment, independent of external illumination conditions. This work focuses on advancing active optical sensing using Light Detection and Ranging (LiDAR), which captures intensity data, enabling the estimation of surface reflectance that remains invariant under varying illumination. Such properties are crucial for robotic perception tasks, including detection, recognition, segmentation, and Simultaneous Localization and Mapping (SLAM). A key challenge with …
Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo
Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo
Research Collection School Of Computing and Information Systems
Penetration testing is essential for identifying vulnerabilities in web applications before real adversaries can exploit them. Recent work has explored automating this process with Large Language Model (LLM)-powered agents, but existing approaches either rely on a single generic agent that struggles in complex scenarios or narrowly specialized agents that cannot adapt to diverse vulnerability types. We therefore introduce PenForge, a framework that dynamically constructs expert agents during testing rather than relying on those prepared beforehand. By integrating automated reconnaissance of potential attack surfaces with agents instantiated on the fly for context-aware exploitation, PenForge achieves a 30.0% exploit success rate (12/40) …
Semat: Semantic Enhanced Natural Image Interactive Matting, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Qianru Sun, Yang Tang, Bo Li, Pan Zhou
Semat: Semantic Enhanced Natural Image Interactive Matting, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Qianru Sun, Yang Tang, Bo Li, Pan Zhou
Research Collection School Of Computing and Information Systems
Recent approaches attempt to adapt powerful interactive segmentation models, such as SAM, to interactive matting and fine-tune the models based on synthetic matting datasets. However, models trained on synthetic data fail to generalize to complex and occlusion scenes. We address this challenge by proposing a new matting dataset based on the COCO dataset, namely COCO-Matting. It selects real-world complex images from COCO and converts semantic segmentation masks to matting labels. The built COCO-Matting comprises an extensive collection of 36,980 human instance-level alpha mattes in complex natural scenarios. Furthermore, existing SAM-based matting methods extract intermediate features and masks from a frozen …
Stacked From One: Multi-Scale Self-Injection For Context Window Extension, Wei Han, Pan Zhou, Shuicheng Yan
Stacked From One: Multi-Scale Self-Injection For Context Window Extension, Wei Han, Pan Zhou, Shuicheng Yan
Research Collection School Of Computing and Information Systems
The limited context window of contemporary large language models (LLMs) remains a primary bottleneck for their broader application across diverse domains. Although continual pre-training on long-context data offers a straightforward solution, it incurs prohibitive data acquisition and computational costs. To address this challenge, we propose SHAREDLLM, a novel framework based on multi-grained context compression and query-aware information acquisition. SHAREDLLM comprises two stacked short-context LLMs: a lower model serving as a compressor and an upper model acting as a decoder. The lower model compresses long inputs into compact, multi-grained representations, which are then forwarded to the upper model for context-aware processing. …
Distributional Vision-Language Alignment By Cauchy-Schwarz Divergence, Wenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu, Jiayi Shen, Jan-Jakob Sonke, Stratis Gavves
Distributional Vision-Language Alignment By Cauchy-Schwarz Divergence, Wenzhe Yin, Zehao Xiao, Pan Zhou, Shujian Yu, Jiayi Shen, Jan-Jakob Sonke, Stratis Gavves
Research Collection School Of Computing and Information Systems
Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairwise samples across modalities while overlooking distributional differences. In addition, InfoNCE has inherent conflict in terms of alignment and uniformity in multimodality, leading to suboptimal alignment with modality gaps. To overcome the limitations, we propose CS-Aligner, a novel framework that performs distributional vision-language alignment by integrating Cauchy-Schwarz (CS) divergence with mutual information. CS-Aligner captures both the global distribution information of each modality and the pairwise semantic relationships. We find that the CS divergence …
Bridging Draft Policy Misalignment: Group Tree Optimization For Speculative Decoding, Shijing Hu, Jingyang Li, Zhihui Lu, Pan Zhou
Bridging Draft Policy Misalignment: Group Tree Optimization For Speculative Decoding, Shijing Hu, Jingyang Li, Zhihui Lu, Pan Zhou
Research Collection School Of Computing and Information Systems
Speculative decoding accelerates large language model (LLM) inference by letting a lightweight draft model propose multiple tokens that the target model verifies in parallel. Yet existing training objectives optimize only a single greedy draft path, while decoding follows a tree policy that re-ranks and verifies multiple branches. This draft policy misalignment limits achievable speedups. We introduce Group Tree Optimization (GTO), which aligns training with the decoding-time tree policy through two components: (i) Draft Tree Reward, a sampling-free objective equal to the expected acceptance length of the draft tree under the target model, directly measuring decoding performance; (ii) Group-based Draft Policy …
Dragging With Geometry: From Pixels To Geometry-Guided Image Editing, Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou
Dragging With Geometry: From Pixels To Geometry-Guided Image Editing, Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou
Research Collection School Of Computing and Information Systems
Interactive point-based image editing serves as a controllable editor, enabling precise and flexible manipulation of image content. However, most drag-based methods operate primarily on the 2D pixel plane with limited use of 3D cues. As a result, they often produce imprecise and inconsistent edits, particularly in geometry-intensive scenarios such as rotations and perspective transformations. To address these limitations, we propose a novel geometry-guided drag-based image editing method—GeoDrag, which addresses three key challenges: 1) incorporating 3D geometric cues into pixel-level editing, 2) mitigating discontinuities caused by geometry-only guidance, and 3) resolving conflicts arising from multi-point dragging. Built upon a unified displacement …
Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou
Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou
Research Collection School Of Computing and Information Systems
While text-to-3D generation has attracted growing interest, existing methods often struggle to produce 3D assets that align well with human preferences. Current preference alignment techniques for 3D content typically rely on hardly-collected preference-paired multi-view 2D images to train 2D reward models, when then guide 3D generation — leading to geometric artifacts, such as the Janus face problem and geometric incompleteness, due to their inherent 2D bias. To address these limitations, we construct 3D-MeshPref, the first large-scale unpaired 3D preference dataset, featuring diverse 3D meshes annotated by a large language model and refined by human evaluators. We then develop RewardCS, the …