Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Physical Sciences and Mathematics (1031)
- Computer Sciences (1030)
- Databases and Information Systems (510)
- Artificial Intelligence and Robotics (316)
- Software Engineering (185)
-
- Graphics and Human Computer Interfaces (150)
- Numerical Analysis and Scientific Computing (142)
- Social and Behavioral Sciences (106)
- Programming Languages and Compilers (103)
- Communication (87)
- Social Media (79)
- Engineering (58)
- Computer Engineering (43)
- Business (30)
- Information Security (30)
- Theory and Algorithms (28)
- Data Storage Systems (25)
- Education (22)
- OS and Networks (18)
- E-Commerce (12)
- Medicine and Health Sciences (11)
- Operations Research, Systems Engineering and Industrial Engineering (11)
- Higher Education (10)
- Arts and Humanities (7)
- Asian Studies (7)
- Digital Communications and Networking (7)
- Health Information Technology (7)
- International and Area Studies (7)
- Educational Assessment, Evaluation, and Research (6)
- Finance and Financial Management (6)
- Keyword
-
- Social media (36)
- Deep learning (25)
- Natural language processing (22)
- Large Language Models (19)
- LLMs (18)
-
- Machine learning (18)
- Large language models (17)
- Twitter (15)
- Sentiment analysis (14)
- Software engineering (13)
- Text mining (13)
- Computational linguistics (11)
- Neural networks (11)
- Natural Language Processing (10)
- Language model (9)
- Data mining (8)
- Deep Learning (8)
- Large Language Model (8)
- Large language model (8)
- Semantics (8)
- Artificial intelligence (7)
- Generative AI (7)
- Information retrieval (7)
- Question answering (7)
- Reinforcement learning (7)
- Transformer (7)
- Clustering (6)
- Graph neural networks (6)
- Multimodal (6)
- Natural language processing systems (6)
- Publication Year
Articles 1 - 30 of 1047
Full-Text Articles in Entire DC Network
Metarag: Identifying Website Owner Using Meta-Path-Guided Dynamic Graph Retrieval-Augmented Generation, Cheng Tu, Yunshan Ma, Bingyang Guo, Qianyu Li, Yang Li, Min Zhang, Fan Shi, Xiang Wang
Metarag: Identifying Website Owner Using Meta-Path-Guided Dynamic Graph Retrieval-Augmented Generation, Cheng Tu, Yunshan Ma, Bingyang Guo, Qianyu Li, Yang Li, Min Zhang, Fan Shi, Xiang Wang
Research Collection School Of Computing and Information Systems
Website owner identification aims to link websites to their real-world owners, which is crucial for credibility assessment and information provenance in information retrieval and vital for applications in cybersecurity, Internet governance, and digital regulation. Existing approaches for website owner identification primarily rely on querying infrastructure registration records or analyzing webpage content. However, these methods often fail due to incomplete or outdated registration records and sparse webpage content. We observe that inter-website relationships, derived from shared infrastructure data such as primary domains, IP blocks, and geolocations, can provide valuable but underutilized ownership cues. To exploit this insight, we propose MetaRAG, a …
Logupdater: Automated Detection And Repair Of Specific Defects In Logging Statements, Renyi Zhong, Yichen Li, Jinxi Kuang, Wenwei Gu, Yintong Huo, R. Michael Lyu
Logupdater: Automated Detection And Repair Of Specific Defects In Logging Statements, Renyi Zhong, Yichen Li, Jinxi Kuang, Wenwei Gu, Yintong Huo, R. Michael Lyu
Research Collection School Of Computing and Information Systems
Developers write logging statements to monitor software runtime behaviors and system state. However, poorly constructed or misleading log messages can inadvertently obfuscate actual program execution patterns, thereby impeding effective software maintenance. Existing research on analyzing issues within logging statements is limited, primarily focusing on detecting a singular type of defect and relying on manual intervention for fixes rather than automated solutions.To address the limitation, we initiate a systematic study that pinpoints four specific types of defects in logging statements (i.e., statement code inconsistency, static dynamic inconsistency, temporal relation inconsistency, and readability issues) through the analysis of real-world log-centric changes. We …
Generative Ai Adoption And Solvers' Popularity On Supply-Driven Crowdsourcing Platforms: The Dual Role Of Price Signals, Zimeng Zhu, Carol Hsu, Fiona Fui-Hoon Nah, Na Liu
Generative Ai Adoption And Solvers' Popularity On Supply-Driven Crowdsourcing Platforms: The Dual Role Of Price Signals, Zimeng Zhu, Carol Hsu, Fiona Fui-Hoon Nah, Na Liu
Research Collection School Of Computing and Information Systems
Purpose – We investigate the effect of solvers’ adoption of Generative AI (GenAI) on their popularity in a supply-driven crowdsourcing platform. We also examine the impact of price signals as well as their heterogeneous impact based on the solvers’ membership duration on the platform. Design/methodology/approach – Our analysis focuses on solvers who adopt GenAI for design-related gigs on the supply-driven crowdsourcing platform. By combining propensity score matching (PSM) with multi-period difference-in-differences (DID), we examine how GenAI adoption impacts solvers’ popularity and how price signals affect this main effect. Findings – Our findings reveal that solvers who adopt GenAI tend to …
Semantic-Structural Decoupling: Disentangling Semantic Attention From Structural Bias In The Attention Manifold, Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-Gang Jiang
Semantic-Structural Decoupling: Disentangling Semantic Attention From Structural Bias In The Attention Manifold, Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual tokens, a phenomenon termed "register" or "Visual Attention Sinks." While existing inference intervention methods attempt to identify these sink tokens and redistribute their attention weights, such approaches typically treat these tokens in isolation and suffer from computational inefficiency. Instead, we reframe this phenomenon as a generalized textual bias exerted over visual features that extends beyond isolated sink tokens. From this perspective, a pervasive structural bias leads to the dilution of the …
Spatialimaginer: Towards Adaptive Visual Imagination For Spatial Reasoning, Yian Li, Yang Jiao, Bin Zhu, Tianwen Qian, Shaoxiang Chen, Jingjing Chen, Yu-Gang Jiang
Spatialimaginer: Towards Adaptive Visual Imagination For Spatial Reasoning, Yian Li, Yang Jiao, Bin Zhu, Tianwen Qian, Shaoxiang Chen, Jingjing Chen, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language models. Despite promising performance, recent multimodal large language models (MLLMs) often exhibit fragile reasoning traces in spatial intelligence tasks that involve consistent spatial state recognition. We argue that these failures stem from a mismatch between the spatial recognition mechanism and the text-only reasoning behavior of these MLLMs. Effective spatial reasoning requires low-level geometric structure to be faithfully preserved and updated throughout the reasoning process, whereas textual representations tend to abstract away precisely these critical details. …
Restoring Linguistic Grounding In Vla Models Via Train-Free Attention Recalibration, Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen
Restoring Linguistic Grounding In Vla Models Via Train-Free Attention Recalibration, Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen
Research Collection School Of Computing and Information Systems
Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generalist robotic policies. However, their reliability under Out-Of-Distribution (OOD) instructions remains underexplored. In this paper, we reveal a critical failure mode in which VLA policies continue executing visually plausible actions even when the language instruction contradicts the scene. We refer to this phenomenon as linguistic blindness, where VLA policies prioritize visual priors over instruction semantics during action generation. To systematically analyze this issue, we introduce ICBench, a diagnostic benchmark constructed from the LIBERO dataset that probes language–action coupling …
Llm-Based Early Rumor Detection With Imitation Agent, Fengzhu Zeng, Qian Shao, Ling Cheng, Wei Gao, Shih-Fen Cheng, Jing Ma, Cheng Niu
Llm-Based Early Rumor Detection With Imitation Agent, Fengzhu Zeng, Qian Shao, Ling Cheng, Wei Gao, Shih-Fen Cheng, Jing Ma, Cheng Niu
Research Collection School Of Computing and Information Systems
Early Rumor Detection (EARD) aims to identify the earliest point at which a claim can be accurately classified based on a sequence of social media posts. This is especially challenging in data-scarce settings. While Large Language Models (LLMs) perform well in few-shot NLP tasks, they are not well-suited for time-series data and are computationally expensive for both training and inference. In this work, we propose a novel EARD framework that combines an autonomous agent and an LLM-based detection model, where the agent acts as a reliable decision-maker for \textit{early time point determination}, while the LLM serves as a powerful \textit{rumor …
Audeter: A Large-Scale Dataset For Deepfake Audio Detection In Open Worlds, Qizhou Wang, Hanxun Huang, Guansong Pang, Sarah Erfani, Christopher Leckie
Audeter: A Large-Scale Dataset For Deepfake Audio Detection In Open Worlds, Qizhou Wang, Hanxun Huang, Guansong Pang, Sarah Erfani, Christopher Leckie
Research Collection School Of Computing and Information Systems
Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving distribution shifts between training and test data, driven by the complexity of human speech and the rapid evolution of synthesis systems. Existing datasets suffer from limited real speech diversity, insufficient coverage of recent synthesis systems, and heterogeneous mixtures of deepfake sources, which hinder systematic evaluation and open-world model training. To address these issues, we introduce AUDETER (AUdio DEepfake TEst Range), a large-scale and highly diverse deepfake audio dataset comprising over 4,500 …
Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Although Large Language Models (LLMs) demonstrate significant capabilities, their reliance on parametric knowledge often leads to inaccuracies. Retrieval Augmented Generation (RAG) mitigates this by incorporating external knowledge, but these methods may introduce irrelevant retrieved documents, leading to inaccurate responses. While the integration methods filter out incorrect answers from multiple responses, but lack external knowledge like RAG methods, and their high costs require balancing overhead with performance gains. To address these issues, we propose an Efficient Test-Time Retrieval-Augmented Generation Framework named ET2RAG to improve the performance of LLMs while maintaining efficiency. Specifically, ET2RAG is a training-free method, that first retrieves the …
Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen
Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen
Research Collection School Of Computing and Information Systems
Large Language Model (LLM) agents are increasingly deployed in practice across a wide range of autonomous applications. Yet current safety mechanisms for LLM agents focus almost exclusively on preventing failures in advance, providing limited capabilities for responding to, containing, or recovering from incidents after they inevitably arise. In this work, we introduce AIR, the first incident response framework for LLM agent systems. AIR defines a domain-specific language for managing the incident response lifecycle autonomously in LLM agent systems, and integrates it into the agent's execution loop to (1) detect incidents via semantic checks grounded in the current environment state and …
Operationalizing Ethics For Ai Agents: How Developers Encode Values Into Repository Context Files, Christoph Treude, Sebastian Baltes, Marc Cheong
Operationalizing Ethics For Ai Agents: How Developers Encode Values Into Repository Context Files, Christoph Treude, Sebastian Baltes, Marc Cheong
Research Collection School Of Computing and Information Systems
As AI coding agents become embedded in software development workflows, developers are beginning to operationalize ethical principles by encoding behavioral rules into repository-level context files for AI agents, such as AGENTS.md files. Rather than examining the ethics of AI agents in the abstract, this vision paper investigates how ethics and values are already being translated for AI agents into actionable instructions that shape agent behavior. Through a preliminary investigation, we find that developers are already embedding guidance related to fairness, accessibility, sustainability, tone, and privacy. These artifacts function as a developer-authored governance layer, translating abstract principles into situated, natural-language directives …
A Dataset Of Agentic Ai Coding Tool Configurations, Matthias Galster, Seyedmoein Mohsenimofidi, Levi Böhme, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes
A Dataset Of Agentic Ai Coding Tool Configurations, Matthias Galster, Seyedmoein Mohsenimofidi, Levi Böhme, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes
Research Collection School Of Computing and Information Systems
Agentic AI coding tools such as Claude Code and OpenAI Codex execute multi-step coding tasks with limited human oversight. To steer these tools, developers create repository-level configuration artifacts (e.g., Markdown files) for configuration mechanisms such as Context Files, Skills, Rules, and Hooks. There is no curated dataset yet that captures these configurations at scale. This dataset, collected from open-source GitHub repositories, fills that gap. We selected 40,585 actively maintained repositories through metadata filtering, classified them using GPT-5.2 to identify 36,710 as belonging to engineered software projects, and systematically detected configuration artifacts in these repositories. The dataset covers 4,738 repositories across …
Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang
Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely underexplored. In this paper, we identify spatiotemporal sycophancy, a failure mode in which Vid-LLMs retract initially correct, visually grounded judgments and conform to misleading user feedback under negation-based gaslighting. Rather than merely changing their answers, the models often fabricate unsupported temporal or spatial explanations to justify incorrect revisions. To systematically investigate this phenomenon, we propose a negation-based gaslighting evaluation framework and introduce GasVideo-1000, a curated benchmark designed to probe spatiotemporal sycophancy with clear visual grounding and temporal reasoning requirements. …
Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen
Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen
Research Collection School Of Computing and Information Systems
Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object’s state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action–object interactions into …
Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo
Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo
Research Collection School Of Computing and Information Systems
Rapid advances in AI-generated image (AIGI) technology enable highly realistic synthesis, threatening public information integrity and security. Recent studies have demonstrated that incorporating texture-level artifact features alongside semantic features into multimodal large language models (MLLMs) can enhance their AIGI detection capability. However, our preliminary analyses reveal that artifact features exhibit high intra-feature similarity, leading to an almost uniform attention map after the softmax operation. This phenomenon causes attention dilution, thereby hindering effective fusion between semantic and artifact features. To overcome this limitation, we propose a lightweight fusion adapter, TranX-Adapter, which integrates a Task-aware Optimal-Transport Fusion that leverages the Jensen-Shannon divergence …
Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun
Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun
Research Collection School Of Computing and Information Systems
Large language models (LLMs) are increasingly trained on massive, heterogeneous text corpora, raising serious concerns about the unauthorised use of proprietary or personal data during model training. In this work, we address the problem of data protection against unwanted model learning in a realistic blackbox setting. We propose Disclaimer Injection, a novel data-level defence that renders text unlearnable to LLMs. Rather than relying on model-side controls or explicit data removal, our approach exploits the models’ own alignment mechanisms: injecting carefully designed alignment-triggers to prevent effective learning. Through layer-wise analysis, we find that finetuning on such protected data induces persistent activation …
Knowledge-State Generative Agents For Pre-Assessment Question Evaluation, Ping Fan Ke, Yi Meng Lau, Siaw Ling Lo
Knowledge-State Generative Agents For Pre-Assessment Question Evaluation, Ping Fan Ke, Yi Meng Lau, Siaw Ling Lo
Research Collection School Of Computing and Information Systems
This paper introduces a Knowledge‑State Generative Agent framework for evaluating the quality of pre‑assessment questions. The framework employs large language model (LLM)–based agents prompted to adopt a teacher persona to simulate the responses of students with and without mastery of targeted knowledge components. A preliminary empirical study using archival data from 424 students enrolled in an Information Systems Management course indicates that the proposed approach yields interpretable metrics under Classical Test Theory. Results further show that agents instantiated with the relevant mastered knowledge components exhibit systematically higher performance than agents lacking such mastery. In addition, the study suggests that teacher-persona …
Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma
Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma
Research Collection School Of Computing and Information Systems
Personalized generative recommender systems have emerged as a promising solution for fashion recommendation. However, existing methods primarily rely on implicit visual embeddings from historical interactions, which often contain preference-irrelevant information and result in insufficient user behavior modeling. Moreover, these models typically generate only item images, providing limited interpretability. To address these limitations, we propose DualFashion, a Dual-Diffusional Generative Fashion Recommendation Architecture that jointly models image and text modalities for personalized and explainable recommendation. DualFashion adopts a dual-diffusion Transformer with image and text branches, where structured attribute-level captions and visual outfit information are jointly used as conditioning signals to model user …
Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma
Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma
Research Collection School Of Computing and Information Systems
Despite the importance of open-ended event forecasting for risk management, current LLM-based methods predominantly target only the most probable outcomes, neglecting the intrinsic uncertainty of real-world events. To bridge this gap, we advance open-ended event forecasting from pinpoint forecasting to scatter forecasting by introducing the proxy task of hypothesis generation. This paradigm aims to generate an inclusive and diverse set of hypotheses that broadly cover the space of plausible future events. To this end, we propose SCATTER, a reinforcement learning framework that jointly optimizes inclusiveness and diversity of the hypothesis. Specifically, we design a novel hybrid reward that consists of …
Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang
Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang
Research Collection School Of Computing and Information Systems
Document Question Answering (DQA) involves generating answers from a document based on a user’s query, representing a key task in document understanding. This task requires interpreting visual layouts, which has prompted recent studies to adopt multimodal Retrieval-Augmented Generation (RAG) that processes page images for answer generation. However, in multimodal RAG, visual DQA struggles to utilize a large number of images effectively, as the retrieval stage often retains only a few candidate pages (e.g., Top-4), causing informative but less visually salient content to be overlooked in favor of common yet low-information pages. To address this issue, we propose a Multi-Armed Bandit–based …
Purifai: Detecting And Fixing Search-Induced Distortions In Web-Augmented Llms, Guoqing Wang, Zhao Zhang, Zeyu Sun, Xiaofei Xie, Yizhou Chen, Yanchao Tan, Dan Hao
Purifai: Detecting And Fixing Search-Induced Distortions In Web-Augmented Llms, Guoqing Wang, Zhao Zhang, Zeyu Sun, Xiaofei Xie, Yizhou Chen, Yanchao Tan, Dan Hao
Research Collection School Of Computing and Information Systems
As Large Language Models (LLMs) increasingly serve as interfaces for proprietary data (e.g., enterprise knowledge bases, legal statutes), ensuring their fidelity to trusted internal information is paramount. While integrating real-time web search can enhance model utility, it introduces a critical vulnerability: the ingestion of conflicting, misleading, or hallucinated content from the open web can override the model's adherence to its verified internal knowledge. We define this failure mode as search-induced distortion, a significant risk in high-stakes domains where the internal knowledge base serves as the absolute ground truth.To address this challenge, we present PurifAI, a proactive, model-agnostic, cache-level purification system …
Verbalizing Lightgcn: Direct Learning Of Textual Representations From User-Item Interaction Graph Via Llms, Manh-Khanh Ngo Huu, Hady Wirawan Lauw
Verbalizing Lightgcn: Direct Learning Of Textual Representations From User-Item Interaction Graph Via Llms, Manh-Khanh Ngo Huu, Hady Wirawan Lauw
Research Collection School Of Computing and Information Systems
In this work, we propose VerbaLightGCN, a novel LLM-based recommendation framework that integrates the semantic understanding of LLMs with user-item interaction modeling. Traditional collaborative filtering (CF) models typically embed user and item IDs into a latent space to capture interaction signals. However, pretrained LLMs cannot natively interpret these learned embeddings. To bridge this gap, VerbaLightGCN adopts a CF-as-text paradigm, in which collaborative signals are encoded in textual form and directly learned from the user–item interaction graph, and are then combined with semantic information to construct user and item profiles that function as latent embeddings. Inspired by LightGCN, our method retains …
Generation-Augmented Video Corpus Moment Retrieval, Mingjin Kuai, Qianyin Xiao, Juncheng Li, Jin Peng, Lizi Liao, Wei Ji
Generation-Augmented Video Corpus Moment Retrieval, Mingjin Kuai, Qianyin Xiao, Juncheng Li, Jin Peng, Lizi Liao, Wei Ji
Research Collection School Of Computing and Information Systems
Video Corpus Moment Retrieval (VCMR) requires models to efficiently retrieve and precisely locate specific moments relevant to natural language queries within a massive, untrimmed video corpus. However, existing discriminative approaches typically rely on shallow visual-textual feature matching mechanisms, which often struggle to capture fine-grained semantic differences. To address this limitation, we propose Video-GAR, a novel framework that reframes the conventional retrieval task from superficial matching to generative understanding, positing that the capability for query reconstruction evidences deep semantic comprehension. Specifically, Video-GAR orchestrates three synergistic components: To overcome the computational efficiency bottleneck, we construct a Bi-Mamba backbone that leverages the linear …
Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou
Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou
Research Collection School Of Computing and Information Systems
Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize single greedy trajectories, decoding involves verifying and ranking multiple sampled draft paths. We propose Variational Speculative Decoding (VSD), formulating draft training as variational inference over latent proposals (draft paths). VSD maximizes the marginal probability of target-model acceptance, yielding an ELBO that promotes high-quality latent proposals while minimizing divergence from the target distribution. To enhance quality and reduce variance, we incorporate a path-level utility and optimize via an Expectation-Maximization procedure. The E-step draws MCMC samples from an oracle-filtered posterior, while the M-step maximizes weighted likelihood using …
Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen
Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen
Research Collection School Of Computing and Information Systems
The widespread availability of large-scale code datasets has accelerated the development of code large language models (CodeLLMs), raising concerns about unauthorized dataset usage. Dataset poisoning offers a proactive defense by reducing the utility of such unauthorized training. However, existing poisoning methods often require full-dataset poisoning and introduce transformations that break code compilability. In this paper, we introduce FunPoison, a functionality-preserving poisoning approach that injects short, compilable weak-use fragments into executed code paths. FunPoison leverages reusable statement-level templates with automatic repair and conservative safety checking to ensure side-effect freedom, while a type-aware synthesis module preserves type correctness, suppresses static-analysis warnings, and …
Itimo: An Llm-Empowered Synthesis Dataset For Travel Itinerary Modification, Zhuoxuan Huang, Yunshan Ma, Hongyu Zhang, Hua Ma, Zhu Sun
Itimo: An Llm-Empowered Synthesis Dataset For Travel Itinerary Modification, Zhuoxuan Huang, Yunshan Ma, Hongyu Zhang, Hua Ma, Zhu Sun
Research Collection School Of Computing and Information Systems
Addressing itinerary modification is crucial for enhancing the travel experience as it is a frequent requirement during traveling. However, existing research mainly focuses on fixed itinerary planning, leaving modification underexplored due to the scarcity of shape need-to-modify itinerary data. To bridge this gap, we formally define the itinerary modification task and propose a general pipeline to construct the corresponding dataset, namely iTIMO. This pipeline frames the generation of shape need-to-modify itinerary data as an intent-driven perturbation task. It instructs large language models to perturb real-world itineraries using three operations: REPLACE, ADD, and DELETE. Each perturbation is grounded in three intents: …
Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang
Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang
Research Collection School Of Computing and Information Systems
Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we …
Frozen Lvlms For Micro-Video Recommendation: A Systematic Study Of Feature Extraction And Fusion, Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du
Frozen Lvlms For Micro-Video Recommendation: A Systematic Study Of Feature Extraction And Fusion, Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du
Research Collection School Of Computing and Information Systems
Frozen Large Video Language Models (LVLMs) are increasingly employed in micro-video recommendation (MVR) due to their strong multimodal understanding. However, existing apporches typically deploy LVLMs as fixed black-box feature extractors without systematically comparing alternative representation strategies. To address this gap, we present the first systematic empirical study on various feature extraction paradigms and integration strategies, along with hierarchical representations from frozen LVLMs for MVR. Extensive experiments on representative LVLMs reveal that hidden states from multiple decoder layers provide richer and more effective representations for MVR. Guided by this insight, we propose the Dual Feature Fusion (DFF) Framework, a lightweight approach …
Co-Designing With Autistic Livestreamers: Care, Constraints, And Trade-Offs In Livestreaming, Terrance Mok, Anthony Tang, Lora Oehlberg
Co-Designing With Autistic Livestreamers: Care, Constraints, And Trade-Offs In Livestreaming, Terrance Mok, Anthony Tang, Lora Oehlberg
Research Collection School Of Computing and Information Systems
Autistic livestreamers use platforms like Twitch for social connection, self-expression, and community, but these spaces also impose ongoing social and emotional demands. Prior work has documented these experiences, but less is known about what autistic creators themselves envision for the tools and platforms they use. We address this gap through a Research through Design (RtD) co-design study with three autistic Twitch streamers, using speculative artefacts as discussion prompts to explore how participants reasoned about potential livestreaming technologies. Across three co-design activities, we identify three overarching tensions shaping autistic streaming practice: Expression versus Misinterpretation and Harm; Public Participation versus Control and …
History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu
History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu
Research Collection School Of Computing and Information Systems
Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution in addressing the drawbacks of poor generalization in the training-based paradigm. However, existing LLM-based works primarily perform naive reasoning for decision-making and lack feedback, e.g., reviewing historical errors and predicting future potentials. Consequently, it may suffer from continuous failure for those initial error tasks. In this paper, we rethink LLM-based zero-shot VLN-CE and propose a new paradigm, named EvoNav, to improve …