Open Access. Powered by Scholars. Published by Universities.®

Digital Commons Network™

Open Access. Powered by Scholars. Published by Universities.®

Research Collection School Of Computing and Information Systems

Discipline
Keyword
Publication Year
File Type

Articles 61 - 90 of 8553

Full-Text Articles in Entire DC Network

Anomaly Management In Unmanned Aerial Vehicles: A Systematic Literature Review, Ivan Tan Wei Han, Christopher M. Poskitt, Lingxiao Jiang, Lwin Khin Shar Jul 2026

Anomaly Management In Unmanned Aerial Vehicles: A Systematic Literature Review, Ivan Tan Wei Han, Christopher M. Poskitt, Lingxiao Jiang, Lwin Khin Shar

Research Collection School Of Computing and Information Systems

Unmanned Aerial Vehicles (UAVs) are increasingly deployed in safety-critical applications such as logistics, surveillance, disaster response, and urban air mobility. While their autonomy enables powerful capabilities, it also introduces vulnerabilities due to hardware faults, software defects, communication failures, and adversarial interference. This survey presents a comprehensive review of research studies closely related to UAV anomalies published between 2015 and 2025, covering 111 papers from academic and industrial sources. We introduce a unified five-pillar taxonomy—anomaly generation, prevention, detection, recovery, and analysis—that organizes existing work across the full anomaly management lifecycle. In contrast to prior surveys that focus primarily on detection algorithms, …


Towards Uniformity And Alignment For Multimodal Representation Learning, Wenzhe Yin, Pan Zhou, Zehao Xiao, Jie Liu, Shujian Yu, Jan-Jakob Sonke, Efstratios Gavves Jul 2026

Towards Uniformity And Alignment For Multimodal Representation Learning, Wenzhe Yin, Pan Zhou, Zehao Xiao, Jie Liu, Shujian Yu, Jan-Jakob Sonke, Efstratios Gavves

Research Collection School Of Computing and Information Systems

Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-based objectives introduce inherent conflicts that yield distribution gaps across modalities. In this work, we identify two conflicts in the multimodal regime, both exacerbated as the number of modalities increases: (i) an alignment–uniformity conflict, whereby the repulsion of uniformity undermines pairwise alignment, and (ii) an intra-alignment conflict, where aligning multiple modalities induces competing alignment directions. To address these issues, we propose a principled decoupling of alignment and uniformity for multimodal representations, providing a conflict-free recipe for multimodal learning that …


Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou Jul 2026

Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou

Research Collection School Of Computing and Information Systems

Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize single greedy trajectories, decoding involves verifying and ranking multiple sampled draft paths. We propose Variational Speculative Decoding (VSD), formulating draft training as variational inference over latent proposals (draft paths). VSD maximizes the marginal probability of target-model acceptance, yielding an ELBO that promotes high-quality latent proposals while minimizing divergence from the target distribution. To enhance quality and reduce variance, we incorporate a path-level utility and optimize via an Expectation-Maximization procedure. The E-step draws MCMC samples from an oracle-filtered posterior, while the M-step maximizes weighted likelihood using …


Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen Jul 2026

Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen

Research Collection School Of Computing and Information Systems

The widespread availability of large-scale code datasets has accelerated the development of code large language models (CodeLLMs), raising concerns about unauthorized dataset usage. Dataset poisoning offers a proactive defense by reducing the utility of such unauthorized training. However, existing poisoning methods often require full-dataset poisoning and introduce transformations that break code compilability. In this paper, we introduce FunPoison, a functionality-preserving poisoning approach that injects short, compilable weak-use fragments into executed code paths. FunPoison leverages reusable statement-level templates with automatic repair and conservative safety checking to ensure side-effect freedom, while a type-aware synthesis module preserves type correctness, suppresses static-analysis warnings, and …


Itimo: An Llm-Empowered Synthesis Dataset For Travel Itinerary Modification, Zhuoxuan Huang, Yunshan Ma, Hongyu Zhang, Hua Ma, Zhu Sun Jul 2026

Itimo: An Llm-Empowered Synthesis Dataset For Travel Itinerary Modification, Zhuoxuan Huang, Yunshan Ma, Hongyu Zhang, Hua Ma, Zhu Sun

Research Collection School Of Computing and Information Systems

Addressing itinerary modification is crucial for enhancing the travel experience as it is a frequent requirement during traveling. However, existing research mainly focuses on fixed itinerary planning, leaving modification underexplored due to the scarcity of shape need-to-modify itinerary data. To bridge this gap, we formally define the itinerary modification task and propose a general pipeline to construct the corresponding dataset, namely iTIMO. This pipeline frames the generation of shape need-to-modify itinerary data as an intent-driven perturbation task. It instructs large language models to perturb real-world itineraries using three operations: REPLACE, ADD, and DELETE. Each perturbation is grounded in three intents: …


Bridging Llm Embeddings And Vae Parameters For Disentangled Recommendation, Nhu-Thuat Tran, Hady Wirawan Lauw Jul 2026

Bridging Llm Embeddings And Vae Parameters For Disentangled Recommendation, Nhu-Thuat Tran, Hady Wirawan Lauw

Research Collection School Of Computing and Information Systems

Disentangled recommendation within the Variational Autoencoder (VAE) framework aims to capture multiple user interests. While effective, these VAEs are fundamentally constrained by their reliance on interaction data alone, lacking the rich external semantic knowledge needed to properly structure and separate latent interests. Meanwhile, Large Language Models (LLMs) excel at deriving profound user preference signals from textual data. Prevailing methods for integrating LLMs into recommendation, however, either focus on single-interest modeling or perform a shallow fusion by aligning LLM and VAE representation spaces. Thus, they fail to fundamentally shape the VAE's latent space for multi-interest learning, hindering recommendation performance. To bridge …


Ppg-Sport: A Dataset For Reliable Heart Rate Monitoring From Wrist Ppg Under Dynamic Sports Conditions, Changshuo Hu, Hung Manh Pham, Yiming Zhang, Guanru Yan, Xiao Ma, Yuezhong Wu, Thivya Kandappu, Archan Misra, Dong Ma Jun 2026

Ppg-Sport: A Dataset For Reliable Heart Rate Monitoring From Wrist Ppg Under Dynamic Sports Conditions, Changshuo Hu, Hung Manh Pham, Yiming Zhang, Guanru Yan, Xiao Ma, Yuezhong Wu, Thivya Kandappu, Archan Misra, Dong Ma

Research Collection School Of Computing and Information Systems

Photoplethysmography (PPG) has become a cornerstone of physiological sensing in wearable devices, enabling non-invasive monitoring of heart rate and related biomarkers. However, its reliability deteriorates sharply under dynamic, high-intensity, or non-periodic motions such as those in sports, where existing datasets fail to capture realistic wrist dynamics. To address this gap, we introduce PPG-Sport, the first large-scale dataset designed for heart rate monitoring from wrist-worn PPG under real sports conditions. The PPG-Sport dataset includes synchronized PPG, inertial measurement unit (IMU), and electrocardiography (ECG) recordings from both wrists of 30 participants across six representative activities: stationary, walking, running, badminton, table tennis, and …


“From Remembering To Shaping”: Narrating Shared Experiences By Co-Designing Cultural Heritage Artifacts In Collaborative Vr, Yushang Yang, Fanxu Meng, Fiona Fui-Hoon Nah, L. C. Ray Jun 2026

“From Remembering To Shaping”: Narrating Shared Experiences By Co-Designing Cultural Heritage Artifacts In Collaborative Vr, Yushang Yang, Fanxu Meng, Fiona Fui-Hoon Nah, L. C. Ray

Research Collection School Of Computing and Information Systems

The ways people remember and recall places reveal an invisible aspect of cultural heritage (CH), reflecting how individuals and communities relate to these places. Heritage is communal, emerging through collaboratively constructed narratives rather than individual records. To probe how people may share collective memories, we designed an immersive two-person workflow for collaboratively co-designing 3D artifacts and environments in virtual heritage locations, using Generative AI (GenAI) to instantiate these intangible memories. Observations of the co-creation process revealed that participants merged prompts and model placements when negotiating different perspectives. They used spatial operations to compose scenes, and also to express personal and …


Context Matters: Auditing Gender Bias In T2i Generation Through Risk-Tiered Use-Case Profiles, Jose Luis Luna Campoverde, Yankun Wu, Xiaofei Xie, Noa Garcia Jun 2026

Context Matters: Auditing Gender Bias In T2i Generation Through Risk-Tiered Use-Case Profiles, Jose Luis Luna Campoverde, Yankun Wu, Xiaofei Xie, Noa Garcia

Research Collection School Of Computing and Information Systems

Text-to-image (T2I) generative models are increasingly used to produce content for education, media, and public-facing communication, and are starting to be integrated into higher-impact pipelines. Since generated images tend to reinforce stereotypes, producing representational erasure via “default” depictions and shaping perceptions of who belongs in certain roles, a growing body of work has proposed metrics to quantify gender bias in T2I outputs. Yet existing evaluations remain fragmented. Metrics are often reported without a shared view of what they measure, what assumptions they entail, or how their results should be interpreted under different deployment contexts. This limits the usefulness of gender …


Patchfuzz: Patch Fuzzing For Javascript Engines, Junjie Wang, Zhihua Xie, Xiaofei Xie, Xiaoning Du, Xiangwei Zhang Jun 2026

Patchfuzz: Patch Fuzzing For Javascript Engines, Junjie Wang, Zhihua Xie, Xiaofei Xie, Xiaoning Du, Xiangwei Zhang

Research Collection School Of Computing and Information Systems

Context: Patch fuzzing is a technique aimed at identifying vulnerabilities that arise from newly patched code. While researchers have made efforts to apply patch fuzzing to testing JavaScript (JS) engines with considerable success, these efforts have been limited to using ordinary test cases or publicly available vulnerability PoCs (Proof of Concepts) as seeds, and the sustainability of these approaches is hindered by the challenges associated with automating the PoC collection. Objective: To address these limitations, we propose an end-to-end sustainable approach for JS engine patch fuzzing, named PatchFuzz. Method: It automates the collection of PoCs of a broader range of …


Hydpn: A Hybrid Deep Reinforcement Learning, Programming, And Neighborhood Operations Framework For Integrated Scheduling On Parallel Batch Processing Machines, Yuqi Wang, He Luo, Guoqiang Wang, Zhaoxia Wang Jun 2026

Hydpn: A Hybrid Deep Reinforcement Learning, Programming, And Neighborhood Operations Framework For Integrated Scheduling On Parallel Batch Processing Machines, Yuqi Wang, He Luo, Guoqiang Wang, Zhaoxia Wang

Research Collection School Of Computing and Information Systems

Batch processing machines (BPMs) are widely used in industries such as semiconductors, metal processing, and healthcare, where jobs are processed in batches. As production, inventory, and distribution become increasingly integrated to improve efficiency, research on their joint scheduling in parallel BPM environments remains scarce. This paper addresses the integrated scheduling problem in parallel BPMs, involving production, inventory, and distribution stages, with the objective of minimizing total costs. A unified cost-based model is first formulated, applicable to both in-facility and external distribution scenarios. A hybrid algorithm framework, HyDPN, combining deep reinforcement learning, dynamic programming, and neighborhood operations is proposed. Extensive experiments …


Extensive And Intensive Margin Labor Supply On Ride-Sourcing Platforms, Hao Sun, Hai Wang, Zhixi Wan Jun 2026

Extensive And Intensive Margin Labor Supply On Ride-Sourcing Platforms, Hao Sun, Hai Wang, Zhixi Wan

Research Collection School Of Computing and Information Systems

The rapid expansion of ride-sourcing platforms has enabled freelance drivers to flexibly determine both their participation and working hours. Understanding this flexible labor supply behavior is essential for managing platform capacity and evaluating the impacts of pricing and incentive policies on driver welfare. This study develops a labor supply model in which drivers optimally choose whether to participate (extensive margin) and how long to work (intensive margin) to maximize their utility from consumption and leisure. The model incorporates heterogeneity in drivers’ other income, idle time, and participation costs, allowing us to analytically characterize equilibrium labor supply decisions. The results show …


When Politics Meets Digital Assets: Gender Identity Salience And Nft Pricing After Roe V. Wade, Xiang Liu, Yao Zhao, Ping Fan Ke Jun 2026

When Politics Meets Digital Assets: Gender Identity Salience And Nft Pricing After Roe V. Wade, Xiang Liu, Yao Zhao, Ping Fan Ke

Research Collection School Of Computing and Information Systems

Major sociopolitical events can reshape public attention toward identity-related issues, potentially influencing valuation patterns in digital markets where identity-related characteristics are embedded in digital assets. Using the overturning of Roe v. Wade as an exogenous policy shock, this paper examines how gender attributes represented in non-fungible token (NFT) avatars affect market outcomes. Using transaction data from six major avatar-based NFT collections traded on Etherscan in 2022, we apply a quasi-experimental design combining propensity score matching and a difference-in-differences model. The results indicate that the policy shock significantly increased the resale prices of NFTs representing female avatars. These findings suggest that …


On-The-Fly Generation-Quality Enhancement Of Deep Code Models Via Model Collaboration, Weifeng Sun, Naiqi Huang, Meng Yan, Zhongxin Liu, Hongyan Li, Yan Lei, David Lo Jun 2026

On-The-Fly Generation-Quality Enhancement Of Deep Code Models Via Model Collaboration, Weifeng Sun, Naiqi Huang, Meng Yan, Zhongxin Liu, Hongyan Li, Yan Lei, David Lo

Research Collection School Of Computing and Information Systems

The growing prominence of deep code models in automating software engineering tasks is undeniable. However, their deployment encounters significant challenges in on-the-fly performance enhancement, which refers to dynamically improving the performance of deep code models during real-time execution. Conventional techniques, such as retraining or fine-tuning, are effective in controlled pre-deployment scenarios but fall short when adapting to on-the-fly adjustments post-deployment. CodeDenoise, a notable on-the-fly performance enhancement technology, leverages uncertainty-based methods to identify misclassified inputs and applies an input modification strategy to rectify classification errors. While effective for classification tasks, this approach is inapplicable to generative tasks due to two key …


Videocreator: An Agentic System For Multi-Turn Video Production, Zhengyang Liang, Yan Shu, Cathal Gurrin, Nicu Sebe, Lizi Liao Jun 2026

Videocreator: An Agentic System For Multi-Turn Video Production, Zhengyang Liang, Yan Shu, Cathal Gurrin, Nicu Sebe, Lizi Liao

Research Collection School Of Computing and Information Systems

Recent advances in video generation models enable visually compelling single clips. However, real-world video creation is inherently continuous and iterative: creators refine content over multiple rounds while maintaining narrative, style, and entity consistency. Existing standalone generators are largely stateless and lack memory of previously generated segments, making it difficult to produce a coherent and consistent video project. To address this gap, we present VideoCreator, a unified video agent that integrates generation and understanding with a project-level memory system. VideoCreator leverages understanding capabilities to perform fine-grained analysis of newly produced content and uses persistent memory to retain and reuse prior context …


Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang Jun 2026

Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang

Research Collection School Of Computing and Information Systems

Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we …


Frozen Lvlms For Micro-Video Recommendation: A Systematic Study Of Feature Extraction And Fusion, Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du Jun 2026

Frozen Lvlms For Micro-Video Recommendation: A Systematic Study Of Feature Extraction And Fusion, Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du

Research Collection School Of Computing and Information Systems

Frozen Large Video Language Models (LVLMs) are increasingly employed in micro-video recommendation (MVR) due to their strong multimodal understanding. However, existing apporches typically deploy LVLMs as fixed black-box feature extractors without systematically comparing alternative representation strategies. To address this gap, we present the first systematic empirical study on various feature extraction paradigms and integration strategies, along with hierarchical representations from frozen LVLMs for MVR. Extensive experiments on representative LVLMs reveal that hidden states from multiple decoder layers provide richer and more effective representations for MVR. Guided by this insight, we propose the Dual Feature Fusion (DFF) Framework, a lightweight approach …


Co-Designing With Autistic Livestreamers: Care, Constraints, And Trade-Offs In Livestreaming, Terrance Mok, Anthony Tang, Lora Oehlberg Jun 2026

Co-Designing With Autistic Livestreamers: Care, Constraints, And Trade-Offs In Livestreaming, Terrance Mok, Anthony Tang, Lora Oehlberg

Research Collection School Of Computing and Information Systems

Autistic livestreamers use platforms like Twitch for social connection, self-expression, and community, but these spaces also impose ongoing social and emotional demands. Prior work has documented these experiences, but less is known about what autistic creators themselves envision for the tools and platforms they use. We address this gap through a Research through Design (RtD) co-design study with three autistic Twitch streamers, using speculative artefacts as discussion prompts to explore how participants reasoned about potential livestreaming technologies. Across three co-design activities, we identify three overarching tensions shaping autistic streaming practice: Expression versus Misinterpretation and Harm; Public Participation versus Control and …


“Grandpa, Can You Speak Nicer?”: Envisioned Chatbot Roles And Design Tensions In Intergenerational Communication Conflicts, Tianyi Zhang, Emran Bin Elias Poh, Yueyue Hou, Yi-Chieh Lee, Renwen Zhang, Jiannan Li, Anthony Tang Jun 2026

“Grandpa, Can You Speak Nicer?”: Envisioned Chatbot Roles And Design Tensions In Intergenerational Communication Conflicts, Tianyi Zhang, Emran Bin Elias Poh, Yueyue Hou, Yi-Chieh Lee, Renwen Zhang, Jiannan Li, Anthony Tang

Research Collection School Of Computing and Information Systems

Intergenerational conversations often break down when differences in tone, language, or expectations lead participants to feel dismissed or misunderstood. In this work, we explore how people envision AI-driven chatbot interventions for addressing communication problems in text-based intergenerational family chat. We conducted a scenario-based design interview with 10 pairs of family members from different generations, in which participants designed chatbot interventions that varied in intervention target and timing. Our findings show that participants expect chatbots to perform multiple themes of intervention, including mediating understanding, providing emotional support, offering evaluative commentary, and guiding interaction through behavioral suggestions. These expectations varied systematically across …


Group Conversational Agents: A Review Of Designs That Support And Shape Group Interaction, Shunyi Yeo, Tianyi Zhang, Scott Bateman, Gary Hsieh, Young-Ho Kim, Simon Tangi Perrault, Jiannan Li, Anthony Tang Jun 2026

Group Conversational Agents: A Review Of Designs That Support And Shape Group Interaction, Shunyi Yeo, Tianyi Zhang, Scott Bateman, Gary Hsieh, Young-Ho Kim, Simon Tangi Perrault, Jiannan Li, Anthony Tang

Research Collection School Of Computing and Information Systems

Conversational agents that participate in or mediate group interaction introduce challenges that extend beyond supporting individual users, raising new questions about how agents participate in and influence groups. To characterise this emerging design space, we present a systematic review of 53 peer-reviewed studies on group conversational agents (GCAs). We analyse how GCAs intervene in group-level processes, including participation regulation, conflict mediation, task alignment, and execution support. Using concepts from group research as an analytic lens, we organise prior GCA work around recurring group interactional challenges (orientation, conflict, alignment, and execution), and examine the roles agents are designed to play in …


The Stars Align: Modeling User Rating Calibration With Sparse Semantic Review Features, Rodrigo Alves, Antoine Ledent Jun 2026

The Stars Align: Modeling User Rating Calibration With Sparse Semantic Review Features, Rodrigo Alves, Antoine Ledent

Research Collection School Of Computing and Information Systems

User ratings are often treated as comparable across users, although identical scores may reflect different experiences. We study whether ratings can be viewed as user-specific discretizations of a shared semantic continuum derived from review text. Our method maps reviews into sparse semantic features with a sparse autoencoder and learns user-specific filters for each rating level. On Amazon Electronics, the learned embeddings align along a shared low-dimensional rating axis. Users differ mainly in how they anchor and partition this continuum, while preserving its overall ordinal structure. These findings support a semantic view of calibration beyond scalar bias correction.


Hide-And-Sweep: Detecting Concealed Cameras Via Led Illumination Sweeps, Jonghyuk Yun, Jaeyoung Moon, Yunseo Park, Sean Rui Xiang Tan, Byunghyun Kim, Rajesh Krishna Balan, Jun Han Jun 2026

Hide-And-Sweep: Detecting Concealed Cameras Via Led Illumination Sweeps, Jonghyuk Yun, Jaeyoung Moon, Yunseo Park, Sean Rui Xiang Tan, Byunghyun Kim, Rajesh Krishna Balan, Jun Han

Research Collection School Of Computing and Information Systems

Hidden cameras have increasingly infiltrated hotel and Airbnb rooms, posing serious privacy risks. Detecting such cameras is challenging because they are visually inconspicuous and often embedded inside everyday objects. Even worse, existing handheld detectors are manual and also rely on single-angle illumination and hence suffer from high false-positive rates. We present SweepLED (pronounced "sweepled")1, a practical hidden camera detection system that operates on a commodity smartphone augmented with an unobtrusive LED-embedded case. SweepLED performs LED sweeping - a controlled sequence of multi-angle illumination - while the user simply holds the phone still by hand, enabling the camera to capture how …


Rc-Nf: Robot-Conditioned Normalizing Flow For Real-Time Anomaly Detection In Robotic Manipulation, Shijie Zhou, Bin Zhu, Jiarui Yang, Xiangyu Zhao, Jingjing Chen, Yu-Gang Jiang Jun 2026

Rc-Nf: Robot-Conditioned Normalizing Flow For Real-Time Anomaly Detection In Robotic Manipulation, Shijie Zhou, Bin Zhu, Jiarui Yang, Xiangyu Zhao, Jingjing Chen, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

Recent advances in Vision-Language-Action (VLA) models have enabled robots to execute increasingly complex tasks. However, VLA models trained through imitation learning struggle to operate reliably in dynamic environments and often fail under Out-of-Distribution (OOD) conditions. To address this issue, we propose Robot-Conditioned Normalizing Flow(RC-NF), a real-time monitoring model for robotic anomaly detection and intervention that ensures the robot's state and the object's motion trajectory align with the task. RC-NF decouples the processing of task-aware robot and object states within the normalizing flow. It requires only positive samples for unsupervised training and calculates accurate robotic anomaly scores during inference through the …


Long-Term Mine Planning: A Survey Of Classical, Hybrid And Artificial Intelligence-Based Methods, Nurul Asyikeen Azhar, Aldy Gunawan, Shih-Fen Cheng, Erwin Leonardi Jun 2026

Long-Term Mine Planning: A Survey Of Classical, Hybrid And Artificial Intelligence-Based Methods, Nurul Asyikeen Azhar, Aldy Gunawan, Shih-Fen Cheng, Erwin Leonardi

Research Collection School Of Computing and Information Systems

The aim of long-term mine planning (LTMP) is two-fold: to maximize the net present value of profits (NPV) and determine how ores are sequentially processed over the lifetime. This scheduling task is computationally complex as it is rife with variables, constraints, periods, uncertainties, and unique operations. In this paper, we present trends in the literature in the recent decade. One trend is the shift from deterministic toward stochastic problems as they reflect real-world complexities. A complexity of growing concern is also in sustainable mine planning. Another trend is the shift from traditional operational research solutions — relying on exact or …


History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu Jun 2026

History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu

Research Collection School Of Computing and Information Systems

Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution in addressing the drawbacks of poor generalization in the training-based paradigm. However, existing LLM-based works primarily perform naive reasoning for decision-making and lack feedback, e.g., reviewing historical errors and predicting future potentials. Consequently, it may suffer from continuous failure for those initial error tasks. In this paper, we rethink LLM-based zero-shot VLN-CE and propose a new paradigm, named EvoNav, to improve …


Task Complexity Matters: An Empirical Study Of Reasoning In Llms For Sentiment Analysis, Donghao Huang, Zhaoxia Wang Jun 2026

Task Complexity Matters: An Empirical Study Of Reasoning In Llms For Sentiment Analysis, Donghao Huang, Zhaoxia Wang

Research Collection School Of Computing and Information Systems

Large language models (LLMs) with reasoning capabilities have fueled a compelling narrative that reasoning universally improves performance across language tasks. We test this claim through a comprehensive evaluation of 504 configurations across seven model families—including adaptive, conditional, and reinforcement learning-based reasoning architectures—on sentiment analysis datasets of varying granularity (binary, five-class, and 27-class emotion). Our findings reveal that reasoning effectiveness is strongly task-dependent, challenging prevailing assumptions: (1) Reasoning shows task-complexity dependence—binary classification degrades up to -19.9 F1% points (pp), while 27-class emotion recognition gains up to  +16.0 pp; (2) Distilled reasoning variants underperform base models by 3–18 pp on simpler tasks, …


Cfalr: Collaborative Filtering-Augmented Large Language Model For Personalized Fashion Outfit Recommendation, Yujuan Ding, Junrong Liao, Yunshan Ma, Yi Bin, Wenqi Fan, Tat-Seng Chua, Qing Li Jun 2026

Cfalr: Collaborative Filtering-Augmented Large Language Model For Personalized Fashion Outfit Recommendation, Yujuan Ding, Junrong Liao, Yunshan Ma, Yi Bin, Wenqi Fan, Tat-Seng Chua, Qing Li

Research Collection School Of Computing and Information Systems

Personalized outfit recommendation poses a significant challenge in e-commerce and social media platforms, requiring systems that balance user preferences with aesthetic compatibility. Collaborative filtering (CF) provides a traditional solution for this, but it struggles with data-sparse scenarios and complex user-item-outfit relationships. Meanwhile, existing template-based approaches are constrained by rigid pre-designed structures. To bridge these research gaps, we introduce CFALR (Collaborative Filtering-Augmented Large Language Model for Recommendation), a novel framework that synergizes collaborative filtering with large language models for personalized outfit recommendation. Specifically, CFALR describes user-outfit interactions in natural language and leverages LLMs to capture fashion semantics while employing CF-enhanced embeddings …


To Wait Or To Transfer? A Three-Level Optimization Framework For Intermodal Transfer Coordination In First Train Timetabling And Bus Bridging Services Management, Hao Li, Liujiang Kang, Norman Weik, Huijun Sun, Qingying Lai, Zhiguang Cao Jun 2026

To Wait Or To Transfer? A Three-Level Optimization Framework For Intermodal Transfer Coordination In First Train Timetabling And Bus Bridging Services Management, Hao Li, Liujiang Kang, Norman Weik, Huijun Sun, Qingying Lai, Zhiguang Cao

Research Collection School Of Computing and Information Systems

This study addresses the integrated optimization of the first train timetabling and bus bridging service design (FTT-BBSD) for morning transfer challenges, two critical but interdependent passenger services in the public transit system. In contrast to most existing studies and conventional approaches, this study explicitly models the influence of passenger path choices and transfer mode selections on FTT-BBSD. Through a novel dual-level network representation that integrates subway and bus systems, we formulate the FTT-BBSD problem as a mixed-integer nonlinear programming model. The model simultaneously determines subway and bus timetables and bridging line deployment to minimize total travel time for all first …


A Pruning-Based Question-Answering For Interactive Video Search: A Simple Baseline, Yu Tong Cheng, Phuong Anh Nguyen, Chong-Wah Ngo Jun 2026

A Pruning-Based Question-Answering For Interactive Video Search: A Simple Baseline, Yu Tong Cheng, Phuong Anh Nguyen, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

There are various factors affecting the performance of video search. An imprecise query will enlarge search space and reduce the discriminative power of ranking functions. This problem is further exacerbated by the presence of numerous visually or semantically similar videos in large datasets. Consequently, users need to painstakingly browse through many highly similar candidates to locate the search target, leading to increased cognitive load and inefficient searching. Ideally, engaging users through interactive questioning to resolve uncertainties in the search process is an effective strategy for progressively narrowing down the search space. However, despite rapid advances in deep learning, generating informative …


Interfold: Learning Interpretable Diffusion Manifolds Beyond Binary Samples, Alexander Vincent Lewi, Rainer Tan, Shengfeng He Jun 2026

Interfold: Learning Interpretable Diffusion Manifolds Beyond Binary Samples, Alexander Vincent Lewi, Rainer Tan, Shengfeng He

Research Collection School Of Computing and Information Systems

We propose InterFold, a framework for learning and applying interpretable semantic manifolds in latent diffusion models, without requiring binary or paired supervision. Existing methods for semantic editing either rely on limited paired data or uncover only coarse, unsupervised directions that fail to capture user-specific, fine-grained attributes. InterFold addresses these limitations by learning a target attribute manifold in the H-space of diffusion models using only a set of positive, unlabeled examples. To edit a new image, InterFold projects its H-space representation toward this learned manifold through test-time optimization, enabling precise, identity-preserving modifications of complex, non-binary concepts. To make these edits effective …