Open Access. Powered by Scholars. Published by Universities.®
Graphics and Human Computer Interfaces Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (313)
- Artificial Intelligence and Robotics (168)
- Software Engineering (77)
- Engineering (70)
- Computer Engineering (69)
-
- Data Storage Systems (61)
- OS and Networks (32)
- Social and Behavioral Sciences (26)
- Theory and Algorithms (22)
- Numerical Analysis and Scientific Computing (20)
- Communication (18)
- Business (17)
- Information Security (13)
- Education (11)
- Medicine and Health Sciences (11)
- Programming Languages and Compilers (10)
- Social Media (9)
- Health Information Technology (8)
- Educational Methods (5)
- Technology and Innovation (5)
- Asian Studies (4)
- Communication Technology and New Media (4)
- Data Science (4)
- E-Commerce (4)
- International and Area Studies (4)
- Computer and Systems Architecture (3)
- Digital Communications and Networking (3)
- Keyword
-
- Visualization (21)
- Computer vision (17)
- Graph Neural Networks (14)
- Deep learning (13)
- Task analysis (12)
-
- Feature extraction (11)
- Training (11)
- Accessibility (10)
- Deep Learning (10)
- Design (10)
- Face recognition (9)
- Graph neural networks (9)
- Semantics (9)
- Categorization (8)
- Gamification (8)
- Virtual reality (8)
- Codes (7)
- Data visualization (7)
- Domain adaptation (7)
- Human-centered computing (7)
- Recipe retrieval (7)
- Video search (7)
- Web video (7)
- Augmented reality (6)
- Cross-modal retrieval (6)
- Few-shot learning (6)
- Human-computer interaction (6)
- Image search (6)
- Knowledge Graph (6)
- Object detection (6)
- Publication Year
- Publication
- Publication Type
Articles 1 - 30 of 938
Full-Text Articles in Graphics and Human Computer Interfaces
Defense-To-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks In Vision-Language Models, Yunhan Zhao, Xiang Zheng, Yige Li, Xingjun Ma
Defense-To-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks In Vision-Language Models, Yunhan Zhao, Xiang Zheng, Yige Li, Xingjun Ma
Research Collection School Of Computing and Information Systems
Despite their superb capabilities, Vision-Language Models (VLMs) have been shown to be vulnerable to jailbreak attacks. While recent jailbreaks have achieved notable progress, their effectiveness and efficiency can still be improved. In this work, we reveal an interesting phenomenon: incorporating weak defense cues into the attack pipeline can significantly enhance both the effectiveness and efficiency of jailbreaks on VLMs. Building on this insight, we propose Defense2Attack, a novel jailbreak method that bypasses the safety guardrails of VLMs by leveraging defensive patterns to guide jailbreak prompt construction. Specifically, Defense2Attack consists of three key components: (1) a visual optimizer that embeds universal …
Neural Symphony Of Flow Experience: Evidence For High-Dimensional Metastable Dynamics, Abdelrahman B. M. Eldaly, Kris Zhangguang Kang, Fiona Fui-Hoon Nah, Leanne Lai-Hang Chan, Keng Siau, Xiao Fan Liu, Richard Huskey, Langtao Chen, Tejaswini Yelamanchili, Rene Weber
Neural Symphony Of Flow Experience: Evidence For High-Dimensional Metastable Dynamics, Abdelrahman B. M. Eldaly, Kris Zhangguang Kang, Fiona Fui-Hoon Nah, Leanne Lai-Hang Chan, Keng Siau, Xiao Fan Liu, Richard Huskey, Langtao Chen, Tejaswini Yelamanchili, Rene Weber
Research Collection School Of Computing and Information Systems
Flow, an optimal experience characterized by deep immersion and engagement in an activity, has been extensively studied in behavioral research. However, its neural dynamic mechanism remains poorly understood. In a within-subject video gaming experiment, we captured neural activity underlying flow, boredom, and anxiety using a 64-channel electroencephalogram (EEG) system. Compared to boredom and anxiety, flow exhibits the highest global functional connectivity, metastability, and dimensionality of dynamic functional connectivity patterns, suggesting that flow is a highly adaptable process that is supported by high-dimensional neural dynamics. Unlike previous studies that focused on identifying static or localized brain activity, we examine the neural …
Restoring Linguistic Grounding In Vla Models Via Train-Free Attention Recalibration, Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen
Restoring Linguistic Grounding In Vla Models Via Train-Free Attention Recalibration, Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen
Research Collection School Of Computing and Information Systems
Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generalist robotic policies. However, their reliability under Out-Of-Distribution (OOD) instructions remains underexplored. In this paper, we reveal a critical failure mode in which VLA policies continue executing visually plausible actions even when the language instruction contradicts the scene. We refer to this phenomenon as linguistic blindness, where VLA policies prioritize visual priors over instruction semantics during action generation. To systematically analyze this issue, we introduce ICBench, a diagnostic benchmark constructed from the LIBERO dataset that probes language–action coupling …
Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo
Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo
Research Collection School Of Computing and Information Systems
Rapid advances in AI-generated image (AIGI) technology enable highly realistic synthesis, threatening public information integrity and security. Recent studies have demonstrated that incorporating texture-level artifact features alongside semantic features into multimodal large language models (MLLMs) can enhance their AIGI detection capability. However, our preliminary analyses reveal that artifact features exhibit high intra-feature similarity, leading to an almost uniform attention map after the softmax operation. This phenomenon causes attention dilution, thereby hindering effective fusion between semantic and artifact features. To overcome this limitation, we propose a lightweight fusion adapter, TranX-Adapter, which integrates a Task-aware Optimal-Transport Fusion that leverages the Jensen-Shannon divergence …
Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen
Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen
Research Collection School Of Computing and Information Systems
Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object’s state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action–object interactions into …
Activity Transition Graph Generation: How Far Are We?, Jiakun Liu, Peixin Zhang, Han Hu, Yonghui Liu, Wei Minn, Ferdian Thung, Shahar Maoz, Eran Toch, Debin Gao, David Lo
Activity Transition Graph Generation: How Far Are We?, Jiakun Liu, Peixin Zhang, Han Hu, Yonghui Liu, Wei Minn, Ferdian Thung, Shahar Maoz, Eran Toch, Debin Gao, David Lo
Research Collection School Of Computing and Information Systems
Android applications (i.e., apps) are indispensable nowadays and are getting bigger and bigger with an increasing number offunctionalities. To understand how to access functionalities in an app, prior studies proposed tools to model the transitionsbetween functionalities with the activity transition graph (ATG). ATG is an important data structure and has been used forvarious Android app analyses, including app design, understanding, and testing. However, there is no benchmarking work onATG generation. It is still unclear whether the transitions identified by tools are correct and how many transitions are missed.To fill this gap, we manually identified all transitions in 98 applications to …
Enhancing Pointing Gestures Of Non-Hmd Users In Asymmetric Collocated Mixed Reality Collaboration, Nam-Dang Vo, Van-Vinh Thai, Anthony Tang, Khanh-Duy Le
Enhancing Pointing Gestures Of Non-Hmd Users In Asymmetric Collocated Mixed Reality Collaboration, Nam-Dang Vo, Van-Vinh Thai, Anthony Tang, Khanh-Duy Le
Research Collection School Of Computing and Information Systems
A common collocated group setting in mixed-reality (MR) collaboration is a person wearing a MR headset (HMD user) and presenting MR contents to audiences who are not provided with such specialized devices (Non-HMD users). In this setting, while Non-HMD users can view the MR environment shown on a large physical display, it still remains challenging for the HMD user to interpret their pointing gesture when they spatially refer to objects in the MR environment. To address this, we designed and evaluated two pointing techniques—SCREEN and SCREEN+SPACE—that support Non-HMD users in referring to MR content. Screen pointing allows users to refer …
Not Too Early, Not All At Once: Design Tensions In Ai-Mediated Self-Disclosure In Online Dating, Pei-Hua Tsai, Tianyi Zhang, Emran Bin Elias Poh, Anthony Tang, Yung-Ju Chang
Not Too Early, Not All At Once: Design Tensions In Ai-Mediated Self-Disclosure In Online Dating, Pei-Hua Tsai, Tianyi Zhang, Emran Bin Elias Poh, Anthony Tang, Yung-Ju Chang
Research Collection School Of Computing and Information Systems
Online dating relies on self-disclosure, yet initial conversations are fragile: users must navigate uncertainty around timing, boundaries, and reciprocity with little shared context. While advances in AI raise the possibility of mediating disclosure, how such support might reshape the experience of early-stage relational disclosure remains underexplored. We conducted 29 semi-structured interviews to examine how daters envision AI-mediated self-disclosure in online dating. Our findings surface recurring design tensions rather than simple opportunities or risks. Participants welcomed guidance that could pace disclosure, support reflection, and reduce social awkwardness, but stressed preserving agency and authorship. They valued interpretive assistance for sense-making of ambiguous …
A Pruning-Based Question-Answering For Interactive Video Search: A Simple Baseline, Yu Tong Cheng, Phuong Anh Nguyen, Chong-Wah Ngo
A Pruning-Based Question-Answering For Interactive Video Search: A Simple Baseline, Yu Tong Cheng, Phuong Anh Nguyen, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
There are various factors affecting the performance of video search. An imprecise query will enlarge search space and reduce the discriminative power of ranking functions. This problem is further exacerbated by the presence of numerous visually or semantically similar videos in large datasets. Consequently, users need to painstakingly browse through many highly similar candidates to locate the search target, leading to increased cognitive load and inefficient searching. Ideally, engaging users through interactive questioning to resolve uncertainties in the search process is an effective strategy for progressively narrowing down the search space. However, despite rapid advances in deep learning, generating informative …
Videocreator: An Agentic System For Multi-Turn Video Production, Zhengyang Liang, Yan Shu, Cathal Gurrin, Nicu Sebe, Lizi Liao
Videocreator: An Agentic System For Multi-Turn Video Production, Zhengyang Liang, Yan Shu, Cathal Gurrin, Nicu Sebe, Lizi Liao
Research Collection School Of Computing and Information Systems
Recent advances in video generation models enable visually compelling single clips. However, real-world video creation is inherently continuous and iterative: creators refine content over multiple rounds while maintaining narrative, style, and entity consistency. Existing standalone generators are largely stateless and lack memory of previously generated segments, making it difficult to produce a coherent and consistent video project. To address this gap, we present VideoCreator, a unified video agent that integrates generation and understanding with a project-level memory system. VideoCreator leverages understanding capabilities to perform fine-grained analysis of newly produced content and uses persistent memory to retain and reuse prior context …
Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang
Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang
Research Collection School Of Computing and Information Systems
Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we …
Co-Designing With Autistic Livestreamers: Care, Constraints, And Trade-Offs In Livestreaming, Terrance Mok, Anthony Tang, Lora Oehlberg
Co-Designing With Autistic Livestreamers: Care, Constraints, And Trade-Offs In Livestreaming, Terrance Mok, Anthony Tang, Lora Oehlberg
Research Collection School Of Computing and Information Systems
Autistic livestreamers use platforms like Twitch for social connection, self-expression, and community, but these spaces also impose ongoing social and emotional demands. Prior work has documented these experiences, but less is known about what autistic creators themselves envision for the tools and platforms they use. We address this gap through a Research through Design (RtD) co-design study with three autistic Twitch streamers, using speculative artefacts as discussion prompts to explore how participants reasoned about potential livestreaming technologies. Across three co-design activities, we identify three overarching tensions shaping autistic streaming practice: Expression versus Misinterpretation and Harm; Public Participation versus Control and …
“From Remembering To Shaping”: Narrating Shared Experiences By Co-Designing Cultural Heritage Artifacts In Collaborative Vr, Yushang Yang, Fanxu Meng, Fiona Fui-Hoon Nah, L. C. Ray
“From Remembering To Shaping”: Narrating Shared Experiences By Co-Designing Cultural Heritage Artifacts In Collaborative Vr, Yushang Yang, Fanxu Meng, Fiona Fui-Hoon Nah, L. C. Ray
Research Collection School Of Computing and Information Systems
The ways people remember and recall places reveal an invisible aspect of cultural heritage (CH), reflecting how individuals and communities relate to these places. Heritage is communal, emerging through collaboratively constructed narratives rather than individual records. To probe how people may share collective memories, we designed an immersive two-person workflow for collaboratively co-designing 3D artifacts and environments in virtual heritage locations, using Generative AI (GenAI) to instantiate these intangible memories. Observations of the co-creation process revealed that participants merged prompts and model placements when negotiating different perspectives. They used spatial operations to compose scenes, and also to express personal and …
History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu
History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu
Research Collection School Of Computing and Information Systems
Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution in addressing the drawbacks of poor generalization in the training-based paradigm. However, existing LLM-based works primarily perform naive reasoning for decision-making and lack feedback, e.g., reviewing historical errors and predicting future potentials. Consequently, it may suffer from continuous failure for those initial error tasks. In this paper, we rethink LLM-based zero-shot VLN-CE and propose a new paradigm, named EvoNav, to improve …
Happycal: Designing Text And Image-Based Supports For Savouring Positive Work Experiences, Molly Stewart, Minghao Cai, Anthony Tang, Sam Liu, Chris Mosunic, Sowmya Somanath
Happycal: Designing Text And Image-Based Supports For Savouring Positive Work Experiences, Molly Stewart, Minghao Cai, Anthony Tang, Sam Liu, Chris Mosunic, Sowmya Somanath
Research Collection School Of Computing and Information Systems
Savouring positive work experiences can promote positive affect and well-being at work, yet there is limited guidance on how digital applications can support workers to engage in savouring. We developed HappyCal, a work-focused savouring application offering two forms of savouring support: text-based, a common modality in workplace reflection tools, and images, a largely unexplored approach in work-related savouring. We conducted an exploratory qualitative study where participants (N=36) used HappyCal over five days and engaged in savouring through either a text-only modality (n=17) or text input paired with image output (n=19). We found that (1) participants in both groups reported heightened …
Teacher-Student Diffusion Model For Text-Driven 3d Hand Motion Generation, Ching Lam Cheng, Bin Zhu, Shengfeng He
Teacher-Student Diffusion Model For Text-Driven 3d Hand Motion Generation, Ching Lam Cheng, Bin Zhu, Shengfeng He
PhD Student’s Publications Collection
Generating realistic 3D hand motion from natural language is vital for VR, robotics, and human-computer interaction. Existing methods either focus on full-body motion, overlooking detailed hand gestures, or require explicit 3D object meshes, limiting generality. We propose TSHaMo, a model-agnostic teacher-student diffusion framework for text-driven hand motion generation. The student model learns to synthesize motions from text alone, while the teacher leverages auxiliary signals (e.g., MANO parameters) to provide structured guidance during training. A co-training strategy enables the student to benefit from the teacher’s intermediate predictions while remaining text-only at inference. Evaluated using two diffusion backbones on GRAB and H2O, …
Portrait Shadow Removal Via Self-Exemplar Illumination Equalization, Qian Huang, Cheng Xu, Guiqing Li, Ziheng Wu, Shengxin Liu, Shengfeng He
Portrait Shadow Removal Via Self-Exemplar Illumination Equalization, Qian Huang, Cheng Xu, Guiqing Li, Ziheng Wu, Shengxin Liu, Shengfeng He
Research Collection School Of Computing and Information Systems
We introduce the Self-Exemplar Illumination Equalization Network, designed specifically for effective portrait shadow removal. The core idea of our method is that partially shadowed portraits can find ideal exemplars within their non-shadowed facial regions. Rather than directly fusing two distinct classes of facial features, our approach utilizes non-shadowed regions as an illumination indicator to equalize the shadowed regions, generating deshadowed results without boundary-merging artifacts. Our network comprises cascaded Self-Exemplar Illumination Equalization Blocks (SExmBlock), each containing two modules: a self-exemplar feature matching module and a feature-level illumination rectification module. The former identifies and applies internal illumination exemplars to shadowed areas, producing …
Teamwise: Exploring Virtually Embodied Ai Facilitation For Video-Based Team Onboarding, Venkata Akhila Rani Obilisetty, Mikkeline Elleby, Anthony Tang, April Yi Wang
Teamwise: Exploring Virtually Embodied Ai Facilitation For Video-Based Team Onboarding, Venkata Akhila Rani Obilisetty, Mikkeline Elleby, Anthony Tang, April Yi Wang
Research Collection School Of Computing and Information Systems
AI-mediated facilitation has emerged as a scalable approach to supporting onboarding and coordination in newly formed remote teams, yet existing systems are predominantly text-based. To explore how video-based, virtually embodied AI facilitators shape team experiences, we present TeamWise, which joins video-based onboarding meetings as an on-screen avatar. TeamWise guides teams through a structured facilitation flow of low-stakes activities to foster rapport, mutual awareness, and shared identity. While the overall sequence of activities and facilitation goals is predefined, the facilitator’s turn-by-turn utterances are generated dynamically by an LLM in response to participant input. We conducted a formative study of TeamWise to …
Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong
Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong
Research Collection School Of Computing and Information Systems
Conventionally, human intuition defines vision as a modality of passive optical sensing, relying on ambient light to perceive the environment. However, active optical sensing, which involves emitting and receiving signals, offers unique advantages by capturing both radiometric and geometric properties of the environment, independent of external illumination conditions. This work focuses on advancing active optical sensing using Light Detection and Ranging (LiDAR), which captures intensity data, enabling the estimation of surface reflectance that remains invariant under varying illumination. Such properties are crucial for robotic perception tasks, including detection, recognition, segmentation, and Simultaneous Localization and Mapping (SLAM). A key challenge with …
Learning Feature Inversion For Multi-Class Anomaly Detection Under General-Purpose Coco-Ad Benchmark, Jiangning Zhang, Chengjie Wang, Xiangtai Li, Guanzhong Tian, Zhucun Xue, Yong Liu, Guansong Pang, Dacheng Tao
Learning Feature Inversion For Multi-Class Anomaly Detection Under General-Purpose Coco-Ad Benchmark, Jiangning Zhang, Chengjie Wang, Xiangtai Li, Guanzhong Tian, Zhucun Xue, Yong Liu, Guansong Pang, Dacheng Tao
Research Collection School Of Computing and Information Systems
Anomaly detection (AD) is often focused on detecting anomaly areas for industrial quality inspection and medical lesion examination. However, due to the specific scenario targets, the data scale for AD is relatively small, and evaluation metrics are still deficient compared to classic vision tasks, such as object detection and semantic segmentation. To fill these gaps, this work first constructs a large-scale and general-purpose COCO-AD dataset by extending COCO to the AD field. This enables fair evaluation and sustainable development for different methods on this challenging benchmark. Moreover, current metrics such as AU-ROC have nearly reached saturation on simple datasets, which …
Weakly Supervised Video Anomaly Detection And Localization With Spatio-Temporal Prompts, Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang, Qingsen Yan, Peng Wang, Yanning Zhang
Weakly Supervised Video Anomaly Detection And Localization With Spatio-Temporal Prompts, Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang, Qingsen Yan, Peng Wang, Yanning Zhang
Research Collection School Of Computing and Information Systems
Current weakly supervised video anomaly detection (WSVAD) task aims to achieve frame-level anomalous event detection with only coarse video-level annotations available. Existing works typically involve extracting global features from full-resolution video frames and training frame-level classifiers to detect anomalies in the temporal dimension. However, most anomalous events tend to occur in localized spatial regions rather than the entire video frames, which implies existing frame-level feature based works may be misled by the dominant background information and lack the interpretation of the detected anomalies. To address this dilemma, this paper introduces a novel method called STPrompt that learns spatio-temporal prompt embeddings …
Semat: Semantic Enhanced Natural Image Interactive Matting, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Qianru Sun, Yang Tang, Bo Li, Pan Zhou
Semat: Semantic Enhanced Natural Image Interactive Matting, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Qianru Sun, Yang Tang, Bo Li, Pan Zhou
Research Collection School Of Computing and Information Systems
Recent approaches attempt to adapt powerful interactive segmentation models, such as SAM, to interactive matting and fine-tune the models based on synthetic matting datasets. However, models trained on synthetic data fail to generalize to complex and occlusion scenes. We address this challenge by proposing a new matting dataset based on the COCO dataset, namely COCO-Matting. It selects real-world complex images from COCO and converts semantic segmentation masks to matting labels. The built COCO-Matting comprises an extensive collection of 36,980 human instance-level alpha mattes in complex natural scenarios. Furthermore, existing SAM-based matting methods extract intermediate features and masks from a frozen …
Dragging With Geometry: From Pixels To Geometry-Guided Image Editing, Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou
Dragging With Geometry: From Pixels To Geometry-Guided Image Editing, Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou
Research Collection School Of Computing and Information Systems
Interactive point-based image editing serves as a controllable editor, enabling precise and flexible manipulation of image content. However, most drag-based methods operate primarily on the 2D pixel plane with limited use of 3D cues. As a result, they often produce imprecise and inconsistent edits, particularly in geometry-intensive scenarios such as rotations and perspective transformations. To address these limitations, we propose a novel geometry-guided drag-based image editing method—GeoDrag, which addresses three key challenges: 1) incorporating 3D geometric cues into pixel-level editing, 2) mitigating discontinuities caused by geometry-only guidance, and 3) resolving conflicts arising from multi-point dragging. Built upon a unified displacement …
Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou
Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou
Research Collection School Of Computing and Information Systems
While text-to-3D generation has attracted growing interest, existing methods often struggle to produce 3D assets that align well with human preferences. Current preference alignment techniques for 3D content typically rely on hardly-collected preference-paired multi-view 2D images to train 2D reward models, when then guide 3D generation — leading to geometric artifacts, such as the Janus face problem and geometric incompleteness, due to their inherent 2D bias. To address these limitations, we construct 3D-MeshPref, the first large-scale unpaired 3D preference dataset, featuring diverse 3D meshes annotated by a large language model and refined by human evaluators. We then develop RewardCS, the …
From Spatial To Actions: Grounding Vision-Language-Action Model In Spatial Foundation Priors, Zhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu, Lei Zhou, Chenchen Liu, Dong Wang, Francis E. H. Tay, Sijin Chen, Ziwei Liu, Yuxiao Liu, Xinghang Li, Pan Zhou
From Spatial To Actions: Grounding Vision-Language-Action Model In Spatial Foundation Priors, Zhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu, Lei Zhou, Chenchen Liu, Dong Wang, Francis E. H. Tay, Sijin Chen, Ziwei Liu, Yuxiao Liu, Xinghang Li, Pan Zhou
Research Collection School Of Computing and Information Systems
Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require specialized sensors and transfer poorly across modalities, or inject weak cues that lack geometry and degrade vision-language alignment. In this work, we introduce FALCON (From Spatial to Action), a novel paradigm that injects rich 3D spatial tokens into the action head. FALCON leverages spatial foundation models to deliver strong geometric priors from RGB alone, and includes an Embodied Spatial Model that can optionally fuse depth, or pose …
Who You Explain To Matters: Learning By Explaining To Conversational Agents With Different Pedagogical Roles, Zhengtao Xu, Junti Zhang, Anthony Tang, Yi-Chieh Lee
Who You Explain To Matters: Learning By Explaining To Conversational Agents With Different Pedagogical Roles, Zhengtao Xu, Junti Zhang, Anthony Tang, Yi-Chieh Lee
Research Collection School Of Computing and Information Systems
Conversational agents are increasingly used in education for learning support. An application is “learning by explaining”, where learners explain their understanding to an agent. However, existing research focuses on single roles, leaving it unclear how different pedagogical roles influence learners’ interaction patterns, learning outcomes and experiences. We conducted a between-subjects study (N=96) comparing agents with three pedagogical roles (Tutee, Peer, Challenger) and a control condition while learning an economics concept. We found that different pedagogical roles shaped learning dynamics, including interaction patterns and experiences. Specifically, the Tutee agent elicited the most cognitive investment but led to high pressure. The Peer …
Invert Your Prompt: Editing-Aware Diffusion Inversion, Yangyang Xu, Wenqi Shao, Yong Du, Haiming Zhu, Yang Zhou, Jiayuan Xie, Ping Luo, Shengfeng He
Invert Your Prompt: Editing-Aware Diffusion Inversion, Yangyang Xu, Wenqi Shao, Yong Du, Haiming Zhu, Yang Zhou, Jiayuan Xie, Ping Luo, Shengfeng He
Research Collection School Of Computing and Information Systems
Recent advancements in text-guided diffusion models have enabled powerful image manipulation capabilities. However, balancing reconstruction fidelity and editability for real images remains a significant challenge. In this work, we introduce Editing Inversion (EditInv), a novel framework that inverts and edits real images for specific editing tasks by optimizing specific prompt embeddings within the extended space. By leveraging distinct embeddings across different U-Net layers and time steps, EditInv seamlessly integrates inversion and editing through reciprocal optimization, ensuring both high fidelity and precise editability. This hierarchical editing mechanism classifies tasks into structure, appearance, and global edits, optimizing only those embeddings that are …
Cylindformer: Image-To-Point Cloud Registration With Cylindrical Transformer, Jingtao Wang, Hao Tang, Yanpeng Sun, Shengfeng He, Zechao Li
Cylindformer: Image-To-Point Cloud Registration With Cylindrical Transformer, Jingtao Wang, Hao Tang, Yanpeng Sun, Shengfeng He, Zechao Li
Research Collection School Of Computing and Information Systems
Accurate correspondence extraction between distinctive pixel-wise and point-wise features is critical for image-to-point cloud (I2P) registration. Recent efforts leveraging Transformers for I2P feature representation have demonstrated potential, primarily by first capturing intra-modality global contextual dependencies via self-attention, and then learning cross-modality correlations via cross-attention. The strength of vanilla Transformers lies in modeling cross-modality global feature correlations. However, such mechanisms often struggle with the structural disparity between dense image pixels and sparse 3D points, hindering the establishment of fine-grained correspondences. Moreover, global attention may introduce ambiguity, as interactions with many inconsistent regions of intra-modality may degrade feature distinctiveness. To address these …
Zero-Shot Video Translation Via Token Warping, Haiming Zhu, Yangyang Xu, Jun Yu, Shengfeng He
Zero-Shot Video Translation Via Token Warping, Haiming Zhu, Yangyang Xu, Jun Yu, Shengfeng He
Research Collection School Of Computing and Information Systems
With the revolution of generative AI, video-related tasks have been widely studied. However, current state-of-the-art video models still lag behind image models in visual quality and user control over generated content. In this paper, we introduce TokenWarping, a novel framework for temporally coherent video translation. Existing diffusion-based video editing approaches rely solely on key and value patches in self-attention to ensure temporal consistency, often sacrificing the preservation of local and structural regions. Critically, these methods overlook the significance of the query patches in achieving accurate feature aggregation and temporal coherence. In contrast, TokenWarping leverages complementary token priors by constructing temporal …
Lagrangian Motion Fields For Long-Term Motion Generation, Yifei Yang, Zikai Huang, Chenshu Xu, Shengfeng He
Lagrangian Motion Fields For Long-Term Motion Generation, Yifei Yang, Zikai Huang, Chenshu Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
Long-term motion generation is a challenging task that requires producing coherent and realistic sequences over extended durations. Current methods primarily rely on framewise motion representations, which capture only static spatial details and overlook temporal dynamics. This approach leads to significant redundancy across the temporal dimension, complicating the generation of effective long-term motion. To overcome these limitations, we introduce the novel concept of Lagrangian Motion Fields, specifically designed for long-term motion generation. By treating each joint as a Lagrangian particle with uniform velocity over short intervals, our approach condenses motion representations into a series of "supermotions" (analogous to superpixels). This method …