Open Access. Powered by Scholars. Published by Universities.®
Graphics and Human Computer Interfaces Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (313)
- Artificial Intelligence and Robotics (168)
- Software Engineering (77)
- Engineering (70)
- Computer Engineering (69)
-
- Data Storage Systems (61)
- OS and Networks (32)
- Social and Behavioral Sciences (26)
- Theory and Algorithms (22)
- Numerical Analysis and Scientific Computing (20)
- Communication (18)
- Business (17)
- Information Security (13)
- Education (11)
- Medicine and Health Sciences (11)
- Programming Languages and Compilers (10)
- Social Media (9)
- Health Information Technology (8)
- Educational Methods (5)
- Technology and Innovation (5)
- Asian Studies (4)
- Communication Technology and New Media (4)
- Data Science (4)
- E-Commerce (4)
- International and Area Studies (4)
- Computer and Systems Architecture (3)
- Digital Communications and Networking (3)
- Keyword
-
- Visualization (21)
- Computer vision (17)
- Graph Neural Networks (14)
- Deep learning (13)
- Task analysis (12)
-
- Feature extraction (11)
- Training (11)
- Accessibility (10)
- Deep Learning (10)
- Design (10)
- Face recognition (9)
- Graph neural networks (9)
- Semantics (9)
- Categorization (8)
- Gamification (8)
- Virtual reality (8)
- Codes (7)
- Data visualization (7)
- Domain adaptation (7)
- Human-centered computing (7)
- Recipe retrieval (7)
- Video search (7)
- Web video (7)
- Augmented reality (6)
- Cross-modal retrieval (6)
- Few-shot learning (6)
- Human-computer interaction (6)
- Image search (6)
- Knowledge Graph (6)
- Object detection (6)
- Publication Year
- Publication
- Publication Type
Articles 91 - 120 of 938
Full-Text Articles in Graphics and Human Computer Interfaces
Efficient Prompt Tuning For Hierarchical Ingredient Recognition, Yinxuan Gui, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Efficient Prompt Tuning For Hierarchical Ingredient Recognition, Yinxuan Gui, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Fine-grained ingredient recognition presents a significant challenge due to the diverse appearances of ingredients, resulting from different cutting and cooking methods. While existing approaches have shown promising results, they still require extensive training costs and focus solely on fine-grained ingredient recognition. In this paper, we address these limitations by introducing an efficient prompt-tuning framework that adapts pretrained visual-language models (VLMs), such as CLIP, to the ingredient recognition task without requiring full model finetuning. Additionally, we introduce three-level ingredient hierarchies to enhance both training performance and evaluation robustness. Specifically, we propose a hierarchical ingredient recognition task, designed to evaluate model performance …
Action Dubber: Timing Audible Actions Via Inflectional Flow, Wenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li, Shengfeng He
Action Dubber: Timing Audible Actions Via Inflectional Flow, Wenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li, Shengfeng He
Research Collection School Of Computing and Information Systems
We introduce the task of Audible Action Temporal Localization, which aims to identify the spatiotemporal coordinates of audible movements. Unlike conventional tasks such as action recognition and temporal action localization, which broadly analyze video content, our task focuses on the distinct kinematic dynamics of audible actions. It is based on the premise that key actions are driven by inflectional movements; for example, collisions that produce sound often involve abrupt changes in motion. To capture this, we propose T A2Net, a novel architecture that estimates inflectional flow using the second derivative of motion to determine collision timings without relying on audio …
Robust Relevance Feedback For Interactive Known-Item Video Search, Zhixin Ma, Chong-Wah Ngo
Robust Relevance Feedback For Interactive Known-Item Video Search, Zhixin Ma, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Known-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to infer user intent-inapplicable. PicHunter addresses this issue by asking users to select the top-k most similar examples to the unique search target from a displayed set. Under ideal conditions, when the user's perception aligns closely with the machine's perception of similarity, consistent and precise judgments can elevate the target to the top position within a few iterations. However, in practical scenarios, expecting users to provide consistent judgments is often unrealistic, especially when the underlying embedding features used for …
Retrieval Augmented Generation For Dynamic Graph Modeling, Yuxia Wu, Lizi Liao, Yuan Fang
Retrieval Augmented Generation For Dynamic Graph Modeling, Yuxia Wu, Lizi Liao, Yuan Fang
Research Collection School Of Computing and Information Systems
Modeling dynamic graphs, such as those found in social networks, recommendation systems, and e-commerce platforms, is crucial for capturing evolving relationships and delivering relevant insights over time. Traditional approaches primarily rely on graph neural networks with temporal components or sequence generation models, which often focus narrowly on the historical context of target nodes. This limitation restricts the ability to adapt to new and emerging patterns in dynamic graphs. To address this challenge, we propose a novel framework, Retrieval-Augmented Generation for Dy namic Graph modeling (RAG4DyG ), which enhances dynamic graph predictions by incorporating contextually and temporally relevant examples from broader …
Nexusgs: Sparse View Synthesis With Epipolar Depth Priors In 3d Gaussian Splatting, Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du
Nexusgs: Sparse View Synthesis With Epipolar Depth Priors In 3d Gaussian Splatting, Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du
Research Collection School Of Computing and Information Systems
Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have noticeably advanced photo-realistic novel view synthesis using images from densely spaced camera viewpoints. However, these methods struggle in few-shot scenarios due to limited supervision. In this paper, we present NexusGS, a 3DGS-based approach that enhances novel view synthesis from sparse-view images by directly embedding depth information into point clouds, without relying on complex manual regularizations. Exploiting the inherent epipolar geometry of 3DGS, our method introduces a novel point cloud densification strategy that initializes 3DGS with a dense point cloud, reducing randomness in point placement while preventing over-smoothing and overfitting. Specifically, …
Hd-Epic: A Highly-Detailed Egocentric Video Dataset, Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Kumar Parida, Kaiting Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rhodri Guerrier, Fahd Abdelazim, Bin Zhu, Davide Moltisanti, Michael Wray, Hazel Doughty, Dima Damen
Hd-Epic: A Highly-Detailed Egocentric Video Dataset, Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Kumar Parida, Kaiting Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rhodri Guerrier, Fahd Abdelazim, Bin Zhu, Davide Moltisanti, Michael Wray, Hazel Doughty, Dima Damen
Research Collection School Of Computing and Information Systems
We present a validation dataset of newly-collected kitchenbased egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional values, moving objects, and audio annotations. Importantly, all annotations are grounded in 3D through digital twinning of the scene, fixtures, object locations, and primed with gaze. Footage is collected from unscripted recordings in diverse home environments, making HDEPIC the first dataset collected in-the-wild but with detailed annotations matching those in controlled lab environments. We show the potential of our highly-detailed annotations through a challenging VQA benchmark of 26K questions assessing the capability to …
Dupin: A Parallel Framework For Densest Subgraph Discovery In Fraud Detection On Massive Graphs, Jiaxin Jiang, Siyuan Yao, Yuchen Li, Qiange Wang, Bingsheng He, Min Chen
Dupin: A Parallel Framework For Densest Subgraph Discovery In Fraud Detection On Massive Graphs, Jiaxin Jiang, Siyuan Yao, Yuchen Li, Qiange Wang, Bingsheng He, Min Chen
Research Collection School Of Computing and Information Systems
Detecting fraudulent activities in financial and e-commerce transaction networks is crucial. One effective method for this is Densest Subgraph Discovery (DSD). However, deploying DSD methods in production systems faces substantial scalability challenges due to the predominantly sequential nature of existing methods, which impedes their ability to handle large-scale transaction networks and results in significant detection delays. To address these challenges, we introduce Dupin, a novel parallel processing framework designed for efficient DSD processing in billion-scale graphs. Dupin is powered by a processing engine that exploits the unique properties of the peeling process, with theoretical guarantees on detection quality and efficiency. …
Community Detection In Heterogeneous Information Networks Without Materialization, Jiaxin Jiang, Siyuan Yao, Yuhang Chen, Bingsheng He, Yudong Niu, Yuchen Li, Shixuan Sun, Yongchao Liu
Community Detection In Heterogeneous Information Networks Without Materialization, Jiaxin Jiang, Siyuan Yao, Yuhang Chen, Bingsheng He, Yudong Niu, Yuchen Li, Shixuan Sun, Yongchao Liu
Research Collection School Of Computing and Information Systems
Community detection in heterogeneous information networks (HINs) poses significant challenges due to the diversity of entity types and the complexity of their interrelations. While traditional algorithms may perform adequately in some scenarios, many struggle with the high memory usage and computational demands of large-scale HINs. To address these challenges, we introduce a novel framework, SCAR, which efficiently uncovers community structures in HINs without requiring network materialization. SCAR leverages insights from meta-paths to interpret multi-relational data through compact vertex-based sketches, significantly reducing computational overhead and materialization overhead. We propose a sketch-based technique for estimating changes in modularity, improving both the precision …
Keep The Balance: A Parameter-Efficient Symmetrical Framework For Rgb+X Semantic Segmentation, Jiaxin Cai, Jingze Su, Qi Li, Wenjie Yang, Shu Wang, Tiesong Zhao, Shengfeng He, Wenxi Liu
Keep The Balance: A Parameter-Efficient Symmetrical Framework For Rgb+X Semantic Segmentation, Jiaxin Cai, Jingze Su, Qi Li, Wenjie Yang, Shu Wang, Tiesong Zhao, Shengfeng He, Wenxi Liu
Research Collection School Of Computing and Information Systems
Multimodal semantic segmentation is a critical challenge in computer vision, with early methods suffering from high computational costs and limited transferability due to full fine-tuning of RGB-based pre-trained parameters. Recent studies, while leveraging additional modalities as supplementary prompts to RGB, still predominantly rely on RGB, which restricts the full potential of other modalities. To address these issues, we propose a novel symmetric parameter-efficient fine-tuning framework for multimodal segmentation, featuring with a modality-aware prompting and adaptation scheme, to simultaneously adapt the capabilities of a powerful pre-trained model to both RGB and X modalities. Furthermore, prevalent approaches use the global cross-modality correlations …
Hvi: A New Color Space For Low-Light Image Enhancement, Qingsen Yan, Yixu Feng, Cheng Zhang, Guansong Pang, Kangbiao Shi, Peng Wu, Wei Dong, Jinqiu Sun, Yanning Zhang
Hvi: A New Color Space For Low-Light Image Enhancement, Qingsen Yan, Yixu Feng, Cheng Zhang, Guansong Pang, Kangbiao Shi, Peng Wu, Wei Dong, Jinqiu Sun, Yanning Zhang
Research Collection School Of Computing and Information Systems
Low-Light Image Enhancement (LLIE) is a crucial computer vision task that aims to restore detailed visual information from corrupted low-light images. Many existing LLIE methods are based on standard RGB (sRGB) space, which often produce color bias and brightness artifacts due to inherent high color sensitivity in sRGB. While converting the images using Hue, Saturation and Value (HSV) color space helps resolve the brightness issue, it introduces significant red and black noise artifacts. To address this issue, we propose a new color space for LLIE, namely Horizontal/Vertical-Intensity (HVI), defined by polarized HS maps and learnable inten sity. The former enforces …
Event-Based Eye Tracking: Event-Based Vision Workshop 2025, Qinyu Chen, Et. Al.
Event-Based Eye Tracking: Event-Based Vision Workshop 2025, Qinyu Chen, Et. Al.
Research Collection School Of Computing and Information Systems
No abstract provided.
Modfinity: Unsupervised Domain Adaptation With Multimodal Information Flow Intertwining, Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He
Modfinity: Unsupervised Domain Adaptation With Multimodal Information Flow Intertwining, Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He
Research Collection School Of Computing and Information Systems
Multimodal unsupervised domain adaptation leverages unlabeled data in the target domain to enhance multimodal systems continuously. While current state-of-the-art methods encourage interaction between sub-models of different modalities through pseudo-labeling and feature-level exchange, varying sample quality across modalities can lead to the propagation of inaccurate information, resulting in error accumulation. To address this, we propose Modal-Affinity Multimodal Domain Adaptation (MODfinity), a method that dynamically manages multimodal information flow through fine-grained control over teacher model selection, guiding information intertwining at both feature and label levels. By treating labels as an independent modality, MODfinity enables balanced performance assessment across modalities, employing a novel …
Building Narratives And Probing Concepts: Preparing Materials For Co-Design With Autistic Livestreamers, Terrance Mok, Tyson Hartley, Anthony Tang, Adam Mccrimmon, Lora Oehlberg
Building Narratives And Probing Concepts: Preparing Materials For Co-Design With Autistic Livestreamers, Terrance Mok, Tyson Hartley, Anthony Tang, Adam Mccrimmon, Lora Oehlberg
Research Collection School Of Computing and Information Systems
Based on ten semi-structured interviews with autistic Twitch streamers, we introduce a series of scenario-based design narratives coupled with technology design concepts as a starting point for co-design discussion about autistic streaming. This work builds on prior thematic analysis of the unique intersection between autism and livestreaming. Our user-centered scenarios highlight the needs, goals, and challenges of autistic individuals in livestreaming contexts. By using evocative narratives, the scenarios serve to facilitate empathy and deeper engagement with the needs of autistic users, and help facilitate and support co-creative dialogues and discussions about new technology designs. We contribute this starting point for …
Guest Editorial: When Multimedia Meets Food: Multimedia Computing For Food Data Analysis And Applications, Weiqing Min, Shuqiang Jiang, Petia Radeva, Vladimir Pavlovic, Chong-Wah Ngo, Kiyoharu Aizawa, Wanqing Li
Guest Editorial: When Multimedia Meets Food: Multimedia Computing For Food Data Analysis And Applications, Weiqing Min, Shuqiang Jiang, Petia Radeva, Vladimir Pavlovic, Chong-Wah Ngo, Kiyoharu Aizawa, Wanqing Li
Research Collection School Of Computing and Information Systems
Food is central in our life for its fundamental role in our survival, health, mood and culture. The deployment of various networks (e.g., IoT and mobile networks), devices (e.g., hyperspectral imaging devices, electronic nose/tongue), databases (e.g., nutrition tables and food compositional databases), recipe-sharing websites (e.g., Yummly and Meishijie) and social media (e.g., Twitter and Weibo) has generated unprecedented volumes of multi-modal food data. Such multi-source multi-modal food data provides new perspectives to analyze and understand food consumption via multimedia computing. Riding on the wave of AI, food-oriented multimedia computing integrates AI, multimedia technology and food science to enable a wide …
Exploring The Potential Of Large Language Models For Heterophilic Graphs, Yuxia Wu, Shujie Li, Yuan Fang, Chuan Shi
Exploring The Potential Of Large Language Models For Heterophilic Graphs, Yuxia Wu, Shujie Li, Yuan Fang, Chuan Shi
Research Collection School Of Computing and Information Systems
Large language models (LLMs) have presented significant opportunities to enhance various machine learning applications, including graph neural networks (GNNs). By leveraging the vast open-world knowledge within LLMs, we can more effectively interpret and utilize textual data to better characterize heterophilic graphs, where neighboring nodes often have different labels. However, existing approaches for heterophilic graphs overlook the rich textual data associated with nodes, which could unlock deeper insights into their heterophilic contexts. In this work, we explore the potential of LLMs for modeling heterophilic graphs and propose a novel two-stage framework: LLM-enhanced edge discriminator and LLM-guided edge reweighting. In the first …
Cyberoception: Finding A Painlessly-Measurable New Sense In The Cyberworld Towards Emotion-Awareness In Computing, Tadashi Okoshi, Zexiong Gao, Yi Zhen Tan, Takumi Karasawa, Takeshi Miki, Wataru Sasaki, Rajesh Krishna Balan
Cyberoception: Finding A Painlessly-Measurable New Sense In The Cyberworld Towards Emotion-Awareness In Computing, Tadashi Okoshi, Zexiong Gao, Yi Zhen Tan, Takumi Karasawa, Takeshi Miki, Wataru Sasaki, Rajesh Krishna Balan
Research Collection School Of Computing and Information Systems
In Affective computing, recognizing users’ emotions accurately is the basis of affective human–computer interaction. Understanding users’ interoception contributes to a better understanding of individually different emotional abilities, which is essential for achieving inter-individually accurate emotion estimation. However, existing interoception measurement methods, such as the heart rate discrimination task, have several limitations, including their dependence on a well-controlled laboratory environment and precision apparatus, making monitoring users’ interoception challenging. This study aims to determine other forms of data that can explain users’ interoceptive or similar states in their real-world lives and propose a novel hypothetical concept “cyberoception,” a new sense (1) which …
Oscar: Object Status And Contextual Awareness For Recipes To Support Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington
Oscar: Object Status And Contextual Awareness For Recipes To Support Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington
Research Collection School Of Computing and Information Systems
Following recipes while cooking is an important but difficult task for visually impaired individuals. We developed OSCAR (Object Status Context Awareness for Recipes), a novel approach that provides recipe progress tracking and context-aware feedback on the completion of cooking tasks through tracking object statuses. OSCAR leverages both Large-Language Models (LLMs) and Vision-Language Models (VLMs) to manipulate recipe steps, extract object status information, align visual frames with object status, and provide cooking progress tracking log. We evaluated OSCAR’s recipe following functionality using 173 YouTube cooking videos and 12 real-world non-visual cooking videos to demonstrate OSCAR’s capability to track cooking steps and …
Sans: Efficient Densest Subgraph Discovery Over Relational Graphs Without Materialization, Yudong Niu, Yuchen Li, Jiaxin Jiang, Laks V. S. Lakshmanan
Sans: Efficient Densest Subgraph Discovery Over Relational Graphs Without Materialization, Yudong Niu, Yuchen Li, Jiaxin Jiang, Laks V. S. Lakshmanan
Research Collection School Of Computing and Information Systems
How can we efficiently identify the densest subgraph over relational graphs? Existing dense subgraph discovery (DSD) approaches assume that a relational graph H is already derived from a heterogeneous data source and they focus on efficient discovery of the densest subgraph on the materialized H. Unfortunately, materializing relational graphs can be resource-intensive, which thus limits the practical usefulness of existing algorithms over large datasets. To mitigate this, we propose a novel Summary-bAsed deNsest Subgraph discovery (SANS) system. Our unique summary-based peeling algorithm forms the core of SANS. Following the peeling paradigm, it utilizes summaries of each node's neighborhood to efficiently …
Few-Shot Learning On Graphs: From Meta-Learning To Llm-Empowered Pre-Training And Beyond, Yuan Fang, Yuxia Wu, Xingtong Yu, Shirui Pan
Few-Shot Learning On Graphs: From Meta-Learning To Llm-Empowered Pre-Training And Beyond, Yuan Fang, Yuxia Wu, Xingtong Yu, Shirui Pan
Research Collection School Of Computing and Information Systems
Graph representation learning has become central to many graph-based tasks, driving advancements in various domains such as web search, recommendation systems, and social network analysis. Traditionally, these methods rely on end-to-end supervised learning paradigms that require abundant labeled data, which can be costly and difficult to obtain. To address this limitation, few-shot learning on graphs has emerged as a promising approach, allowing models to generalize with minimal supervision and overcome data scarcity in real-world applications. This tutorial offers an in-depth exploration of recent advancements in few-shot learning for graphs, providing a comparative analysis of state-of-the-art methods and identifying future research …
“I Can Run At Night!”: Using Augmented Reality To Support Nighttime Guided Running For Low-Vision Runners, Yuki Abe, Keisuke Matsushima, Kotaro Hara, Daisuke Sakamoto, Tetsuo Ono
“I Can Run At Night!”: Using Augmented Reality To Support Nighttime Guided Running For Low-Vision Runners, Yuki Abe, Keisuke Matsushima, Kotaro Hara, Daisuke Sakamoto, Tetsuo Ono
Research Collection School Of Computing and Information Systems
Dark environment challenges low-vision (LV) individuals to engage in running by following sighted guide—a Caller-style guided running—due to insufficient illumination, because it prevents them from using their residual vision to follow the guide and be aware about their environment. We design, develop, and evaluate RunSight, an augmented reality (AR)-based assistive tool to support LV individuals to run at night. RunSight combines see-through HMD and image processing to enhance one’s visual awareness of the surrounding environment (e.g., potential hazard) and visualize the guide’s position with AR-based visualization. To demonstrate RunSight’s efficacy, we conducted a user study with 8 LV runners. The …
Worldcuisines: A Massive-Scale Benchmark For Multilingual And Multicultural Visual Question Answering On Global Cuisines, Genta Indra Winata, Et. Al
Worldcuisines: A Massive-Scale Benchmark For Multilingual And Multicultural Visual Question Answering On Global Cuisines, Genta Indra Winata, Et. Al
Research Collection School Of Computing and Information Systems
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate their understanding of such knowledge, we introduce WorldCuisines, a massive-scale benchmark for multilingual and multicultural, visually grounded language understanding. This benchmark includes a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects, spanning 9 language families and featuring over 1 million data points, making it the largest multicultural VQA benchmark to date. It includes tasks for identifying dish names and their origins. We provide evaluation datasets in two sizes (12k and 60k instances) alongside …
Towards Real-World Unsupervised Anomaly Detection For Images, Zhonghang Liu
Towards Real-World Unsupervised Anomaly Detection For Images, Zhonghang Liu
Dissertations and Theses Collection (Open Access)
In the era of big data, data quality plays a critical role in computer vision, where the reliability and purity of training images are essential for optimal performance. When training models such as image classifiers and object detectors, the quality of the training data directly influences the success of the model. In other words, if the training dataset is contaminated, the model’s performance might accordingly decrease.
To address this challenge, unsupervised anomaly detection (UAD) has become an attractive research area. By automatically removing these anomalous data points, UAD can help improve the accuracy and robustness of machine learning models in …
Temporal Relational Graph Convolutional Networks For Financial Applications, Brindha Priyadarshini Jeyaraman
Temporal Relational Graph Convolutional Networks For Financial Applications, Brindha Priyadarshini Jeyaraman
Dissertations and Theses Collection (Open Access)
The financial industry operates within a highly dynamic and interconnected ecosystem, presenting unique challenges for predictive modeling and decision-making. Accurately forecasting financial performance, assessing credit risk, detecting fraud, and ensuring compliance require methodologies that can capture complex temporal, relational, and contextual dependencies within financial data. This thesis investigates the use of Temporal Relational Graph Convolutional Networks (TRGCNs) combined with financial knowledge graphs (FKGs) to address these challenges and enable advanced analytics in the financial domain. We introduce FintechKG, a financial knowledge graph constructed through a threedimensional information extraction process, incorporating entities, temporal dimensions, and domain-specific financial relationships. A TRGCN-based framework …
Samgpt: Text-Free Graph Foundation Model For Multi-Domain Pre-Training And Cross-Domain Adaptation, Xingtong Yu, Zechuan Gong, Chang Zhou, Yuan Fang, Hui Zhang
Samgpt: Text-Free Graph Foundation Model For Multi-Domain Pre-Training And Cross-Domain Adaptation, Xingtong Yu, Zechuan Gong, Chang Zhou, Yuan Fang, Hui Zhang
Research Collection School Of Computing and Information Systems
Graphs are able to model interconnected entities in many online services, supporting a wide range of applications on the Web. This raises an important question: How can we train a graph foundational model on multiple source domains and adapt to an unseen target domain? A major obstacle is that graphs from different domains often exhibit divergent characteristics. Some studies leverage large language models to align multiple domains based on textual descriptions associated with the graphs, limiting their applicability to text-attributed graphs. For text-free graphs, a few recent works attempt to align different feature distributions across domains, while generally neglecting structural …
One-For-All: Towards Universal Domain Translation With A Single Stylegan, Yong Du, Jiahui Zhan, Xinzhe Li, Junyu Dong, Sheng Chen, Ming-Hsuan Yang, Shengfeng He
One-For-All: Towards Universal Domain Translation With A Single Stylegan, Yong Du, Jiahui Zhan, Xinzhe Li, Junyu Dong, Sheng Chen, Ming-Hsuan Yang, Shengfeng He
Research Collection School Of Computing and Information Systems
In this paper, we propose a novel translation model, UniTranslator, for transforming representations between visually distinct domains under conditions of limited training data and significant visual differences. The main idea behind our approach is leveraging the domain-neutral capabilities of CLIP as a bridging mechanism, while utilizing a separate module to extract abstract, domain-agnostic semantics from the embeddings of both the source and target realms. Fusing these abstract semantics with target-specific semantics results in a transformed embedding within the CLIP space. To bridge the gap between the disparate worlds of CLIP and StyleGAN, we introduce a new non-linear mapper, the CLIP2P …
Sita: Structurally Imperceptible And Transferable Adversarial Attacks For Stylized Image Generation, Jingdan Kang, Haoxin Yang, Yan Cai, Huaidong Zhang, Xuemiao Xu, Yong Du, Shengfeng He
Sita: Structurally Imperceptible And Transferable Adversarial Attacks For Stylized Image Generation, Jingdan Kang, Haoxin Yang, Yan Cai, Huaidong Zhang, Xuemiao Xu, Yong Du, Shengfeng He
Research Collection School Of Computing and Information Systems
Image generation technology has brought significant advancements across various fields but has also raised concerns about data misuse and potential rights infringements, particularly with respect to creating visual artworks. Current methods aimed at safeguarding artworks often employ adversarial attacks. However, these methods face challenges such as poor transferability, high computational costs, and the introduction of noticeable noise, which compromises the aesthetic quality of the original artwork. To address these limitations, we propose a Structurally Imperceptible and Transferable Adversarial (SITA) attacks. SITA leverages a CLIP-based destylization loss, which decouples and disrupts the robust style representation of the image. This disruption hinders …
Open-Set Graph Anomaly Detection Via Normal Structure Regularisation, Qizhou Wang, Guansong Pang, Mahsa Salehi, Xiaokun Xia, Christopher Leckie
Open-Set Graph Anomaly Detection Via Normal Structure Regularisation, Qizhou Wang, Guansong Pang, Mahsa Salehi, Xiaokun Xia, Christopher Leckie
Research Collection School Of Computing and Information Systems
This paper considers an important Graph Anomaly Detection (GAD) task, namely open-set GAD, which aims to train a detection model using a small number of normal and anomaly nodes (referred to as *seen anomalies*) to detect both seen anomalies and *unseen anomalies* (*i.e*., anomalies that cannot be illustrated the training anomalies). Those labelled training data provide crucial prior knowledge about abnormalities for GAD models, enabling substantially reduced detection errors. However, current supervised GAD methods tend to over-emphasise fitting the seen anomalies, leading to many errors of detecting the unseen anomalies as normal nodes. Further, existing open-set AD models were introduced …
Minimum Multi-Service Fleet Size Problem: Shareability Graph And Network Flow Approach, Dingtong Yang, Yubin Liu, Hai Wang, Jinhua Zhao, Hamsa Balakrishnan
Minimum Multi-Service Fleet Size Problem: Shareability Graph And Network Flow Approach, Dingtong Yang, Yubin Liu, Hai Wang, Jinhua Zhao, Hamsa Balakrishnan
Research Collection School Of Computing and Information Systems
On-demand, vehicle-based services—such as ride-hailing, food, grocery, and parcel delivery—have become ubiquitous over the past decade. These services can be categorized into four types (Sun et al., 2023): passenger mobility, goods delivery, information acquisition (e.g., probe vehicle for traffic conditions), and mobile server (e.g., vehicle displaying advertisements). Passenger mobility and goods delivery are typically fulfilled by separate fleets, each dedicated to a single service. However, if various services can be pooled and handled simultaneously by a multi-functional fleet while maintaining service quality, the total number of required vehicles and overall vehicle mileage could be significantly reduced. This exciting potential motivates …
Node-Time Conditional Prompt Learning In Dynamic Graphs, Xingtong Yu, Zhenghao Liu, Xinming Zhang, Yuan Fang
Node-Time Conditional Prompt Learning In Dynamic Graphs, Xingtong Yu, Zhenghao Liu, Xinming Zhang, Yuan Fang
Research Collection School Of Computing and Information Systems
Dynamic graphs capture evolving interactions between entities, such as in social networks, online learning platforms, and crowdsourcing projects. For dynamic graph modeling, dynamic graph neural networks (DGNNs) have emerged as a mainstream technique. However, they are generally pre-trained on the link prediction task, leaving a significant gap from the objectives of downstream tasks such as node classification. To bridge the gap, prompt-based learning has gained traction on graphs, but most existing efforts focus on static graphs and neglect the evolution of dynamic graphs. In this paper, we propose DYGPROMPT, a novel pre-training and prompt learning framework for dynamic graph modeling. …
Imageinthat: Manipulating Images To Convey User Instructions To Robots, Karthik Mahadevan, Blaine Lewis, Jiannan Li, Bilge Mutlu, Anthony Tang, Tovi Grossman
Imageinthat: Manipulating Images To Convey User Instructions To Robots, Karthik Mahadevan, Blaine Lewis, Jiannan Li, Bilge Mutlu, Anthony Tang, Tovi Grossman
Research Collection School Of Computing and Information Systems
Foundation models are rapidly improving the capability of robots in performing everyday tasks autonomously such as meal preparation, yet robots will still need to be instructed by humans due to model performance, the difficulty of capturing user preferences, and the need for user agency. Robots can be instructed using various methods---natural language conveys immediate instructions but can be abstract or ambiguous, whereas end-user programming supports longer-horizon tasks but interfaces face difficulties in capturing user intent. In this work, we propose using direct manipulation of images as an alternative paradigm to instruct robots, and introduce a specific instantiation called ImageInThat which …