Open Access. Powered by Scholars. Published by Universities.®

Graphics and Human Computer Interfaces Commons™

Open Access. Powered by Scholars. Published by Universities.®

2,378 Full-Text Articles 4,459 Authors 1,194,747 Downloads 165 Institutions

All Articles in Graphics and Human Computer Interfaces

Faceted Search

2,378 full-text articles. Page 37 of 101.

Dualformer: Local-Global Stratified Transformer For Efficient Video Recognition, Yuxuan LIANG, Pan ZHOU, Roger ZIMMERMANN, Shuicheng YAN 2022 Singapore Management University

Dualformer: Local-Global Stratified Transformer For Efficient Video Recognition, Yuxuan Liang, Pan Zhou, Roger Zimmermann, Shuicheng Yan

Research Collection School Of Computing and Information Systems

While transformers have shown great potential on video recognition with their strong capability of capturing long-range dependencies, they often suffer high computational costs induced by the self-attention to the huge number of 3D tokens. In this paper, we present a new transformer architecture termed DualFormer, which can efficiently perform space-time attention for video recognition. Concretely, DualFormer stratifies the full space-time attention into dual cascaded levels, i.e., to first learn fine-grained local interactions among nearby 3D tokens, and then to capture coarse-grained global dependencies between the query token and global pyramid contexts. Different from existing methods that apply space-time factorization or …


Self-Promoted Supervision For Few-Shot Transformer, Bowen DONG, Pan ZHOU, Shuicheng YAN, Wangmeng ZUO 2022 Singapore Management University

Self-Promoted Supervision For Few-Shot Transformer, Bowen Dong, Pan Zhou, Shuicheng Yan, Wangmeng Zuo

Research Collection School Of Computing and Information Systems

The few-shot learning ability of vision transformers (ViTs) is rarely investigated though heavily desired. In this work, we empirically find that with the same few-shot learning frameworks, e.g. MetaBaseline, replacing the widely used CNN feature extractor with a ViT model often severely impairs few-shot classification performance. Moreover, our empirical study shows that in the absence of inductive bias, ViTs often learn the low-qualified token dependencies under few-shot learning regime where only a few labeled training data are available, which largely contributes to the above performance degradation. To alleviate this issue, for the first time, we propose a simple yet effective …


Video Graph Transformer For Video Question Answering, Junbin XIAO, Pan ZHOU, Tat-Seng CHUA, Shuicheng YAN 2022 Singapore Management University

Video Graph Transformer For Video Question Answering, Junbin Xiao, Pan Zhou, Tat-Seng Chua, Shuicheng Yan

Research Collection School Of Computing and Information Systems

This paper proposes a Video Graph Transformer (VGT) model for Video Quetion Answering (VideoQA). VGT’s uniqueness are two-fold: 1) it designs a dynamic graph transformer module which encodes video by explicitly capturing the visual objects, their relations, and dynamics for complex spatio-temporal reasoning; and 2) it exploits disentangled video and text Transformers for relevance comparison between the video and text to perform QA, instead of entangled crossmodal Transformer for answer classification. Vision-text communication is done by additional cross-modal interaction modules. With more reasonable video encoding and QA solution, we show that VGT can achieve much better performances on VideoQA tasks …


Unsupervised Video Hashing With Multi-Granularity Contextualization And Multi-Structure Preservation, Yanbin HAO, Jingru DUAN, Hao ZHANG, Bin ZHU, Pengyuan ZHOU, Xiangnan HE 2022 Singapore Management University

Unsupervised Video Hashing With Multi-Granularity Contextualization And Multi-Structure Preservation, Yanbin Hao, Jingru Duan, Hao Zhang, Bin Zhu, Pengyuan Zhou, Xiangnan He

Research Collection School Of Computing and Information Systems

Unsupervised video hashing typically aims to learn a compact binary vector to represent complex video content without using manual annotations. Existing unsupervised hashing methods generally suffer from incomplete exploration of various perspective dependencies (e.g., long-range and short-range) and data structures that exist in visual contents, resulting in less discriminative hash codes. In this paper, we propose aMulti-granularity Contextualized and Multi-Structure preserved Hashing (MCMSH) method, exploring multiple axial contexts for discriminative video representation generation and various structural information for unsupervised learning simultaneously. Specifically, we delicately design three self-gating modules to separately model three granularities of dependencies (i.e., long/middle/short-range dependencies) and densely …


Mix-Dann And Dynamic-Modal-Distillation For Video Domain Adaptation, Yuehao YIN, Bin ZHU, Jingjing CHEN, Lechao CHENG, Yu-Gang JIANG 2022 Singapore Management University

Mix-Dann And Dynamic-Modal-Distillation For Video Domain Adaptation, Yuehao Yin, Bin Zhu, Jingjing Chen, Lechao Cheng, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

Video domain adaptation is non-trivial due to video is inherently involved with multi-dimensional and multi-modal information. Existing works mainly adopt adversarial learning and self-supervised tasks to align features. Nevertheless, the explicit interaction between source and target in the temporal dimension, as well as the adaptation between modalities, are unexploited. In this paper, we propose Mix-Domain-Adversarial Neural Network and Dynamic-Modal-Distillation (MD-DMD), a novel multi-modal adversarial learning framework for unsupervised video domain adaptation. Our approach incorporates the temporal information between source and target domains, as well as the diversity of adaptability between modalities. On the one hand, for every single modality, we …


Locally Varying Distance Transform For Unsupervised Visual Anomaly Detection, Wen-yan LIN, Zhonghang LIU, Siying LIU 2022 Singapore Management University

Locally Varying Distance Transform For Unsupervised Visual Anomaly Detection, Wen-Yan Lin, Zhonghang Liu, Siying Liu

Research Collection School Of Computing and Information Systems

Unsupervised anomaly detection on image data is notoriously unstable. We believe this is because many classical anomaly detectors implicitly assume data is low dimensional. However, image data is always high dimensional. Images can be projected to a low dimensional embedding but such projections rely on global transformations that truncate minor variations. As anomalies are rare, the final embedding often lacks the key variations needed to distinguish anomalies from normal instances. This paper proposes a new embedding using a set of locally varying data projections, with each projection responsible for persevering the variations that distinguish a local cluster of instances from …


Ergo: Event Relational Graph Transformer For Document-Level Event Causality Identification, Meiqi CHEN, Yixin CAO, Kunquan DENG, Mukai LI, Kun WANG, Jing SHAO, Yan ZHANG 2022 Singapore Management University

Ergo: Event Relational Graph Transformer For Document-Level Event Causality Identification, Meiqi Chen, Yixin Cao, Kunquan Deng, Mukai Li, Kun Wang, Jing Shao, Yan Zhang

Research Collection School Of Computing and Information Systems

Document-level Event Causality Identification (DECI) aims to identify event-event causal relations in a document. Existing works usually build an event graph for global reasoning across multiple sentences. However, the edges between events have to be carefully designed through heuristic rules or external tools. In this paper, we propose a novel Event Relational Graph TransfOrmer (ERGO) framework1 for DECI, to ease the graph construction and improve it over the noisy edge issue. Different from conventional event graphs, we define a pair of events as a node and build a complete event relational graph without any prior knowledge or tools. This naturally …


Interactive Contrastive Learning For Self-Supervised Entity Alignment, Kaisheng ZENG, Zhenhao DONG, Lei HOU, Yixin CAO, Minghao HU, Jifan YU, Xin LV, Lei CAO, Xin WANG, Haozhuang LIU, Yi HUANG, Jing WAN, Juanzi LI 2022 Singapore Management University

Interactive Contrastive Learning For Self-Supervised Entity Alignment, Kaisheng Zeng, Zhenhao Dong, Lei Hou, Yixin Cao, Minghao Hu, Jifan Yu, Xin Lv, Lei Cao, Xin Wang, Haozhuang Liu, Yi Huang, Jing Wan, Juanzi Li

Research Collection School Of Computing and Information Systems

Self-supervised entity alignment (EA) aims to link equivalent entities across different knowledge graphs (KGs) without the use of pre-aligned entity pairs. The current state-of-the-art (SOTA) selfsupervised EA approach draws inspiration from contrastive learning, originally designed in computer vision based on instance discrimination and contrastive loss, and suffers from two shortcomings. Firstly, it puts unidirectional emphasis on pushing sampled negative entities far away rather than pulling positively aligned pairs close, as is done in the well-established supervised EA. Secondly, it advocates the minimum information requirement for self-supervised EA, while we argue that self-described KG’s side information (e.g., entity name, relation name, …


Tgdm: Target Guided Dynamic Mixup For Cross-Domain Few-Shot Learning, Linhai ZHUO, Yuqian FU, Jingjing CHEN, Yixin CAO, Yu-Gang JIANG 2022 Singapore Management University

Tgdm: Target Guided Dynamic Mixup For Cross-Domain Few-Shot Learning, Linhai Zhuo, Yuqian Fu, Jingjing Chen, Yixin Cao, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

Given sufficient training data on the source domain, cross-domain few-shot learning (CD-FSL) aims at recognizing new classes with a small number of labeled examples on the target domain. The key to addressing CD-FSL is to narrow the domain gap and transferring knowledge of a network trained on the source domain to the target domain. To help knowledge transfer, this paper introduces an intermediate domain generated by mixing images in the source and the target domain. Specifically, to generate the optimal intermediate domain for different target data, we propose a novel target guided dynamic mixup (TGDM) framework that leverages the target …


Investigating Accessibility Challenges And Opportunities For Users With Low Vision Disabilities In Customer-To-Customer (C2c) Marketplaces, Bektur RYSKELDIEV, Kotaro HARA, Mariko KOBAYASHI, Koki KUSANO 2022 University of Tsukuba

Investigating Accessibility Challenges And Opportunities For Users With Low Vision Disabilities In Customer-To-Customer (C2c) Marketplaces, Bektur Ryskeldiev, Kotaro Hara, Mariko Kobayashi, Koki Kusano

Research Collection School Of Computing and Information Systems

Inaccessible e-commerce websites and mobile applications exclude people with visual impairments (PVI) from online shopping. Customer-to-customer (C2C) marketplaces, a form of e-commerce where trading happens not between businesses and customers but between customers, could pose a unique set of challenges in the interactions that the platform brings about. Through online questionnaire and remote interviews, we investigate problems experienced by people with low vision disabilities in common C2C scenarios. Our study with low vision participants (N = 12) reveal both previously known general accessibility issues (e.g., web and mobile interface accessibility) and C2C specific accessibility issues (e.g., inability to confirm item …


Interactive Video Corpus Moment Retrieval Using Reinforcement Learning, Zhixin MA, Chong-wah NGO 2022 Singapore Management University

Interactive Video Corpus Moment Retrieval Using Reinforcement Learning, Zhixin Ma, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Known-item video search is effective with human-in-the-loop to interactively investigate the search result and refine the initial query. Nevertheless, when the first few pages of results are swamped with visually similar items, or the search target is hidden deep in the ranked list, finding the know-item target usually requires a long duration of browsing and result inspection. This paper tackles the problem by reinforcement learning, aiming to reach a search target within a few rounds of interaction by long-term learning from user feedbacks. Specifically, the system interactively plans for navigation path based on feedback and recommends a potential target that …


Long-Term Leap Attention, Short-Term Periodic Shift For Video Classification, Hao ZHANG, Lechao CHENG, Yanbin HAO, Chong-wah NGO 2022 Singapore Management University

Long-Term Leap Attention, Short-Term Periodic Shift For Video Classification, Hao Zhang, Lechao Cheng, Yanbin Hao, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Video transformer naturally incurs a heavier computation burden than a static vision transformer, as the former processes �� times longer sequence than the latter under the current attention of quadratic complexity (�� 2�� 2 ). The existing works treat the temporal axis as a simple extension of spatial axes, focusing on shortening the spatio-temporal sequence by either generic pooling or local windowing without utilizing temporal redundancy. However, videos naturally contain redundant information between neighboring frames; thereby, we could potentially suppress attention on visually similar frames in a dilated manner. Based on this hypothesis, we propose the LAPS, a long-term “Leap …


Wave-Vit: Unifying Wavelet And Transformers For Visual Representation Learning, Ting YAO, Yingwei PAN, Yehao LI, Chong-wah NGO, Tao MEI 2022 JD Explore Academy

Wave-Vit: Unifying Wavelet And Transformers For Visual Representation Learning, Ting Yao, Yingwei Pan, Yehao Li, Chong-Wah Ngo, Tao Mei

Research Collection School Of Computing and Information Systems

Multi-scale Vision Transformer (ViT) has emerged as a powerful backbone for computer vision tasks, while the self-attention computation in Transformer scales quadratically w.r.t. the input patch number. Thus, existing solutions commonly employ down-sampling operations (e.g., average pooling) over keys/values to dramatically reduce the computational cost. In this work, we argue that such over-aggressive down-sampling design is not invertible and inevitably causes information dropping especially for high-frequency components in objects (e.g., texture details). Motivated by the wavelet theory, we construct a new Wavelet Vision Transformer (Wave-ViT) that formulates the invertible down-sampling with wavelet transforms and self-attention learning in a unified way. …


Dynamic Temporal Filtering In Video Models, Fuchen LONG, Zhaofan QIU, Yingwei PAN, Ting YAO, Chong-wah NGO, Tao MEI 2022 JD Explore Academy

Dynamic Temporal Filtering In Video Models, Fuchen Long, Zhaofan Qiu, Yingwei Pan, Ting Yao, Chong-Wah Ngo, Tao Mei

Research Collection School Of Computing and Information Systems

Video temporal dynamics is conventionally modeled with 3D spatial-temporal kernel or its factorized version comprised of 2D spatial kernel and 1D temporal kernel. The modeling power, nevertheless, is limited by the fixed window size and static weights of a kernel along the temporal dimension. The pre-determined kernel size severely limits the temporal receptive fields and the fixed weights treat each spatial location across frames equally, resulting in sub-optimal solution for longrange temporal modeling in natural scenes. In this paper, we present a new recipe of temporal feature learning, namely Dynamic Temporal Filter (DTF), that novelly performs spatial-aware temporal modeling in …


On Mitigating Hard Clusters For Face Clustering, Yingjie CHEN, Huasong ZHONG, Chong CHEN, Chen SHEN, Jianqiang HUANG, Tao WANG, Yun LIANG, Qianru SUN 2022 Singapore Management University

On Mitigating Hard Clusters For Face Clustering, Yingjie Chen, Huasong Zhong, Chong Chen, Chen Shen, Jianqiang Huang, Tao Wang, Yun Liang, Qianru Sun

Research Collection School Of Computing and Information Systems

Face clustering is a promising way to scale up face recognition systems using large-scale unlabeled face images. It remains challenging to identify small or sparse face image clusters that we call hard clusters, which is caused by the heterogeneity, i.e., high variations in size and sparsity, of the clusters. Consequently, the conventional way of using a uniform threshold (to identify clusters) often leads to a terrible misclassification for the samples that should belong to hard clusters. We tackle this problem by leveraging the neighborhood information of samples and inferring the cluster memberships (of samples) in a probabilistic way. We introduce …


Equivariance And Invariance Inductive Bias For Learning From Insufficient Data, Tan WANG, Qianru SUN, Sugiri PRANATA, Karlekar JAYASHREE, Hanwang ZHANG 2022 Nanyang Technological University

Equivariance And Invariance Inductive Bias For Learning From Insufficient Data, Tan Wang, Qianru Sun, Sugiri Pranata, Karlekar Jayashree, Hanwang Zhang

Research Collection School Of Computing and Information Systems

We are interested in learning robust models from insufficient data, without the need for any externally pre-trained model checkpoints. First, compared to sufficient data, we show why insufficient data renders the model more easily biased to the limited training environments that are usually different from testing. For example, if all the training "swan" samples are "white", the model may wrongly use the "white" environment to represent the intrinsic class "swan". Then, we justify that equivariance inductive bias can retain the class feature while invariance inductive bias can remove the environmental feature, leaving only the class feature that generalizes to any …


Cvfnet: Real-Time 3d Object Detection By Learning Cross View Features, Jiaqi GU, Zhiyu XIANG, Pan ZHAO, Tingming BAI, Lingxuan WANG, Xijun ZHAO, Zhiyuan ZHANG 2022 Singapore Management University

Cvfnet: Real-Time 3d Object Detection By Learning Cross View Features, Jiaqi Gu, Zhiyu Xiang, Pan Zhao, Tingming Bai, Lingxuan Wang, Xijun Zhao, Zhiyuan Zhang

Research Collection School Of Computing and Information Systems

In recent years 3D object detection from LiDAR point clouds has made great progress thanks to the development of deep learning technologies. Although voxel or point based methods are popular in 3D object detection, they usually involve time-consuming operations such as 3D convolutions on voxels or ball query among points, making the resulting network inappropriate for time critical applications. On the other hand, 2D view-based methods feature high computing efficiency while usually obtaining inferior performance than the voxel or point based methods. In this work, we present a real-time view-based single stage 3D object detector, namely CVFNet to fulfill this …


Learning Discriminative Representations Via Variational Self-Distillation For Cross-View Geo-Localization, Qian HU, Wansi LI, Xing XU, Ning LIU, Lei WANG 2022 Singapore Management University

Learning Discriminative Representations Via Variational Self-Distillation For Cross-View Geo-Localization, Qian Hu, Wansi Li, Xing Xu, Ning Liu, Lei Wang

Research Collection School Of Computing and Information Systems

Cross-view geo-localization is to localize the same geographic target in images from different perspectives, e.g., satellite-view and drone-view. The primary challenge faced by existing methods is the large visual appearance changes across views. Most previous work utilizes the deep neural network to obtain the discriminative representations and directly uses them to accomplish the geo-localization task. However, these approaches ignore that the redundancy retained in the extracted features negatively impacts the result. In this paper, we argue that the information bottleneck (IB) can retain the most relevant information while removing as much redundancy as possible. The variational self-distillation (VSD) strategy provides …


A Roller Coaster For The Mind: Virtual Reality Sickness Modes, Metrics, And Mitigation, Dalton C. Sparks 2022 University of Louisville

A Roller Coaster For The Mind: Virtual Reality Sickness Modes, Metrics, And Mitigation, Dalton C. Sparks

The Cardinal Edge

Understanding and preventing virtual reality sickness(VRS), or cybersickness, is vital in removing barriers for the technology's adoption. Thus, this article aims to synthesize a variety of academic sources to demonstrate the modes by which VRS occurs, the metrics by which it is judged, and the methods to mitigate it. The predominant theories on the biological origins of VRS are discussed, as well as the individual factors which increase the likelihood of a user developing VRS. Moreover, subjective and physiological measurements of VRS are discussed in addition to the development of a predictive model and conceptual framework. Finally, several methodologies of …


On The Effectiveness Of Using Graphics Interrupt As A Side Channel For User Behavior Snooping, Haoyu MA, Jianwen TIAN, Debin GAO, Chunfu JIA 2022 Singapore Management University

On The Effectiveness Of Using Graphics Interrupt As A Side Channel For User Behavior Snooping, Haoyu Ma, Jianwen Tian, Debin Gao, Chunfu Jia

Research Collection School Of Computing and Information Systems

Graphics Processing Units (GPUs) are now a key component of many devices and systems, including those in the cloud and data centers, thus are also subject to side-channel attacks. Existing side-channel attacks on GPUs typically leak information from graphics libraries like OpenGL and CUDA, which require creating contentions within the GPU resource space and are being mitigated with software patches. This paper evaluates potential side channels exposed at a lower-level interface between GPUs and CPUs, namely the graphics interrupts. These signals could indicate unique signatures of GPU workload, allowing a spy process to infer the behavior of other processes. We …


Digital Commons powered by bepress