Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Research Collection School Of Computing and Information Systems

Discipline
Keyword
Publication Year
File Type

Articles 2581 - 2610 of 8479

Full-Text Articles in Computer Sciences

Target-Guided Emotion-Aware Chat Machine, Wei Wei, Jiayi Liu, Xianling Mao, Guibing Guo, Feida Zhu, Pan Zhou, Yuchong Hu, Shanshan Feng Oct 2021

Target-Guided Emotion-Aware Chat Machine, Wei Wei, Jiayi Liu, Xianling Mao, Guibing Guo, Feida Zhu, Pan Zhou, Yuchong Hu, Shanshan Feng

Research Collection School Of Computing and Information Systems

The consistency of a response to a given post at the semantic level and emotional level is essential for a dialogue system to deliver humanlike interactions. However, this challenge is not well addressed in the literature, since most of the approaches neglect the emotional information conveyed by a post while generating responses. This article addresses this problem and proposes a unified end-to-end neural architecture, which is capable of simultaneously encoding the semantics and the emotions in a post and leveraging target information to generate more intelligent responses with appropriately expressed emotions. Extensive experiments on real-world data demonstrate that the proposed …


Revocable Policy-Based Chameleon Hash, Shengmin Xu, Jianting Ning, Jinhua Ma, Guowen Xu, Jiaming Yuan, Robert H. Deng Oct 2021

Revocable Policy-Based Chameleon Hash, Shengmin Xu, Jianting Ning, Jinhua Ma, Guowen Xu, Jiaming Yuan, Robert H. Deng

Research Collection School Of Computing and Information Systems

Policy-based chameleon hash (PCH) is a cryptographic building block which finds increasing practical applications. Given a message and an access policy, for any chameleon hash generated by a PCH scheme, a chameleon trapdoor holder whose rewriting privileges satisfy the access policy can amend the underlying message without affecting the hash value. In practice, it is necessary to revoke the rewriting privileges of a trapdoor holder due to various reasons, such as change of positions, compromise of credentials, or malicious behaviours. In this paper, we introduce the notion of revocable PCH (RPCH) and formally define its security. We instantiate a concrete …


Prediction Of Synthetic Lethal Interactions In Human Cancers Using Multi-View Graph Auto-Encoder, Zhifeng Hao, Di Wu, Yuan Fang, Min Wu, Ruichu Cai, Xiaoli Li Oct 2021

Prediction Of Synthetic Lethal Interactions In Human Cancers Using Multi-View Graph Auto-Encoder, Zhifeng Hao, Di Wu, Yuan Fang, Min Wu, Ruichu Cai, Xiaoli Li

Research Collection School Of Computing and Information Systems

Synthetic lethality (SL) is a very important concept for the development of targeted anticancer drugs. However, experimental methods for SL detection often suffer from various issues like high cost and low consistency across cell lines. Hence, computational methods for predicting novel SLs have recently emerged as complements for wet-lab experiments. In addition, SL data can be represented as a graph where nodes are genes and edges are the SL interactions. It is thus motivated to design advanced graph-based machine learning algorithms for SL prediction. In this paper, we propose a novel SL prediction method using Multi-view Graph Auto-Encoder (SLMGAE). We …


The Efficacy Of Collaborative Authoring Of Video Scene Descriptions, Rosiana Natalie, Jolene Kar Inn Loh, Huei Suen Tan, Joshua Shi-Hao Tseng, Ian Luke Yi-Ren Chan, Ebrima H. Jarjue, Hernisa Kacorri, Kotaro Hara Oct 2021

The Efficacy Of Collaborative Authoring Of Video Scene Descriptions, Rosiana Natalie, Jolene Kar Inn Loh, Huei Suen Tan, Joshua Shi-Hao Tseng, Ian Luke Yi-Ren Chan, Ebrima H. Jarjue, Hernisa Kacorri, Kotaro Hara

Research Collection School Of Computing and Information Systems

The majority of online video contents remain inaccessible to people with visual impairments due to the lack of audio descriptions to depict the video scenes. Content creators have traditionally relied on professionals to author audio descriptions, but their service is costly and not readily-available. We investigate the feasibility of creating more cost-effective audio descriptions that are also of high quality by involving novices. Specifically, we designed, developed, and evaluated ViScene, a web-based collaborative audio description authoring tool that enables a sighted novice author and a reviewer either sighted or blind to interact and contribute to scene descriptions (SDs)—text that can …


Visionary Caption: Improving The Accessibility Of Presentation Slides Through Highlighting Visualization, Carmen Ji Yan Yip, Jie Mi Chong, Sin Yee Kwek, Yong Wang, Kotaro Hara Oct 2021

Visionary Caption: Improving The Accessibility Of Presentation Slides Through Highlighting Visualization, Carmen Ji Yan Yip, Jie Mi Chong, Sin Yee Kwek, Yong Wang, Kotaro Hara

Research Collection School Of Computing and Information Systems

Presentation slides are widely used in occasions such as academic talks and business meetings. Captions placed on slides support deaf and hard of hearing (DHH) people to understand spoken contents, but simultaneously comprehending and associating visual contents on slides and caption text could be challenging. In this paper, we design and develop a visualization technique to highlight and associate chart on a slide and numerical data in caption. We first conduct a small formative study with people with and without hearing impairments to assess the value of the visualization technique using a lo-fidelity video prototype. We then develop Visionary Caption, …


Conquer: Contextual Query-Aware Ranking For Video Corpus Moment Retrieval, Zhijian Hou, Chong-Wah Ngo, W. K. Chan Oct 2021

Conquer: Contextual Query-Aware Ranking For Video Corpus Moment Retrieval, Zhijian Hou, Chong-Wah Ngo, W. K. Chan

Research Collection School Of Computing and Information Systems

This paper tackles a recently proposed Video Corpus Moment Retrieval task. This task is essential because advanced video retrieval applications should enable users to retrieve a precise moment from a large video corpus. We propose a novel CONtextual QUery-awarE Ranking (CONQUER) model for effective moment localization and ranking. CONQUER explores query context for multi-modal fusion and representation learning in two different steps. The first step derives fusion weights for the adaptive combination of multi-modal video content. The second step performs bi-directional attention to tightly couple video and query as a single joint representation for moment localization. As query context is …


Token Shift Transformer For Video Classification, Zhang Hao, Yanbin. Hao, Chong-Wah Ngo Oct 2021

Token Shift Transformer For Video Classification, Zhang Hao, Yanbin. Hao, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Transformer achieves remarkable successes in understanding 1 and 2-dimensional signals (e.g., NLP and Image Content Understanding). As a potential alternative to convolutional neural networks, it shares merits of strong interpretability, high discriminative power on hyper-scale data, and flexibility in processing varying length inputs. However, its encoders naturally contain computational intensive operations such as pair-wise self-attention, incurring heavy computational burden when being applied on the complex 3-dimensional video signals. This paper presents Token Shift Module (i.e., TokShift), a novel, zero-parameter, zero-FLOPs operator, for modeling temporal relations within each transformer encoder. Specifically, the TokShift barely temporally shifts partial [Class] token features back-and-forth …


Cloud, Edge And Fog Computing: Trends And Case Studies, Eng Lieh Ouh, Stanislaw Jarzabek, Geok Shan Lim, Masayoshi Ogawa Oct 2021

Cloud, Edge And Fog Computing: Trends And Case Studies, Eng Lieh Ouh, Stanislaw Jarzabek, Geok Shan Lim, Masayoshi Ogawa

Research Collection School Of Computing and Information Systems

As it is done today, an informal – solely based on experts’ intuition – evaluation of profitability of adopting cloud services is undependable and not scalable as there are many conflicting factors and constraints such evaluation should account for. The revenue from service tenants and the cost of implementing the service architecture are the leading service factors that drive profitability. Cloud service architectures also need to handle a growing number of tenants with increasingly diverse requirements which must be weighed against the capabilities and costs of various service architectures, particularly single- versus multi-tenanted models. We believe a conceptual model enumerating …


Integrated Discourse Analysis & Learning Skills Framework For Class Conversations, Devyn Wei Hung Tan, Gottipati Swapna, Kyong Jin Shim, Shankararaman, Venky Oct 2021

Integrated Discourse Analysis & Learning Skills Framework For Class Conversations, Devyn Wei Hung Tan, Gottipati Swapna, Kyong Jin Shim, Shankararaman, Venky

Research Collection School Of Computing and Information Systems

Constructive interactions through discussion forums allow students to open their horizons and thought processes to acquire more knowledge and develop skills. Thus, discussion forums play an important role in supporting learning. Additionally, the discussion forum provides the content for creating a knowledge repository. It contains discussion threads related to key course topics that are debated by the students. One approach to understanding the student learning experience is through the analysis of the discussion threads. This research proposes the application of discourse analysis and collaborative learning frameworks to discussion forums to gain further insights into the student’s learning in a classroom. …


Interactive Probing Of Multivariate Time Series Prediction Models: A Case Of Freight Rate Analysis, Haonan Xu, Haotian Li, Yong Wang Oct 2021

Interactive Probing Of Multivariate Time Series Prediction Models: A Case Of Freight Rate Analysis, Haonan Xu, Haotian Li, Yong Wang

Research Collection School Of Computing and Information Systems

We present an interactive probing tool to create, modify and analyze what-if scenarios for multivariate time series models. The solution is applied to freight trading, where analysts can carry out sensitivity analysis on freight rates by changing demand and supply-related econometric variables and observing their resultant effects on freight indexes. We utilize various visualization techniques to enable intuitive scenario creation, alteration, and comprehension of time series inputs and model predictions. Our tool proved to be useful to the industry practitioners, demonstrated by a case study where freight traders are given hypothetical market scenarios and successfully generated quantitative freight index projection …


Visilence: An Interactive Visualization Tool For Error Resilience Analysis, Shaolun Ruan, Yong Wang, Qiang Guan Oct 2021

Visilence: An Interactive Visualization Tool For Error Resilience Analysis, Shaolun Ruan, Yong Wang, Qiang Guan

Research Collection School Of Computing and Information Systems

Soft errors have become one of the major concerns for HPC applications, as those errors can result in seriously corrupted outcomes, such as silent data corruptions (SDCs). Prior studies on error resilience have studied the robustness of HPC applications. However, it is still difficult for program developers to identify potential vulnerability to soft errors. In this paper, we present Visilence, a novel visualization tool to visually analyze error vulnerability based on the control-flow graph generated from HPC applications. Visilence efficiently visualizes the affected program states under injected errors and presents the visual analysis of the most vulnerable parts of an …


Quantum Computing: Computational Excellence For Society 5.0, Paul R. Griffin, Michael Boguslavsky, Junye Huang, Robert J. Kauffman, Brian R. Tan Oct 2021

Quantum Computing: Computational Excellence For Society 5.0, Paul R. Griffin, Michael Boguslavsky, Junye Huang, Robert J. Kauffman, Brian R. Tan

Research Collection School Of Computing and Information Systems

In this chapter, we consider which general business problems may be suitable for exploring the utilization of quantum computing and provide a framework for applying quantum computing. The characteristics of quantum computing systems are mapped into business problems to show the potential advantages of quantum computing. The framework shows how quantum computing can be applied in general, and a use case is offered for quantum machine learning (QML) related to the credit ratings of small and medium-size enterprises (SMEs).


Cloudnplay: Resource Optimization For A Cloud-Native Gaming System, Angelus Wibowo, Nguyen Binh Duong Ta Oct 2021

Cloudnplay: Resource Optimization For A Cloud-Native Gaming System, Angelus Wibowo, Nguyen Binh Duong Ta

Research Collection School Of Computing and Information Systems

Cloud gaming enables people playing graphically intensive games from their less powerful, or even outdated computing devices. It is challenging to realize cloud gaming as it requires minimal latency in server-side processing, rendering and streaming, which are expensive in terms of resource requirements, e.g., powerful GPU servers. Commercial gaming providers, e.g., Google Stadia, Amazon Luna, etc., hardly disclose any information on how they optimize gaming performance and cloud cost. In this work, we aim to investigate resource cost optimization for such cloud gaming systems. In contrast to previous work which have been focusing more on theoretical approaches, we deliver a …


Assessing Generalizability Of Codebert, Xin Zhou, Donggyun Han, David Lo Oct 2021

Assessing Generalizability Of Codebert, Xin Zhou, Donggyun Han, David Lo

Research Collection School Of Computing and Information Systems

Pre-trained models like BERT have achieved strong improvements on many natural language processing (NLP) tasks, showing their great generalizability. The success of pre-trained models in NLP inspires pre-trained models for programming language. Recently, CodeBERT, a model for both natural language (NL) and programming language (PL), pre-trained on code search dataset, is proposed. Although promising, CodeBERT has not been evaluated beyond its pre-trained dataset for NL-PL tasks. Also, it has only been shown effective on two tasks that are close in nature to its pre-trained data. This raises two questions: Can CodeBERT generalize beyond its pre-trained data? Can it generalize to …


Design And Supervision Model Of Group Projects For Active Learning, Yi Meng Lau, Kyong Jin Shim, Swapna Gottipati Oct 2021

Design And Supervision Model Of Group Projects For Active Learning, Yi Meng Lau, Kyong Jin Shim, Swapna Gottipati

Research Collection School Of Computing and Information Systems

This research paper presents a group project framework for a second-year programming course, which was conducted during the COVID-19 pandemic. The framework offers well defined stages of the group project which allow students to work on their choice of a real-world problem, integrate their learnings from previous courses, and present a working solution. In the group project, students actively participate, reflect, and contribute to achieving the goals set in the learning objectives of the course. Our framework incorporates key features from Kolb’s Experiential Learning Theory (1984) and principles of active learning from Barnes (1989) to achieve active and experiential learning …


Latent Class Analysis For Identifying Subclasses Of Depression Using Jmp Pro 16, Karishma Yadav, Fei Fei Sue-Ann Seet, Tin Seong Kam, Tin Seong Kam Oct 2021

Latent Class Analysis For Identifying Subclasses Of Depression Using Jmp Pro 16, Karishma Yadav, Fei Fei Sue-Ann Seet, Tin Seong Kam, Tin Seong Kam

Research Collection School Of Computing and Information Systems

According to WHO, “Depression is a leading cause of disability worldwide and is a major contributor to the overall global burden of disease”. A major stumbling block in the care of depressed patients remains the accurate diagnosis of the severity of depression. Patient Health Questionnaire (PHQ-9), a 9-question instrument is widely used for diagnosing and determining the severity of depression. However, the popularly used 5-Category of depression severity based on the sum of responses to the 9 questions was overly subjective. In view of this limitation, our paper aims to demonstrate how Latent Class Analysis of JMP Pro can be …


Disambiguating Mentions Of Api Methods In Stack Overflow Via Type Scoping, Kien Luong, Ferdian Thung, David Lo Oct 2021

Disambiguating Mentions Of Api Methods In Stack Overflow Via Type Scoping, Kien Luong, Ferdian Thung, David Lo

Research Collection School Of Computing and Information Systems

Stack Overflow is one of the most popular venues for developers to find answers to their API-related questions. However, API mentions in informal text content of Stack Overflow are often ambiguous and thus it could be difficult to find the APIs and learn their usages. Disambiguating these API mentions is not trivial, as an API mention can match with names of APIs from different libraries or even the same one. In this paper, we propose an approach called DATYS to disambiguate API mentions in informal text content of Stack Overflow using type scoping. With type scoping, we consider API methods …


Condensing A Sequence To One Informative Frame For Video Recognition, Qiu. Zhaofan, Ting Yao, Yan Shu, Chong-Wah Ngo, Tao Mei Oct 2021

Condensing A Sequence To One Informative Frame For Video Recognition, Qiu. Zhaofan, Ting Yao, Yan Shu, Chong-Wah Ngo, Tao Mei

Research Collection School Of Computing and Information Systems

Video is complex due to large variations in motion and rich content in fine-grained visual details. Abstracting useful information from such information-intensive media requires exhaustive computing resources. This paper studies a two-step alternative that first condenses the video sequence to an informative" frame" and then exploits off-the-shelf image recognition system on the synthetic frame. A valid question is how to define" useful information" and then distill it from a video sequence down to one synthetic frame. This paper presents a novel Informative Frame Synthesis (IFS) architecture that incorporates three objective tasks, ie, appearance reconstruction, video categorization, motion estimation, and two …


Can Differential Testing Improve Automatic Speech Recognition Systems?, Muhammad Hilmi Asyrofi, Zhou Yang, Jieke Shi, Chu Wei Quan, David Lo Oct 2021

Can Differential Testing Improve Automatic Speech Recognition Systems?, Muhammad Hilmi Asyrofi, Zhou Yang, Jieke Shi, Chu Wei Quan, David Lo

Research Collection School Of Computing and Information Systems

Due to the widespread adoption of Automatic Speech Recognition (ASR) systems in many critical domains, ensuring the quality of recognized transcriptions is of great importance. A recent work, CrossASR++, can automatically uncover many failures in ASR systems by taking advantage of the differential testing technique. It employs a Text-To-Speech (TTS) system to synthesize audios from texts and then reveals failed test cases by feeding them to multiple ASR systems for cross-referencing. However, no prior work tries to utilize the generated test cases to enhance the quality of ASR systems. In this paper, we explore the subsequent improvements brought by leveraging …


A First Look At Accessibility Issues In Popular Github Projects, Tingting Bi, Xin Xia, David Lo, Aldeida Aleti Oct 2021

A First Look At Accessibility Issues In Popular Github Projects, Tingting Bi, Xin Xia, David Lo, Aldeida Aleti

Research Collection School Of Computing and Information Systems

Accessibility design elements allow people to access software products and services independent of their different abilities. However, accessibility is challenging to handle and whether accessibility is widely considered in software projects is unclear. In this work, we aim to understand if accessibility is a prevalent consideration in practice, what accessibility issues are discussed in GitHub projects, what potential reasons cause accessibility issues, and what solutions (e.g., tools and standards) are applied for addressing accessibility issues. In this work, we collect 11,820 accessibility issues and their threads discussed by developers in popular GitHub projects. We manually analyzed and grouped the collected …


Deep Learning For Image Super-Resolution: A Survey, Zhihao Wang, Jian Chen, Steven C. H. Hoi Oct 2021

Deep Learning For Image Super-Resolution: A Survey, Zhihao Wang, Jian Chen, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

Image Super-Resolution (SR) is an important class of image processing techniqueso enhance the resolution of images and videos in computer vision. Recent years have witnessed remarkable progress of image super-resolution using deep learning techniques. This article aims to provide a comprehensive survey on recent advances of image super-resolution using deep learning approaches. In general, we can roughly group the existing studies of SR techniques into three major categories: supervised SR, unsupervised SR, and domain-specific SR. In addition, we also cover some other important issues, such as publicly available benchmark datasets and performance evaluation metrics. Finally, we conclude this survey by …


Online Learning: A Comprehensive Survey, Steven C. H. Hoi, Doyen Sahoo, Jing Lu, Peilin Zhao Oct 2021

Online Learning: A Comprehensive Survey, Steven C. H. Hoi, Doyen Sahoo, Jing Lu, Peilin Zhao

Research Collection School Of Computing and Information Systems

Online learning represents a family of machine learning methods, where a learner attempts to tackle some predictive (or any type of decision-making) task by learning from a sequence of data instances one by one at each time. The goal of online learning is to maximize the accuracy/correctness for the sequence of predictions/decisions made by the online learner given the knowledge of correct answers to previous prediction/learning tasks and possibly additional information. This is in contrast to traditional batch or offline machine learning methods that are often designed to learn a model from the entire training data set at once. Online …


Eargate: Gait-Based User Identification With In-Ear Microphones, Andrea Ferlini, Dong Ma, Cecilia Mascolo Oct 2021

Eargate: Gait-Based User Identification With In-Ear Microphones, Andrea Ferlini, Dong Ma, Cecilia Mascolo

Research Collection School Of Computing and Information Systems

Human gait is a widely used biometric trait for user identification and recognition. Given the wide-spreading, steady diffusion of earworn wearables (Earables) as the new frontier of wearable devices, we investigate the feasibility of earable-based gait identification. Specifically, we look at gait-based identification from the sounds induced by walking and propagated through the musculoskeletal system in the body. Our system, EarGate, leverages an in-ear facing microphone which exploits the earable’s occlusion effect to reliably detect the user’s gait from inside the ear canal, without impairing the general usage of earphones. With data collected from 31 subjects, we show that EarGate …


Solarslam: Battery-Free Loop Closure For Indoor Localisation, Bo Wei, Weitao Xu, Chengwen Luo, Guillaume Zoppi, Dong Ma, Sen Wang Oct 2021

Solarslam: Battery-Free Loop Closure For Indoor Localisation, Bo Wei, Weitao Xu, Chengwen Luo, Guillaume Zoppi, Dong Ma, Sen Wang

Research Collection School Of Computing and Information Systems

In this paper, we propose SolarSLAM, a batteryfree loop closure method for indoor localisation. Inertial Measurement Unit (IMU) based indoor localisation method has been widely used due to its ubiquity in mobile devices, such as mobile phones, smartwatches and wearable bands. However, it suffers from the unavoidable long term drift. To mitigate the localisation error, many loop closure solutions have been proposed using sophisticated sensors, such as cameras, laser, etc. Despite achieving high-precision localisation performance, these sensors consume a huge amount of energy. Different from those solutions, the proposed SolarSLAM takes advantage of an energy harvesting solar cell as a …


Weakly-Supervised Video Anomaly Detection With Contrastive Learning Of Long And Short-Range Temporal Features, Yu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh, Johan W. Verjans, Gustavo Carneiro Oct 2021

Weakly-Supervised Video Anomaly Detection With Contrastive Learning Of Long And Short-Range Temporal Features, Yu Tian, Guansong Pang, Yuanhong Chen, Rajvinder Singh, Johan W. Verjans, Gustavo Carneiro

Research Collection School Of Computing and Information Systems

Anomaly detection with weakly supervised video-level labels is typically formulated as a multiple instance learning (MIL) problem, in which we aim to identify snippets containing abnormal events, with each video represented as a bag of video snippets. Although current methods show effective detection performance, their recognition of the positive instances, i.e., rare abnormal snippets in the abnormal videos, is largely biased by the dominant negative instances, especially when the abnormal events are subtle anomalies that exhibit only small differences compared with normal events. This issue is exacerbated in many methods that ignore important video temporal dependencies. To address this issue, …


Constrained Contrastive Distribution Learning For Unsupervised Anomaly Detection And Localisation In Medical Images, Yu Tian, Guansong Pang, Fengbei Liu, Yuanhong Chen, Seon Ho Shin, Johan W. Verjans, Rajvinder Singh Oct 2021

Constrained Contrastive Distribution Learning For Unsupervised Anomaly Detection And Localisation In Medical Images, Yu Tian, Guansong Pang, Fengbei Liu, Yuanhong Chen, Seon Ho Shin, Johan W. Verjans, Rajvinder Singh

Research Collection School Of Computing and Information Systems

Unsupervised anomaly detection (UAD) learns one-class classifiers exclusively with normal (i.e., healthy) images to detect any abnormal (i.e., unhealthy) samples that do not conform to the expected normal patterns. UAD has two main advantages over its fully supervised counterpart. Firstly, it is able to directly leverage large datasets available from health screening programs that contain mostly normal image samples, avoiding the costly manual labelling of abnormal samples and the subsequent issues involved in training with extremely class-imbalanced data. Further, UAD approaches can potentially detect and localise any type of lesions that deviate from the normal patterns. One significant challenge faced …


Learning To Adversarially Blur Visual Object Tracking, Qing Guo, Ziyi Cheng, Felix Juefei-Xu, Lei Ma, Xiaofei Xie, Yang Liu, Jianjun Zhao Oct 2021

Learning To Adversarially Blur Visual Object Tracking, Qing Guo, Ziyi Cheng, Felix Juefei-Xu, Lei Ma, Xiaofei Xie, Yang Liu, Jianjun Zhao

Research Collection School Of Computing and Information Systems

Motion blur caused by the moving of the object or camera during the exposure can be a key challenge for visual object tracking, affecting tracking accuracy significantly. In this work, we explore the robustness of visual object trackers against motion blur from a new angle, i.e., adversarial blur attack (ABA). Our main objective is to online transfer input frames to their natural motion-blurred counterparts while misleading the state-of-the-art trackers during the tracking process. To this end, we first design the motion blur synthesizing method for visual tracking based on the generation principle of motion blur, considering the motion information and …


Causal Attention For Unbiased Visual Recognition, Tan Wang, Chang Zhou, Qianru Sun, Hanwang Zhang Oct 2021

Causal Attention For Unbiased Visual Recognition, Tan Wang, Chang Zhou, Qianru Sun, Hanwang Zhang

Research Collection School Of Computing and Information Systems

Attention module does not always help deep models learn causal features that are robust in any confounding context, e.g., a foreground object feature is invariant to different backgrounds. This is because the confounders trick the attention to capture spurious correlations that benefit the prediction when the training and testing data are IID (identical & independent distribution); while harm the prediction when the data are OOD (out-of-distribution). The sole fundamental solution to learn causal attention is by causal intervention, which requires additional annotations of the confounders, e.g., a “dog” model is learned within “grass+dog” and “road+dog” respectively, so the “grass” and …


Transporting Causal Mechanisms For Unsupervised Domain Adaptation, Zhongqi Yue, Qianru Sun, Xian-Sheng Hua, Hanwang Zhang Oct 2021

Transporting Causal Mechanisms For Unsupervised Domain Adaptation, Zhongqi Yue, Qianru Sun, Xian-Sheng Hua, Hanwang Zhang

Research Collection School Of Computing and Information Systems

Existing Unsupervised Domain Adaptation (UDA) literature adopts the covariate shift and conditional shift assumptions, which essentially encourage models to learn common features across domains. However, due to the lack of supervision in the target domain, they suffer from the semantic loss: the feature will inevitably lose nondiscriminative semantics in source domain, which is however discriminative in target domain. We use a causal view—transportability theory [41]—to identify that such loss is in fact a confounding effect, which can only be removed by causal intervention. However, the theoretical solution provided by transportability is far from practical for UDA, because it requires the …


Self-Regulation For Semantic Segmentation, Dong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua, Qianru Sun Oct 2021

Self-Regulation For Semantic Segmentation, Dong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua, Qianru Sun

Research Collection School Of Computing and Information Systems

In this paper, we seek reasons for the two major failure cases in Semantic Segmentation (SS): 1) missing small objects or minor object parts, and 2) mislabeling minor parts of large objects as wrong classes. We have an interesting finding that Failure-1 is due to the underuse of detailed features and Failure-2 is due to the underuse of visual contexts. To help the model learn a better trade-off, we introduce several Self-Regulation (SR) losses for training SS neural networks. By “self”, we mean that the losses are from the model per se without using any additional data or supervision. By …