Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 511 - 540 of 2362

Full-Text Articles in Graphics and Human Computer Interfaces

Vrmovian - An Immersive Data Annotation Tool For Visual Analysis Of Human Interactions In Vr, Isaac Browen Nov 2023

Vrmovian - An Immersive Data Annotation Tool For Visual Analysis Of Human Interactions In Vr, Isaac Browen

Student Scholar Symposium Abstracts and Posters

Understanding human behavior in virtual reality (VR) is a key component for developing intelligent systems to enhance human focused VR experiences. The ability to annotate human motion data proves to be a very useful way to analyze and understand human behavior. However, due to the complexity and multi-dimensionality of human activity data, it is necessary to develop software that can display the data in a comprehensible way and can support intuitive data annotation for developing machine learning models able recognize and assist human motions in VR (e.g., remote physical therapy). Although past research has been done to improve VR data …


User Feedback On Celebratory Technology Model For Reducing Stigma, Evelyn Lawrie, Daniel Dinh, Sav Avalos, Jack De Bruyn, Spencer Au, Christian Lopez, Ray Tan, Cyrus Fa'amafoe Nov 2023

User Feedback On Celebratory Technology Model For Reducing Stigma, Evelyn Lawrie, Daniel Dinh, Sav Avalos, Jack De Bruyn, Spencer Au, Christian Lopez, Ray Tan, Cyrus Fa'amafoe

Student Scholar Symposium Abstracts and Posters

Social stigma is a complex manifestation that affects humanity, particularly individuals with disabilities and other marginalized groups, including those with physical, cognitive, and emotional conditions. Society often judges these individuals' interactions with the world, and many technologies designed to assist those with disabilities attempt to change their daily interactions and behaviors. Nonetheless, when the emphasis is placed on validating disabled identities, there is a potential for it to be seen as "inspiration porn." This approach might inadvertently reduce inclusivity and do little to challenge negative stereotypes; it can also lead to the objectification of individuals with disabilities. Therefore, this project …


Bridging Domain Gaps For Cross-Spectrum And Long-Range Face Recognition Using Domain Adaptive Machine Learning, Cedric Armel Nimpa Fondje Nov 2023

Bridging Domain Gaps For Cross-Spectrum And Long-Range Face Recognition Using Domain Adaptive Machine Learning, Cedric Armel Nimpa Fondje

Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–

Face recognition technology has witnessed significant advancements in recent decades, enabling its widespread adoption in various applications such as security, surveillance, and biometrics applications. However, one of the primary challenges faced by existing face recognition systems is their limited performance when presented with images from different modalities or domains( such as infrared to visible, long range to close range, nighttime to daytime, profile to f rontal, etc.) Additionally, advancements in camera sensors, analytics beyond the visible spectrum, and the increasing size of cross-modal datasets have led to a particular interest in cross-modal learning for face recognition in the biometrics and …


Enhancing Search Engine Results: A Comparative Study Of Graph And Timeline Visualizations For Semantic And Temporal Relationship Discovery, Muhammad Shahiq Qureshi Nov 2023

Enhancing Search Engine Results: A Comparative Study Of Graph And Timeline Visualizations For Semantic And Temporal Relationship Discovery, Muhammad Shahiq Qureshi

Electronic Theses and Dissertations

In today’s digital age, search engines have become indispensable tools for finding information among the corpus of billions of webpages. The standard that most search engines follow is to display search results in a list-based format arranged according to a ranking algorithm. Although this format is good for presenting the most relevant results to users, it fails to represent the underlying relations between different results. These relations, among others, can generally be of either a temporal or semantic nature. A user who wants to explore the results that are connected by those relations would have to make a manual effort …


Designing Secure Mental Healthcare Chatbots For Older Adults, Aishwarya Surani Nov 2023

Designing Secure Mental Healthcare Chatbots For Older Adults, Aishwarya Surani

Electronic Theses and Dissertations

The landscape of mental health support has evolved as a result of the rising demand for digital mental healthcare services. Users now have an opportunity to seek mental health support online due to the growth of digital platforms. For those looking for mental health treatments, chatbots have evolved as user-friendly, accessible platforms that provide remote access and convenience. However, for chatbots to be effective, users must divulge personal and sensitive information, such as demographics, insurance information, and a history of mental illness. While chatbots offer services to a variety of demographic users, older adults face unique challenges related to usability, …


Optimizing E-Payment Applications For Older Adults: User-Centered Solutions To Improve Security, Privacy, Usability, And Accessibility, Urvashi Kishnani Nov 2023

Optimizing E-Payment Applications For Older Adults: User-Centered Solutions To Improve Security, Privacy, Usability, And Accessibility, Urvashi Kishnani

Electronic Theses and Dissertations

In an increasingly digital world, older adults are rapidly becoming a vital demographic in the realm of electronic financial transactions. It is imperative to address their unique needs and challenges to ensure their financial well-being. Older adults can be more vulnerable to various online threats, making security and privacy paramount. As they adapt to the digital age, understanding their specific privacy concerns and preferences is crucial for creating trustworthy e-payment systems. Moreover, enhancing the usability of e-payment applications for older adults promotes financial independence and inclusion, contributing to their overall quality of life. By focusing on these critical dimensions, we …


Turn-It-Up: Rendering Resistance For Knobs In Virtual Reality Through Undetectable Pseudo-Haptics, Martin Feick, Andre Zenner, Oscar Ariza, Anthony Tang, Cihan Biyikli, Antonio Kruger Nov 2023

Turn-It-Up: Rendering Resistance For Knobs In Virtual Reality Through Undetectable Pseudo-Haptics, Martin Feick, Andre Zenner, Oscar Ariza, Anthony Tang, Cihan Biyikli, Antonio Kruger

Research Collection School Of Computing and Information Systems

Rendering haptic feedback for interactions with virtual objects is an essential part of effective virtual reality experiences. In this work, we explore providing haptic feedback for rotational manipulations, e.g., through knobs. We propose the use of a Pseudo-Haptic technique alongside a physical proxy knob to simulate various physical resistances. In a psychophysical experiment with 20 participants, we found that designers can introduce unnoticeable offsets between real and virtual rotations of the knob, and we report the corresponding detection thresholds. Based on these, we present the Pseudo-Haptic Resistance technique to convey physical resistance while applying only unnoticeable pseudo-haptic manipulation. Additionally, we …


Metaformer Baselines For Vision, Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, Xinchao Wang Nov 2023

Metaformer Baselines For Vision, Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, Xinchao Wang

Research Collection School Of Computing and Information Systems

Abstract—MetaFormer, the abstracted architecture of Transformer, has been found to play a significant role in achieving competitive performance. In this paper, we further explore the capacity of MetaFormer, again, by migrating our focus away from the token mixer design: we introduce several baseline models under MetaFormer using the most basic or common mixers, and demonstrate their gratifying performance. We summarize our observations as follows: (1) MetaFormer ensures solid lower bound of performance. By merely adopting identity mapping as the token mixer, the MetaFormer model, termed IdentityFormer, achieves >80% accuracy on ImageNet-1K. (2) MetaFormer works well with arbitrary token mixers. When …


Pro-Cap: Leveraging A Frozen Vision-Language Model For Hateful Meme Detection, Rui Cao, Ming Shan Hee, Adriel Kuek, Wen Haw Chong, Roy Ka-Wei Lee, Jing Jiang Nov 2023

Pro-Cap: Leveraging A Frozen Vision-Language Model For Hateful Meme Detection, Rui Cao, Ming Shan Hee, Adriel Kuek, Wen Haw Chong, Roy Ka-Wei Lee, Jing Jiang

Research Collection School Of Computing and Information Systems

Hateful meme detection is a challenging multimodal task that requires comprehension of both vision and language, as well as cross-modal interactions. Recent studies have tried to fine-tune pre-trained vision-language models (PVLMs) for this task. However, with increasing model sizes, it becomes important to leverage powerful PVLMs more efficiently, rather than simply fine-tuning them. Recently, researchers have attempted to convert meme images into textual captions and prompt language models for predictions. This approach has shown good performance but suffers from non-informative image captions. Considering the two factors mentioned above, we propose a probing-based captioning approach to leverage PVLMs in a zero-shot …


Revisiting Disentanglement And Fusion On Modality And Context In Conversational Multimodal Emotion Recognition, Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li Nov 2023

Revisiting Disentanglement And Fusion On Modality And Context In Conversational Multimodal Emotion Recognition, Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li

Research Collection School Of Computing and Information Systems

It has been a hot research topic to enable machines to understand human emotions in multimodal contexts under dialogue scenarios, which is tasked with multimodal emotion analysis in conversation (MM-ERC). MM-ERC has received consistent attention in recent years, where a diverse range of methods has been proposed for securing better task performance. Most existing works treat MM-ERC as a standard multimodal classification problem and perform multimodal feature disentanglement and fusion for maximizing feature utility. Yet after revisiting the characteristic of MM-ERC, we argue that both the feature multimodality and conversational contextualization should be properly modeled simultaneously during the feature disentanglement …


Editanything: Empowering Unparalleled Flexibility In Image Editing And Generation, Shanghua Gao, Zhijie Lin, Xingyu Xie, Pan Zhou, Ming-Ming Cheng, Shuicheng Yan Nov 2023

Editanything: Empowering Unparalleled Flexibility In Image Editing And Generation, Shanghua Gao, Zhijie Lin, Xingyu Xie, Pan Zhou, Ming-Ming Cheng, Shuicheng Yan

Research Collection School Of Computing and Information Systems

Image editing plays a vital role in computer vision field, aiming to realistically manipulate images while ensuring seamless integration. It finds numerous applications across various fields. In this work, we present EditAnything, a novel approach that empowers users with unparalleled flexibility in editing and generating image content. EditAnything introduces an array of advanced features, including crossimage dragging (e.g., try-on), region-interactive editing, controllable layout generation, and virtual character replacement. By harnessing these capabilities, users can engage in interactive and flexible editing, giving captivating outcomes that uphold the integrity of the original image. With its diverse range of tools, EditAnything caters to …


Demo Abstract: Vgglass - Demonstrating Visual Grounding And Localization Synergy With A Lidar-Enabled Smart-Glass, Darshana Rathnayake, Dulanga Weerakoon, Meeralakshmi Radhakrishnan, Vigneshwaran Subbaraju, Inseok Hwang, Archan Misra Nov 2023

Demo Abstract: Vgglass - Demonstrating Visual Grounding And Localization Synergy With A Lidar-Enabled Smart-Glass, Darshana Rathnayake, Dulanga Weerakoon, Meeralakshmi Radhakrishnan, Vigneshwaran Subbaraju, Inseok Hwang, Archan Misra

Research Collection School Of Computing and Information Systems

This work demonstrates the VGGlass system, which simultaneously interprets human instructions for a target acquisition task and determines the precise 3D positions of both user and the target object. This is achieved by utilizing LiDARs mounted in the infrastructure and a smart glass device worn by the user. Key to our system is the union of LiDAR-based localization termed LiLOC and a multi-modal visual grounding approach termed RealG(2)In-Lite. To demonstrate the system, we use Intel RealSense L515 cameras and a Microsoft HoloLens 2, as the user devices. VGGlass is able to: a) track the user in real-time in a global …


Disentangling Multi-View Representations Beyond Inductive Bias, Guanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang, Chenyang Xu, Shengfeng He Nov 2023

Disentangling Multi-View Representations Beyond Inductive Bias, Guanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang, Chenyang Xu, Shengfeng He

Research Collection School Of Computing and Information Systems

Multi-view (or -modality) representation learning aims to understand the relationships between different view representations. Existing methods disentangle multi-view representations into consistent and view-specific representations by introducing strong inductive biases, which can limit their generalization ability. In this paper, we propose a novel multi-view representation disentangling method that aims to go beyond inductive biases, ensuring both interpretability and generalizability of the resulting representations. Our method is based on the observation that discovering multi-view consistency in advance can determine the disentangling information boundary, leading to a decoupled learning objective. We also found that the consistency can be easily extracted by maximizing the …


Matk: The Meme Analytical Tool Kit, Ming Shan Hee, Aditi Kumaresan, Nguyen Khoi Hoang, Nirmalendu Prakash, Rui Cao, Roy Ka-Wei Lee Nov 2023

Matk: The Meme Analytical Tool Kit, Ming Shan Hee, Aditi Kumaresan, Nguyen Khoi Hoang, Nirmalendu Prakash, Rui Cao, Roy Ka-Wei Lee

Research Collection School Of Computing and Information Systems

The rise of social media platforms has brought about a new digital culture called memes. Memes, which combine visuals and text, can strongly influence public opinions on social and cultural issues. As a result, people have become interested in categorizing memes, leading to the development of various datasets and multimodal models that show promising results in this field. However, there is currently a lack of a single library that allows for the reproduction, evaluation, and comparison of these models using fair benchmarks and settings. To fill this gap, we introduce the Meme Analytical Tool Kit (MATK), an open-source toolkit specifically …


Cgt-Gan: Clip-Guided Text Gan For Image Captioning, Jiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu, Tong Xu, Xiangnan He Nov 2023

Cgt-Gan: Clip-Guided Text Gan For Image Captioning, Jiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu, Tong Xu, Xiangnan He

Research Collection School Of Computing and Information Systems

The large-scale visual-language pre-trained model, Contrastive Language-Image Pre-training (CLIP), has significantly improved image captioning for scenarios without human-annotated image-caption pairs. Recent advanced CLIP-based image captioning without human annotations follows a text-only training paradigm, i.e., reconstructing text from shared embedding space. Nevertheless, these approaches are limited by the training/inference gap or huge storage requirements for text embeddings. Given that it is trivial to obtain images in the real world, we propose CLIP-guided text GAN (CgT-GAN), which incorporates images into the training process to enable the model to "see" real visual modality. Particularly, we use adversarial training to teach CgT-GAN to mimic …


Constructing Holistic Spatio-Temporal Scene Graph For Video Semantic Role Labeling, Yu Zhao, Hao Fei, Yixin Cao, Bobo Li, Meishan Zhang, Jianguo Wei, Min Zhang, Tat-Seng Chua Nov 2023

Constructing Holistic Spatio-Temporal Scene Graph For Video Semantic Role Labeling, Yu Zhao, Hao Fei, Yixin Cao, Bobo Li, Meishan Zhang, Jianguo Wei, Min Zhang, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

As one of the core video semantic understanding tasks, Video Semantic Role Labeling (VidSRL) aims to detect the salient events from given videos, by recognizing the predict-argument event structures and the interrelationships between events. While recent endeavors have put forth methods for VidSRL, they can be mostly subject to two key drawbacks, including the lack of fine-grained spatial scene perception and the insufficiently modeling of video temporality. Towards this end, this work explores a novel holistic spatio-temporal scene graph (namely HostSG) representation based on the existing dynamic scene graph structures, which well model both the fine-grained spatial semantics and temporal …


Voxelhap: A Toolkit For Constructing Proxies Providing Tactile And Kinesthetic Haptic Feedback In Virtual Reality, M. Feick, C. Biyikli, K. Gani, A. Wittig, Anthony Tang, A. Krüger Nov 2023

Voxelhap: A Toolkit For Constructing Proxies Providing Tactile And Kinesthetic Haptic Feedback In Virtual Reality, M. Feick, C. Biyikli, K. Gani, A. Wittig, Anthony Tang, A. Krüger

Research Collection School Of Computing and Information Systems

Experiencing virtual environments is often limited to abstract interactions with objects. Physical proxies allow users to feel virtual objects, but are often inaccessible. We present the VoxelHap toolkit which enables users to construct highly functional proxy objects using Voxels and Plates. Voxels are blocks with special functionalities that form the core of each physical proxy. Plates increase a proxy’s haptic resolution, such as its shape, texture or weight. Beyond providing physical capabilities to realize haptic sensations, VoxelHap utilizes VR illusion techniques to expand its haptic resolution. We evaluated the capabilities of the VoxelHap toolkit through the construction of a range …


Npf-200: A Multi-Modal Eye Fixation Dataset And Method For Non-Photorealistic Videos, Ziyu Yang, Sucheng Ren, Zongwei Wu, Nanxuan Zhao, Junle Wang, Jing Qin, Shengfeng He Nov 2023

Npf-200: A Multi-Modal Eye Fixation Dataset And Method For Non-Photorealistic Videos, Ziyu Yang, Sucheng Ren, Zongwei Wu, Nanxuan Zhao, Junle Wang, Jing Qin, Shengfeng He

Research Collection School Of Computing and Information Systems

Non-photorealistic videos are in demand with the wave of the metaverse, but lack of sufficient research studies. This work aims to take a step forward to understand how humans perceive nonphotorealistic videos with eye fixation (i.e., saliency detection), which is critical for enhancing media production, artistic design, and game user experience. To fill in the gap of missing a suitable dataset for this research line, we present NPF-200, the first largescale multi-modal dataset of purely non-photorealistic videos with eye fixations. Our dataset has three characteristics: 1) it contains soundtracks that are essential according to vision and psychological studies; 2) it …


Improving Human-Automation Collaboration In Motion Planning, Torin J. Adamson Oct 2023

Improving Human-Automation Collaboration In Motion Planning, Torin J. Adamson

Computer Science ETDs

Human-automation collaboration is becoming a part of everyday life as AI helps us drive, make decisions, and solve a variety of other tasks. However, safe and effective collaboration systems depend on factors in trust, communication, and more. Existing studies to explore these are typically carried out in laboratory settings, providing robust data under tight environmental control. However, human behavior evolves over time, driven by external factors that cannot be fully captured in single participation sessions. These factors form the "human context", contextualizing the behavioral data for a more complete understanding. In this thesis, video game adaptations upon conventional subject studies …


Evocative And Provocative Image-Making In The Age Of Generative Ai, Julian Kilker Oct 2023

Evocative And Provocative Image-Making In The Age Of Generative Ai, Julian Kilker

Tradition Innovations in Arts, Design, and Media Higher Education

Editorial for inaugural AI-focused special issue of Tradition-Innovations in Arts, Design, and Media Higher Education, published under the auspices of the Alliance for the Arts in Research Universities (a2ru). Discusses three articles by five authors in this issue: (1) Choreographing Shadows: Interdisciplinary collaboration to orchestrate ethical image-making by Mark Burchick and Diana Pasulka; (2) Giving Up Control: Hybrid AI-augmented workflows for image-making by Joshua Vermillion; and (3) Hands are Hard: Unlearning how we talk about machine learning in the arts by Adam Hyland and Oscar Keyes.

Editing this special issue explored several key questions: What does “innovation” mean when …


Exploring Approaches To Engage K-12 Students In Learning Computational Thinking Using Collaborative Robots, Zoila Anuri Kanu Oct 2023

Exploring Approaches To Engage K-12 Students In Learning Computational Thinking Using Collaborative Robots, Zoila Anuri Kanu

College of Engineering Summer Undergraduate Research Program

Minority students are largely underrepresented in the STEM field. The goal for this project was to develop a program which promotes the inclusion of computation skills among students and help them work collaboratively with the use of human – robot interaction. Robots are such a strong tool that can be used to enhance computational thinking and engage students towards a technical field. Through workshops and readings about computational thinking we worked on building a block-based program that introduces the uses of robots as teaching tool for computational thinking.


Supporting Artefact Awareness In Partially-Replicated Workspaces, Emran Poh, Anthony Tang, Jenanie S. Lee, Zhao Shengdong Oct 2023

Supporting Artefact Awareness In Partially-Replicated Workspaces, Emran Poh, Anthony Tang, Jenanie S. Lee, Zhao Shengdong

Research Collection School Of Computing and Information Systems

Using Cross Reality (CR) approaches for remote collaboration will often result in partially-replicated workspaces. Here, workspace artefacts are not equally accessible - i.e. a physical artefact may only be manipulated by one collaborator - and in general, the artefacts become desynchronised over time. In this paper, we introduce a framework for artefact awareness that can help collaborators maintain an understanding of each others' manipulations with workspace artefacts. We illustrate our design explorations through sketches, and outline how we aim to study the effectiveness and utility of artefact awareness in cross reality remote collaboration. In our work, we expect to show …


Ciri: Curricular Inactivation For Residue-Aware One-Shot Video Inpainting, Weiying Zheng, Cheng Xu, Xuemiao Xu, Wenxi Liu, Shengfeng He Oct 2023

Ciri: Curricular Inactivation For Residue-Aware One-Shot Video Inpainting, Weiying Zheng, Cheng Xu, Xuemiao Xu, Wenxi Liu, Shengfeng He

Research Collection School Of Computing and Information Systems

Video inpainting aims at filling in missing regions of a video. However, when dealing with dynamic scenes with camera or object movements, annotating the inpainting target becomes laborious and impractical. In this paper, we resolve the one-shot video inpainting problem in which only one annotated first frame is provided. A naive solution is to propagate the initial target to the other frames with techniques like object tracking. In this context, the main obstacles are the unreliable propagation and the partially inpainted artifacts due to the inaccurate mask. For the former problem, we propose curricular inactivation to replace the hard masking …


Masked Diffusion Transformer Is A Strong Image Synthesizer, Shanghua Gao, Pan Zhou, Ming-Ming Cheng, Shuicheng Yan Oct 2023

Masked Diffusion Transformer Is A Strong Image Synthesizer, Shanghua Gao, Pan Zhou, Ming-Ming Cheng, Shuicheng Yan

Research Collection School Of Computing and Information Systems

Despite its success in image synthesis, we observe that diffusion probabilistic models (DPMs) often lack contextual reasoning ability to learn the relations among object parts in an image, leading to a slow learning process. To solve this issue, we propose a Masked Diffusion Transformer (MDT) that introduces a mask latent modeling scheme to explicitly enhance the DPMs’ ability to contextual relation learning among object semantic parts in an image. During training, MDT operates in the latent space to mask certain tokens. Then, an asymmetric masking diffusion transformer is designed to predict masked tokens from unmasked ones while maintaining the diffusion …


Underwater Image Translation Via Multi-Scale Generative Adversarial Network, Dongmei Yang, Tianzi Zhang, Boquan Li, Menghao Li, Weijing Chen, Xiaoqing Li, Xingmei Wang Oct 2023

Underwater Image Translation Via Multi-Scale Generative Adversarial Network, Dongmei Yang, Tianzi Zhang, Boquan Li, Menghao Li, Weijing Chen, Xiaoqing Li, Xingmei Wang

Research Collection School Of Computing and Information Systems

The role that underwater image translation plays assists in generating rare images for marine applications. However, such translation tasks are still challenging due to data lacking, insufficient feature extraction ability, and the loss of content details. To address these issues, we propose a novel multi-scale image translation model based on style-independent discriminators and attention modules (SID-AM-MSITM), which learns the mapping relationship between two unpaired images for translation. We introduce Convolution Block Attention Modules (CBAM) to the generators and discriminators of SID-AM-MSITM to improve its feature extraction ability. Moreover, we construct style-independent discriminators that enable the discriminant results of SID-AM-MSITM to …


Stprivacy: Spatio-Temporal Privacy-Preserving Action Recognition, Ming Li, Xiangyu Xu, Hehe Fan, Pan Zhou, Jun Liu, Jia-Wei Liu, Jiahe Li, Jussi Keppo, Mike Zheng Shou, Shuicheng Yan Oct 2023

Stprivacy: Spatio-Temporal Privacy-Preserving Action Recognition, Ming Li, Xiangyu Xu, Hehe Fan, Pan Zhou, Jun Liu, Jia-Wei Liu, Jiahe Li, Jussi Keppo, Mike Zheng Shou, Shuicheng Yan

Research Collection School Of Computing and Information Systems

Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks. First, they may compromise temporal dynamics in input videos, which are critical for accurate action recognition. Second, they are vulnerable to practical attacking scenarios where attackers probe for privacy from an entire video rather than individual frames. To address these issues, we propose a novel framework STPrivacy to perform video-level PPAR. For the first time, we introduce vision Transformers into PPAR by treating a video as a tubelet sequence, and accordingly design two complementary mechanisms, i.e., sparsification …


Feature Prediction Diffusion Model For Video Anomaly Detection, Cheng Yan, Shiyu Zhang, Yang Liu, Guansong Pang, Wenjun Wang Oct 2023

Feature Prediction Diffusion Model For Video Anomaly Detection, Cheng Yan, Shiyu Zhang, Yang Liu, Guansong Pang, Wenjun Wang

Research Collection School Of Computing and Information Systems

Anomaly detection in the video is an important research area and a challenging task in real applications. Due to the unavailability of large-scale annotated anomaly events, most existing video anomaly detection (VAD) methods focus on learning the distribution of normal samples to detect the substantially deviated samples as anomalies. To well learn the distribution of normal motion and appearance, many auxiliary networks are employed to extract foreground object or action information. These high-level semantic features effectively filter the noise from the background to decrease its influence on detection models. However, the capability of these extra semantic models heavily affects the …


Rigid: Recurrent Gan Inversion And Editing Of Real Face Videos, Yangyang Xu, Shengfeng He, Kwan-Yee K. Wong, Pingluo Luo Oct 2023

Rigid: Recurrent Gan Inversion And Editing Of Real Face Videos, Yangyang Xu, Shengfeng He, Kwan-Yee K. Wong, Pingluo Luo

Research Collection School Of Computing and Information Systems

GAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired inconsistent results over time. In this paper, we propose a unified recurrent framework, named Recurrent vIdeo GAN Inversion and eDiting (RIGID), to explicitly and simultaneously enforce temporally coherent GAN inversion and facial editing of real videos. Our approach models the temporal relations between current and previous frames from three aspects. To enable a faithful real video reconstruction, we first maximize the inversion fidelity and consistency by learning a temporal compensated latent code. Second, we observe …


Deep Video Demoireing Via Compact Invertible Dyadic Decomposition, Yuhui Quan, Haoran Huang, Shengfeng He, Ruotao Xu Oct 2023

Deep Video Demoireing Via Compact Invertible Dyadic Decomposition, Yuhui Quan, Haoran Huang, Shengfeng He, Ruotao Xu

Research Collection School Of Computing and Information Systems

Removing moire patterns from videos recorded on screens or complex textures is known as video demoireing. It is a challenging task as both structures and textures of an image usually exhibit strong periodic patterns, which thus are easily confused with moire patterns and can be significantly erased in the removal process. By interpreting video demoireing as a multi-frame decomposition problem, we propose a compact invertible dyadic network called CIDNet that progressively decouples latent frames and the moire patterns from an input video sequence. Using a dyadic cross-scale coupling structure with coupling layers tailored for multi-scale processing, CIDNet aims at disentangling …


Diffuse3d: Wide-Angle 3d Photography Via Bilateral Diffusion, Yutao Jiang, Yang Zhou, Yuan Liang, Wenxi Liu, Jianbo Jiao, Yuhui Quan, Shengfeng He Oct 2023

Diffuse3d: Wide-Angle 3d Photography Via Bilateral Diffusion, Yutao Jiang, Yang Zhou, Yuan Liang, Wenxi Liu, Jianbo Jiao, Yuhui Quan, Shengfeng He

Research Collection School Of Computing and Information Systems

This paper aims to resolve the challenging problem of wide-angle novel view synthesis from a single image, a.k.a. wide-angle 3D photography. Existing approaches rely on local context and treat them equally to inpaint occluded RGB and depth regions, which fail to deal with large-region occlusion (i.e., observing from an extreme angle) and foreground layers might blend into background inpainting. To address the above issues, we propose Diffuse3D which employs a pre-trained diffusion model for global synthesis, while amending the model to activate depth-aware inference. Our key insight is to alter the convolution mechanism in the denoising process. We inject depth …