Open Access. Powered by Scholars. Published by Universities.®

2023

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 31 - 60 of 191

Full-Text Articles in Graphics and Human Computer Interfaces

Revisiting Disentanglement And Fusion On Modality And Context In Conversational Multimodal Emotion Recognition, Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li Nov 2023

Revisiting Disentanglement And Fusion On Modality And Context In Conversational Multimodal Emotion Recognition, Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li

Research Collection School Of Computing and Information Systems

It has been a hot research topic to enable machines to understand human emotions in multimodal contexts under dialogue scenarios, which is tasked with multimodal emotion analysis in conversation (MM-ERC). MM-ERC has received consistent attention in recent years, where a diverse range of methods has been proposed for securing better task performance. Most existing works treat MM-ERC as a standard multimodal classification problem and perform multimodal feature disentanglement and fusion for maximizing feature utility. Yet after revisiting the characteristic of MM-ERC, we argue that both the feature multimodality and conversational contextualization should be properly modeled simultaneously during the feature disentanglement …


Editanything: Empowering Unparalleled Flexibility In Image Editing And Generation, Shanghua Gao, Zhijie Lin, Xingyu Xie, Pan Zhou, Ming-Ming Cheng, Shuicheng Yan Nov 2023

Editanything: Empowering Unparalleled Flexibility In Image Editing And Generation, Shanghua Gao, Zhijie Lin, Xingyu Xie, Pan Zhou, Ming-Ming Cheng, Shuicheng Yan

Research Collection School Of Computing and Information Systems

Image editing plays a vital role in computer vision field, aiming to realistically manipulate images while ensuring seamless integration. It finds numerous applications across various fields. In this work, we present EditAnything, a novel approach that empowers users with unparalleled flexibility in editing and generating image content. EditAnything introduces an array of advanced features, including crossimage dragging (e.g., try-on), region-interactive editing, controllable layout generation, and virtual character replacement. By harnessing these capabilities, users can engage in interactive and flexible editing, giving captivating outcomes that uphold the integrity of the original image. With its diverse range of tools, EditAnything caters to …


Demo Abstract: Vgglass - Demonstrating Visual Grounding And Localization Synergy With A Lidar-Enabled Smart-Glass, Darshana Rathnayake, Dulanga Weerakoon, Meeralakshmi Radhakrishnan, Vigneshwaran Subbaraju, Inseok Hwang, Archan Misra Nov 2023

Demo Abstract: Vgglass - Demonstrating Visual Grounding And Localization Synergy With A Lidar-Enabled Smart-Glass, Darshana Rathnayake, Dulanga Weerakoon, Meeralakshmi Radhakrishnan, Vigneshwaran Subbaraju, Inseok Hwang, Archan Misra

Research Collection School Of Computing and Information Systems

This work demonstrates the VGGlass system, which simultaneously interprets human instructions for a target acquisition task and determines the precise 3D positions of both user and the target object. This is achieved by utilizing LiDARs mounted in the infrastructure and a smart glass device worn by the user. Key to our system is the union of LiDAR-based localization termed LiLOC and a multi-modal visual grounding approach termed RealG(2)In-Lite. To demonstrate the system, we use Intel RealSense L515 cameras and a Microsoft HoloLens 2, as the user devices. VGGlass is able to: a) track the user in real-time in a global …


Disentangling Multi-View Representations Beyond Inductive Bias, Guanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang, Chenyang Xu, Shengfeng He Nov 2023

Disentangling Multi-View Representations Beyond Inductive Bias, Guanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang, Chenyang Xu, Shengfeng He

Research Collection School Of Computing and Information Systems

Multi-view (or -modality) representation learning aims to understand the relationships between different view representations. Existing methods disentangle multi-view representations into consistent and view-specific representations by introducing strong inductive biases, which can limit their generalization ability. In this paper, we propose a novel multi-view representation disentangling method that aims to go beyond inductive biases, ensuring both interpretability and generalizability of the resulting representations. Our method is based on the observation that discovering multi-view consistency in advance can determine the disentangling information boundary, leading to a decoupled learning objective. We also found that the consistency can be easily extracted by maximizing the …


Matk: The Meme Analytical Tool Kit, Ming Shan Hee, Aditi Kumaresan, Nguyen Khoi Hoang, Nirmalendu Prakash, Rui Cao, Roy Ka-Wei Lee Nov 2023

Matk: The Meme Analytical Tool Kit, Ming Shan Hee, Aditi Kumaresan, Nguyen Khoi Hoang, Nirmalendu Prakash, Rui Cao, Roy Ka-Wei Lee

Research Collection School Of Computing and Information Systems

The rise of social media platforms has brought about a new digital culture called memes. Memes, which combine visuals and text, can strongly influence public opinions on social and cultural issues. As a result, people have become interested in categorizing memes, leading to the development of various datasets and multimodal models that show promising results in this field. However, there is currently a lack of a single library that allows for the reproduction, evaluation, and comparison of these models using fair benchmarks and settings. To fill this gap, we introduce the Meme Analytical Tool Kit (MATK), an open-source toolkit specifically …


Cgt-Gan: Clip-Guided Text Gan For Image Captioning, Jiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu, Tong Xu, Xiangnan He Nov 2023

Cgt-Gan: Clip-Guided Text Gan For Image Captioning, Jiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu, Tong Xu, Xiangnan He

Research Collection School Of Computing and Information Systems

The large-scale visual-language pre-trained model, Contrastive Language-Image Pre-training (CLIP), has significantly improved image captioning for scenarios without human-annotated image-caption pairs. Recent advanced CLIP-based image captioning without human annotations follows a text-only training paradigm, i.e., reconstructing text from shared embedding space. Nevertheless, these approaches are limited by the training/inference gap or huge storage requirements for text embeddings. Given that it is trivial to obtain images in the real world, we propose CLIP-guided text GAN (CgT-GAN), which incorporates images into the training process to enable the model to "see" real visual modality. Particularly, we use adversarial training to teach CgT-GAN to mimic …


Constructing Holistic Spatio-Temporal Scene Graph For Video Semantic Role Labeling, Yu Zhao, Hao Fei, Yixin Cao, Bobo Li, Meishan Zhang, Jianguo Wei, Min Zhang, Tat-Seng Chua Nov 2023

Constructing Holistic Spatio-Temporal Scene Graph For Video Semantic Role Labeling, Yu Zhao, Hao Fei, Yixin Cao, Bobo Li, Meishan Zhang, Jianguo Wei, Min Zhang, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

As one of the core video semantic understanding tasks, Video Semantic Role Labeling (VidSRL) aims to detect the salient events from given videos, by recognizing the predict-argument event structures and the interrelationships between events. While recent endeavors have put forth methods for VidSRL, they can be mostly subject to two key drawbacks, including the lack of fine-grained spatial scene perception and the insufficiently modeling of video temporality. Towards this end, this work explores a novel holistic spatio-temporal scene graph (namely HostSG) representation based on the existing dynamic scene graph structures, which well model both the fine-grained spatial semantics and temporal …


Voxelhap: A Toolkit For Constructing Proxies Providing Tactile And Kinesthetic Haptic Feedback In Virtual Reality, M. Feick, C. Biyikli, K. Gani, A. Wittig, Anthony Tang, A. Krüger Nov 2023

Voxelhap: A Toolkit For Constructing Proxies Providing Tactile And Kinesthetic Haptic Feedback In Virtual Reality, M. Feick, C. Biyikli, K. Gani, A. Wittig, Anthony Tang, A. Krüger

Research Collection School Of Computing and Information Systems

Experiencing virtual environments is often limited to abstract interactions with objects. Physical proxies allow users to feel virtual objects, but are often inaccessible. We present the VoxelHap toolkit which enables users to construct highly functional proxy objects using Voxels and Plates. Voxels are blocks with special functionalities that form the core of each physical proxy. Plates increase a proxy’s haptic resolution, such as its shape, texture or weight. Beyond providing physical capabilities to realize haptic sensations, VoxelHap utilizes VR illusion techniques to expand its haptic resolution. We evaluated the capabilities of the VoxelHap toolkit through the construction of a range …


Npf-200: A Multi-Modal Eye Fixation Dataset And Method For Non-Photorealistic Videos, Ziyu Yang, Sucheng Ren, Zongwei Wu, Nanxuan Zhao, Junle Wang, Jing Qin, Shengfeng He Nov 2023

Npf-200: A Multi-Modal Eye Fixation Dataset And Method For Non-Photorealistic Videos, Ziyu Yang, Sucheng Ren, Zongwei Wu, Nanxuan Zhao, Junle Wang, Jing Qin, Shengfeng He

Research Collection School Of Computing and Information Systems

Non-photorealistic videos are in demand with the wave of the metaverse, but lack of sufficient research studies. This work aims to take a step forward to understand how humans perceive nonphotorealistic videos with eye fixation (i.e., saliency detection), which is critical for enhancing media production, artistic design, and game user experience. To fill in the gap of missing a suitable dataset for this research line, we present NPF-200, the first largescale multi-modal dataset of purely non-photorealistic videos with eye fixations. Our dataset has three characteristics: 1) it contains soundtracks that are essential according to vision and psychological studies; 2) it …


Improving Human-Automation Collaboration In Motion Planning, Torin J. Adamson Oct 2023

Improving Human-Automation Collaboration In Motion Planning, Torin J. Adamson

Computer Science ETDs

Human-automation collaboration is becoming a part of everyday life as AI helps us drive, make decisions, and solve a variety of other tasks. However, safe and effective collaboration systems depend on factors in trust, communication, and more. Existing studies to explore these are typically carried out in laboratory settings, providing robust data under tight environmental control. However, human behavior evolves over time, driven by external factors that cannot be fully captured in single participation sessions. These factors form the "human context", contextualizing the behavioral data for a more complete understanding. In this thesis, video game adaptations upon conventional subject studies …


Evocative And Provocative Image-Making In The Age Of Generative Ai, Julian Kilker Oct 2023

Evocative And Provocative Image-Making In The Age Of Generative Ai, Julian Kilker

Tradition Innovations in Arts, Design, and Media Higher Education

Editorial for inaugural AI-focused special issue of Tradition-Innovations in Arts, Design, and Media Higher Education, published under the auspices of the Alliance for the Arts in Research Universities (a2ru). Discusses three articles by five authors in this issue: (1) Choreographing Shadows: Interdisciplinary collaboration to orchestrate ethical image-making by Mark Burchick and Diana Pasulka; (2) Giving Up Control: Hybrid AI-augmented workflows for image-making by Joshua Vermillion; and (3) Hands are Hard: Unlearning how we talk about machine learning in the arts by Adam Hyland and Oscar Keyes.

Editing this special issue explored several key questions: What does “innovation” mean when …


Exploring Approaches To Engage K-12 Students In Learning Computational Thinking Using Collaborative Robots, Zoila Anuri Kanu Oct 2023

Exploring Approaches To Engage K-12 Students In Learning Computational Thinking Using Collaborative Robots, Zoila Anuri Kanu

College of Engineering Summer Undergraduate Research Program

Minority students are largely underrepresented in the STEM field. The goal for this project was to develop a program which promotes the inclusion of computation skills among students and help them work collaboratively with the use of human – robot interaction. Robots are such a strong tool that can be used to enhance computational thinking and engage students towards a technical field. Through workshops and readings about computational thinking we worked on building a block-based program that introduces the uses of robots as teaching tool for computational thinking.


Supporting Artefact Awareness In Partially-Replicated Workspaces, Emran Poh, Anthony Tang, Jenanie S. Lee, Zhao Shengdong Oct 2023

Supporting Artefact Awareness In Partially-Replicated Workspaces, Emran Poh, Anthony Tang, Jenanie S. Lee, Zhao Shengdong

Research Collection School Of Computing and Information Systems

Using Cross Reality (CR) approaches for remote collaboration will often result in partially-replicated workspaces. Here, workspace artefacts are not equally accessible - i.e. a physical artefact may only be manipulated by one collaborator - and in general, the artefacts become desynchronised over time. In this paper, we introduce a framework for artefact awareness that can help collaborators maintain an understanding of each others' manipulations with workspace artefacts. We illustrate our design explorations through sketches, and outline how we aim to study the effectiveness and utility of artefact awareness in cross reality remote collaboration. In our work, we expect to show …


Ciri: Curricular Inactivation For Residue-Aware One-Shot Video Inpainting, Weiying Zheng, Cheng Xu, Xuemiao Xu, Wenxi Liu, Shengfeng He Oct 2023

Ciri: Curricular Inactivation For Residue-Aware One-Shot Video Inpainting, Weiying Zheng, Cheng Xu, Xuemiao Xu, Wenxi Liu, Shengfeng He

Research Collection School Of Computing and Information Systems

Video inpainting aims at filling in missing regions of a video. However, when dealing with dynamic scenes with camera or object movements, annotating the inpainting target becomes laborious and impractical. In this paper, we resolve the one-shot video inpainting problem in which only one annotated first frame is provided. A naive solution is to propagate the initial target to the other frames with techniques like object tracking. In this context, the main obstacles are the unreliable propagation and the partially inpainted artifacts due to the inaccurate mask. For the former problem, we propose curricular inactivation to replace the hard masking …


Masked Diffusion Transformer Is A Strong Image Synthesizer, Shanghua Gao, Pan Zhou, Ming-Ming Cheng, Shuicheng Yan Oct 2023

Masked Diffusion Transformer Is A Strong Image Synthesizer, Shanghua Gao, Pan Zhou, Ming-Ming Cheng, Shuicheng Yan

Research Collection School Of Computing and Information Systems

Despite its success in image synthesis, we observe that diffusion probabilistic models (DPMs) often lack contextual reasoning ability to learn the relations among object parts in an image, leading to a slow learning process. To solve this issue, we propose a Masked Diffusion Transformer (MDT) that introduces a mask latent modeling scheme to explicitly enhance the DPMs’ ability to contextual relation learning among object semantic parts in an image. During training, MDT operates in the latent space to mask certain tokens. Then, an asymmetric masking diffusion transformer is designed to predict masked tokens from unmasked ones while maintaining the diffusion …


Underwater Image Translation Via Multi-Scale Generative Adversarial Network, Dongmei Yang, Tianzi Zhang, Boquan Li, Menghao Li, Weijing Chen, Xiaoqing Li, Xingmei Wang Oct 2023

Underwater Image Translation Via Multi-Scale Generative Adversarial Network, Dongmei Yang, Tianzi Zhang, Boquan Li, Menghao Li, Weijing Chen, Xiaoqing Li, Xingmei Wang

Research Collection School Of Computing and Information Systems

The role that underwater image translation plays assists in generating rare images for marine applications. However, such translation tasks are still challenging due to data lacking, insufficient feature extraction ability, and the loss of content details. To address these issues, we propose a novel multi-scale image translation model based on style-independent discriminators and attention modules (SID-AM-MSITM), which learns the mapping relationship between two unpaired images for translation. We introduce Convolution Block Attention Modules (CBAM) to the generators and discriminators of SID-AM-MSITM to improve its feature extraction ability. Moreover, we construct style-independent discriminators that enable the discriminant results of SID-AM-MSITM to …


Stprivacy: Spatio-Temporal Privacy-Preserving Action Recognition, Ming Li, Xiangyu Xu, Hehe Fan, Pan Zhou, Jun Liu, Jia-Wei Liu, Jiahe Li, Jussi Keppo, Mike Zheng Shou, Shuicheng Yan Oct 2023

Stprivacy: Spatio-Temporal Privacy-Preserving Action Recognition, Ming Li, Xiangyu Xu, Hehe Fan, Pan Zhou, Jun Liu, Jia-Wei Liu, Jiahe Li, Jussi Keppo, Mike Zheng Shou, Shuicheng Yan

Research Collection School Of Computing and Information Systems

Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks. First, they may compromise temporal dynamics in input videos, which are critical for accurate action recognition. Second, they are vulnerable to practical attacking scenarios where attackers probe for privacy from an entire video rather than individual frames. To address these issues, we propose a novel framework STPrivacy to perform video-level PPAR. For the first time, we introduce vision Transformers into PPAR by treating a video as a tubelet sequence, and accordingly design two complementary mechanisms, i.e., sparsification …


Feature Prediction Diffusion Model For Video Anomaly Detection, Cheng Yan, Shiyu Zhang, Yang Liu, Guansong Pang, Wenjun Wang Oct 2023

Feature Prediction Diffusion Model For Video Anomaly Detection, Cheng Yan, Shiyu Zhang, Yang Liu, Guansong Pang, Wenjun Wang

Research Collection School Of Computing and Information Systems

Anomaly detection in the video is an important research area and a challenging task in real applications. Due to the unavailability of large-scale annotated anomaly events, most existing video anomaly detection (VAD) methods focus on learning the distribution of normal samples to detect the substantially deviated samples as anomalies. To well learn the distribution of normal motion and appearance, many auxiliary networks are employed to extract foreground object or action information. These high-level semantic features effectively filter the noise from the background to decrease its influence on detection models. However, the capability of these extra semantic models heavily affects the …


Rigid: Recurrent Gan Inversion And Editing Of Real Face Videos, Yangyang Xu, Shengfeng He, Kwan-Yee K. Wong, Pingluo Luo Oct 2023

Rigid: Recurrent Gan Inversion And Editing Of Real Face Videos, Yangyang Xu, Shengfeng He, Kwan-Yee K. Wong, Pingluo Luo

Research Collection School Of Computing and Information Systems

GAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired inconsistent results over time. In this paper, we propose a unified recurrent framework, named Recurrent vIdeo GAN Inversion and eDiting (RIGID), to explicitly and simultaneously enforce temporally coherent GAN inversion and facial editing of real videos. Our approach models the temporal relations between current and previous frames from three aspects. To enable a faithful real video reconstruction, we first maximize the inversion fidelity and consistency by learning a temporal compensated latent code. Second, we observe …


Deep Video Demoireing Via Compact Invertible Dyadic Decomposition, Yuhui Quan, Haoran Huang, Shengfeng He, Ruotao Xu Oct 2023

Deep Video Demoireing Via Compact Invertible Dyadic Decomposition, Yuhui Quan, Haoran Huang, Shengfeng He, Ruotao Xu

Research Collection School Of Computing and Information Systems

Removing moire patterns from videos recorded on screens or complex textures is known as video demoireing. It is a challenging task as both structures and textures of an image usually exhibit strong periodic patterns, which thus are easily confused with moire patterns and can be significantly erased in the removal process. By interpreting video demoireing as a multi-frame decomposition problem, we propose a compact invertible dyadic network called CIDNet that progressively decouples latent frames and the moire patterns from an input video sequence. Using a dyadic cross-scale coupling structure with coupling layers tailored for multi-scale processing, CIDNet aims at disentangling …


Diffuse3d: Wide-Angle 3d Photography Via Bilateral Diffusion, Yutao Jiang, Yang Zhou, Yuan Liang, Wenxi Liu, Jianbo Jiao, Yuhui Quan, Shengfeng He Oct 2023

Diffuse3d: Wide-Angle 3d Photography Via Bilateral Diffusion, Yutao Jiang, Yang Zhou, Yuan Liang, Wenxi Liu, Jianbo Jiao, Yuhui Quan, Shengfeng He

Research Collection School Of Computing and Information Systems

This paper aims to resolve the challenging problem of wide-angle novel view synthesis from a single image, a.k.a. wide-angle 3D photography. Existing approaches rely on local context and treat them equally to inpaint occluded RGB and depth regions, which fail to deal with large-region occlusion (i.e., observing from an extreme angle) and foreground layers might blend into background inpainting. To address the above issues, we propose Diffuse3D which employs a pre-trained diffusion model for global synthesis, while amending the model to activate depth-aware inference. Our key insight is to alter the convolution mechanism in the denoising process. We inject depth …


Experiences Of Autistic Twitch Livestreamers: “I Have Made Easily The Most Meaningful And Impactful Relationships”, Terrance Mok, Anthony Tang, Adam Mccrimmon, Lora Oehlberg Oct 2023

Experiences Of Autistic Twitch Livestreamers: “I Have Made Easily The Most Meaningful And Impactful Relationships”, Terrance Mok, Anthony Tang, Adam Mccrimmon, Lora Oehlberg

Research Collection School Of Computing and Information Systems

We present perspectives from 10 autistic Twitch streamers regarding their experiences as livestreamers and how autism uniquely colors their experiences. Livestreaming offers a social online experience distinct from in-person, face-to-face communication, where autistic people tend to encounter challenges. Our reflexive thematic analysis of interviews with 10 participants showcases autistic livestreamers’ perspectives in their own words. Our findings center on the importance of having streamers establishing connections with other, sharing autistic identities, controlling a space for social interaction, personal growth, and accessibility challenges. In our discussion, we highlight the crucial value of having a medium for autistic representation, as well as …


Unsupervised Anomaly Detection In Medical Images With A Memory-Augmented Multi-Level Cross-Attentional Masked Autoencoder, Yu Tian, Guansong Pang, Yuyuan Liu, Chong Wang, Yuanhong Chen, Fengbei Liu, Rajvinder Singh, Johan W. Verjans, Mengyu Wang, Gustavo Carneiro Oct 2023

Unsupervised Anomaly Detection In Medical Images With A Memory-Augmented Multi-Level Cross-Attentional Masked Autoencoder, Yu Tian, Guansong Pang, Yuyuan Liu, Chong Wang, Yuanhong Chen, Fengbei Liu, Rajvinder Singh, Johan W. Verjans, Mengyu Wang, Gustavo Carneiro

Research Collection School Of Computing and Information Systems

Unsupervised anomaly detection (UAD) aims to find anomalous images by optimising a detector using a training set that contains only normal images. UAD approaches can be based on reconstruction methods, self-supervised approaches, and Imagenet pre-trained models. Reconstruction methods, which detect anomalies from image reconstruction errors, are advantageous because they do not rely on the design of problem-specific pretext tasks needed by self-supervised approaches, and on the unreliable translation of models pre-trained from non-medical datasets. However, reconstruction methods may fail because they can have low reconstruction errors even for anomalous images. In this paper, we introduce a new reconstruction-based UAD approach …


Ubisurface: A Robotic Touch Surface For Supporting Mid-Air Planar Interactions In Room-Scale Vr, Ryota Gomi, Kazuki Takashima, Yuki Onishi, Kazuyuki Fujita, Yoshifumi Kitamura Oct 2023

Ubisurface: A Robotic Touch Surface For Supporting Mid-Air Planar Interactions In Room-Scale Vr, Ryota Gomi, Kazuki Takashima, Yuki Onishi, Kazuyuki Fujita, Yoshifumi Kitamura

Research Collection School Of Computing and Information Systems

Room-scale VR has been considered an alternative to physical office workspaces. For office activities, users frequently require planar input methods, such as typing or handwriting, to quickly record annotations to virtual content. However, current off-The-shelf VR HMD setups rely on mid-Air interactions, which can cause arm fatigue and decrease input accuracy. To address this issue, we propose UbiSurface, a robotic touch surface that can automatically reposition itself to physically present a virtual planar input surface (VR whiteboard, VR canvas, etc.) to users and to permit them to achieve accurate and fatigue-less input while walking around a virtual room. We design …


Balanced Blended Space: Proposing A Universal Theoretical Framework For Combinative Reality, David Smith, Frederick Bianchi Oct 2023

Balanced Blended Space: Proposing A Universal Theoretical Framework For Combinative Reality, David Smith, Frederick Bianchi

Publications and Research

In today's fragmented societies, a unified framework for communication and collaboration across different realities is crucial. We introduce Balanced Blended Space (BBS) as a framework for describing combinative reality, encompassing virtual, physical, and conceptual realms, all intrinsically connected. Interactions within these environments shape our perceptual space. This paper outlines key axiomatic assumptions, criteria for a universal framework, and fundamental terminology. We identify deep symmetries enabling the BBS framework, including Cognitive and Computational Symmetry, Physical and Virtual Symmetry, Mediation Pathway Symmetry, Space-Time Symmetry, and Sensory Symmetry. We propose tests to determine its viability, emphasizing virtual intelligence as a collaborative partner. We …


Ai Vs. Ai: Can Ai Detect Ai-Generated Images?, Samah S. Baraheem, Tam Van Nguyen Sep 2023

Ai Vs. Ai: Can Ai Detect Ai-Generated Images?, Samah S. Baraheem, Tam Van Nguyen

Computer Science Faculty Publications

The proliferation of Artificial Intelligence (AI) models such as Generative Adversarial Net- works (GANs) has shown impressive success in image synthesis. Artificial GAN-based synthesized images have been widely spread over the Internet with the advancement in generating naturalistic and photo-realistic images. This might have the ability to improve content and media; however, it also constitutes a threat with regard to legitimacy, authenticity, and security. Moreover, implementing an automated system that is able to detect and recognize GAN-generated images is significant for image synthesis models as an evaluation tool, regardless of the input modality. To this end, we propose a framework …


Edge Distraction-Aware Salient Object Detection, Sucheng Ren, Wenxi Liu, Jianbo Jiao, Guoqiang Han, Shengfeng He Sep 2023

Edge Distraction-Aware Salient Object Detection, Sucheng Ren, Wenxi Liu, Jianbo Jiao, Guoqiang Han, Shengfeng He

Research Collection School Of Computing and Information Systems

Integrating low-level edge features has been proven to be effective in preserving clear boundaries of salient objects. However, the locality of edge features makes it difficult to capture globally salient edges, leading to distraction in the final predictions. To address this problem, we propose to produce distraction-free edge features by incorporating cross-scale holistic interdependencies between high-level features. In particular, we first formulate our edge features extraction process as a boundary-filling problem. In this way, we enforce edge features to focus on closed boundaries instead of those disconnected background edges. Second, we propose to explore cross-scale holistic contextual connections between every …


Graph-Level Anomaly Detection Via Hierarchical Memory Networks, Chaoxi Niu, Guansong Pang, Ling Chen Sep 2023

Graph-Level Anomaly Detection Via Hierarchical Memory Networks, Chaoxi Niu, Guansong Pang, Ling Chen

Research Collection School Of Computing and Information Systems

Graph-level anomaly detection aims to identify abnormal graphs that exhibit deviant structures and node attributes compared to the majority in a graph set. One primary challenge is to learn normal patterns manifested in both fine-grained and holistic views of graphs for identifying graphs that are abnormal in part or in whole. To tackle this challenge, we propose a novel approach called Hierarchical Memory Networks (HimNet), which learns hierarchical memory modules---node and graph memory modules---via a graph autoencoder network architecture. The node-level memory module is trained to model fine-grained, internal graph interactions among nodes for detecting locally abnormal graphs, while the …


One Font Doesn’T Fit All: The Influence Of Digital Text Personalization On Comprehension In Child And Adolescent Readers, Shannon M. Sheppard, Susanne L. Nobles, Anton Palma, Sophie Kajfez, Marjorie Jordan, Kathy Crowley, Sofie Beier Aug 2023

One Font Doesn’T Fit All: The Influence Of Digital Text Personalization On Comprehension In Child And Adolescent Readers, Shannon M. Sheppard, Susanne L. Nobles, Anton Palma, Sophie Kajfez, Marjorie Jordan, Kathy Crowley, Sofie Beier

Communication Sciences and Disorders Faculty Articles and Research

Reading comprehension is an essential skill. It is unclear whether and to what degree typography and font personalization may impact reading comprehension in younger readers. With advancements in technology, it is now feasible to personalize digital reading formats in general technology tools, but this feature is not yet available for many educational tools. The current study aimed to investigate the effect of character width and inter-letter spacing on reading speed and comprehension. We enrolled 94 children (kindergarten–8th grade) and compared performance with six font variations on a word-level semantic decision task (Experiment 1) and a passage-level comprehension task (Experiment 2). …


Human Recognition Theory And Facial Recognition Technology: A Topic Modeling Approach To Understanding The Ethical Implication Of A Developing Algorithmic Technologies Landscape On How We View Ourselves And Are Viewed By Others, Hajer Albalawi Aug 2023

Human Recognition Theory And Facial Recognition Technology: A Topic Modeling Approach To Understanding The Ethical Implication Of A Developing Algorithmic Technologies Landscape On How We View Ourselves And Are Viewed By Others, Hajer Albalawi

Electronic Theses and Dissertations, 2020-2023

The emergence of algorithmic-driven technology has significantly impacted human life in the current century. Algorithms, as versatile constructs, hold different meanings across various disciplines, including computer science, mathematics, social science, and human-artificial intelligence studies. This study defines algorithms from an ethical perspective as the foundation of an information society and focuses on their implications in the context of human recognition. Facial recognition technology, driven by algorithms, has gained widespread use, raising important ethical questions regarding privacy, bias, and accuracy. This dissertation aims to explore the impact of algorithms on machine perception of human individuals and how humans perceive one another …