Empathyear : An Open-Source Avatar Multimodal Empathetic Chatbot,
2024
Singapore Management University
Empathyear : An Open-Source Avatar Multimodal Empathetic Chatbot, Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, Erik Cambria
Research Collection School Of Computing and Information Systems
This paper introduces EmpathyEar, a pioneering open-source, avatar-based multimodal empathetic chatbot, to fill the gap in traditional text-only empathetic response generation (ERG) systems. Leveraging the advancements of a large language model, combined with multimodal encoders and generators, EmpathyEar supports user inputs in any combination of text, sound, and vision, and produces multimodal empathetic responses, offering users, not just textual responses but also digital avatars with talking faces and synchronized speeches. A series of emotion-aware instruction-tuning is performed for comprehensive emotional understanding and generation capabilities. In this way, EmpathyEar provides users with responses that achieve a deeper emotional resonance, closely emulating …
Unifying Global-Local Representations In Salient Object Detection With Transformers,
2024
Singapore Management University
Unifying Global-Local Representations In Salient Object Detection With Transformers, Sucheng Ren, Nanxuan Zhao, Qiang Wen, Guoqiang Han, Shengfeng He
Research Collection School Of Computing and Information Systems
The fully convolutional network (FCN) has dominated salient object detection for a long period. However, the locality of CNN requires the model deep enough to have a global receptive field and such a deep model always leads to the loss of local details. In this paper, we introduce a new attention-based encoder, vision transformer, into salient object detection to ensure the globalization of the representations from shallow to deep layers. With the global view in very shallow layers, the transformer encoder preserves more local representations to recover the spatial details in final saliency maps. Besides, as each layer can capture …
Exploring A Multimodal Fusion-Based Deep Learning Network For Detecting Facial Palsy,
2024
Singapore Management University
Exploring A Multimodal Fusion-Based Deep Learning Network For Detecting Facial Palsy, Heng Yim Nicole Oo, Min Hun Lee, J. H. Lim
Research Collection School Of Computing and Information Systems
Algorithmic detection of facial palsy offers the potential to improve current practices, which usually involve labor-intensive and subjective assessment by clinicians. In this paper, we present a multimodal fusion-based deep learning model that utilizes unstructured data (i.e. an image frame with facial line segments) and structured data (i.e. features of facial expressions) to detect facial palsy. We then contribute to a study to analyze the effect of different data modalities and the benefits of a multimodal fusion-based approach using videos of 21 facial palsy patients. Our experimental results show that among various data modalities (i.e. unstructured data - RGB images …
How To Avoid Jumping To Conclusions: Measuring The Robustness Of Outstanding Facts In Knowledge Graphs,
2024
Singapore Management University
How To Avoid Jumping To Conclusions: Measuring The Robustness Of Outstanding Facts In Knowledge Graphs, Hanhua Xiao, Yuchen Li, Yanhao Wang, Panagiotis Karras, Kyriakos Mouratidis, Natalia Rozalia Avlona
Research Collection School Of Computing and Information Systems
An outstanding fact (OF) is a striking claim by which some entities stand out from their peers on someattribute. OFs serve data journalism, fact checking, and recommendation. However, one could jump to conclusions by selecting truthful OFs while intentionally or inadvertently ignoring lateral contexts and data that render them less striking. This jumping conclusion bias from unstable OFs may disorient the public, including voters and consumers, raising concerns about fairness and transparency in political and business competition. It is thus ethically imperative for several stakeholders to measure the robustness of OFs with respect to lateral contexts and data. Unfortunately, a …
Heterogeneous Graph Transformer With Poly-Tokenization,
2024
Singapore Management University
Heterogeneous Graph Transformer With Poly-Tokenization, Zhiyuan Lu, Yuan Fang, Cheng Yang, Chuan Shi
Research Collection School Of Computing and Information Systems
Graph neural networks have shown widespread success for learning on graphs, but they still face fundamental drawbacks, such as limited expressive power, over-smoothing, and over-squashing. Meanwhile, the transformer architecture offers a potential solution to these issues. However, existing graph transformers primarily cater to homogeneous graphs and are unable to model the intricate semantics of heterogeneous graphs. Moreover, unlike small molecular graphs where the entire graph can be considered as the receptive field in graph transformers, real-world heterogeneous graphs comprise a significantly larger number of nodes and cannot be entirely treated as such. Consequently, existing graph transformers struggle to capture the …
Human Centered Approaches And Taxonomies For Explainable Artificial Intelligence,
2024
Technological University Dublin
Human Centered Approaches And Taxonomies For Explainable Artificial Intelligence, Helen Sheridan, Emma Murphy, Dympna O'Sullivan
Conference papers
Recent interest within the research community related to explainable artificial intelligence (XAI) has led to a profuse amount of literature on the subject. Those who wish to tackle the domain from an HCI focus may be presented with overwhelming material, most of which does not pertain to human aspects of XAI. Taxonomies can serve to categorize a subject into topic areas and distill content into an overview of the field. This late breaking work intends to help those within the HCI community with a focus on XAI to understand relevant aspects of human centered XAI. We also present a taxonomy …
Hierarchical Damage Correlations For Old Photo Restoration,
2024
Singapore Management University
Hierarchical Damage Correlations For Old Photo Restoration, Weiwei Cai, Xuemiao Xu, Jiajia Xu, Huaidong Zhang, Haoxin Yang, Kun Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
Restoring old photographs can preserve cherished memories. Previous methods handled diverse damages within the same network structure, which proved impractical. In addition, these methods cannot exploit correlations among artifacts, especially in scratches versus patch-misses issues. Hence, a tailored network is particularly crucial. In light of this, we propose a unified framework consisting of two key components: ScratchNet and PatchNet. In detail, ScratchNet employs the parallel Multi-scale Partial Convolution Module to effectively repair scratches, learning from multi-scale local receptive fields. In contrast, the patch-misses necessitate the network to emphasize global information. To this end, we incorporate a transformer-based encoder and decoder …
How People Prompt Generative Ai To Create Interactive Vr Scenes,
2024
University of Calgary
How People Prompt Generative Ai To Create Interactive Vr Scenes, Setareh Aghel Manesh, Tianyi Zhang, Yuki Onishi, Kotaro Hara, Scott Bateman, Jiannan Li, Anthony Tang
Research Collection School Of Computing and Information Systems
Generative AI tools can provide people with the ability to create virtual environments and scenes with natural language prompts. Yet, how people will formulate such prompts is unclear---particularly when they inhabit the environment that they are designing. For instance, it is likely that a person might say, "Put a chair here,'' while pointing at a location. If such linguistic and embodied features are common to people's prompts, we need to tune models to accommodate them. In this work, we present a Wizard of Oz elicitation study with 22 participants, where we studied people's implicit expectations when verbally prompting such programming …
Certified Robust Accuracy Of Neural Networks Are Bounded Due To Bayes Errors,
2024
Singapore Management University
Certified Robust Accuracy Of Neural Networks Are Bounded Due To Bayes Errors, Ruihan Zhang, Jun Sun
Research Collection School Of Computing and Information Systems
Adversarial examples pose a security threat to many critical systems built on neural networks. While certified training improves robustness, it also decreases accuracy noticeably. Despite various proposals for addressing this issue, the significant accuracy drop remains. More importantly, it is not clear whether there is a certain fundamental limit on achieving robustness whilst maintaining accuracy. In this work, we offer a novel perspective based on Bayes errors. By adopting Bayes error to robustness analysis, we investigate the limit of certified robust accuracy, taking into account data distribution uncertainties. We first show that the accuracy inevitably decreases in the pursuit of …
Jigsaw: Edge-Based Streaming Perception Over Spatially Overlapped Multi-Camera Deployments,
2024
Singapore Management University
Jigsaw: Edge-Based Streaming Perception Over Spatially Overlapped Multi-Camera Deployments, Ila Gokarn, Yigong Hu, Tarek Abdelzaher, Archan Misra
Research Collection School Of Computing and Information Systems
We present JIGSAW, a novel system that performs edge-based streaming perception over multiple video streams, while additionally factoring in the redundancy offered by the spatial overlap often exhibited in urban, multi-camera deployments. To assure high streaming throughput, JIGSAW extracts and spatially multiplexes multiple regions-of-interest from different camera frames into a smaller canvas frame. Moreover, to ensure that perception stays abreast of evolving object kinematics, JIGSAW includes a utility-based weighted scheduler to preferentially prioritize and even skip object-specific tiles extracted from an incoming stream of camera frames. Using the CityflowV2 traffic surveillance dataset, we show that JIGSAW can simultaneously process 25 …
My Ai Companion: An Examination Of The Removal Of Erotic Role Play From Replika Through User Discussion On Reddit,
2024
University of Nebraska-Lincoln
My Ai Companion: An Examination Of The Removal Of Erotic Role Play From Replika Through User Discussion On Reddit, Chelsee M. Allen
Department of Sociology: Dissertations, Theses, and Student Research
The development of artificial intelligence (AI) software has expanded rapidly in recent years, and thus has emerged the importance of exploring human relationships with AI chatbots. Replika, an app which uses AI to mimic human conversation, removed a function called Erotic Role Play (ERP) that allowed for sexual conversation with users’ customizable chatbots in February of 2023. This exploratory qualitative study examines the aftermath of ERP’s removal through an analysis of user interactions on Reddit. Five overarching themes emerged through the analysis of top posts to a Replika-specific subreddit, encompassing topics around mental health, stigma, coping, sex work and gendered …
Creating And Delivering Audio Descriptions For Videos,
2024
Singapore Management University
Creating And Delivering Audio Descriptions For Videos, Rosiana Natalie
Dissertations and Theses Collection (Open Access)
Despite anti-discrimination regulations mandating the provision of audio descriptions (ADs), the majority of online video content remains inaccessible to blind and low-vision (BLV) individuals. This is because these ADs are either absent or fail to adequately address the diverse and unique needs of the audience. Traditionally, content creators have relied on professionals to author ADs. However, this gold standard may not be accessible for some content creators because this method is still costly and has a long turnaround time. Moreover, when ADs are available, they tend to be static and unalterable, failing to cater to the unique preferences of BLV …
An Exploratory Study Of Conventional Machine Learning And Large Language Models For Sentiment Analysis,
2024
Singapore Management University
An Exploratory Study Of Conventional Machine Learning And Large Language Models For Sentiment Analysis, Cui Zou, Jingyuan Cai, Langtao Chen, Fiona Fui-Hoon Nah
Research Collection School Of Computing and Information Systems
Sentiment analysis is the use of natural language processing to identify affective states and determine people’s opinions in various analytical applications such as customer reviews and social media analyses. Large language models (LLMs) such as GPT-4o demonstrate impressive performance in text generation tasks. Despite numerous studies in the extant literature, few have compared the performance of conventional machine learning models with LLMs for sentiment analysis. This study aims to fill this gap by conducting an evaluation of these models using a balanced dataset of 2,000 IMDb movie reviews. Our study shows that GPT-4o achieves the highest performance, while GPT-3.5 and …
A Computational Aesthetic Design Science Study On Online Video Based On Triple-Dimensional Multimodal Analysis,
2024
Singapore Management University
A Computational Aesthetic Design Science Study On Online Video Based On Triple-Dimensional Multimodal Analysis, Zhangguang Kang, Fiona Fui-Hoon Nah, Keng Siau
Research Collection School Of Computing and Information Systems
Computational video aesthetic prediction refers to using models that automatically evaluate the features of videos to produce their aesthetic scores. Current video aesthetic prediction models are designed based on bimodal frameworks. To address their limitations, we developed the Triple-Dimensional Multimodal Temporal Video Aesthetic neural network (TMTVA-net) model. The Long Short-Term Memory (LSTM) forms the conceptual foundation for the design framework. In the multimodal transformer layer, we employed two distinct transformers: the multimodal transformer and the feature transformer, enabling the acquisition of modality-specific patterns and representational features uniquely adapted to each modality. The fusion layer has also been redesigned to compute …
Understanding And Fighting Scams: Media, Language, Appeals And Effects,
2024
City University of Hong Kong
Understanding And Fighting Scams: Media, Language, Appeals And Effects, Shuhua Zhou, Xiao Fan Liu, Fiona Fui-Hoon Nah, S. Harrison, X. Zhang, S. Zhen, D. Yeung, J. Hsiao, R. Lc, A. Chan, X. Wang, C. Jiang, F. Lin, J. Li, A. Wong, L. Chan, B. George, P. Li
Research Collection School Of Computing and Information Systems
Scams are fraudulent activities aiming to deceive individuals into relinquishing money, property, or rights, and they have proliferated in the context of widespread misinformation and disinformation. In this paper, we propose strategies and a research plan to address key questions about the exploitation of new communication technologies by scammers, the prevalence and nature of different scam types, and the language characteristics and appeals used in scamming content. We aim to develop a comprehensive taxonomy of scams and identify factors that contribute to their persuasiveness. Additionally, we propose the use of advanced technologies, including artificial intelligence, physiological measures, and brain mapping, …
Evilscreen Attack: Smart Tv Hijacking Via Multi-Channel Remote Control Mimicry,
2024
Singapore Management University
Evilscreen Attack: Smart Tv Hijacking Via Multi-Channel Remote Control Mimicry, Yiwei Zhang, Siqi Ma, Tiancheng Chen, Juanru Li, Robert H. Deng, Elisa Bertino
Research Collection School Of Computing and Information Systems
Modern smart TVs often communicate with their remote controls (including the smartphone simulated ones) using multiple wireless channels (e.g., Infrared, Bluetooth, and Wi-Fi). However, this multi-channel remote control communication introduces a new attack surface. An inherent security flaw is that remote controls of most smart TVs are designed to work in a benign environment rather than an adversarial one, and thus wireless communications between a smart TV and its remote controls are not strongly protected. Attackers can leverage such a flaw to abuse the remote control communication and compromise smart TV systems. In this paper, we propose EvilScreen, a novel …
Learning Topological Representations With Bidirectional Graph Attention Network For Solving Job Shop Scheduling Problem,
2024
Singapore Management University
Learning Topological Representations With Bidirectional Graph Attention Network For Solving Job Shop Scheduling Problem, Cong Zhang, Zhiguang Cao, Yaoxin Wu, Wen Song, Jing Sun
Research Collection School Of Computing and Information Systems
Existing learning-based methods for solving job shop scheduling problems (JSSP) usually use off-the-shelf GNN models tailored to undirected graphs and neglect the rich and meaningful topological structures of disjunctive graphs (DGs). This paper proposes the topology-aware bidirectional graph attention network (TBGAT), a novel GNN architecture based on the attention mechanism, to embed the DG for solving JSSP in a local search framework. Specifically, TBGAT embeds the DG from a forward and a backward view, respectively, where the messages are propagated by following the different topologies of the views and aggregated via graph attention. Then, we propose a novel operator based …
Creative Insights Into Motion: Enhancing Human Activity Understanding With 3d Data Visualization And Annotation,
2024
Chapman University
Creative Insights Into Motion: Enhancing Human Activity Understanding With 3d Data Visualization And Annotation, Isaac Browen, Hector M. Camarillo-Abad, Franceli L. Cibrian, Trudi Di Qi
Engineering Faculty Articles and Research
This paper presents a novel 3D system for human motion analysis - Motion Data Visualization and Annotation (MoViAn). Designed to provide a comprehensive visual representation of 3D human motion data, MoViAn incorporates detailed visualization of gaze direction, hand movements, and object interactions, alongside an interactive interface for efficient data annotation. A user study involving eight participants indicates that MoViAn enables users to thoroughly explore and annotate human motion data, with System Usability Scale (SUS) results demonstrating a satisfactory usability level. The contribution of this paper lies in the development of an interactive and usable data analytics tool aimed at deepening …
Technology And Homelessness: How Website Design And Blockchain Technology Could Impact The Unhoused,
2024
University of Denver
Technology And Homelessness: How Website Design And Blockchain Technology Could Impact The Unhoused, Casey Pratt
Undergraduate Theses, Capstones, and Recitals
Although technology could be used to combat inequality, it is instead increasing it. This paper discusses how the unhoused population suffers at the hand of technological inequality despite being relatively offline. It presents theories on how this would change if we reapproached how technology is used to assist the unhoused. It suggests implementing blockchain as a resource as well as modifying the websites built to assist in accessing benefits. Employees at shelters are interviewed for this paper about their experiences with using digital resources to rehouse and restabilize the vulnerable. They are asked how the sites can be improved for …
Curating Familiarity Within The Unfamiliar: Exploring Non-Native Mobile App Experiences To Create Cross-Cultural Design Frameworks,
2024
Dartmouth College
Curating Familiarity Within The Unfamiliar: Exploring Non-Native Mobile App Experiences To Create Cross-Cultural Design Frameworks, Hanna Hong
Computer Science Senior Theses
Global mobility and markets are expanding, and as a result, countries are becoming less and less monocultural. With multiple cultural affinity groups to cater towards, companies often will deploy different versions of a website or app based on the country a user is accessing it from. This strategy of catering to geographic location results in a lack of accommodation for people living within a culture that is different from their native one. In order to increase accessibility and equal ease-of-use for all audiences, designers should understand and work towards the needs of a multicultural user base. This study investigates how …
