Graph Contrastive Learning With Stable And Scalable Spectral Encoding,
2023
Singapore Management University
Graph Contrastive Learning With Stable And Scalable Spectral Encoding, Deyu Bo, Yuan Fang, Yang Liu, Chuan Shi
Research Collection School Of Computing and Information Systems
Graph contrastive learning (GCL) aims to learn representations by capturing the agreements between different graph views. Traditional GCL methods generate views in the spatial domain, but it has been recently discovered that the spectral domain also plays a vital role in complementing spatial views. However, existing spectral-based graph views either ignore the eigenvectors that encode valuable positional information, or suffer from high complexity when trying to address the instability of spectral features. To tackle these challenges, we first design an informative, stable, and scalable spectral encoder, termed EigenMLP, to learn effective representations from the spectral features. Theoretically, EigenMLP is invariant …
Clueless: Revolutionizing Sustainable Fashion And Combating Overconsumption,
2023
California Polytechnic State University, San Luis Obispo
Clueless: Revolutionizing Sustainable Fashion And Combating Overconsumption, Tanya Ravichandran
Graphic Communication
“Clueless” revolutionizes sustainable fashion by combating wardrobe overconsumption and the industry’s carbon footprint, using AI to suggest personalized outfits from existing wardrobes tailored to weather and wear history. It enhances user engagement through features like outfit ‘shuffle’ and provides insights into wardrobe utilization and carbon impact.
It’s more than an app; it’s a step towards a greener wardrobe and a healthier planet.
Leveraging Artificial Intelligence For Team Cognition In Human-Ai Teams,
2023
Clemson University
Leveraging Artificial Intelligence For Team Cognition In Human-Ai Teams, Beau Schelble
All Dissertations
Advances in artificial intelligence (AI) technologies have enabled AI to be applied across a wide variety of new fields like cryptography, art, and data analysis. Several of these fields are social in nature, including decision-making and teaming, which introduces a new set of challenges for AI research. While each of these fields has its unique challenges, the area of human-AI teaming is beset with many that center around the expectations and abilities of AI teammates. One such challenge is understanding team cognition in these human-AI teams and AI teammates' ability to contribute towards, support, and encourage it. Team cognition is …
Mermaid: A Dataset And Framework For Multimodal Meme Semantic Understanding,
2023
Singapore Management University
Mermaid: A Dataset And Framework For Multimodal Meme Semantic Understanding, Shaun Toh, Adriel Kuek, Wen Haw Chong, Roy Ka Wei Lee
Research Collection School Of Computing and Information Systems
Memes are widely used to convey cultural and societal issues and have a significant impact on public opinion. However, little work has been done on understanding and explaining the semantics expressed in multimodal memes. To fill this research gap, we introduce MERMAID, a dataset consisting of 3,633 memes annotated with their entities and relations, and propose a novel MERF pipeline that extracts entities and their relationships in memes. Our framework combines state-of-the-art techniques from natural language processing and computer vision to extract text and image features and infer relationships between entities in memes. We evaluate the proposed framework on a …
Video Sentiment Analysis For Child Safety,
2023
Singapore Management University
Video Sentiment Analysis For Child Safety, Yee Sen Tan, Nicole Anne Huiying Teo, Ezekiel En Zhe Ghe, Jolie Zhi Yi Fong, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
The proliferation of online video content underscores the critical need for effective sentiment analysis, particularly in safeguarding children from potentially harmful material. This research addresses this concern by presenting a multimodal analysis method for assessing video sentiment, categorizing it as either positive (child-friendly) or negative (potentially harmful). This method leverages three key components: text analysis, facial expression analysis, and audio analysis, including music mood analysis, resulting in a comprehensive sentiment assessment. Our evaluation results validate the effectiveness of this approach, making significant contributions to the field of video sentiment analysis and bolstering child safety measures. This research serves as a …
Depwignn: A Depth-Wise Graph Neural Network For Multi-Hop Spatial Reasoning In Text,
2023
Singapore Management University
Depwignn: A Depth-Wise Graph Neural Network For Multi-Hop Spatial Reasoning In Text, Shuaiyi Li, Yang Deng, Wai Lam
Research Collection School Of Computing and Information Systems
Spatial reasoning in text plays a crucial role in various real-world applications. Existing approaches for spatial reasoning typically infer spatial relations from pure text, which overlook the gap between natural language and symbolic structures. Graph neural networks (GNNs) have showcased exceptional proficiency in inducing and aggregating symbolic structures. However, classical GNNs face challenges in handling multi-hop spatial reasoning due to the over-smoothing issue, i.e., the performance decreases substantially as the number of graph layers increases. To cope with these challenges, we propose a novel Depth-Wise Graph Neural Network (DepWiGNN). Specifically, we design a novel node memory scheme and aggregate the …
Unifying Text, Tables, And Images For Multimodal Question Answering,
2023
Singapore Management University
Unifying Text, Tables, And Images For Multimodal Question Answering, Haohao Luo, Ying Shen, Yang Deng
Research Collection School Of Computing and Information Systems
Multimodal question answering (MMQA), which aims to derive the answer from multiple knowledge modalities (e.g., text, tables, and images), has received increasing attention due to its board applications. Current approaches to MMQA often rely on single-modal or bi-modal QA models, which limits their ability to effectively integrate information across all modalities and leverage the power of pre-trained language models. To address these limitations, we propose a novel framework called UniMMQA, which unifies three different input modalities into a text-to-text format by employing position-enhanced table linearization and diversified image captioning techniques. Additionally, we enhance cross-modal reasoning by incorporating a multimodal rationale …
Developing Detection And Mapping Of Roads Within Various Forms Of Media Using Opencv,
2023
University of Arkansas, Fayetteville
Developing Detection And Mapping Of Roads Within Various Forms Of Media Using Opencv, Jordan C. Lyle
Computer Science and Computer Engineering Undergraduate Honors Theses
OpenCV, and Computer Vision in general, has been a Computer Science topic that has interested me for a long time while completing my Bachelor’s degree at the University of Arkansas. As a result of this, I ended up choosing to utilize OpenCV in order to complete the task of detecting road-lines and mapping roads when given a wide variety of images. The purpose of my Honors research and this thesis is to detail the process of creating an algorithm to detect the road-lines such that the results are effective and instantaneous, as well as detail how Computer Vision can be …
The Propagation And Execution Of Malware In Images,
2023
Christopher Newport University
The Propagation And Execution Of Malware In Images, Piper Hall
Cybersecurity Undergraduate Research Showcase
Malware has become increasingly prolific and severe in its consequences as information systems mature and users become more reliant on computing in their daily lives. As cybercrime becomes more complex in its strategies, an often-overlooked manner of propagation is through images. In recent years, several high-profile vulnerabilities in image libraries have opened the door for threat actors to steal money and information from unsuspecting users. This paper will explore the mechanisms by which these exploits function and how they can be avoided.
Performative Mixing For Immersive Audio,
2023
Louisiana State University and Agricultural and Mechanical College
Performative Mixing For Immersive Audio, Brian A. Elizondo
LSU Doctoral Dissertations
Immersive multichannel audio can be produced with specialized setups of loudspeakers, often surrounding the audience. These setups can feature as few as four loudspeakers or more than 300. Performative mixing in these environments requires a bespoke solution offering intuitive gestural control. Beyond the usual faders for gain control, advancements in multichannel sound demand interfaces capable of quickly positioning sounds between channels. The Quad Cartesian Positioner is such a solution in the form of a Eurorack module for surround mixing for use in live or studio performances.
Diffusion/mixing methods for live multichannel immersive music often rely on the repurposing of hardware …
Vrmovian - An Immersive Data Annotation Tool For Visual Analysis Of Human Interactions In Vr,
2023
Chapman University
Vrmovian - An Immersive Data Annotation Tool For Visual Analysis Of Human Interactions In Vr, Isaac Browen
Student Scholar Symposium Abstracts and Posters
Understanding human behavior in virtual reality (VR) is a key component for developing intelligent systems to enhance human focused VR experiences. The ability to annotate human motion data proves to be a very useful way to analyze and understand human behavior. However, due to the complexity and multi-dimensionality of human activity data, it is necessary to develop software that can display the data in a comprehensible way and can support intuitive data annotation for developing machine learning models able recognize and assist human motions in VR (e.g., remote physical therapy). Although past research has been done to improve VR data …
User Feedback On Celebratory Technology Model For Reducing Stigma,
2023
Chapman University
User Feedback On Celebratory Technology Model For Reducing Stigma, Evelyn Lawrie, Daniel Dinh, Sav Avalos, Jack De Bruyn, Spencer Au, Christian Lopez, Ray Tan, Cyrus Fa'amafoe
Student Scholar Symposium Abstracts and Posters
Social stigma is a complex manifestation that affects humanity, particularly individuals with disabilities and other marginalized groups, including those with physical, cognitive, and emotional conditions. Society often judges these individuals' interactions with the world, and many technologies designed to assist those with disabilities attempt to change their daily interactions and behaviors. Nonetheless, when the emphasis is placed on validating disabled identities, there is a potential for it to be seen as "inspiration porn." This approach might inadvertently reduce inclusivity and do little to challenge negative stereotypes; it can also lead to the objectification of individuals with disabilities. Therefore, this project …
Bridging Domain Gaps For Cross-Spectrum And Long-Range Face Recognition Using Domain Adaptive Machine Learning,
2023
University of Nebraska-Lincoln
Bridging Domain Gaps For Cross-Spectrum And Long-Range Face Recognition Using Domain Adaptive Machine Learning, Cedric Armel Nimpa Fondje
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
Face recognition technology has witnessed significant advancements in recent decades, enabling its widespread adoption in various applications such as security, surveillance, and biometrics applications. However, one of the primary challenges faced by existing face recognition systems is their limited performance when presented with images from different modalities or domains( such as infrared to visible, long range to close range, nighttime to daytime, profile to f rontal, etc.) Additionally, advancements in camera sensors, analytics beyond the visible spectrum, and the increasing size of cross-modal datasets have led to a particular interest in cross-modal learning for face recognition in the biometrics and …
Enhancing Search Engine Results: A Comparative Study Of Graph And Timeline Visualizations For Semantic And Temporal Relationship Discovery,
2023
University of Denver
Enhancing Search Engine Results: A Comparative Study Of Graph And Timeline Visualizations For Semantic And Temporal Relationship Discovery, Muhammad Shahiq Qureshi
Electronic Theses and Dissertations
In today’s digital age, search engines have become indispensable tools for finding information among the corpus of billions of webpages. The standard that most search engines follow is to display search results in a list-based format arranged according to a ranking algorithm. Although this format is good for presenting the most relevant results to users, it fails to represent the underlying relations between different results. These relations, among others, can generally be of either a temporal or semantic nature. A user who wants to explore the results that are connected by those relations would have to make a manual effort …
Designing Secure Mental Healthcare Chatbots For Older Adults,
2023
University of Denver
Designing Secure Mental Healthcare Chatbots For Older Adults, Aishwarya Surani
Electronic Theses and Dissertations
The landscape of mental health support has evolved as a result of the rising demand for digital mental healthcare services. Users now have an opportunity to seek mental health support online due to the growth of digital platforms. For those looking for mental health treatments, chatbots have evolved as user-friendly, accessible platforms that provide remote access and convenience. However, for chatbots to be effective, users must divulge personal and sensitive information, such as demographics, insurance information, and a history of mental illness. While chatbots offer services to a variety of demographic users, older adults face unique challenges related to usability, …
Optimizing E-Payment Applications For Older Adults: User-Centered Solutions To Improve Security, Privacy, Usability, And Accessibility,
2023
University of Denver
Optimizing E-Payment Applications For Older Adults: User-Centered Solutions To Improve Security, Privacy, Usability, And Accessibility, Urvashi Kishnani
Electronic Theses and Dissertations
In an increasingly digital world, older adults are rapidly becoming a vital demographic in the realm of electronic financial transactions. It is imperative to address their unique needs and challenges to ensure their financial well-being. Older adults can be more vulnerable to various online threats, making security and privacy paramount. As they adapt to the digital age, understanding their specific privacy concerns and preferences is crucial for creating trustworthy e-payment systems. Moreover, enhancing the usability of e-payment applications for older adults promotes financial independence and inclusion, contributing to their overall quality of life. By focusing on these critical dimensions, we …
Turn-It-Up: Rendering Resistance For Knobs In Virtual Reality Through Undetectable Pseudo-Haptics,
2023
Universitat des Saarlandes
Turn-It-Up: Rendering Resistance For Knobs In Virtual Reality Through Undetectable Pseudo-Haptics, Martin Feick, Andre Zenner, Oscar Ariza, Anthony Tang, Cihan Biyikli, Antonio Kruger
Research Collection School Of Computing and Information Systems
Rendering haptic feedback for interactions with virtual objects is an essential part of effective virtual reality experiences. In this work, we explore providing haptic feedback for rotational manipulations, e.g., through knobs. We propose the use of a Pseudo-Haptic technique alongside a physical proxy knob to simulate various physical resistances. In a psychophysical experiment with 20 participants, we found that designers can introduce unnoticeable offsets between real and virtual rotations of the knob, and we report the corresponding detection thresholds. Based on these, we present the Pseudo-Haptic Resistance technique to convey physical resistance while applying only unnoticeable pseudo-haptic manipulation. Additionally, we …
Metaformer Baselines For Vision,
2023
Singapore Management University
Metaformer Baselines For Vision, Weihao Yu, Chenyang Si, Pan Zhou, Mi Luo, Yichen Zhou, Jiashi Feng, Shuicheng Yan, Xinchao Wang
Research Collection School Of Computing and Information Systems
Abstract—MetaFormer, the abstracted architecture of Transformer, has been found to play a significant role in achieving competitive performance. In this paper, we further explore the capacity of MetaFormer, again, by migrating our focus away from the token mixer design: we introduce several baseline models under MetaFormer using the most basic or common mixers, and demonstrate their gratifying performance. We summarize our observations as follows: (1) MetaFormer ensures solid lower bound of performance. By merely adopting identity mapping as the token mixer, the MetaFormer model, termed IdentityFormer, achieves >80% accuracy on ImageNet-1K. (2) MetaFormer works well with arbitrary token mixers. When …
Pro-Cap: Leveraging A Frozen Vision-Language Model For Hateful Meme Detection,
2023
Singapore Management University
Pro-Cap: Leveraging A Frozen Vision-Language Model For Hateful Meme Detection, Rui Cao, Ming Shan Hee, Adriel Kuek, Wen Haw Chong, Roy Ka-Wei Lee, Jing Jiang
Research Collection School Of Computing and Information Systems
Hateful meme detection is a challenging multimodal task that requires comprehension of both vision and language, as well as cross-modal interactions. Recent studies have tried to fine-tune pre-trained vision-language models (PVLMs) for this task. However, with increasing model sizes, it becomes important to leverage powerful PVLMs more efficiently, rather than simply fine-tuning them. Recently, researchers have attempted to convert meme images into textual captions and prompt language models for predictions. This approach has shown good performance but suffers from non-informative image captions. Considering the two factors mentioned above, we propose a probing-based captioning approach to leverage PVLMs in a zero-shot …
Revisiting Disentanglement And Fusion On Modality And Context In Conversational Multimodal Emotion Recognition,
2023
Singapore Management University
Revisiting Disentanglement And Fusion On Modality And Context In Conversational Multimodal Emotion Recognition, Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, Fei Li
Research Collection School Of Computing and Information Systems
It has been a hot research topic to enable machines to understand human emotions in multimodal contexts under dialogue scenarios, which is tasked with multimodal emotion analysis in conversation (MM-ERC). MM-ERC has received consistent attention in recent years, where a diverse range of methods has been proposed for securing better task performance. Most existing works treat MM-ERC as a standard multimodal classification problem and perform multimodal feature disentanglement and fusion for maximizing feature utility. Yet after revisiting the characteristic of MM-ERC, we argue that both the feature multimodality and conversational contextualization should be properly modeled simultaneously during the feature disentanglement …
