Image De‑Photobombing Benchmark,
2024
University of Dayton
Image De‑Photobombing Benchmark, Vatsa S. Patel, Kunal Agrawal, Samah Baraheem, Amira Yousif, Tam Nguyen
Computer Science Faculty Publications
Removing photobombing elements from images is a challenging task that requires sophisticated image inpainting techniques. Despite the availability of various methods, their effectiveness depends on the complexity of the image and the nature of the distracting element. To address this issue, we conducted a benchmark study to evaluate 10 state-of-the-art photobombing removal methods on a dataset of over 300 images. Our study focused on identifying the most effective image inpainting techniques for removing unwanted regions from images. We annotated the photobombed regions that require removal and evaluated the performance of each method using peak signal-to-noise ratio (PSNR), structural similarity index …
Improving Implicit Communication In Remote Collaboration Through Augmented Reality And Digital Twins,
2024
Louisiana State University and Agricultural and Mechanical College
Improving Implicit Communication In Remote Collaboration Through Augmented Reality And Digital Twins, Nicholas Levergne
LSU Master's Theses
Large scale digital twinning projects are beginning to emerge across the tech industry. Within these projects is a desire to integrate augmented reality capabilities into industrial workflows. However, research on augmented reality technology for remote collaboration lacks ecologically valid studies of real world scenarios. Additionally, prior remote collaboration literature is focused on white-collar applications instead of blue-collar field work. Prior AR collaboration software is similarly limited, with most software allowing mixed camera views and annotation that requires participants to be stationary. This thesis introduces SpectAR, an augmented reality and desktop remote collaboration software suite developed in Unreal Engine 5.1.1. SpectAR …
Visualizing Routes With Ai-Discovered Street-View Patterns,
2024
Kent State University
Visualizing Routes With Ai-Discovered Street-View Patterns, Tsung Heng Wu, Md Amiruzzaman, Ye Zhao, Deepshikha Bhati, Jing Yang
Computer Science Faculty Publications
Street-level visual appearances play an important role in studying social systems, such as understanding the built environment, driving routes, and associated social and economic factors. It has not been integrated into a typical geographical visualization interface (e.g., map services) for planning driving routes. In this article, we study this new visualization task with several new contributions. First, we experiment with a set of AI techniques and propose a solution of using semantic latent vectors for quantifying visual appearance features. Second, we calculate image similarities among a large set of street-view images and then discover spatial imagery patterns. Third, we integrate …
Coca: Improving And Explaining Graph Neural Network-Based Vulnerability Detection Systems,
2024
Singapore Management University
Coca: Improving And Explaining Graph Neural Network-Based Vulnerability Detection Systems, Sicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo, Lili Bo, Bin Li, Wei Liu
Research Collection School Of Computing and Information Systems
Recently, Graph Neural Network (GNN)-based vulnerability detection systems have achieved remarkable success. However, the lack of explainability poses a critical challenge to deploy black-box models in security-related domains. For this reason, several approaches have been proposed to explain the decision logic of the detection model by providing a set of crucial statements positively contributing to its predictions. Unfortunately, due to the weakly-robust detection models and suboptimal explanation strategy, they have the danger of revealing spurious correlations and redundancy issue.In this paper, we propose Coca, a general framework aiming to 1) enhance the robustness of existing GNN-based vulnerability detection models to …
Lyraquist: Language Learning Via Music App,
2024
University of South Carolina - Columbia
Lyraquist: Language Learning Via Music App, Vivian D'Souza, Siri Avula, Mahi Patel, Tanvi Singh, Ashley Bickham
Senior Theses
Lyraquist is a new language learning mobile app that encourages the practice of a foreign language through music. Language learners can connect their Spotify Premium account to Lyraquist to listen to music in their target languages and utilize tools such as translation and vocabulary lists to facilitate language practice. Through integration with Spotify, users can import and create Spotify playlists and search the service’s entire catalog. By combining daily listening habits with several tasks associated with language learning in one place, Lyraquist hopes to be a useful language learning tool.
Elevating Academic Administration: A Comprehensive Faculty Dashboard For Tracking Student Evaluations And Research,
2024
University of South Carolina
Elevating Academic Administration: A Comprehensive Faculty Dashboard For Tracking Student Evaluations And Research, Musa M. Azeem
Senior Theses
The USC Faculty Dashboard is a web application designed to revolutionize how department heads, professors, and instructors monitor progress and make decisions, providing a centralized hub for efficient data storage and analysis. Currently, there’s a gap in tools tailored for department heads to concisely manage the performance of their department, which our platform aims to fill. The USC Faculty Dashboard offers easy access to upload and view student evaluation and research information, empowering department heads to evaluate the performance of faculty members and seamlessly track their research grants, publications, and expenditures. Furthermore, professors and instructors gain personalized performance analysis tools, …
Hop‑Based Heterogeneous Graph Transformer,
2024
Singapore Management University
Hop‑Based Heterogeneous Graph Transformer, Zixuan Yang, Xiao Wang, Yanhua Yu, Yuling Wang, Kangkang Lu, Zirui Guo, Xiting Qin, Yunshan Ma, Tat‑Seng Chua
Research Collection School Of Computing and Information Systems
The Graph Transformer (GT) has shown significant ability in processing graph-structured data, addressing limitations in graph neural networks, such as over-smoothing and over-squashing. However, the implementation of GT in real-world heterogeneous graphs (HGs) with complex topology continues to present numerous challenges. Firstly, a challenge arises in designing a tokenizer that is compatible with heterogeneity. Secondly, the complexity of the transformer hampers the acquisition of high-order neighbor information in HGs. In this paper, we propose a novel Hop-basedHeterogeneous Graph Transformer (H2Gormer) framework, paving a promising path for HGs to benefit from the capabilities of Transformers. We propose a Heterogeneous Hop-based Token …
Exploring Neural Networks For Developing A Chess Learning Platform With Integrated Ai Agent,
2024
St. Mary's University
Exploring Neural Networks For Developing A Chess Learning Platform With Integrated Ai Agent, Lauren Escobedo
Posters - 2024
Chess is a highly strategic, complex, and long-form game that has been popular for many centuries. Due to the aforementioned complexities of this game, new players often have a hard time learning how to effectively and successfully play. With the recent developments in machine learning algorithms, new opportunities arise to create artificially intelligent tutors - not only for chess, but for all subjects. This project aims to develop a product which investigates the integration of an artificially intelligent “coach”, named Chesster, to train the player, which is trained on a neural network machine learning algorithm.
Nookipedia,
2024
St. Mary's University
Nookipedia, Shyann Francis
Posters - 2024
Animal Crossing involves managing around 9000 items within a gameplay environment. This abundance of items makes it difficult for players to effectively track progress and navigate through their possessions. The primary motivation of Nookipedia is to enhance overall gameplay experience by streamlining achievement management, thereby reducing unnecessary interactions with non-playable characters (NPCs) and enabling players to focus more on core game activities. By improving inventory organization and reducing clutter, the aim is to create a smoother and more immersive gaming experience for players.
Terry Riley's "In C" For Mobile Ensemble,
2024
Loyola University Chicago
Terry Riley's "In C" For Mobile Ensemble, David B. Wetzel, Griffin Moe, George K. Thiruvathukal
Computer Science: Faculty Publications and Other Works
This workshop presents a mobile-friendly Web Audio application for a “technology ensemble play-along” of Terry Riley’s 1964 composition In C. Attendees will join in a reading of In C using available web-enabled devices as musical instruments. We hope to demonstrate an accessible music-technology experience that relies on face-to-face interaction within a shared space. In this all-electronic implementation, no special musical or technical expertise is required.
Towards Understanding Convergence And Generalization Of Adamw,
2024
Singapore Management University
Towards Understanding Convergence And Generalization Of Adamw, Pan Zhou, Xingyu Xie, Zhouchen Lin, Shuicheng Yan
Research Collection School Of Computing and Information Systems
AdamW modifies Adam by adding a decoupled weight decay to decay network weights per training iteration. For adaptive algorithms, this decoupled weight decay does not affect specific optimization steps, and differs from the widely used ℓ2-regularizer which changes optimization steps via changing the first- and second-order gradient moments. Despite its great practical success, for AdamW, its convergence behavior and generalization improvement over Adam and ℓ2-regularized Adam (ℓ2-Adam) remain absent yet. To solve this issue, we prove the convergence of AdamW and justify its generalization advantages over Adam and ℓ2-Adam. Specifically, AdamW provably converges but minimizes a dynamically regularized loss that …
Test-Time Augmentation For 3d Point Cloud Classification And Segmentation,
2024
Singapore Management University
Test-Time Augmentation For 3d Point Cloud Classification And Segmentation, Tuan-Anh Vu, Srinjay Sarkar, Zhiyuan Zhang, Binh-Son Hua, Sai-Kit Yeung
Research Collection School Of Computing and Information Systems
Data augmentation is a powerful technique to enhance the performance of a deep learning task but has received less attention in 3D deep learning. It is well known that when 3D shapes are sparsely represented with low point density, the performance of the downstream tasks drops significantly. This work explores test-time augmentation (TTA) for 3D point clouds. We are inspired by the recent revolution of learning implicit representation and point cloud upsampling, which can produce high-quality 3D surface reconstruction and proximity-to-surface, respectively. Our idea is to leverage the implicit field reconstruction or point cloud upsampling techniques as a systematic way …
Iterative Graph Self-Distillation,
2024
Singapore Management University
Iterative Graph Self-Distillation, Hanlin Zhang, Shuai Lin, Weiyang Liu, Pan Zhou, Jian Tang, Xiaodan Liang, Eric Xing
Research Collection School Of Computing and Information Systems
Recently, there has been increasing interest in the challenge of how to discriminatively vectorize graphs. To address this, we propose a method called Iterative Graph Self-Distillation (IGSD) which learns graph-level representation in an unsupervised manner through instance discrimination using a self-supervised contrastive learning approach. IGSD involves a teacher-student distillation process that uses graph diffusion augmentations and constructs the teacher model using an exponential moving average of the student model. The intuition behind IGSD is to predict the teacher network representation of the graph pairs under different augmented views. As a natural extension, we also apply IGSD to semi-supervised scenarios by …
Transiam: Aggregating Multi-Modal Visual Features With Locality For Medical Image Segmentation,
2024
Central South University
Transiam: Aggregating Multi-Modal Visual Features With Locality For Medical Image Segmentation, Xuejian Li, Shiqiang Ma, Junhai Xu, Jijun Tang, Shengfeng He, Fei Guo
Research Collection School Of Computing and Information Systems
Automatic segmentation of medical images plays an important role in the diagnosis of diseases. On single-modal data, convolutional neural networks have demonstrated satisfactory performance. However, multi-modal data encompasses a greater amount of information rather than single-modal data. Multi-modal data can be effectively used to improve the segmentation accuracy of regions of interest by analyzing both spatial and temporal information. In this study, we propose a dual-path segmentation model for multi-modal medical images, named TranSiam. Taking into account that there is a significant diversity between the different modalities, TranSiam employs two parallel CNNs to extract the features which are specific to …
Hgprompt: Bridging Homogeneous And Heterogeneous Graphs For Few-Shot Prompt Learning,
2024
Singapore Management University
Hgprompt: Bridging Homogeneous And Heterogeneous Graphs For Few-Shot Prompt Learning, Xingtong Yu, Yuan Fang, Zemin Liu, Xinming Zhang
Research Collection School Of Computing and Information Systems
Graph neural networks (GNNs) and heterogeneous graph neural networks (HGNNs) are prominent techniques for homogeneous and heterogeneous graph representation learning, yet their performance in an end-to-end supervised framework greatly depends on the availability of task-specific supervision. To reduce the labeling cost, pre-training on selfsupervised pretext tasks has become a popular paradigm, but there is often a gap between the pre-trained model and downstream tasks, stemming from the divergence in their objectives. To bridge the gap, prompt learning has risen as a promising direction especially in few-shot settings, without the need to fully fine-tune the pre-trained model. While there has been …
Hackles: Simulating And Visually Representing The Anxiety Of Walking Alone,
2024
Singapore Management University
Hackles: Simulating And Visually Representing The Anxiety Of Walking Alone, Sydney Pratte, Anthony Tang, Shannon Hoover, Lora Oehlberg
Research Collection School Of Computing and Information Systems
In this work, we compare the designs of two fashion-tech garments that communicate the anxiety felt when walking alone. While the two garments share a common vision, they are designed to be worn in two radically different settings and to communicate to different audiences: one directly communicates an empathetic experience to its wearer; the other a model wears at a runway show and must share its story to a general audience. We used Research Through Design (RtD) methods to design both fashion-tech garments. Then, we recorded and analyzed the design process for both garments via an annotated portfolio to compare …
Foodmask: Real-Time Food Instance Counting, Segmentation And Recognition,
2024
Singapore Management University
Foodmask: Real-Time Food Instance Counting, Segmentation And Recognition, Huu-Thanh Nguyen, Yu Cao, Chong-Wah Ngo, Wing-Kwong Chan
Research Collection School Of Computing and Information Systems
Food computing has long been studied and deployed to several applications. Understanding a food image at the instance level, including recognition, counting and segmentation, is essential to quantifying nutrition and calorie consumption. Nevertheless, existing techniques are limited to either category-specific instance detection, which does not reflect precisely the instance size at the pixel level, or category-agnostic instance segmentation, which is insufficient for dish recognition. This paper presents a compact and fast multi-task network, namely FoodMask, for clustering-based food instance counting, segmentation and recognition. The network learns a semantic space simultaneously encoding food category distribution and instance height at pixel basis. …
Leveraging Llms And Generative Models For Interactive Known-Item Video Search,
2024
Singapore Management University
Leveraging Llms And Generative Models For Interactive Known-Item Video Search, Zhixin Ma, Jiaxin Wu, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
While embedding techniques such as CLIP have considerably boosted search performance, user strategies in interactive video search still largely operate on a trial-and-error basis. Users are often required to manually adjust their queries and carefully inspect the search results, which greatly rely on the users’ capability and proficiency. Recent advancements in large language models (LLMs) and generative models offer promising avenues for enhancing interactivity in video retrieval and reducing the personal bias in query interpretation, particularly in the known-item search. Specifically, LLMs can expand and diversify the semantics of the queries while avoiding grammar mistakes or the language barrier. In …
What Does One Billion Dollars Look Like?: Visualizing Extreme Wealth,
2024
CUNY Graduate Center
What Does One Billion Dollars Look Like?: Visualizing Extreme Wealth, William Mahoney Luckman
Dissertations, Theses, and Capstone Projects
The word “billion” is a mathematical abstraction related to “big,” but it is difficult to understand the vast difference in value between one million and one billion; even harder to understand the vast difference in purchasing power between one billion dollars, and the average U.S. yearly income. Perhaps most difficult to conceive of is what that purchasing power and huge mass of capital translates to in terms of power. This project blends design, text, facts, and figures into an interactive narrative website that helps the user better understand their position in relation to extreme wealth: https://whatdoesonebilliondollarslooklike.website/
The site incorporates …
Catnet: Cross-Modal Fusion For Audio-Visual Speech Recognition,
2024
Singapore Management University
Catnet: Cross-Modal Fusion For Audio-Visual Speech Recognition, Xingmei Wang, Jianchen Mi, Boquan Li, Yixu Zhao, Jiaxiang Meng
Research Collection School Of Computing and Information Systems
Automatic speech recognition (ASR) is a typical pattern recognition technology that converts human speeches into texts. With the aid of advanced deep learning models, the performance of speech recognition is significantly improved. Especially, the emerging Audio–Visual Speech Recognition (AVSR) methods achieve satisfactory performance by combining audio-modal and visual-modal information. However, various complex environments, especially noises, limit the effectiveness of existing methods. In response to the noisy problem, in this paper, we propose a novel cross-modal audio–visual speech recognition model, named CATNet. First, we devise a cross-modal bidirectional fusion model to analyze the close relationship between audio and visual modalities. Second, …
