Zero-Shot Object Counting With Good Exemplars,
2024
Singapore Management University
Zero-Shot Object Counting With Good Exemplars, Huilin Zhu, Jingling Yuan, Zhengwei Yang, Yu Guo, Zheng Wang, Xian Zhong, Shengfeng He
Research Collection School Of Computing and Information Systems
Zero-shot object counting (ZOC) aims to enumerate objects in images using only the names of object classes during testing, without the need for manual annotations. However, a critical challenge in current ZOC methods lies in their inability to identify high-quality exemplars effectively. This deficiency hampers scalability across diverse classes and undermines the development of strong visual associations between the identified classes and image content. To this end, we propose the Visual Association-based Zero-shot Object Counting (VA-Count) framework. VACount consists of an Exemplar Enhancement Module (EEM) and a Noise Suppression Module (NSM) that synergistically refine the process of class exemplar identification …
Beat-It : Beat-Synchronized Multi-Condition 3d Dance Generation,
2024
Singapore Management University
Beat-It : Beat-Synchronized Multi-Condition 3d Dance Generation, Zikai Huang, Xuemiao Xu, Cheng Xu, Huaidong Zhang, Chenxi Zheng, Jing Qin, Shengfeng He
Research Collection School Of Computing and Information Systems
Dance, as an art form, fundamentally hinges on the precise synchronization with musical beats. However, achieving aesthetically pleasing dance sequences from music is challenging, with existing methods often falling short in controllability and beat alignment. To address these shortcomings, this paper introduces Beat-It, a novel framework for beat-specific, key pose-guided dance generation. Unlike prior approaches, Beat-It uniquely integrates explicit beat awareness and key pose guidance, effectively resolving two main issues: the misalignment of generated dance motions with musical beats, and the inability to map key poses to specific beats, critical for practical choreography. Our approach disentangles beat conditions from music …
Onerestore : A Universal Restoration Framework For Composite Degradation,
2024
Singapore Management University
Onerestore : A Universal Restoration Framework For Composite Degradation, Yu Guo, Yuan Gao, Yuxu Lu, Huilin Zhu, Ryan Wen Liu, Shengfeng He
Research Collection School Of Computing and Information Systems
In real-world scenarios, image impairments often manifest as composite degradations, presenting a complex interplay of elements such as low light, haze, rain, and snow. Despite this reality, existing restoration methods typically target isolated degradation types, thereby falling short in environments where multiple degrading factors coexist. To bridge this gap, our study proposes a versatile imaging model that consolidates four physical corruption paradigms to accurately represent complex, composite degradation scenarios. In this context, we propose OneRestore, a novel transformer-based framework designed for adaptive, controllable scene restoration. The proposed framework leverages a unique cross-attention mechanism, merging degraded scene descriptors with image features, …
Gradualreality : Enhancing Physical Object Interaction In Virtual Reality Via Interaction State-Aware Blending,
2024
Singapore Management University
Gradualreality : Enhancing Physical Object Interaction In Virtual Reality Via Interaction State-Aware Blending, Hyuna Seo, Juheon Yi, Rajesh Krishna Balan, Youngki Lee
Research Collection School Of Computing and Information Systems
We present GradualReality, a novel interface enabling a Cross Reality experience that includes gradual interaction with physical objects in a virtual environment and supports both presence and usability. Daily Cross Reality interaction is challenging as the user’s physical object interaction state is continuously changing over time, causing their attention to frequently shift between the virtual and physical worlds. As such, presence in the virtual environment and seamless usability for interacting with physical objects should be maintained at a high level. To address this issue, we present an Interaction State-Aware Blending approach that (i) balances immersion and interaction capability and (ii) …
Graph Continual Learning With Debiased Lossless Memory Replay,
2024
Singapore Management University
Graph Continual Learning With Debiased Lossless Memory Replay, Chaoxi Niu, Guansong Pang, Ling Chen
Research Collection School Of Computing and Information Systems
Real-life graph data often expands continually, rendering the learning of graph neural networks (GNNs) on static graph data impractical. Graph continual learning (GCL) tackles this problem by continually adapting GNNs to the expanded graph of the current task while maintaining the performance over the graph of previous tasks. Memory replay-based methods, which aim to replay data of previous tasks when learning new tasks, have been explored as one principled approach to mitigate the forgetting of the knowledge learned from the previous tasks. In this paper we extend this methodology with a novel framework, called Debiased Lossless Memory replay (DeLoMe). Unlike …
Pvp-Ssd: Point-Voxel Fusion With Partitioned Point Cloud Sampling For Anchor-Free Single-Stage Small 3d Object Detection,
2024
Singapore Management University
Pvp-Ssd: Point-Voxel Fusion With Partitioned Point Cloud Sampling For Anchor-Free Single-Stage Small 3d Object Detection, Xinlin Wu, Yibin Tian, Yin Pan, Zhiyuan Zhang, Xuesong Wu, Ruisheng Wang, Zhi Zeng
Research Collection School Of Computing and Information Systems
Single-stage object detection from 3D point clouds in autonomous driving faces significant challenges, particularly in accurately detecting small objects. To address this issue, we propose a novel method called Point-Voxel dual-branch feature extraction with Partitioned point cloud sampling for anchor-free Single-Stage Detection of 3D objects (PVP-SSD). The network comprises two branches: a point branch and a voxel branch. In the point branch, a partitioned point cloud sampling strategy leverages axial features to divide the point cloud. Then, it assigns different sampling weights to various segments to enhance the sampling accuracy. Additionally, a local feature enhancement module explicitly calculates the correlation …
Video Editing For Video Retrieval,
2024
Singapore Management University
Video Editing For Video Retrieval, Bin Zhu, Kevin Flanagan, Adriano Fragomeni, Michael Wray, Dima Damen
Research Collection School Of Computing and Information Systems
Though pre-training vision-language models have demonstrated significant benefits in boosting video-text retrieval performance from large-scale web videos, fine-tuning still plays a critical role with manually annotated clips with start and end times, which requires considerable human effort. To address this issue, we explore an alternative cheaper source of annotations, single timestamps, for video-text retrieval. We initialise clips from timestamps in a heuristic way to warm up a retrieval model. Then a video clip editing method is proposed to refine the initial rough boundaries to improve retrieval performance. A student-teacher network is introduced for video clip editing: the teacher model is …
Getting To The Point: Contrasting Directness And Warmth In Motivational Embodied Conversational Agents,
2024
Technological University Dublin
Getting To The Point: Contrasting Directness And Warmth In Motivational Embodied Conversational Agents, Michael O'Mahony, Cathy Ennis, Robert Ross
Conference papers
Enhancing long-term engagement with conversational agents remains a significant challenge. Controlling the perceived warmth or directness of an agent’s personality through the style of its generated text could be used to increase user likeability. This paper reports an investigation of a Wizard-of-Oz (WoZ) mediated study of two variants of a motivational embodied conversational agent to measure user perception of and attitudes towards warmth in interaction style. Results show a significant effect of users preferring an agent with a "more direct" personality for this scenario, though this effect is in many ways nuanced.
Imbalanced Graph Classification With Multi-Scale Oversampling Graph Neural Networks,
2024
Singapore Management University
Imbalanced Graph Classification With Multi-Scale Oversampling Graph Neural Networks, Rongrong Ma, Guansong Pang, Ling Chen
Research Collection School Of Computing and Information Systems
One main challenge in imbalanced graph classification is to learn expressive representations of the graphs in under-represented (minority) classes. Existing generic imbalanced learning methods, such as oversampling and imbalanced learning loss functions, can be adopted for enabling graph representation learning models to cope with this challenge. However, these methods often directly operate on the graph representations, ignoring rich discriminative information within the graphs and their interactions. To tackle this issue, we introduce a novel multi-scale oversampling graph neural network (MOSGNN) that learns expressive minority graph representations based on intra- and inter-graph semantics resulting from oversampled graphs at multiple scales - …
Efficient Cascaded Multiscale Adaptive Network For Image Restoration,
2024
Singapore Management University
Efficient Cascaded Multiscale Adaptive Network For Image Restoration, Yichen Zhou, Pan Zhou, Teck Khim Ng
Research Collection School Of Computing and Information Systems
Image restoration, encompassing tasks such as deblurring, denoising, and super-resolution, remains a pivotal area in computer vision. However, efficiently addressing the spatially varying artifacts of various low-quality images with local adaptiveness and handling their degradations at different scales poses significant challenges. To efficiently tackle these issues, we propose the novel Efficient Cascaded Multiscale Adaptive (ECMA) Network. ECMA employs Local Adaptive Module, LAM, which dynamically adjusts convolution kernels across local image regions to efficiently handle varying artifacts. Thus, LAM addresses the local adaptiveness challenge more efficiently than costlier mechanisms like self-attention, due to its less computationally intensive convolutions. To construct a …
Robust Image Classification System Via Cloud Computing, Aligned Multimodal Embeddings, Centroids And Neighbours,
2024
Singapore Management University
Robust Image Classification System Via Cloud Computing, Aligned Multimodal Embeddings, Centroids And Neighbours, Wei Lun Koh, Boon Yong Koh, Bing Tian Dai
Research Collection School Of Computing and Information Systems
We propose a framework for a cloud-based application of an image classification system that is highly accessible, maintains data confidentiality, and robust to incorrect training labels. The end-to-end system is implemented using Amazon Web Services (AWS), with a detailed guide provided for replication, enhancing the ways which researchers can collaborate with a community of users for mutual benefits. A front-end web application allows users across the world to securely log in, contribute labelled training images conveniently via a drag-and-drop approach, and use that same application to query an up-to-date model that has knowledge of images from the community of users. …
Text-Driven Video Prediction,
2024
Fudan University
Text-Driven Video Prediction, Xue Song, Jingjing Chen, Bin Zhu, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
Current video generation models usually convert signals indicating appearance and motion received from inputs (e.g., image and text) or latent spaces (e.g., noise vectors) into consecutive frames, fulfilling a stochastic generation process for the uncertainty introduced by latent code sampling. However, this generation pattern lacks deterministic constraints for both appearance and motion, leading to uncontrollable and undesirable outcomes. To this end, we propose a new task called Text-driven Video Prediction (TVP). Taking the first frame and text caption as inputs, this task aims to synthesize the following frames. Specifically, appearance and motion components are provided by the image and caption …
Granular3d: Delving Into Multi-Granularity 3d Scene Graph Prediction,
2024
Singapore Management University
Granular3d: Delving Into Multi-Granularity 3d Scene Graph Prediction, Kaixiang Huang, Jingru Yang, Jin Wang, Shengfeng He, Zhan Wang, Haiyan He, Qifeng Zhang, Guodong Lu
Research Collection School Of Computing and Information Systems
This paper addresses the significant challenges in 3D Semantic Scene Graph (3DSSG) prediction, essential for understanding complex 3D environments. Traditional approaches, primarily using PointNet and Graph Convolutional Networks, struggle with effectively extracting multi-grained features from intricate 3D scenes, largely due to a focus on global scene processing and single-scale feature extraction. To overcome these limitations, we introduce Granular3D, a novel approach that shifts the focus towards multi-granularity analysis by predicting relation triplets from specific sub-scenes. One key is the Adaptive Instance Enveloping Method (AIEM), which establishes an approximate envelope structure around irregular instances, providing shape-adaptive local point cloud sampling, thereby …
Adavis: Adaptive And Explainable Visualization Recommendation For Tabular Data,
2024
Singapore Management University
Adavis: Adaptive And Explainable Visualization Recommendation For Tabular Data, Songheng Zhang, Haotian Li, Huamin Qu, Yong Wang
Research Collection School Of Computing and Information Systems
Automated visualization recommendation facilitates the rapid creation of effective visualizations, which is especially beneficial for users with limited time and limited knowledge of data visualization. There is an increasing trend in leveraging machine learning (ML) techniques to achieve an end-to-end visualization recommendation. However, existing ML-based approaches implicitly assume that there is only one appropriate visualization for a specific dataset, which is often not true for real applications. Also, they often work like a black box, and are difficult for users to understand the reasons for recommending specific visualizations. To fill the research gap, we propose AdaVis, an adaptive and explainable …
Integrating Authentication Schemes In Augmented And Virtual Reality Classrooms,
2024
University of Denver
Integrating Authentication Schemes In Augmented And Virtual Reality Classrooms, Naheem Noah
Electronic Theses and Dissertations
Augmented Reality and Virtual Reality (AR/VR) technologies are revolutionizing educational experiences, but their widespread adoption hinges on addressing critical security and usability challenges, particularly in the domain of user authentication. This research presents an investigation into the security landscape of AR/VR and explores a graphical authentication scheme called “Things” that enhances both security and usability in immersive learning environments. Through a systematic evaluation of popular AR/VR devices and applications, potential vulnerabilities and limitations were identified, such as high usage of pin/passwords which are susceptible to shoulder-surfing attacks, lack of multi-factor authentication, and unclear data-sharing practices. A review of existing knowledge-based …
Changes In Near-Field Perception And Reaching Behavior In Virtual Environments Over Time,
2024
Clemson University
Changes In Near-Field Perception And Reaching Behavior In Virtual Environments Over Time, Kristopher C. Kohm
All Dissertations
Near-field perception and reaching capabilities are fundamental for most interactions in immersive virtual environments (IVEs). To perform actions in IVEs accurately and efficiently, virtual reality (VR) users need to be able to adapt to changes in their perception. Some of these perceptual differences may be inherent to virtual environments, such as the difference in depth perception between the virtual and non-virtual worlds. Others may be deliberate alterations to the user's action capabilities or to their surroundings to make interactions easier. Both the alterations and the user's ability to adjust to them may change over time as they gain experience in …
Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study,
2024
Clemson University
Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny
All Theses
High blood pressure, also known as hypertension, significantly increases the risk of heart disease and stroke, which are leading causes of death in the United States. While contributing to over 691,000 deaths in 2021 alone in the United States (U.S.), it also imposes immense economic burden on the healthcare system, costing approximately $131 billion annually. One way to address this issue is for increased self-care behaviors and medication adherence, both of which require sufficient health literacy. Despite the importance of health literacy, 90% of U.S. adults struggle with health-related subjects. Overcoming the issues associated with health literacy requires addressing the …
We Train Ai, Why Not Humans, Too? An Exploration Of Human-Ai Team Training For Future Workplace Viability,
2024
Clemson University
We Train Ai, Why Not Humans, Too? An Exploration Of Human-Ai Team Training For Future Workplace Viability, Caitlin M. Lancaster
All Dissertations
The integration of Artificial Intelligence (AI) in the workforce is transforming team dynamics, leading to the emergence of Human-AI Teams (HATs). These teams offer opportunities to capitalize on human strengths with AI's prowess, offering significant opportunities for innovation and efficiency. Effective HAT functioning requires aligning human expectations with AI capabilities and bridging knowledge gaps between teammates. Despite this potential, key integration challenges remain, such as developing shared mental models, addressing skill limitations, and overcoming negative AI perceptions. Existing training efforts often apply human-human teaming principles directly to HATs, overlooking AI's role as a teammate and limiting the development of HAT-specific …
Empathyear : An Open-Source Avatar Multimodal Empathetic Chatbot,
2024
Singapore Management University
Empathyear : An Open-Source Avatar Multimodal Empathetic Chatbot, Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, Erik Cambria
Research Collection School Of Computing and Information Systems
This paper introduces EmpathyEar, a pioneering open-source, avatar-based multimodal empathetic chatbot, to fill the gap in traditional text-only empathetic response generation (ERG) systems. Leveraging the advancements of a large language model, combined with multimodal encoders and generators, EmpathyEar supports user inputs in any combination of text, sound, and vision, and produces multimodal empathetic responses, offering users, not just textual responses but also digital avatars with talking faces and synchronized speeches. A series of emotion-aware instruction-tuning is performed for comprehensive emotional understanding and generation capabilities. In this way, EmpathyEar provides users with responses that achieve a deeper emotional resonance, closely emulating …
Unifying Global-Local Representations In Salient Object Detection With Transformers,
2024
Singapore Management University
Unifying Global-Local Representations In Salient Object Detection With Transformers, Sucheng Ren, Nanxuan Zhao, Qiang Wen, Guoqiang Han, Shengfeng He
Research Collection School Of Computing and Information Systems
The fully convolutional network (FCN) has dominated salient object detection for a long period. However, the locality of CNN requires the model deep enough to have a global receptive field and such a deep model always leads to the loss of local details. In this paper, we introduce a new attention-based encoder, vision transformer, into salient object detection to ensure the globalization of the representations from shallow to deep layers. With the global view in very shallow layers, the transformer encoder preserves more local representations to recover the spatial details in final saliency maps. Besides, as each layer can capture …
