Open Access. Powered by Scholars. Published by Universities.®

2024

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 61 - 90 of 199

Full-Text Articles in Graphics and Human Computer Interfaces

Empathyear : An Open-Source Avatar Multimodal Empathetic Chatbot, Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, Erik Cambria Aug 2024

Empathyear : An Open-Source Avatar Multimodal Empathetic Chatbot, Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, Erik Cambria

Research Collection School Of Computing and Information Systems

This paper introduces EmpathyEar, a pioneering open-source, avatar-based multimodal empathetic chatbot, to fill the gap in traditional text-only empathetic response generation (ERG) systems. Leveraging the advancements of a large language model, combined with multimodal encoders and generators, EmpathyEar supports user inputs in any combination of text, sound, and vision, and produces multimodal empathetic responses, offering users, not just textual responses but also digital avatars with talking faces and synchronized speeches. A series of emotion-aware instruction-tuning is performed for comprehensive emotional understanding and generation capabilities. In this way, EmpathyEar provides users with responses that achieve a deeper emotional resonance, closely emulating …


Human Centered Approaches And Taxonomies For Explainable Artificial Intelligence, Helen Sheridan, Emma Murphy, Dympna O'Sullivan Jul 2024

Human Centered Approaches And Taxonomies For Explainable Artificial Intelligence, Helen Sheridan, Emma Murphy, Dympna O'Sullivan

Conference papers

Recent interest within the research community related to explainable artificial intelligence (XAI) has led to a profuse amount of literature on the subject. Those who wish to tackle the domain from an HCI focus may be presented with overwhelming material, most of which does not pertain to human aspects of XAI. Taxonomies can serve to categorize a subject into topic areas and distill content into an overview of the field. This late breaking work intends to help those within the HCI community with a focus on XAI to understand relevant aspects of human centered XAI. We also present a taxonomy …


Creating And Delivering Audio Descriptions For Videos, Rosiana Natalie Jul 2024

Creating And Delivering Audio Descriptions For Videos, Rosiana Natalie

Dissertations and Theses Collection (Open Access)

Despite anti-discrimination regulations mandating the provision of audio descriptions (ADs), the majority of online video content remains inaccessible to blind and low-vision (BLV) individuals. This is because these ADs are either absent or fail to adequately address the diverse and unique needs of the audience. Traditionally, content creators have relied on professionals to author ADs. However, this gold standard may not be accessible for some content creators because this method is still costly and has a long turnaround time. Moreover, when ADs are available, they tend to be static and unalterable, failing to cater to the unique preferences of BLV …


My Ai Companion: An Examination Of The Removal Of Erotic Role Play From Replika Through User Discussion On Reddit, Chelsee M. Allen Jul 2024

My Ai Companion: An Examination Of The Removal Of Erotic Role Play From Replika Through User Discussion On Reddit, Chelsee M. Allen

Department of Sociology: Dissertations, Theses, and Student Research

The development of artificial intelligence (AI) software has expanded rapidly in recent years, and thus has emerged the importance of exploring human relationships with AI chatbots. Replika, an app which uses AI to mimic human conversation, removed a function called Erotic Role Play (ERP) that allowed for sexual conversation with users’ customizable chatbots in February of 2023. This exploratory qualitative study examines the aftermath of ERP’s removal through an analysis of user interactions on Reddit. Five overarching themes emerged through the analysis of top posts to a Replika-specific subreddit, encompassing topics around mental health, stigma, coping, sex work and gendered …


An Exploratory Study Of Conventional Machine Learning And Large Language Models For Sentiment Analysis, Cui Zou, Jingyuan Cai, Langtao Chen, Fiona Fui-Hoon Nah Jul 2024

An Exploratory Study Of Conventional Machine Learning And Large Language Models For Sentiment Analysis, Cui Zou, Jingyuan Cai, Langtao Chen, Fiona Fui-Hoon Nah

Research Collection School Of Computing and Information Systems

Sentiment analysis is the use of natural language processing to identify affective states and determine people’s opinions in various analytical applications such as customer reviews and social media analyses. Large language models (LLMs) such as GPT-4o demonstrate impressive performance in text generation tasks. Despite numerous studies in the extant literature, few have compared the performance of conventional machine learning models with LLMs for sentiment analysis. This study aims to fill this gap by conducting an evaluation of these models using a balanced dataset of 2,000 IMDb movie reviews. Our study shows that GPT-4o achieves the highest performance, while GPT-3.5 and …


A Computational Aesthetic Design Science Study On Online Video Based On Triple-Dimensional Multimodal Analysis, Zhangguang Kang, Fiona Fui-Hoon Nah, Keng Siau Jul 2024

A Computational Aesthetic Design Science Study On Online Video Based On Triple-Dimensional Multimodal Analysis, Zhangguang Kang, Fiona Fui-Hoon Nah, Keng Siau

Research Collection School Of Computing and Information Systems

Computational video aesthetic prediction refers to using models that automatically evaluate the features of videos to produce their aesthetic scores. Current video aesthetic prediction models are designed based on bimodal frameworks. To address their limitations, we developed the Triple-Dimensional Multimodal Temporal Video Aesthetic neural network (TMTVA-net) model. The Long Short-Term Memory (LSTM) forms the conceptual foundation for the design framework. In the multimodal transformer layer, we employed two distinct transformers: the multimodal transformer and the feature transformer, enabling the acquisition of modality-specific patterns and representational features uniquely adapted to each modality. The fusion layer has also been redesigned to compute …


Understanding And Fighting Scams: Media, Language, Appeals And Effects, Shuhua Zhou, Xiao Fan Liu, Fiona Fui-Hoon Nah, S. Harrison, X. Zhang, S. Zhen, D. Yeung, J. Hsiao, R. Lc, A. Chan, X. Wang, C. Jiang, F. Lin, J. Li, A. Wong, L. Chan, B. George, P. Li Jul 2024

Understanding And Fighting Scams: Media, Language, Appeals And Effects, Shuhua Zhou, Xiao Fan Liu, Fiona Fui-Hoon Nah, S. Harrison, X. Zhang, S. Zhen, D. Yeung, J. Hsiao, R. Lc, A. Chan, X. Wang, C. Jiang, F. Lin, J. Li, A. Wong, L. Chan, B. George, P. Li

Research Collection School Of Computing and Information Systems

Scams are fraudulent activities aiming to deceive individuals into relinquishing money, property, or rights, and they have proliferated in the context of widespread misinformation and disinformation. In this paper, we propose strategies and a research plan to address key questions about the exploitation of new communication technologies by scammers, the prevalence and nature of different scam types, and the language characteristics and appeals used in scamming content. We aim to develop a comprehensive taxonomy of scams and identify factors that contribute to their persuasiveness. Additionally, we propose the use of advanced technologies, including artificial intelligence, physiological measures, and brain mapping, …


Certified Robust Accuracy Of Neural Networks Are Bounded Due To Bayes Errors, Ruihan Zhang, Jun Sun Jul 2024

Certified Robust Accuracy Of Neural Networks Are Bounded Due To Bayes Errors, Ruihan Zhang, Jun Sun

Research Collection School Of Computing and Information Systems

Adversarial examples pose a security threat to many critical systems built on neural networks. While certified training improves robustness, it also decreases accuracy noticeably. Despite various proposals for addressing this issue, the significant accuracy drop remains. More importantly, it is not clear whether there is a certain fundamental limit on achieving robustness whilst maintaining accuracy. In this work, we offer a novel perspective based on Bayes errors. By adopting Bayes error to robustness analysis, we investigate the limit of certified robust accuracy, taking into account data distribution uncertainties. We first show that the accuracy inevitably decreases in the pursuit of …


Jigsaw: Edge-Based Streaming Perception Over Spatially Overlapped Multi-Camera Deployments, Ila Gokarn, Yigong Hu, Tarek Abdelzaher, Archan Misra Jul 2024

Jigsaw: Edge-Based Streaming Perception Over Spatially Overlapped Multi-Camera Deployments, Ila Gokarn, Yigong Hu, Tarek Abdelzaher, Archan Misra

Research Collection School Of Computing and Information Systems

We present JIGSAW, a novel system that performs edge-based streaming perception over multiple video streams, while additionally factoring in the redundancy offered by the spatial overlap often exhibited in urban, multi-camera deployments. To assure high streaming throughput, JIGSAW extracts and spatially multiplexes multiple regions-of-interest from different camera frames into a smaller canvas frame. Moreover, to ensure that perception stays abreast of evolving object kinematics, JIGSAW includes a utility-based weighted scheduler to preferentially prioritize and even skip object-specific tiles extracted from an incoming stream of camera frames. Using the CityflowV2 traffic surveillance dataset, we show that JIGSAW can simultaneously process 25 …


Learning Topological Representations With Bidirectional Graph Attention Network For Solving Job Shop Scheduling Problem, Cong Zhang, Zhiguang Cao, Yaoxin Wu, Wen Song, Jing Sun Jul 2024

Learning Topological Representations With Bidirectional Graph Attention Network For Solving Job Shop Scheduling Problem, Cong Zhang, Zhiguang Cao, Yaoxin Wu, Wen Song, Jing Sun

Research Collection School Of Computing and Information Systems

Existing learning-based methods for solving job shop scheduling problems (JSSP) usually use off-the-shelf GNN models tailored to undirected graphs and neglect the rich and meaningful topological structures of disjunctive graphs (DGs). This paper proposes the topology-aware bidirectional graph attention network (TBGAT), a novel GNN architecture based on the attention mechanism, to embed the DG for solving JSSP in a local search framework. Specifically, TBGAT embeds the DG from a forward and a backward view, respectively, where the messages are propagated by following the different topologies of the views and aggregated via graph attention. Then, we propose a novel operator based …


Evilscreen Attack: Smart Tv Hijacking Via Multi-Channel Remote Control Mimicry, Yiwei Zhang, Siqi Ma, Tiancheng Chen, Juanru Li, Robert H. Deng, Elisa Bertino Jul 2024

Evilscreen Attack: Smart Tv Hijacking Via Multi-Channel Remote Control Mimicry, Yiwei Zhang, Siqi Ma, Tiancheng Chen, Juanru Li, Robert H. Deng, Elisa Bertino

Research Collection School Of Computing and Information Systems

Modern smart TVs often communicate with their remote controls (including the smartphone simulated ones) using multiple wireless channels (e.g., Infrared, Bluetooth, and Wi-Fi). However, this multi-channel remote control communication introduces a new attack surface. An inherent security flaw is that remote controls of most smart TVs are designed to work in a benign environment rather than an adversarial one, and thus wireless communications between a smart TV and its remote controls are not strongly protected. Attackers can leverage such a flaw to abuse the remote control communication and compromise smart TV systems. In this paper, we propose EvilScreen, a novel …


Hierarchical Damage Correlations For Old Photo Restoration, Weiwei Cai, Xuemiao Xu, Jiajia Xu, Huaidong Zhang, Haoxin Yang, Kun Zhang, Shengfeng He Jul 2024

Hierarchical Damage Correlations For Old Photo Restoration, Weiwei Cai, Xuemiao Xu, Jiajia Xu, Huaidong Zhang, Haoxin Yang, Kun Zhang, Shengfeng He

Research Collection School Of Computing and Information Systems

Restoring old photographs can preserve cherished memories. Previous methods handled diverse damages within the same network structure, which proved impractical. In addition, these methods cannot exploit correlations among artifacts, especially in scratches versus patch-misses issues. Hence, a tailored network is particularly crucial. In light of this, we propose a unified framework consisting of two key components: ScratchNet and PatchNet. In detail, ScratchNet employs the parallel Multi-scale Partial Convolution Module to effectively repair scratches, learning from multi-scale local receptive fields. In contrast, the patch-misses necessitate the network to emphasize global information. To this end, we incorporate a transformer-based encoder and decoder …


How People Prompt Generative Ai To Create Interactive Vr Scenes, Setareh Aghel Manesh, Tianyi Zhang, Yuki Onishi, Kotaro Hara, Scott Bateman, Jiannan Li, Anthony Tang Jul 2024

How People Prompt Generative Ai To Create Interactive Vr Scenes, Setareh Aghel Manesh, Tianyi Zhang, Yuki Onishi, Kotaro Hara, Scott Bateman, Jiannan Li, Anthony Tang

Research Collection School Of Computing and Information Systems

Generative AI tools can provide people with the ability to create virtual environments and scenes with natural language prompts. Yet, how people will formulate such prompts is unclear---particularly when they inhabit the environment that they are designing. For instance, it is likely that a person might say, "Put a chair here,'' while pointing at a location. If such linguistic and embodied features are common to people's prompts, we need to tune models to accommodate them. In this work, we present a Wizard of Oz elicitation study with 22 participants, where we studied people's implicit expectations when verbally prompting such programming …


Creative Insights Into Motion: Enhancing Human Activity Understanding With 3d Data Visualization And Annotation, Isaac Browen, Hector M. Camarillo-Abad, Franceli L. Cibrian, Trudi Di Qi Jun 2024

Creative Insights Into Motion: Enhancing Human Activity Understanding With 3d Data Visualization And Annotation, Isaac Browen, Hector M. Camarillo-Abad, Franceli L. Cibrian, Trudi Di Qi

Engineering Faculty Articles and Research

This paper presents a novel 3D system for human motion analysis - Motion Data Visualization and Annotation (MoViAn). Designed to provide a comprehensive visual representation of 3D human motion data, MoViAn incorporates detailed visualization of gaze direction, hand movements, and object interactions, alongside an interactive interface for efficient data annotation. A user study involving eight participants indicates that MoViAn enables users to thoroughly explore and annotate human motion data, with System Usability Scale (SUS) results demonstrating a satisfactory usability level. The contribution of this paper lies in the development of an interactive and usable data analytics tool aimed at deepening …


Technology And Homelessness: How Website Design And Blockchain Technology Could Impact The Unhoused, Casey Pratt Jun 2024

Technology And Homelessness: How Website Design And Blockchain Technology Could Impact The Unhoused, Casey Pratt

Undergraduate Theses, Capstones, and Recitals

Although technology could be used to combat inequality, it is instead increasing it. This paper discusses how the unhoused population suffers at the hand of technological inequality despite being relatively offline. It presents theories on how this would change if we reapproached how technology is used to assist the unhoused. It suggests implementing blockchain as a resource as well as modifying the websites built to assist in accessing benefits. Employees at shelters are interviewed for this paper about their experiences with using digital resources to rehouse and restabilize the vulnerable. They are asked how the sites can be improved for …


Curating Familiarity Within The Unfamiliar: Exploring Non-Native Mobile App Experiences To Create Cross-Cultural Design Frameworks, Hanna Hong Jun 2024

Curating Familiarity Within The Unfamiliar: Exploring Non-Native Mobile App Experiences To Create Cross-Cultural Design Frameworks, Hanna Hong

Computer Science Senior Theses

Global mobility and markets are expanding, and as a result, countries are becoming less and less monocultural. With multiple cultural affinity groups to cater towards, companies often will deploy different versions of a website or app based on the country a user is accessing it from. This strategy of catering to geographic location results in a lack of accommodation for people living within a culture that is different from their native one. In order to increase accessibility and equal ease-of-use for all audiences, designers should understand and work towards the needs of a multicultural user base. This study investigates how …


Combinatorial Creativity: Knowledge Graphs And Idea Generation In Crowdsourcing Innovation, Zhi Wei Vincent Mack Jun 2024

Combinatorial Creativity: Knowledge Graphs And Idea Generation In Crowdsourcing Innovation, Zhi Wei Vincent Mack

Dissertations and Theses Collection (Open Access)

This dissertation explores the dynamic interplay between combinatorial creativity and technology-driven innovation within various knowledge-intensive fields. It critically examines the role of combinatorial creativity in generating groundbreaking innovations by amalgamating existing ideas and technologies. This research incorporates a detailed examination of how knowledge, whether tacit or explicit, can be transformed into actionable data to foster innovation in crowdsourcing contexts. Chapter 2 provides an overview of the relevant literature on how Artificial Intelligence and Knowledge Management Systems can support combinatorial creativity. The study further delves into the transformative impact of knowledge management systems, particularly focusing on crowdsourcing platforms that leverage collective …


Impact Of Similarities In Gender And Physical Appearance Between User And Embodied Conversational Agents On Trustworthiness, Empathy, And Service Evaluation, Sookyoung Park Jun 2024

Impact Of Similarities In Gender And Physical Appearance Between User And Embodied Conversational Agents On Trustworthiness, Empathy, And Service Evaluation, Sookyoung Park

Dartmouth College Master’s Theses

Embodied conversational agents (ECAs) have significantly enhanced human-machine interactions and show considerable potential in various industries such as customer service, education, healthcare, entertainment, and finance [1, 2]. This study explores the impact of similarities in gender and physical appearance between ECAs and users on the perceptions of trustworthiness, empathy, and service evaluation within the context of counselor ECAs. We conducted a within-subject experiment (n=50), using a 2x2 factorial arrangement, that varied the gender and the physical appearance of four distinct AI avatars. Participants interacted with each avatar, completing a post-experiment survey and participating in semi-structured interviews. Our findings indicate that …


Balancing Darkness And Visibility: An Algorithmic Approach To Light Placement In Low-Light, Ray-Traced Scenes, Briana Kuo Jun 2024

Balancing Darkness And Visibility: An Algorithmic Approach To Light Placement In Low-Light, Ray-Traced Scenes, Briana Kuo

Master's Theses

In recent years, digital media has seen incredible advancements in rendering visually stunning computer graphics scenes. Photo-realistic games, animated films, and more leave viewers blown away by the sheer beauty of their graphics. However, challenges arise when depicting dark scenes, often resulting in visual monotony and difficulty in comprehension due to insufficient detail within the scene. In order to enhance readability and visual interest of a scene, additional, artificial lights can be placed throughout a scene to enhance the aesthetic. These lights, however, must be strategically placed in order to retain an essence of darkness and maintain the delicate balance …


D3still : Decoupled Differential Distillation For Asymmetric Image Retrieval, Yi Xie, Yihong Lin, Wenjie Cai, Xuemiao Xu, Huaidong Zhang, Yong Du, Shengfeng He Jun 2024

D3still : Decoupled Differential Distillation For Asymmetric Image Retrieval, Yi Xie, Yihong Lin, Wenjie Cai, Xuemiao Xu, Huaidong Zhang, Yong Du, Shengfeng He

Research Collection School Of Computing and Information Systems

Existing methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However, these oneto-one constraint approaches often fail to maintain retrieval order consistency, especially when the query network has limited representational capacity. To overcome this problem, we introduce the Decoupled Differential Distillation (D3still) framework. This framework shifts from absolute one-to-one supervision to optimizing the relational differences in pairwise similarities produced by the query and gallery networks, thereby preserving a consistent retrieval order across both networks. Our method involves computing a pairwise similarity differential matrix within the gallery domain, which is …


Inceptionnext: When Inception Meets Convnext, Weihao Yu, Pan Zhou, Shuicheng Yan, Xinchao Wang Jun 2024

Inceptionnext: When Inception Meets Convnext, Weihao Yu, Pan Zhou, Shuicheng Yan, Xinchao Wang

Research Collection School Of Computing and Information Systems

Inspired by the long-range modeling ability of ViTs, large-kernel convolutions are widely studied and adopted recently to enlarge the receptive field and improve model performance, like the remarkable work ConvNeXt which employs 7×7 depthwise convolution. Although such depthwise operator only consumes a few FLOPs, it largely harms the model efficiency on powerful computing devices due to the high memory access costs. For example, ConvNeXtT has similar FLOPs with ResNet-50 but only achieves ∼ 60% throughputs when trained on A100 GPUs with full precision. Although reducing the kernel size of ConvNeXt can improve speed, it results in significant performance degradation, which …


Efficient Cross-Modal Video Retrieval With Meta-Optimized Frames, Ning Han, Xun Yang, Ee-Peng Lim, Hao Chen, Qianru Sun Jun 2024

Efficient Cross-Modal Video Retrieval With Meta-Optimized Frames, Ning Han, Xun Yang, Ee-Peng Lim, Hao Chen, Qianru Sun

Research Collection School Of Computing and Information Systems

Cross-modal video retrieval aims to retrieve semantically relevant videos when given a textual query, and is one of the fundamental multimedia tasks. Most top-performing methods primarily leverage Vision Transformer (ViT) to extract video features [1]-[3]. However, they suffer from the high computational complexity of ViT, especially when encoding long videos. A common and simple solution is to uniformly sample a small number (e.g., 4 or 8) of frames from the target video (instead of using the whole video) as ViT inputs. The number of frames has a strong influence on the performance of ViT, e.g., using 8 frames yields better …


Open-Vocabulary Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun, Jing Liu, Peng Wang, Yanning Zhang Jun 2024

Open-Vocabulary Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun, Jing Liu, Peng Wang, Yanning Zhang

Research Collection School Of Computing and Information Systems

Current video anomaly detection (VAD) approaches with weak supervisions are inherently limited to a closed-set setting and may struggle in open-world applications where there can be anomaly categories in the test data unseen during training. A few recent studies attempt to tackle a more realistic setting, open-set VAD, which aims to de-tect unseen anomalies given seen anomalies and normal videos. However, such a setting focuses on predicting frame anomaly scores, having no ability to recognize the specific categories of anomalies, despite the fact that this ability is essential for building more informed video surveillance systems. This paper takes a step …


Beyond Textual Constraints : Learning Novel Diffusion Conditions With Fewer Examples, Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He Jun 2024

Beyond Textual Constraints : Learning Novel Diffusion Conditions With Fewer Examples, Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He

Research Collection School Of Computing and Information Systems

In this paper, we delve into a novel aspect of learning novel diffusion conditions with datasets an order of magnitude smaller. The rationale behind our approach is the elimination of textual constraints during the few-shot learning process. To that end, we implement two optimization strategies. The first, prompt-free conditional learning, utilizes a prompt-free encoder derived from a pre-trained Stable Diffusion model. This strategy is designed to adapt new conditions to the diffusion process by minimizing the textual-visual cor-relation, thereby ensuring a more precise alignment between the generated content and the specified conditions. The second strategy entails condition-specific negative rectification, which …


Learning With Unreliability : Fast Few-Shot Voxel Radiance Fields With Relative Geometric Consistency, Yingjie Xu, Bangzhen Liu, Hao Tang, Bailin Deng, Shengfeng He Jun 2024

Learning With Unreliability : Fast Few-Shot Voxel Radiance Fields With Relative Geometric Consistency, Yingjie Xu, Bangzhen Liu, Hao Tang, Bailin Deng, Shengfeng He

Research Collection School Of Computing and Information Systems

We propose a voxel-based optimization framework, Re VoRF, for few-shot radiance fields that strategically ad-dress the unreliability in pseudo novel view synthesis. Our method pivots on the insight that relative depth relationships within neighboring regions are more reliable than the ab-solute color values in disoccluded areas. Consequently, we devise a bilateral geometric consistency loss that carefully navigates the trade-off between color fidelity and geometric accuracy in the context of depth consistency for uncertain regions. Moreover, we present a reliability-guided learning strategy to discern and utilize the variable quality across syn-thesized views, complemented by a reliability-aware voxel smoothing algorithm that smoothens …


Violet: Visual Analytics For Explainable Quantum Neural Networks, Shaolun Ruan, Zhiding Liang, Qiang Guan, Paul Robert Griffin, Xiaolin Wen, Yanna Lin, Yong Wang Jun 2024

Violet: Visual Analytics For Explainable Quantum Neural Networks, Shaolun Ruan, Zhiding Liang, Qiang Guan, Paul Robert Griffin, Xiaolin Wen, Yanna Lin, Yong Wang

Research Collection School Of Computing and Information Systems

With the rapid development of Quantum Machine Learning, quantum neural networks (QNN) have experienced great advancement in the past few years, harnessing the advantages of quantum computing to significantly speed up classical machine learning tasks. Despite their increasing popularity, the quantum neural network is quite counter-intuitive and difficult to understand, due to their unique quantum-specific layers (e.g., data encoding and measurement) in their architecture. It prevents QNN users and researchers from effectively understanding its inner workings and exploring the model training status. To fill the research gap, we propose VIOLET , a novel visual analytics approach to improve the explainability …


Diffusion Time-Step Curriculum For One Image To 3d Generation, Xuanyu Yi, Zike Wu, Qingshan Xu, Pan Zhou, Joo Hwee Lim, Hanwang Zhang Jun 2024

Diffusion Time-Step Curriculum For One Image To 3d Generation, Xuanyu Yi, Zike Wu, Qingshan Xu, Pan Zhou, Joo Hwee Lim, Hanwang Zhang

Research Collection School Of Computing and Information Systems

Score distillation sampling (SDS) has been widely adopted to overcome the absence of unseen views in reconstructing 3D objects from a single image. It leverages pretrained 2D diffusion models as teacher to guide the reconstruction of student 3D models. Despite their remarkable success, SDS-based methods often encounter geometric artifacts and texture saturation. We find out the crux is the overlooked indiscriminate treatment of diffusion time-steps during optimization: it unreasonably treats the studentteacher knowledge distillation to be equal at all time-steps and thus entangles coarse-grained and fine-grained modeling. Therefore, we propose the Diffusion Time-step Curriculum one-image-to-3D pipeline (DTC123), which involves both …


Jollygesture: Exploring Dual-Purpose Gestures In Vr Presentations, Gun Woo Warren Park, Anthony Tang, Fanny Chevalier Jun 2024

Jollygesture: Exploring Dual-Purpose Gestures In Vr Presentations, Gun Woo Warren Park, Anthony Tang, Fanny Chevalier

Research Collection School Of Computing and Information Systems

Virtual reality (VR) offers new opportunities for presenters to use expressive body language to engage their audience. Yet, most VR presentation systems have adopted control mechanisms that mimic those found in face-to-face presentation systems. We explore the use of gestures that have dual-purpose: first, for the audience, a communicative purpose; second, for the presenter, a control purpose to alter content in slides. To support presenters, we provide guidance on what gestures are available and their effects. We realize our design approach in JollyGesture, a VR technology probe that recognizes dual-purpose gestures in a presentation scenario. We evaluate our approach through …


Posmlp-Video: Spatial And Temporal Relative Position Encoding For Efficient Video Recognition, Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Xiangnan He, Meng Wang Jun 2024

Posmlp-Video: Spatial And Temporal Relative Position Encoding For Efficient Video Recognition, Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Xiangnan He, Meng Wang

Research Collection School Of Computing and Information Systems

In recent years, vision Transformers and MLPs have demonstrated remarkable performance in image understanding tasks. However, their inherently dense computational operators, such as self-attention and token-mixing layers, pose significant challenges when applied to spatio-temporal video data. To address this gap, we propose PosMLP-Video, a lightweight yet powerful MLP-like backbone for video recognition. Instead of dense operators, we use efficient relative positional encoding (RPE) to build pairwise token relations, leveraging small-sized parameterized relative position biases to obtain each relation score. Specifically, to enable spatio-temporal modeling, we extend the image PosMLP’s positional gating unit to temporal, spatial, and spatio-temporal variants, namely PoTGU, …


Crime Prediction Using Agent-Based Modeling, Yifei Gong Jun 2024

Crime Prediction Using Agent-Based Modeling, Yifei Gong

Dissertations, Theses, and Capstone Projects

Crime risk evaluation and crime prediction using agent-based modeling (ABM) have gained popularity in the field of computational criminology in recent years. Traditionally, researchers rely on statistical methods and machine learning models to predict crimes using historical data. ABM generates macro-level crime patterns in a bottom-up fashion by simulating the daily behaviors of autonomous entities, such as citizens and offenders. ABM takes into consideration the non-linear interactions between agents under complex social contexts. Currently, the comprehensive usage of ABM for criminological theory testing and urban policy evaluations calls for a unified software framework. In this research, we introduce CARESim, an …