Open Access. Powered by Scholars. Published by Universities.®
Graphics and Human Computer Interfaces Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (473)
- Artificial Intelligence and Robotics (400)
- Engineering (335)
- Software Engineering (307)
- Social and Behavioral Sciences (264)
-
- Other Computer Sciences (241)
- Computer Engineering (178)
- Arts and Humanities (142)
- Theory and Algorithms (133)
- Education (126)
- OS and Networks (97)
- Medicine and Health Sciences (93)
- Numerical Analysis and Scientific Computing (93)
- Systems Architecture (81)
- Art and Design (79)
- Business (78)
- Programming Languages and Compilers (78)
- Psychology (77)
- Communication (74)
- Electrical and Computer Engineering (73)
- Data Storage Systems (71)
- Information Security (69)
- Life Sciences (58)
- Data Science (51)
- Educational Technology (48)
- Library and Information Science (40)
- Communication Technology and New Media (39)
- Institution
-
- Singapore Management University (938)
- University of Dayton (114)
- Air Force Institute of Technology (98)
- Old Dominion University (97)
- California Polytechnic State University, San Luis Obispo (96)
-
- University of Arkansas, Fayetteville (89)
- University of Nebraska - Lincoln (51)
- City University of New York (CUNY) (48)
- Technological University Dublin (48)
- University of Malaya (42)
- San Jose State University (37)
- Dartmouth College (34)
- Embry-Riddle Aeronautical University (24)
- Clemson University (23)
- Purdue University (23)
- Rochester Institute of Technology (23)
- The University of Akron (22)
- Chapman University (20)
- Edith Cowan University (20)
- University of Kentucky (18)
- Michigan Technological University (16)
- University of Central Florida (15)
- Southern Adventist University (13)
- California State University, San Bernardino (12)
- Kennesaw State University (12)
- St. Mary's University (12)
- Nova Southeastern University (11)
- University of Minnesota Morris Digital Well (11)
- Louisiana State University (10)
- University of Nevada, Las Vegas (10)
- Keyword
-
- Virtual reality (62)
- Visualization (46)
- Computer graphics (38)
- Computer vision (37)
- Accessibility (36)
-
- Human-computer interaction (35)
- Augmented reality (33)
- Usability (31)
- Machine learning (29)
- Computer Science (25)
- Data visualization (25)
- Machine Learning (25)
- Artificial intelligence (24)
- Deep learning (24)
- Virtual Reality (23)
- HCI (22)
- Computer science (20)
- Eye tracking (20)
- Human computer interaction (20)
- User experience (20)
- Design (19)
- Deep Learning (16)
- Education (16)
- Feature extraction (15)
- Graph Neural Networks (15)
- Graphics (15)
- VR (15)
- Gamification (14)
- Image processing (14)
- Applied sciences (13)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (912)
- Computer Science Faculty Publications (135)
- Theses and Dissertations (98)
- Master's Theses (50)
- Graduate Theses and Dissertations (43)
-
- Student Works (2000-2009) (33)
- Computer Science and Computer Engineering Undergraduate Honors Theses (31)
- 3-D Printed Model Structural Files (29)
- Publications and Research (28)
- Dartmouth College Master’s Theses (24)
- Williams Honors College, Honors Research Projects (22)
- Master's Projects (20)
- Conference papers (19)
- Frameless (19)
- Computer Science and Software Engineering (18)
- All Dissertations (17)
- Dissertations and Theses Collection (Open Access) (16)
- Dissertations, Master's Theses and Master's Reports (16)
- H-Workload 2017: Models and Applications (Works in Progress) (15)
- Computer Engineering (14)
- Electronic Theses and Dissertations (14)
- Theses : Honours (14)
- Honors Theses (13)
- MAICS: The Modern Artificial Intelligence and Cognitive Science Conference (12)
- AFIT Patents (11)
- CCAC Theses and Dissertations (11)
- Engineering Faculty Articles and Research (11)
- Scholarly Horizons: University of Minnesota, Morris Undergraduate Journal (10)
- Inquiry: The University of Arkansas Undergraduate Research Journal (9)
- Publications (9)
- Publication Type
- File Type
Articles 361 - 390 of 2362
Full-Text Articles in Graphics and Human Computer Interfaces
Evilscreen Attack: Smart Tv Hijacking Via Multi-Channel Remote Control Mimicry, Yiwei Zhang, Siqi Ma, Tiancheng Chen, Juanru Li, Robert H. Deng, Elisa Bertino
Evilscreen Attack: Smart Tv Hijacking Via Multi-Channel Remote Control Mimicry, Yiwei Zhang, Siqi Ma, Tiancheng Chen, Juanru Li, Robert H. Deng, Elisa Bertino
Research Collection School Of Computing and Information Systems
Modern smart TVs often communicate with their remote controls (including the smartphone simulated ones) using multiple wireless channels (e.g., Infrared, Bluetooth, and Wi-Fi). However, this multi-channel remote control communication introduces a new attack surface. An inherent security flaw is that remote controls of most smart TVs are designed to work in a benign environment rather than an adversarial one, and thus wireless communications between a smart TV and its remote controls are not strongly protected. Attackers can leverage such a flaw to abuse the remote control communication and compromise smart TV systems. In this paper, we propose EvilScreen, a novel …
Hierarchical Damage Correlations For Old Photo Restoration, Weiwei Cai, Xuemiao Xu, Jiajia Xu, Huaidong Zhang, Haoxin Yang, Kun Zhang, Shengfeng He
Hierarchical Damage Correlations For Old Photo Restoration, Weiwei Cai, Xuemiao Xu, Jiajia Xu, Huaidong Zhang, Haoxin Yang, Kun Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
Restoring old photographs can preserve cherished memories. Previous methods handled diverse damages within the same network structure, which proved impractical. In addition, these methods cannot exploit correlations among artifacts, especially in scratches versus patch-misses issues. Hence, a tailored network is particularly crucial. In light of this, we propose a unified framework consisting of two key components: ScratchNet and PatchNet. In detail, ScratchNet employs the parallel Multi-scale Partial Convolution Module to effectively repair scratches, learning from multi-scale local receptive fields. In contrast, the patch-misses necessitate the network to emphasize global information. To this end, we incorporate a transformer-based encoder and decoder …
How People Prompt Generative Ai To Create Interactive Vr Scenes, Setareh Aghel Manesh, Tianyi Zhang, Yuki Onishi, Kotaro Hara, Scott Bateman, Jiannan Li, Anthony Tang
How People Prompt Generative Ai To Create Interactive Vr Scenes, Setareh Aghel Manesh, Tianyi Zhang, Yuki Onishi, Kotaro Hara, Scott Bateman, Jiannan Li, Anthony Tang
Research Collection School Of Computing and Information Systems
Generative AI tools can provide people with the ability to create virtual environments and scenes with natural language prompts. Yet, how people will formulate such prompts is unclear---particularly when they inhabit the environment that they are designing. For instance, it is likely that a person might say, "Put a chair here,'' while pointing at a location. If such linguistic and embodied features are common to people's prompts, we need to tune models to accommodate them. In this work, we present a Wizard of Oz elicitation study with 22 participants, where we studied people's implicit expectations when verbally prompting such programming …
Creative Insights Into Motion: Enhancing Human Activity Understanding With 3d Data Visualization And Annotation, Isaac Browen, Hector M. Camarillo-Abad, Franceli L. Cibrian, Trudi Di Qi
Creative Insights Into Motion: Enhancing Human Activity Understanding With 3d Data Visualization And Annotation, Isaac Browen, Hector M. Camarillo-Abad, Franceli L. Cibrian, Trudi Di Qi
Engineering Faculty Articles and Research
This paper presents a novel 3D system for human motion analysis - Motion Data Visualization and Annotation (MoViAn). Designed to provide a comprehensive visual representation of 3D human motion data, MoViAn incorporates detailed visualization of gaze direction, hand movements, and object interactions, alongside an interactive interface for efficient data annotation. A user study involving eight participants indicates that MoViAn enables users to thoroughly explore and annotate human motion data, with System Usability Scale (SUS) results demonstrating a satisfactory usability level. The contribution of this paper lies in the development of an interactive and usable data analytics tool aimed at deepening …
Technology And Homelessness: How Website Design And Blockchain Technology Could Impact The Unhoused, Casey Pratt
Technology And Homelessness: How Website Design And Blockchain Technology Could Impact The Unhoused, Casey Pratt
Undergraduate Theses, Capstones, and Recitals
Although technology could be used to combat inequality, it is instead increasing it. This paper discusses how the unhoused population suffers at the hand of technological inequality despite being relatively offline. It presents theories on how this would change if we reapproached how technology is used to assist the unhoused. It suggests implementing blockchain as a resource as well as modifying the websites built to assist in accessing benefits. Employees at shelters are interviewed for this paper about their experiences with using digital resources to rehouse and restabilize the vulnerable. They are asked how the sites can be improved for …
Curating Familiarity Within The Unfamiliar: Exploring Non-Native Mobile App Experiences To Create Cross-Cultural Design Frameworks, Hanna Hong
Computer Science Senior Theses
Global mobility and markets are expanding, and as a result, countries are becoming less and less monocultural. With multiple cultural affinity groups to cater towards, companies often will deploy different versions of a website or app based on the country a user is accessing it from. This strategy of catering to geographic location results in a lack of accommodation for people living within a culture that is different from their native one. In order to increase accessibility and equal ease-of-use for all audiences, designers should understand and work towards the needs of a multicultural user base. This study investigates how …
Combinatorial Creativity: Knowledge Graphs And Idea Generation In Crowdsourcing Innovation, Zhi Wei Vincent Mack
Combinatorial Creativity: Knowledge Graphs And Idea Generation In Crowdsourcing Innovation, Zhi Wei Vincent Mack
Dissertations and Theses Collection (Open Access)
This dissertation explores the dynamic interplay between combinatorial creativity and technology-driven innovation within various knowledge-intensive fields. It critically examines the role of combinatorial creativity in generating groundbreaking innovations by amalgamating existing ideas and technologies. This research incorporates a detailed examination of how knowledge, whether tacit or explicit, can be transformed into actionable data to foster innovation in crowdsourcing contexts. Chapter 2 provides an overview of the relevant literature on how Artificial Intelligence and Knowledge Management Systems can support combinatorial creativity. The study further delves into the transformative impact of knowledge management systems, particularly focusing on crowdsourcing platforms that leverage collective …
Impact Of Similarities In Gender And Physical Appearance Between User And Embodied Conversational Agents On Trustworthiness, Empathy, And Service Evaluation, Sookyoung Park
Dartmouth College Master’s Theses
Embodied conversational agents (ECAs) have significantly enhanced human-machine interactions and show considerable potential in various industries such as customer service, education, healthcare, entertainment, and finance [1, 2]. This study explores the impact of similarities in gender and physical appearance between ECAs and users on the perceptions of trustworthiness, empathy, and service evaluation within the context of counselor ECAs. We conducted a within-subject experiment (n=50), using a 2x2 factorial arrangement, that varied the gender and the physical appearance of four distinct AI avatars. Participants interacted with each avatar, completing a post-experiment survey and participating in semi-structured interviews. Our findings indicate that …
D3still : Decoupled Differential Distillation For Asymmetric Image Retrieval, Yi Xie, Yihong Lin, Wenjie Cai, Xuemiao Xu, Huaidong Zhang, Yong Du, Shengfeng He
D3still : Decoupled Differential Distillation For Asymmetric Image Retrieval, Yi Xie, Yihong Lin, Wenjie Cai, Xuemiao Xu, Huaidong Zhang, Yong Du, Shengfeng He
Research Collection School Of Computing and Information Systems
Existing methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However, these oneto-one constraint approaches often fail to maintain retrieval order consistency, especially when the query network has limited representational capacity. To overcome this problem, we introduce the Decoupled Differential Distillation (D3still) framework. This framework shifts from absolute one-to-one supervision to optimizing the relational differences in pairwise similarities produced by the query and gallery networks, thereby preserving a consistent retrieval order across both networks. Our method involves computing a pairwise similarity differential matrix within the gallery domain, which is …
Inceptionnext: When Inception Meets Convnext, Weihao Yu, Pan Zhou, Shuicheng Yan, Xinchao Wang
Inceptionnext: When Inception Meets Convnext, Weihao Yu, Pan Zhou, Shuicheng Yan, Xinchao Wang
Research Collection School Of Computing and Information Systems
Inspired by the long-range modeling ability of ViTs, large-kernel convolutions are widely studied and adopted recently to enlarge the receptive field and improve model performance, like the remarkable work ConvNeXt which employs 7×7 depthwise convolution. Although such depthwise operator only consumes a few FLOPs, it largely harms the model efficiency on powerful computing devices due to the high memory access costs. For example, ConvNeXtT has similar FLOPs with ResNet-50 but only achieves ∼ 60% throughputs when trained on A100 GPUs with full precision. Although reducing the kernel size of ConvNeXt can improve speed, it results in significant performance degradation, which …
Efficient Cross-Modal Video Retrieval With Meta-Optimized Frames, Ning Han, Xun Yang, Ee-Peng Lim, Hao Chen, Qianru Sun
Efficient Cross-Modal Video Retrieval With Meta-Optimized Frames, Ning Han, Xun Yang, Ee-Peng Lim, Hao Chen, Qianru Sun
Research Collection School Of Computing and Information Systems
Cross-modal video retrieval aims to retrieve semantically relevant videos when given a textual query, and is one of the fundamental multimedia tasks. Most top-performing methods primarily leverage Vision Transformer (ViT) to extract video features [1]-[3]. However, they suffer from the high computational complexity of ViT, especially when encoding long videos. A common and simple solution is to uniformly sample a small number (e.g., 4 or 8) of frames from the target video (instead of using the whole video) as ViT inputs. The number of frames has a strong influence on the performance of ViT, e.g., using 8 frames yields better …
Open-Vocabulary Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun, Jing Liu, Peng Wang, Yanning Zhang
Open-Vocabulary Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun, Jing Liu, Peng Wang, Yanning Zhang
Research Collection School Of Computing and Information Systems
Current video anomaly detection (VAD) approaches with weak supervisions are inherently limited to a closed-set setting and may struggle in open-world applications where there can be anomaly categories in the test data unseen during training. A few recent studies attempt to tackle a more realistic setting, open-set VAD, which aims to de-tect unseen anomalies given seen anomalies and normal videos. However, such a setting focuses on predicting frame anomaly scores, having no ability to recognize the specific categories of anomalies, despite the fact that this ability is essential for building more informed video surveillance systems. This paper takes a step …
Beyond Textual Constraints : Learning Novel Diffusion Conditions With Fewer Examples, Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Beyond Textual Constraints : Learning Novel Diffusion Conditions With Fewer Examples, Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
In this paper, we delve into a novel aspect of learning novel diffusion conditions with datasets an order of magnitude smaller. The rationale behind our approach is the elimination of textual constraints during the few-shot learning process. To that end, we implement two optimization strategies. The first, prompt-free conditional learning, utilizes a prompt-free encoder derived from a pre-trained Stable Diffusion model. This strategy is designed to adapt new conditions to the diffusion process by minimizing the textual-visual cor-relation, thereby ensuring a more precise alignment between the generated content and the specified conditions. The second strategy entails condition-specific negative rectification, which …
Learning With Unreliability : Fast Few-Shot Voxel Radiance Fields With Relative Geometric Consistency, Yingjie Xu, Bangzhen Liu, Hao Tang, Bailin Deng, Shengfeng He
Learning With Unreliability : Fast Few-Shot Voxel Radiance Fields With Relative Geometric Consistency, Yingjie Xu, Bangzhen Liu, Hao Tang, Bailin Deng, Shengfeng He
Research Collection School Of Computing and Information Systems
We propose a voxel-based optimization framework, Re VoRF, for few-shot radiance fields that strategically ad-dress the unreliability in pseudo novel view synthesis. Our method pivots on the insight that relative depth relationships within neighboring regions are more reliable than the ab-solute color values in disoccluded areas. Consequently, we devise a bilateral geometric consistency loss that carefully navigates the trade-off between color fidelity and geometric accuracy in the context of depth consistency for uncertain regions. Moreover, we present a reliability-guided learning strategy to discern and utilize the variable quality across syn-thesized views, complemented by a reliability-aware voxel smoothing algorithm that smoothens …
Violet: Visual Analytics For Explainable Quantum Neural Networks, Shaolun Ruan, Zhiding Liang, Qiang Guan, Paul Robert Griffin, Xiaolin Wen, Yanna Lin, Yong Wang
Violet: Visual Analytics For Explainable Quantum Neural Networks, Shaolun Ruan, Zhiding Liang, Qiang Guan, Paul Robert Griffin, Xiaolin Wen, Yanna Lin, Yong Wang
Research Collection School Of Computing and Information Systems
With the rapid development of Quantum Machine Learning, quantum neural networks (QNN) have experienced great advancement in the past few years, harnessing the advantages of quantum computing to significantly speed up classical machine learning tasks. Despite their increasing popularity, the quantum neural network is quite counter-intuitive and difficult to understand, due to their unique quantum-specific layers (e.g., data encoding and measurement) in their architecture. It prevents QNN users and researchers from effectively understanding its inner workings and exploring the model training status. To fill the research gap, we propose VIOLET , a novel visual analytics approach to improve the explainability …
Diffusion Time-Step Curriculum For One Image To 3d Generation, Xuanyu Yi, Zike Wu, Qingshan Xu, Pan Zhou, Joo Hwee Lim, Hanwang Zhang
Diffusion Time-Step Curriculum For One Image To 3d Generation, Xuanyu Yi, Zike Wu, Qingshan Xu, Pan Zhou, Joo Hwee Lim, Hanwang Zhang
Research Collection School Of Computing and Information Systems
Score distillation sampling (SDS) has been widely adopted to overcome the absence of unseen views in reconstructing 3D objects from a single image. It leverages pretrained 2D diffusion models as teacher to guide the reconstruction of student 3D models. Despite their remarkable success, SDS-based methods often encounter geometric artifacts and texture saturation. We find out the crux is the overlooked indiscriminate treatment of diffusion time-steps during optimization: it unreasonably treats the studentteacher knowledge distillation to be equal at all time-steps and thus entangles coarse-grained and fine-grained modeling. Therefore, we propose the Diffusion Time-step Curriculum one-image-to-3D pipeline (DTC123), which involves both …
Jollygesture: Exploring Dual-Purpose Gestures In Vr Presentations, Gun Woo Warren Park, Anthony Tang, Fanny Chevalier
Jollygesture: Exploring Dual-Purpose Gestures In Vr Presentations, Gun Woo Warren Park, Anthony Tang, Fanny Chevalier
Research Collection School Of Computing and Information Systems
Virtual reality (VR) offers new opportunities for presenters to use expressive body language to engage their audience. Yet, most VR presentation systems have adopted control mechanisms that mimic those found in face-to-face presentation systems. We explore the use of gestures that have dual-purpose: first, for the audience, a communicative purpose; second, for the presenter, a control purpose to alter content in slides. To support presenters, we provide guidance on what gestures are available and their effects. We realize our design approach in JollyGesture, a VR technology probe that recognizes dual-purpose gestures in a presentation scenario. We evaluate our approach through …
Posmlp-Video: Spatial And Temporal Relative Position Encoding For Efficient Video Recognition, Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Xiangnan He, Meng Wang
Posmlp-Video: Spatial And Temporal Relative Position Encoding For Efficient Video Recognition, Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Xiangnan He, Meng Wang
Research Collection School Of Computing and Information Systems
In recent years, vision Transformers and MLPs have demonstrated remarkable performance in image understanding tasks. However, their inherently dense computational operators, such as self-attention and token-mixing layers, pose significant challenges when applied to spatio-temporal video data. To address this gap, we propose PosMLP-Video, a lightweight yet powerful MLP-like backbone for video recognition. Instead of dense operators, we use efficient relative positional encoding (RPE) to build pairwise token relations, leveraging small-sized parameterized relative position biases to obtain each relation score. Specifically, to enable spatio-temporal modeling, we extend the image PosMLP’s positional gating unit to temporal, spatial, and spatio-temporal variants, namely PoTGU, …
Crime Prediction Using Agent-Based Modeling, Yifei Gong
Crime Prediction Using Agent-Based Modeling, Yifei Gong
Dissertations, Theses, and Capstone Projects
Crime risk evaluation and crime prediction using agent-based modeling (ABM) have gained popularity in the field of computational criminology in recent years. Traditionally, researchers rely on statistical methods and machine learning models to predict crimes using historical data. ABM generates macro-level crime patterns in a bottom-up fashion by simulating the daily behaviors of autonomous entities, such as citizens and offenders. ABM takes into consideration the non-linear interactions between agents under complex social contexts. Currently, the comprehensive usage of ABM for criminological theory testing and urban policy evaluations calls for a unified software framework. In this research, we introduce CARESim, an …
Community Discovery Over Attributed Graphs, Yudong Niu
Community Discovery Over Attributed Graphs, Yudong Niu
Dissertations and Theses Collection (Open Access)
Community discovery, as a fundamental problem in graph mining, finds applications in various domains such as biological analysis, system optimization and fraud detection. Although many efforts have been made to address community discovery based on graph topology, few works have been devoted to community discovery over attributed graphs, where graphs are equipped with attribute information such as node and edge types. Thus, this thesis is devoted to designing innovative solutions that can utilize the attribute information together with graph topology for community discovery. In particular, we study novel problems with efficient algorithms for both homogeneous and heterogeneous attributed graphs and …
Improving Interpretable Embeddings For Ad-Hoc Video Search With Generative Captions And Multi-Word Concept Bank, Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan
Improving Interpretable Embeddings For Ad-Hoc Video Search With Generative Captions And Multi-Word Concept Bank, Jiaxin Wu, Chong-Wah Ngo, Wing-Kwong Chan
Research Collection School Of Computing and Information Systems
Aligning a user query and video clips in cross-modal latent space and that with semantic concepts are two mainstream approaches for ad-hoc video search (AVS). However, the effectiveness of existing approaches is bottlenecked by the small sizes of available video-text datasets and the low quality of concept banks, which results in the failures of unseen queries and the out-of-vocabulary problem. This paper addresses these two problems by constructing a new dataset and developing a multi-word concept bank. Specifically, capitalizing on a generative model, we construct a new dataset consisting of 7 million generated text and video pairs for pre-training. To …
Consistent3d: Towards Consistent High-Fidelity Text-To-3d Generation With Deterministic Sampling Prior, Zike Wu, Pan Zhou, Xuanyu Yi, Xiaoding Yuan, Hanwang Zhang
Consistent3d: Towards Consistent High-Fidelity Text-To-3d Generation With Deterministic Sampling Prior, Zike Wu, Pan Zhou, Xuanyu Yi, Xiaoding Yuan, Hanwang Zhang
Research Collection School Of Computing and Information Systems
Score distillation sampling (SDS) and its variants have greatly boosted the development of text-to-3D generation, but are vulnerable to geometry collapse and poor textures yet. To solve this issue, we first deeply analyze the SDS and find that its distillation sampling process indeed corresponds to the trajectory sampling of a stochastic differential equation (SDE): SDS samples along an SDE trajectory to yield a less noisy sample which then serves as a guidance to optimize a 3D model. However, the randomness in SDE sampling often leads to a diverse and unpredictable sample which is not always less noisy, and thus is …
Let’S Think Outside The Box: Exploring Leap-Of-Thought In Large Language Models With Multimodal Humor Generation, Shanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen, Liang Lin, Marinka Zitnik, Pan Zhou
Let’S Think Outside The Box: Exploring Leap-Of-Thought In Large Language Models With Multimodal Humor Generation, Shanshan Zhong, Zhongzhan Huang, Shanghua Gao, Wushao Wen, Liang Lin, Marinka Zitnik, Pan Zhou
Research Collection School Of Computing and Information Systems
Chain-of-Thought (CoT) [2, 3] guides large language models (LLMs) to reason step-by-step, and can motivate their logical reasoning ability. While effective for logical tasks, CoT is not conducive to creative problem-solving which often requires out-of-box thoughts and is crucial for innovation advancements. In this paper, we explore the Leap-of-Thought (LoT) abilities within LLMs — a nonsequential, creative paradigm involving strong associations and knowledge leaps. To this end, we study LLMs on the popular Oogiri game which needs participants to have good creativity and strong associative thinking for responding unexpectedly and humorously to the given image, text, or both, and thus …
Few-Shot Learner Parameterization By Diffusion Time-Steps, Zhongqi Yue, Pan Zhou, Richang Hong, Hanwang Zhang, Sun Qianru
Few-Shot Learner Parameterization By Diffusion Time-Steps, Zhongqi Yue, Pan Zhou, Richang Hong, Hanwang Zhang, Sun Qianru
Research Collection School Of Computing and Information Systems
Even when using large multi-modal foundation models, few-shot learning is still challenging—if there is no proper inductive bias, it is nearly impossible to keep the nuanced class attributes while removing the visually prominent attributes that spuriously correlate with class labels. To this end, we find an inductive bias that the time-steps of a Diffusion Model (DM) can isolate the nuanced class attributes, i.e., as the forward diffusion adds noise to an image at each time-step, nuanced attributes are usually lost at an earlier time-step than the spurious attributes that are visually prominent. Building on this, we propose Time-step Few-shot (TiF) …
Generalized Graph Prompt: Toward A Unification Of Pre-Training And Downstream Tasks On Graphs, Xingtong Yu, Zhenghao Liu, Yuan Fang, Et Al.
Generalized Graph Prompt: Toward A Unification Of Pre-Training And Downstream Tasks On Graphs, Xingtong Yu, Zhenghao Liu, Yuan Fang, Et Al.
Research Collection School Of Computing and Information Systems
Graphs can model complex relationships between objects, enabling a myriad of Web applications such as online page/article classification and social recommendation. While graph neural networks (GNNs) have emerged as a powerful tool for graph representation learning, in an end-to-end supervised setting, their performance heavily relies on a large amount of task-specific supervision. To reduce labeling requirement, the 'pre-train, fine-tune' and 'pre-train, prompt' paradigms have become increasingly common. In particular, prompting is a popular alternative to fine-tuning in natural language processing, which is designed to narrow the gap between pre-training and downstream objectives in a task-specific manner. However, existing study of …
Drag Your Noise: Interactive Point-Based Editing Via Diffusion Semantic Propagation, Haofeng Liu, Chenshu Xu, Yifei Yang, Lihua Zeng, Shengfeng He
Drag Your Noise: Interactive Point-Based Editing Via Diffusion Semantic Propagation, Haofeng Liu, Chenshu Xu, Yifei Yang, Lihua Zeng, Shengfeng He
Research Collection School Of Computing and Information Systems
Point-based interactive editing serves as an essential tool to complement the controllability of existing generative models. A concurrent work, DragDiffusion, updates the diffusion latent map in response to user inputs, causing global latent map alterations. This results in imprecise preservation of the original content and unsuccessful editing due to gradient vanishing. In contrast, we present DragNoise, offering robust and accelerated editing without retracing the latent map. The core rationale of DragNoise lies in utilizing the predicted noise output of each U-Net as a semantic editor. This approach is grounded in two critical observations: firstly, the bottleneck features of U-Net inherently …
Rethinking Multi-View Representation Learning Via Distilled Disentangling, Guanzhou Ke, Bo Wang, Xiaoli Wang, Shengfeng He
Rethinking Multi-View Representation Learning Via Distilled Disentangling, Guanzhou Ke, Bo Wang, Xiaoli Wang, Shengfeng He
Research Collection School Of Computing and Information Systems
Multi-view representation learning aims to derive robust representations that are both view-consistent and view-specific from diverse data sources. This paper presents an in-depth analysis of existing approaches in this domain, highlighting a commonly overlooked aspect: the redundancy between view-consistent and view-specific representations. To this end, we propose an innovative framework for multi-view representation learning, which incorporates a technique we term 'distilled disentangling'. Our method introduces the concept of masked cross-view prediction, enabling the extraction of compact, high-quality view-consistent representations from various sources without incurring extra computational overhead. Additionally, we develop a distilled disentangling module that efficiently filters out consistency-related information …
More Human-Likeness, Less Self-Disclosure? Avatars' Form Realism And Job Applicants' Self-Disclosure In Ai Interviews, Yamin Xu, Keng Siau, Fiona Fui-Hoon Nah
More Human-Likeness, Less Self-Disclosure? Avatars' Form Realism And Job Applicants' Self-Disclosure In Ai Interviews, Yamin Xu, Keng Siau, Fiona Fui-Hoon Nah
Research Collection School Of Computing and Information Systems
The rise of AI in recruitment promises to revolutionize how organizations evaluate job candidates. The quality of AI evaluations is determined by the input data, which depends on job applicants' self-disclosure. However, little is known about how the design elements of AI interview systems, particularly avatar interviewers, influence job applicants' self-disclosure during these interactions. This study aims to address this gap by specifically focusing on how the form realism of avatar interviewers affects job applicants' self-disclosure through their perceptions. In addition, the study will examine the effects of job type as a moderator. Drawing on the Stimulus-Organism-Response (S-O-R) model, this …
Balancing Darkness And Visibility: An Algorithmic Approach To Light Placement In Low-Light, Ray-Traced Scenes, Briana Kuo
Master's Theses
In recent years, digital media has seen incredible advancements in rendering visually stunning computer graphics scenes. Photo-realistic games, animated films, and more leave viewers blown away by the sheer beauty of their graphics. However, challenges arise when depicting dark scenes, often resulting in visual monotony and difficulty in comprehension due to insufficient detail within the scene. In order to enhance readability and visual interest of a scene, additional, artificial lights can be placed throughout a scene to enhance the aesthetic. These lights, however, must be strategically placed in order to retain an essence of darkness and maintain the delicate balance …
Embodied Visions: Interactive Installations That Reimagine Bodily Presence In Digital Imaging Apparatuses As Shadows, Yunzi Shi
Dartmouth College Master’s Theses
Contextualized within a history of technological development, the evolution of imaging devices and technologies is accompanied by the abstraction of spatial relationships between the body of the observer, the apparatus, and physical reality, which leads to disembodying experiences for the observing subject. Compared with devices and interactive experiences, critical reflection on the epistemological impact of digital imaging devices has less priority in computational imaging and human-computer interaction research. Taking an artistic approach, this thesis describes Embodied Visions, an exhibition featuring three interactive installations exploring the technical infrastructure for imaging and reflecting on the (dis)embodied experiences in the digital age. …