Open Access. Powered by Scholars. Published by Universities.®
Graphics and Human Computer Interfaces Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (473)
- Artificial Intelligence and Robotics (400)
- Engineering (335)
- Software Engineering (307)
- Social and Behavioral Sciences (264)
-
- Other Computer Sciences (241)
- Computer Engineering (178)
- Arts and Humanities (142)
- Theory and Algorithms (133)
- Education (126)
- OS and Networks (97)
- Medicine and Health Sciences (93)
- Numerical Analysis and Scientific Computing (93)
- Systems Architecture (81)
- Art and Design (79)
- Business (78)
- Programming Languages and Compilers (78)
- Psychology (77)
- Communication (74)
- Electrical and Computer Engineering (73)
- Data Storage Systems (71)
- Information Security (69)
- Life Sciences (58)
- Data Science (51)
- Educational Technology (48)
- Library and Information Science (40)
- Communication Technology and New Media (39)
- Institution
-
- Singapore Management University (938)
- University of Dayton (114)
- Air Force Institute of Technology (98)
- Old Dominion University (97)
- California Polytechnic State University, San Luis Obispo (96)
-
- University of Arkansas, Fayetteville (89)
- University of Nebraska - Lincoln (51)
- City University of New York (CUNY) (48)
- Technological University Dublin (48)
- University of Malaya (42)
- San Jose State University (37)
- Dartmouth College (34)
- Embry-Riddle Aeronautical University (24)
- Clemson University (23)
- Purdue University (23)
- Rochester Institute of Technology (23)
- The University of Akron (22)
- Chapman University (20)
- Edith Cowan University (20)
- University of Kentucky (18)
- Michigan Technological University (16)
- University of Central Florida (15)
- Southern Adventist University (13)
- California State University, San Bernardino (12)
- Kennesaw State University (12)
- St. Mary's University (12)
- Nova Southeastern University (11)
- University of Minnesota Morris Digital Well (11)
- Louisiana State University (10)
- University of Nevada, Las Vegas (10)
- Keyword
-
- Virtual reality (62)
- Visualization (46)
- Computer graphics (38)
- Computer vision (37)
- Accessibility (36)
-
- Human-computer interaction (35)
- Augmented reality (33)
- Usability (31)
- Machine learning (29)
- Computer Science (25)
- Data visualization (25)
- Machine Learning (25)
- Artificial intelligence (24)
- Deep learning (24)
- Virtual Reality (23)
- HCI (22)
- Computer science (20)
- Eye tracking (20)
- Human computer interaction (20)
- User experience (20)
- Design (19)
- Deep Learning (16)
- Education (16)
- Feature extraction (15)
- Graph Neural Networks (15)
- Graphics (15)
- VR (15)
- Gamification (14)
- Image processing (14)
- Applied sciences (13)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (912)
- Computer Science Faculty Publications (135)
- Theses and Dissertations (98)
- Master's Theses (50)
- Graduate Theses and Dissertations (43)
-
- Student Works (2000-2009) (33)
- Computer Science and Computer Engineering Undergraduate Honors Theses (31)
- 3-D Printed Model Structural Files (29)
- Publications and Research (28)
- Dartmouth College Master’s Theses (24)
- Williams Honors College, Honors Research Projects (22)
- Master's Projects (20)
- Conference papers (19)
- Frameless (19)
- Computer Science and Software Engineering (18)
- All Dissertations (17)
- Dissertations and Theses Collection (Open Access) (16)
- Dissertations, Master's Theses and Master's Reports (16)
- H-Workload 2017: Models and Applications (Works in Progress) (15)
- Computer Engineering (14)
- Electronic Theses and Dissertations (14)
- Theses : Honours (14)
- Honors Theses (13)
- MAICS: The Modern Artificial Intelligence and Cognitive Science Conference (12)
- AFIT Patents (11)
- CCAC Theses and Dissertations (11)
- Engineering Faculty Articles and Research (11)
- Scholarly Horizons: University of Minnesota, Morris Undergraduate Journal (10)
- Inquiry: The University of Arkansas Undergraduate Research Journal (9)
- Publications (9)
- Publication Type
- File Type
Articles 301 - 330 of 2362
Full-Text Articles in Graphics and Human Computer Interfaces
Replay-And-Forget-Free Graph Class-Incremental Learning: A Task Profiling And Prompting Approach, Chaoxi Niu, Guansong Pang, Ling Chen, Bing Liu
Replay-And-Forget-Free Graph Class-Incremental Learning: A Task Profiling And Prompting Approach, Chaoxi Niu, Guansong Pang, Ling Chen, Bing Liu
Research Collection School Of Computing and Information Systems
Class-incremental learning (CIL) aims to continually learn a sequence of tasks, with each task consisting of a set of unique classes. Graph CIL (GCIL) follows the same setting but needs to deal with graph tasks (e.g., node classification in a graph). The key characteristic of CIL lies in the absence of task identifiers (IDs) during inference, which causes a significant challenge in separating classes from different tasks (i.e., inter-task class separation). Being able to accurately predict the task IDs can help address this issue, but it is a challenging problem. In this paper, we show theoretically that accurate task ID …
Interactive Known-Item Search In Large Video Corpora, Zhixin Ma
Interactive Known-Item Search In Large Video Corpora, Zhixin Ma
Dissertations and Theses Collection (Open Access)
The surge in video volume makes it challenging to locate a specific target with a single query using automatic video retrieval systems. The interactive video retrieval offers a solution by enabling users to iteratively refine a search. Nevertheless, existing systems often present users with an overwhelming number of similar videos, which can lead to mental fatigue while inspecting results and increase difficulty in providing feedback. This dissertation studies known-item video search and addresses four key challenges. First and foremost, as the link between users and the system, the interaction must be both efficient and effective. To ensure effectiveness, the user’s …
Triadic Temporal-Semantic Alignment For Weakly-Supervised Video Moment Retrieval, Jin Liu, Jialong Xie, Fengyu Zhou, Shengfeng He
Triadic Temporal-Semantic Alignment For Weakly-Supervised Video Moment Retrieval, Jin Liu, Jialong Xie, Fengyu Zhou, Shengfeng He
Research Collection School Of Computing and Information Systems
Video Moment Retrieval (VMR) aims to identify specific event moments within untrimmed videos based on natural language queries. Existing VMR methods have been criticized for relying heavily on moment annotation bias rather than true multi-modal alignment reasoning. Weakly supervised VMR approaches inherently overcome this issue by training without precise temporal location information. However, they struggle with fine-grained semantic alignment and often yield multiple speculative predictions with prolonged video spans. In this paper, we take a step forward in the context of weakly supervised VMR by proposing a triadic temporalsemantic alignment model. Our proposed approach augments weak supervision by comprehensively addressing …
Unsupervised Modality Adaptation With Text-To-Image Diffusion Models For Semantic Segmentation, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Bo Li, Yang Tang, Pan Zhou
Unsupervised Modality Adaptation With Text-To-Image Diffusion Models For Semantic Segmentation, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Bo Li, Yang Tang, Pan Zhou
Research Collection School Of Computing and Information Systems
Despite their success, unsupervised domain adaptation methods for semantic segmentation primarily focus on adaptation between image domains and do not utilize other abundant visual modalities like depth, infrared and event. This limitation hinders their performance and restricts their application in real-world multimodal scenarios. To address this issue, we propose Modality Adaptation with text-toimage Diffusion Models (MADM) for semantic segmentation task which utilizes text-to-image diffusion models pre-trained on extensive image-text pairs to enhance the model’s cross-modality capabilities. Specifically, MADM comprises two key complementary components to tackle major challenges. First, due to the large modality gap, using one modal data to generate …
3d Snapshot: Invertible Embedding Of 3d Neural Representations In A Single Image, Yuqin Lu, Bailin Deng, Zhixuan Zhong, Tianle Zhang, Yuhui Quan, Hongmin Cai, Shengfeng He
3d Snapshot: Invertible Embedding Of 3d Neural Representations In A Single Image, Yuqin Lu, Bailin Deng, Zhixuan Zhong, Tianle Zhang, Yuhui Quan, Hongmin Cai, Shengfeng He
Research Collection School Of Computing and Information Systems
3D neural rendering enables photo-realistic reconstruction of a specific scene by encoding discontinuous inputs into a neural representation. Despite the remarkable rendering results, the storage of network parameters is not transmission-friendly and not extendable to metaverse applications. In this paper, we propose an invertible neural rendering approach that enables generating an interactive 3D model from a single image (i.e., 3D Snapshot). Our idea is to distill a pre-trained neural rendering model (e.g., NeRF) into a visualizable image form that can then be easily inverted back to a neural network. To this end, we first present a neural image distillation method …
User Acceptance Of Advice By Ai Agents: Expectation-System Fit Perspective, Jingyuan Cai, Fiona Fui-Hoon Nah
User Acceptance Of Advice By Ai Agents: Expectation-System Fit Perspective, Jingyuan Cai, Fiona Fui-Hoon Nah
Research Collection School Of Computing and Information Systems
Algorithms have increasing influence on our daily decisions, especially when the recommendations are presented by human-like AI agents. This study applies the Theory of Effective Use to investigate how the fit between the user’s role expectation for an AI agent and the agent’s interaction style impacts AI advice adoption. We proposed a new concept termed Perceived Expectation-System Fit (PESF) and empirically examined its impact on user perceptions and advice acceptance. We found that low PESF reduces advice acceptance by diminishing cognitive and affective trust in the AI agent. Furthermore, increased algorithm transparency increases PESF's impact on decision-making. Our findings provide …
Habit Coach: Customising Rag-Based Chatbots To Support Behavior Change, Arian Fooroogh Mand Arabi, Cansu Koyuturk, Michael O'Mahony, Raffaella Calati, Dimitri Ognibene
Habit Coach: Customising Rag-Based Chatbots To Support Behavior Change, Arian Fooroogh Mand Arabi, Cansu Koyuturk, Michael O'Mahony, Raffaella Calati, Dimitri Ognibene
Conference papers
This paper presents the iterative development of Habit Coach, a GPT-based chatbot designed to support users in habit change through personalized interaction. Employing a user-centered design approach, we developed the chatbot using a Retrieval-Augmented Generation (RAG) system, which enables behavior personalization without retraining the underlying language model (GPT-4). The system leverages document retrieval and specialized prompts to tailor interactions, drawing from Cognitive Behavioral Therapy (CBT) and narrative therapy techniques. A key challenge in the development process was the difficulty of translating declarative knowledge into effective interaction behaviors. In the initial phase, the chatbot was provided with declarative knowledge about CBT …
Smartphone Haptics Can Uncover Differences In Touch Interactions Between Asd And Neurotypicals, Ivonne Monarca, Franceli L. Cibrian, Isabel López Hurtado, Monica Tentori
Smartphone Haptics Can Uncover Differences In Touch Interactions Between Asd And Neurotypicals, Ivonne Monarca, Franceli L. Cibrian, Isabel López Hurtado, Monica Tentori
Engineering Faculty Articles and Research
Utilizing touch interactions from smartphones for gathering data and identifying digital markers for screening and monitoring neurological disorders, such as Autism Spectrum Disorder (ASD), is an emerging area of research. Smartphones provide multiple benefits for this kind of study, including unobtrusive data collection via built-in sensors, integrated haptic feedback systems, and the capability to create specialized applications. Acknowledging the significant yet understudied presence of tactile processing differences in individuals with ASD, we designed and developed Feel and Touch, a mobile game that leverages the haptic capabilities of smartphones. This game provides vibrotactile feedback in response to touch interactions and collects …
Rubrics Informed By The Cognitive Theory Of Multimedia Learning That Support Research On Personalized Learning Paths, Sean A. Mochocki, Mark G. Reith, Jonathan Zemmer
Rubrics Informed By The Cognitive Theory Of Multimedia Learning That Support Research On Personalized Learning Paths, Sean A. Mochocki, Mark G. Reith, Jonathan Zemmer
AFIT Documents
Personalized Learning Paths (PLP)s are a popular area of research in E-Learning where sequences of Learning Materials (LM)s and activities are returned based on a learner profile, the LM metadata, and a knowledge structure that describes the relationship between the underlying topics. Unfortunately, PLP researchers tend to not use an empirically supported cognitive science framework for their research, instead relying on such unsupported theories as learning styles or developing their own ad hoc approaches. While many of these researchers present and solve challenging PLP problems using a variety of algorithmic approaches, the PLP community in general would benefit from a …
Eyetraes : Fine-Grained, Low-Latency Eye Tracking Via Adaptive Event Slicing, Argha Sen, Panahetipola Mudiyanselage Nuwan Bandara, Ila Gokarn, Thivya Kandappu, Archan Misra
Eyetraes : Fine-Grained, Low-Latency Eye Tracking Via Adaptive Event Slicing, Argha Sen, Panahetipola Mudiyanselage Nuwan Bandara, Ila Gokarn, Thivya Kandappu, Archan Misra
Research Collection School Of Computing and Information Systems
Eye-tracking technology has gained significant attention in recent years due to its wide range of applications in humancomputer interaction, virtual and augmented reality, and wearable health. Traditional RGB camera-based eye-tracking systems often struggle with poor temporal resolution and computational constraints, limiting their effectiveness in capturing rapid eye movements. To address these limitations, we propose EyeTrAES, a novel approach using neuromorphic event cameras for high-fidelity tracking of natural pupillary movement that shows significant kinematic variance. One of EyeTrAES’s highlights is the use of a novel adaptive windowing/slicing algorithm that ensures just the right amount of descriptive asynchronous event data accumulation within …
Hi3d: Pursuing High-Resolution Image-To-3d Generation With Video Diffusion Models, Haibo Yang, Yang Chen, Yingwei Pan, Ting Yao, Zhineng Chen, Chong-Wah Ngo, Tao Mei
Hi3d: Pursuing High-Resolution Image-To-3d Generation With Video Diffusion Models, Haibo Yang, Yang Chen, Yingwei Pan, Ting Yao, Zhineng Chen, Chong-Wah Ngo, Tao Mei
Research Collection School Of Computing and Information Systems
Despite having tremendous progress in image-to-3D generation, existing methods still struggle to produce multi-view consistent images with high-resolution textures in detail, especially in the paradigm of 2D diffusion that lacks 3D awareness. In this work, we present High-resolution Image-to-3D model (Hi3D), a new video diffusion based paradigm that redefines a single image to multi-view images as 3D-aware sequential image generation (i.e., orbital video generation). This methodology delves into the underlying temporal consistency knowledge in video diffusion model that generalizes well to geometry consistency across multiple views in 3D generation. Technically, Hi3D first empowers the pre-trained video diffusion model with 3D-aware …
Unlocking Potential: Analyzing The Content, Style, Structure, And Interactivity Of Mesonets As Operational Dashboards, Savannah Olivas, Jeannette Sutton, Michele K. Olson
Unlocking Potential: Analyzing The Content, Style, Structure, And Interactivity Of Mesonets As Operational Dashboards, Savannah Olivas, Jeannette Sutton, Michele K. Olson
Emergency Preparedness, Homeland Security, and Cybersecurity Faculty Scholarship
Emergency managers need data and information to make life-saving decisions on behalf of the public. Operational dashboards, if designed appropriately, can provide this information in a central location and reduce cognitive demands during decision-making. Mesonet websites can serve as a type of operational dashboard that has the potential to provide the meteorological data necessary for emergency managers to make decisions. In this study, we use quantitative content analysis to examine the content, style, structure, and interactivity of 18 Mesonet websites from across the contiguous United States. We find that Mesonet websites vary in the type and amount of content they …
Cirp: Cross‑Item Relational Pre‑Training For Multimodal Product Bundling, Yunshan Ma, Yingzhi He, Wenjun Zhong, Xiang Wang, Roger Zimmermann, Tat-Seng Chua
Cirp: Cross‑Item Relational Pre‑Training For Multimodal Product Bundling, Yunshan Ma, Yingzhi He, Wenjun Zhong, Xiang Wang, Roger Zimmermann, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Product bundling has been a prevailing marketing strategy that is beneficial in the online shopping scenario. Effective product bundling methods depend on high-quality item representations capturing both the individual items' semantics and cross-item relations. However, previous item representation learning methods, either feature fusion or graph learning, suffer from inadequate cross-modal alignment and struggle to capture the cross-item relations for cold-start items. Multimodal pre-train models could be the potential solutions given their promising performance on various multimodal downstream tasks. However, the cross-item relations have been under-explored in the current multimodal pre-train models.To bridge this gap, we propose a novel and simple …
Eyegraph : Modularity-Aware Spatio Temporal Graph Clustering For Continuous Event-Based Eye Tracking, Panahetipola Mudiyanselage Nuwan Bandara, Thivya Kandappu, Archan Misra, Ila Gokarn, Archan Misra
Eyegraph : Modularity-Aware Spatio Temporal Graph Clustering For Continuous Event-Based Eye Tracking, Panahetipola Mudiyanselage Nuwan Bandara, Thivya Kandappu, Archan Misra, Ila Gokarn, Archan Misra
Research Collection School Of Computing and Information Systems
Continuous tracking of eye movement dynamics plays a significant role in developing a broad spectrum of human-centered applications, such as cognitive skills (visual attention and working memory) modeling, human-machine interaction, biometric user authentication, and foveated rendering. Recently neuromorphic cameras have garnered significant interest in the eye-tracking research community, owing to their sub-microsecond latency in capturing intensity changes resulting from eye movements. Nevertheless, the existing approaches for event-based eye tracking suffer from several limitations: dependence on RGB frames, label sparsity, and training on datasets collected in controlled lab environments that do not adequately reflect real-world scenarios. To address these limitations, in …
Transitioning Our Website To Libguides Cms, Samantha Duncan, Eric Resnis
Transitioning Our Website To Libguides Cms, Samantha Duncan, Eric Resnis
Library Faculty Presentations
In this presentation, we describe how we used data from rapid and in-depth student usability testing to assist with the redesign of the library’s website as we finally transitioned to LibGuides CMS. Using the Springy tools; LibGuides, LibGuides CMS, LibCal, and LibWizard we outlined how we were able to create and carry out this highly effective testing, resulting in a better understanding of how our students navigate our site and how to improve it. During this journey, attendees were provided with the detailed and some might say lengthy process that was undertaken to achieve our goals. We did this by …
The Psychological Impacts Of Algorithmic And Ai-Driven Social Media On Teenagers: A Call To Action, Sunil Arora, Sahil Arora, John Hastings
The Psychological Impacts Of Algorithmic And Ai-Driven Social Media On Teenagers: A Call To Action, Sunil Arora, Sahil Arora, John Hastings
Research & Publications
This study investigates the meta-issues surrounding social media, which, while theoretically designed to enhance social interactions and improve our social lives by facilitating the sharing of personal experiences and life events, often results in adverse psychological impacts. Our investigation reveals a paradoxical outcome: rather than fostering closer relationships and improving social lives, the algorithms and structures that underlie social media platforms inadvertently contribute to a profound psychological impact on individuals, influencing them in unforeseen ways. This phenomenon is particularly pronounced among teenagers, who are disproportionately affected by curated online personas, peer pressure to present a perfect digital image, and the …
Construction And Application Of Multi-Dimensional Portrait System For Green Technology Innovation Enterprises In China: Taking The Green Transportation Technology Field As An Example, Wenke Hao, Jianlin Yang, Lei Miao
Construction And Application Of Multi-Dimensional Portrait System For Green Technology Innovation Enterprises In China: Taking The Green Transportation Technology Field As An Example, Wenke Hao, Jianlin Yang, Lei Miao
Journal of Scientific Information Research
[Purpose/significance]By constructing and applying the multi-dimensional portrait system of green technology innovation enterprises in China, this paper aims to comprehensively understand the status quo, advantages and obstacles of enterprises in specific fields in green technology innovation, so as to provide scientific references and suggestions for relevant government departments and decision makers of enterprises. [Method/process]Based on resource based view and environmental dependence theory, we select the internal and external labels of enterprises, and designs a multi-dimensional label system to objectively describe the performance of enterprises in terms of profitability,scientific research and innovation, public opinion and environmental responsibility. Then, the green technology …
An Efficient Fourier Caching Algorithm For Walk On Spheres, Zihong Zhou
An Efficient Fourier Caching Algorithm For Walk On Spheres, Zihong Zhou
Dartmouth College Master’s Theses
Walk on Spheres (WoS) is a grid-free Monte Carlo method for solving elliptic partial differential equations (PDEs).
Rather than discretizing the domain, WoS leverages the mean-value principle to obtain Monte Carlo estimates by recursively averaging the solution over the largest contained sphere, terminating upon reaching the boundary.
Unfortunately, WoS requires many independent estimates to achieve noise-free results.
We propose an acceleration technique for WoS, inspired by irradiance caching methods, that computes the solution at a sparse set of locations, and extrapolates these cached values to local neighborhoods. A key insight is that WoS can be extended to compute not only …
Fast And High-Resolution View Synthesis From A Single Input Panorama, Nam Nguyen, Angela V. Chen, Theresa Zhu, Seth Johnson, Pranav Dumpa, Benjamin Geil
Fast And High-Resolution View Synthesis From A Single Input Panorama, Nam Nguyen, Angela V. Chen, Theresa Zhu, Seth Johnson, Pranav Dumpa, Benjamin Geil
College of Engineering Summer Undergraduate Research Program
We introduce a novel method to convert a single input panorama into a 3D colored mesh representation of the scene. Unlike recent methods based on neural rendering, which are limited to low-resolution inputs and offline rendering, our approach supports 4k resolution inputs and real time rendering in a virtual reality headset. We first estimate a depth map and produce an initial layered depth image (LDI) representation. We fill unseen regions behind objects by iteratively cutting and inpainting the LDI. We then convert the LDI into an optimized, texture mapped mesh to achieve a compact representation
Digital Twin For Shelf Intelligence: Ai-Driven Inventory Management For Minimizing Food Waste, Charlotte Maples, Marvin Velazquez
Digital Twin For Shelf Intelligence: Ai-Driven Inventory Management For Minimizing Food Waste, Charlotte Maples, Marvin Velazquez
College of Engineering Summer Undergraduate Research Program
This project aims to develop a solution for improving grocery store inventory management by leveraging AI-driven image recognition. Traditional inventory methods, which rely on manual counting or barcode scanning, are inefficient, labor-intensive, and prone to human error. Over an 8-week period, we designed and developed a basic iPad app capable of identifying specific types of fruit and automatically updating inventory records in real time. By utilizing the iPad’s camera and machine learning algorithms, the app demonstrates the potential to streamline inventory tracking, reduce manual labor, and improve accuracy in managing perishable goods. Future work will focus on expanding the app’s …
Enhancing Place-Based Interaction With Emotion Ai And Augmented Reality, Jake Maier, Ivan Martinez
Enhancing Place-Based Interaction With Emotion Ai And Augmented Reality, Jake Maier, Ivan Martinez
College of Engineering Summer Undergraduate Research Program
This project explores the integration of augmented reality (AR) and Emotion AI technologies to enhance user experiences in physical environments. By seamlessly merging virtual elements with real-world contexts, we aim to deepen individuals’ interactions and perceptions of their surroundings. Leveraging AR technology enables users to access contextual information, engage with interactive content, and navigate spaces with heightened immersion and understanding. Additionally, Emotion AI enhances these experiences by detecting and responding to users’ emotional states, fostering personalized and emotionally resonant interactions. We aim to integrate digital content within physical environments using mixed-reality headsets equipped with eye-tracking capabilities and consumer-grade wireless EEG …
Leveraging Tradespace-Exploration For A Senior Project Team Formation Application, Miguel Saenz
Leveraging Tradespace-Exploration For A Senior Project Team Formation Application, Miguel Saenz
College of Engineering Summer Undergraduate Research Program
This project revolves around the development of an app in MATLAB that leverages the VASSAR rule-based system and a genetic algorithm to form groups of teams for the Mechanical Engineering Senior Design project class. We leveraged the iterative design process to eventually attain a functional app with a reasonable runtime that works provided correctly formatted rulesheets describing student project preference and member preference.
Gradualreality : Enhancing Physical Object Interaction In Virtual Reality Via Interaction State-Aware Blending, Hyuna Seo, Juheon Yi, Rajesh Krishna Balan, Youngki Lee
Gradualreality : Enhancing Physical Object Interaction In Virtual Reality Via Interaction State-Aware Blending, Hyuna Seo, Juheon Yi, Rajesh Krishna Balan, Youngki Lee
Research Collection School Of Computing and Information Systems
We present GradualReality, a novel interface enabling a Cross Reality experience that includes gradual interaction with physical objects in a virtual environment and supports both presence and usability. Daily Cross Reality interaction is challenging as the user’s physical object interaction state is continuously changing over time, causing their attention to frequently shift between the virtual and physical worlds. As such, presence in the virtual environment and seamless usability for interacting with physical objects should be maintained at a high level. To address this issue, we present an Interaction State-Aware Blending approach that (i) balances immersion and interaction capability and (ii) …
Improving Out-Of-Distribution Detection With Disentangled Foreground And Background Features, Choubo Ding, Guansong Pang
Improving Out-Of-Distribution Detection With Disentangled Foreground And Background Features, Choubo Ding, Guansong Pang
Research Collection School Of Computing and Information Systems
Detecting out-of-distribution (OOD) inputs is a principal task for ensuring the safety of deploying deep-neural-network classifiers in open-set scenarios. OOD samples can be drawn from arbitrary distributions and exhibit deviations from in-distribution (ID) data in various dimensions, such as foreground features (e.g., objects in CIFAR100 images vs. those in CIFAR10 images) and background features (e.g., textural images vs. objects in CIFAR10). Existing methods can confound foreground and background features in training, failing to utilize the background features for OOD detection. This paper considers the importance of feature disentanglement in out-of-distribution detection and proposes the simultaneous exploitation of both foreground and …
Enhancing Recipe Retrieval With Foundation Models: A Data Augmentation Perspective, Fangzhou Song, Bin Zhu, Yanbin Hao, Shuo Wang
Enhancing Recipe Retrieval With Foundation Models: A Data Augmentation Perspective, Fangzhou Song, Bin Zhu, Yanbin Hao, Shuo Wang
Research Collection School Of Computing and Information Systems
Learning recipe and food image representation in common embedding space is non-trivial but crucial for cross-modal recipe retrieval. In this paper, we propose a new perspective for this problem by utilizing foundation models for data augmentation. Leveraging on the remarkable capabilities of foundation models (i.e., Llama2 and SAM), we propose to augment recipe and food image by extracting alignable information related to the counterpart. Specifically, Llama2 is employed to generate a textual description from the recipe, aiming to capture the visual cues of a food image, and SAM is used to produce image segments that correspond to key ingredients in …
Onerestore : A Universal Restoration Framework For Composite Degradation, Yu Guo, Yuan Gao, Yuxu Lu, Huilin Zhu, Ryan Wen Liu, Shengfeng He
Onerestore : A Universal Restoration Framework For Composite Degradation, Yu Guo, Yuan Gao, Yuxu Lu, Huilin Zhu, Ryan Wen Liu, Shengfeng He
Research Collection School Of Computing and Information Systems
In real-world scenarios, image impairments often manifest as composite degradations, presenting a complex interplay of elements such as low light, haze, rain, and snow. Despite this reality, existing restoration methods typically target isolated degradation types, thereby falling short in environments where multiple degrading factors coexist. To bridge this gap, our study proposes a versatile imaging model that consolidates four physical corruption paradigms to accurately represent complex, composite degradation scenarios. In this context, we propose OneRestore, a novel transformer-based framework designed for adaptive, controllable scene restoration. The proposed framework leverages a unique cross-attention mechanism, merging degraded scene descriptors with image features, …
An End-To-End Bi-Objective Approach To Deep Graph Partitioning, Pengcheng Wei, Yuan Fang, Zhihao Wen, Zheng Xiao, Binbin Chen
An End-To-End Bi-Objective Approach To Deep Graph Partitioning, Pengcheng Wei, Yuan Fang, Zhihao Wen, Zheng Xiao, Binbin Chen
Research Collection School Of Computing and Information Systems
Graphs are ubiquitous in real-world applications, such as computation graphs and social networks. Partitioning large graphs into smaller, balanced partitions is often essential, with the biobjective graph partitioning problem aiming to minimize both the“cut” across partitions and the imbalance in partition sizes. However, existing heuristic methods face scalability challenges or overlook partition balance, leading to suboptimal results. Recent deep learning approaches, while promising, typically focus only on node-level features and lack a truly end-to-end framework, resulting in limited performance. In this paper, we introduce a novel method based on graph neural networks (GNNs) that leverages multilevel graph features and addresses …
Video Editing For Video Retrieval, Bin Zhu, Kevin Flanagan, Adriano Fragomeni, Michael Wray, Dima Damen
Video Editing For Video Retrieval, Bin Zhu, Kevin Flanagan, Adriano Fragomeni, Michael Wray, Dima Damen
Research Collection School Of Computing and Information Systems
Though pre-training vision-language models have demonstrated significant benefits in boosting video-text retrieval performance from large-scale web videos, fine-tuning still plays a critical role with manually annotated clips with start and end times, which requires considerable human effort. To address this issue, we explore an alternative cheaper source of annotations, single timestamps, for video-text retrieval. We initialise clips from timestamps in a heuristic way to warm up a retrieval model. Then a video clip editing method is proposed to refine the initial rough boundaries to improve retrieval performance. A student-teacher network is introduced for video clip editing: the teacher model is …
Densetrack : Drone-Based Crowd Tracking Via Density-Aware Motion-Appearance Synergy, Yi Lei, Huilin Zhu, Jingling Yuan, Guangli Xiang, Xian Zhong, Shengfeng He
Densetrack : Drone-Based Crowd Tracking Via Density-Aware Motion-Appearance Synergy, Yi Lei, Huilin Zhu, Jingling Yuan, Guangli Xiang, Xian Zhong, Shengfeng He
Research Collection School Of Computing and Information Systems
Drone-based crowd tracking faces difficulties in accurately identifying and monitoring objects from an aerial perspective, largely due to their small size and close proximity to each other, which complicates both localization and tracking. To address these challenges, we present the Density-aware Tracking (DenseTrack) framework. DenseTrack capitalizes on crowd counting to precisely determine object locations, blending visual and motion cues to improve the tracking of small-scale objects. It specifically addresses the problem of cross-frame motion to enhance tracking accuracy and dependability. DenseTrack employs crowd density estimates as anchors for exact object localization within video frames. These estimates are merged with motion …
Zero-Shot Object Counting With Good Exemplars, Huilin Zhu, Jingling Yuan, Zhengwei Yang, Yu Guo, Zheng Wang, Xian Zhong, Shengfeng He
Zero-Shot Object Counting With Good Exemplars, Huilin Zhu, Jingling Yuan, Zhengwei Yang, Yu Guo, Zheng Wang, Xian Zhong, Shengfeng He
Research Collection School Of Computing and Information Systems
Zero-shot object counting (ZOC) aims to enumerate objects in images using only the names of object classes during testing, without the need for manual annotations. However, a critical challenge in current ZOC methods lies in their inability to identify high-quality exemplars effectively. This deficiency hampers scalability across diverse classes and undermines the development of strong visual associations between the identified classes and image content. To this end, we propose the Visual Association-based Zero-shot Object Counting (VA-Count) framework. VACount consists of an Exemplar Enhancement Module (EEM) and a Noise Suppression Module (NSM) that synergistically refine the process of class exemplar identification …