Open Access. Powered by Scholars. Published by Universities.®
Graphics and Human Computer Interfaces Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Artificial Intelligence and Robotics (63)
- Software Engineering (28)
- Engineering (22)
- Social and Behavioral Sciences (20)
- Other Computer Sciences (14)
-
- Arts and Humanities (12)
- Computer Engineering (10)
- Psychology (10)
- Databases and Information Systems (8)
- Theory and Algorithms (8)
- Art and Design (7)
- OS and Networks (7)
- Education (6)
- Data Science (5)
- Programming Languages and Compilers (5)
- Communication (4)
- Human Factors Psychology (4)
- Medicine and Health Sciences (4)
- Systems Architecture (4)
- Aerospace Engineering (3)
- Applied Mathematics (3)
- Biomedical Engineering and Bioengineering (3)
- Business (3)
- Cognition and Perception (3)
- Communication Technology and New Media (3)
- Digital Humanities (3)
- Educational Assessment, Evaluation, and Research (3)
- Institution
-
- Singapore Management University (98)
- Old Dominion University (15)
- City University of New York (CUNY) (7)
- Dartmouth College (6)
- Clemson University (5)
-
- Chapman University (4)
- California Polytechnic State University, San Luis Obispo (3)
- Embry-Riddle Aeronautical University (3)
- The University of Akron (3)
- University of Texas at Arlington (3)
- Eastern Washington University (2)
- University of South Alabama (2)
- Air Force Institute of Technology (1)
- Bellarmine University (1)
- Binghamton University (1)
- California State University, San Bernardino (1)
- Colby College (1)
- DePauw University (1)
- Extension Journal Inc (1)
- Florida Institute of Technology (1)
- Fort Hays State University (1)
- Hunan Provincial Institute of Scientific and Technology Information (1)
- Illinois State University (1)
- Kennesaw State University (1)
- Lindenwood University (1)
- Louisiana State University (1)
- Mississippi State University (1)
- Murray State University (1)
- SUNY Geneseo (1)
- Southern Adventist University (1)
- Keyword
-
- Accessibility (9)
- Graph Neural Networks (7)
- Machine learning (6)
- Large Language Models (4)
- Machine Learning (4)
-
- AI (3)
- Anomaly Detection (3)
- Artificial Intelligence (3)
- Artificial intelligence (3)
- Assistive technology (3)
- Automation (3)
- Blind (3)
- Catering industry (3)
- Handicapped aids (3)
- Human-computer interaction (3)
- Large language model (3)
- Training (3)
- Website (3)
- Applied computing (2)
- Augmented reality (2)
- Autism (2)
- Benchmark (2)
- Computer graphics (2)
- Computer vision (2)
- Context Awareness (2)
- Cooking (2)
- Decision making (2)
- Deep Learning (2)
- Densest Subgraph Discovery (2)
- Diffusion (2)
- Publication
-
- Research Collection School Of Computing and Information Systems (93)
- Computer Science Faculty Publications (8)
- All Dissertations (5)
- Dartmouth College Master’s Theses (5)
- Dissertations and Theses Collection (Open Access) (5)
-
- Publications and Research (5)
- Honors Theses (3)
- Williams Honors College, Honors Research Projects (3)
- Doctoral Dissertations and Master's Theses (2)
- Master's Theses (2)
- Psychology Faculty Publications (2)
- Shelby Hall Graduate Research Forum Posters (2)
- Student Scholar Symposium Abstracts and Posters (2)
- Theses and Dissertations (2)
- 2025 Spring Honors Capstone Projects - Archive (1)
- 2025 Symposium (1)
- AFIT Patents (1)
- African Conference on Information Systems and Technology (1)
- Civil & Environmental Engineering Faculty Publications (1)
- Computer Science Theses (1)
- Computer Science and Engineering Theses - Archive (1)
- Dartmouth College Ph.D Dissertations (1)
- EWU Masters Thesis Collection (1)
- Electrical Engineering and Computer Science (MS) Theses (1)
- Electrical Engineering and Computer Science Undergraduate Honors Theses (1)
- Electronic Theses and Dissertations (1)
- Electronic Theses, Projects, and Dissertations (1)
- Engineering Faculty Articles and Research (1)
- Faculty Publications - Information Technology (1)
- Faculty Scholarship (1)
- Publication Type
Articles 1 - 30 of 179
Full-Text Articles in Graphics and Human Computer Interfaces
Tests Without Borders: A Global Approach To Measuring Visualization Literacy, Olivia A. Guess
Tests Without Borders: A Global Approach To Measuring Visualization Literacy, Olivia A. Guess
McKelvey School of Engineering Graduate Student Theses & Dissertations
Visualization literacy assessments shape how we understand people's ability to interpret data, yet most existing instruments embed Western datasets and assumptions that limit their relevance for global audiences. This thesis argues that because data is personal, assessments must also be culturally grounded. We introduce a unified framework for adapting the Mini-VLAT into 22 regionally responsive short-form assessments, each retaining the structure of the original test while incorporating datasets and scenarios tailored to specific regions around the world. To demonstrate how such adaptations can be customized and validated, we present a detailed case study of a Ghana-adapted Mini-VLAT, developed in collaboration …
Developing An Ai-Assisted Grading System Using Large Language Models, Andrei Modiga
Developing An Ai-Assisted Grading System Using Large Language Models, Andrei Modiga
MS in Computer Science Project Reports
We present a grading system that accelerates evaluation of open-ended student work across scanned and digital workflows. The system crops answer regions from PDFs, assigns submissions via OCR on identity regions only, and groups answers by visual semantics using a vision LLM. Instructors review and edit groups, apply rubric items once per group, and export grades from an on-screen table. The solution integrates Ghostscript rasterization, PdfPig page orchestration, SkiaSharp region extraction, Tesseract identity OCR, and GPT-4o Vision for grouping. We detail the architecture, token-budgeted batching strategy, and persistence design, then describe testing results for grouping quality, time-on-task, and usability. The …
My First Conversation With Chatgpt (February 22, 2023): Origins Of A Generative Dialogue, David Smith
My First Conversation With Chatgpt (February 22, 2023): Origins Of A Generative Dialogue, David Smith
Publications and Research
This working paper presents the first recorded interaction between the author and the generative AI system ChatGPT, written on February 22, 2023 during the initial weeks of a faculty sabbatical in Boston. The document preserves a complete and unedited transcript of an exploratory conversation conducted without predetermined research aims, marking the author’s first encounter with a large-language-model conversational interface. Although the exchange includes creative experimentation—including musical and poetic prompts—the discussion remains informal and wide-ranging, and no theoretical framework is articulated at this stage. Rather, this transcript is published as primary-source material documenting the moment of discovery and experimentation that precedes …
Decoding The Chameleon Game, Tri Dang '25, Hieu Tran, Brian T. Howard, Sutthirut Charoenphon, Dat Nguyen '25
Decoding The Chameleon Game, Tri Dang '25, Hieu Tran, Brian T. Howard, Sutthirut Charoenphon, Dat Nguyen '25
Student Research
The Chameleon game is a challenging word association activity where players are given a secret word and must respond with words relevant to that secret word. It requires strategic thinking and deduction. The Chameleon must cleverly guess the secret keyword in this game while avoiding suspicion. Our research aims to create an advanced artificial intelligence (AI) model that can play the Chameleon game from both perspectives: as the Chameleon and as a Human. This AI is designed to guess secret keywords based on the information provided by the players, choose the best strategies to avoid detection as the Chameleon, identify …
Developing Accessible Narrative-Based Stem Learning Software For K-6 Braille Display Users, Dylan Ravel, Daniel Tsivkovski, Brandon Foley, Maryam Etezad, Franceli Cibrian, Ariel Han, Rajeev Joshi
Developing Accessible Narrative-Based Stem Learning Software For K-6 Braille Display Users, Dylan Ravel, Daniel Tsivkovski, Brandon Foley, Maryam Etezad, Franceli Cibrian, Ariel Han, Rajeev Joshi
Student Scholar Symposium Abstracts and Posters
This research develops a free, accessible web application that enables K-6 students who are blind or visually impaired (BVI) to learn STEM concepts using refreshable braille displays. Currently, most online learning tools are not designed for BVI students, creating a significant educational barrier.
The application interfaces with commercial braille displays and uses narrative-based learning to make STEM content approachable and engaging. By presenting material as interactive stories, students can connect with concepts while developing braille reading skills. The curriculum design prioritizes accessibility through the Accessible Rich Internet Applications (ARIA) standards and screen reader support.
The goal is to provide BVI …
Visionglow: Evaluating Minimal-Disruption Smart-Home Control In Apple Vision Pro, Hongxiao Zheng
Visionglow: Evaluating Minimal-Disruption Smart-Home Control In Apple Vision Pro, Hongxiao Zheng
Dartmouth College Master’s Theses
Smart-home control in mixed-reality environments like Apple Vision Pro often relies on disruptive, application-based paradigms, such as using a smartphone or a windowed virtual interface. These methods create a “mode switch” that imposes cognitive load and pulls users from their primary tasks. We present VisionGlow, a minimal-disruption spatial interaction technique for Vision Pro. VisionGlow represents devices as spatially-anchored “orbs.” To control a device, the user looks at its orb and performs a pinch gesture, which invokes a compact, contextual control panel. We conducted a within-subjects study (N=18) comparing VisionGlow against two baselines: the standard Apple Home app on a smartphone …
The Future Is Now: Empowering Society Through Ai Literacy, Jason S. Wrench, Sanae Elmoudden
The Future Is Now: Empowering Society Through Ai Literacy, Jason S. Wrench, Sanae Elmoudden
Milne Open Textbooks
Artificial Intelligence (AI) is no longer a futuristic concept—it is the reality of the present. From the algorithms shaping our social media feeds to the generative tools transforming our workplaces, AI has permeated every aspect of modern life. The Future is Now moves beyond the hype to provide a comprehensive roadmap for understanding, navigating, and shaping this technological revolution.
Demystifying the Machine
This textbook serves as a user-friendly guide to the “black box” of AI. It breaks down complex technical concepts—from machine learning and neural networks to large language models—making them accessible to students across all disciplines. By establishing a …
Digital Reflections: Evaluating Body Dissatisfaction In Xr Through Eye- And Body-Tracked Virtual Humans, Deyrel Diaz
Digital Reflections: Evaluating Body Dissatisfaction In Xr Through Eye- And Body-Tracked Virtual Humans, Deyrel Diaz
All Dissertations
In an era where digital and physical realities increasingly intertwine, the perception of body image is undergoing a significant transformation. Traditional understandings of body dissatisfaction, long studied in relation to psychological distress and eating disorders, are now being reshaped by technologies such as Virtual Reality (VR), Augmented Reality (AR), and Artificially Intelligent (AI)- generated media. These technologies have introduced novel ways of experiencing and interacting with the human form, raising critical questions about their impact on self-perception and internalization of beauty standards.
As virtual representations become more prevalent in entertainment, social media, and interactive platforms, it is becoming more crucial …
Instance-Level Video Depth In Groups Beyond Occlusions, Yuan Liang, Yang Zhou, Ziming Sun, Tianyi Xiang, Guiqing Li, Shengfeng He
Instance-Level Video Depth In Groups Beyond Occlusions, Yuan Liang, Yang Zhou, Ziming Sun, Tianyi Xiang, Guiqing Li, Shengfeng He
Research Collection School Of Computing and Information Systems
Depth estimation in dynamic, multi-object scenes remains a major challenge, especially under severe occlusions. Existing monocular models, including foundation models, struggle with instance-wise depth consistency due to their reliance on global regression. We tackle this problem from two key aspects: data and methodology. First, we introduce the Group Instance Depth (GID) dataset, the first large-scale video depth dataset with instance-level annotations, featuring 101,500 frames from real-world activity scenes. GID bridges the gap between synthetic and real-world depth data by providing high-fidelity depth supervision for multi-object interactions. Second, we propose InstanceDepth, the first occlusion-aware depth estimation framework for multi-object environments. Our …
Examining The Roles Of Embodiment And Theory Of Mind In Shaping User Perceptions Of Llm-Driven Conversational Agents, Elizabeth A. Schlesener
Examining The Roles Of Embodiment And Theory Of Mind In Shaping User Perceptions Of Llm-Driven Conversational Agents, Elizabeth A. Schlesener
All Dissertations
Large Language Models (LLMs) have advanced conversational agents, enabling natural, human-like interactions in domains such as education, programming, and workplace collaboration. Yet, user distrust persists over privacy, accuracy, and bias. As developers work to mitigate these issues and human-AI collaboration expands, reinforcing trust in LLM-driven systems is essential. To address this problem, this dissertation explores the role of anthropomorphic form in LLM-driven conversational agents and its impact on user perception.
According to the familiarity thesis, humans attribute human-like characteristics to nonhuman entities — a process known as anthropomorphism — to better comprehend unfamiliar phenomena, based on the assumption that they …
Cssa-Fusion: Channel Selective And Spatial Alignment Infrared-Visible Image Fusion, Zhen Li, Zhi Zeng, Zhongrui Xiao, Ming Wen, Zhiyuan Zhang, Yibin Tian
Cssa-Fusion: Channel Selective And Spatial Alignment Infrared-Visible Image Fusion, Zhen Li, Zhi Zeng, Zhongrui Xiao, Ming Wen, Zhiyuan Zhang, Yibin Tian
Research Collection School Of Computing and Information Systems
Infrared-visible image fusion aims to integrate complementary information from two modalities to generate images with enriched semantic content. However, existing methods often neglect two critical aspects: the design of a local–global feature enhancement architecture and spatial alignment. To address these challenges, we propose Channel Selective and Spatial Alignment Fusion (CSSA-Fusion), a novel framework composed of two synergistic modules. The first is a selective channel and redundancy suppression module, which introduces a dual-branch selective channel attention mechanism to jointly capture local saliency and global channel importance for enhanced feature representation, and an informativeness–redundancy separation strategy to suppress redundant information while preserving …
Registration Is A Powerful Rotation-Invariance Learner For 3d Anomaly Detection, Yuyang Yu, Zhengwei Chen, Xuemiao Xu, Lei Zhang, Haoxin Yang, Yongwei Nie, Shengfeng He
Registration Is A Powerful Rotation-Invariance Learner For 3d Anomaly Detection, Yuyang Yu, Zhengwei Chen, Xuemiao Xu, Lei Zhang, Haoxin Yang, Yongwei Nie, Shengfeng He
Research Collection School Of Computing and Information Systems
3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based methods often suffer from inconsistent feature transformations and limited discriminative capacity, particularly in capturing local geometric details and achieving rotation invariance. These limitations become more pronounced when registration fails, leading to unreliable detection results. We argue that point-cloud registration plays an essential role not only in aligning geometric structures but also in guiding feature extraction toward rotation-invariant and locally discriminative representations. To this end, we propose a registration-induced, rotation-invariant feature extraction framework that integrates the objectives …
Jury-And-Judge Chain-Of-Thought For Uncovering Toxic Data In 3d Visual Grounding, Kaixiang Huang, Qifeng Zhang, Jin Wang, Jingru Yang, Yang Zhou, Huan Yu, Guodong Lu, Shengfeng He
Jury-And-Judge Chain-Of-Thought For Uncovering Toxic Data In 3d Visual Grounding, Kaixiang Huang, Qifeng Zhang, Jin Wang, Jingru Yang, Yang Zhou, Huan Yu, Guodong Lu, Shengfeng He
Research Collection School Of Computing and Information Systems
3D Visual Grounding (3DVG) faces persistent challenges due to coarse scene-level observations and logically inconsistent annotations, which introduce ambiguities that compromise data quality and hinder effective model supervision. To address these challenges, we introduce Refer-Judge, a novel framework that harnesses the reasoning capabilities of Multimodal Large Language Models (MLLMs) to identify and mitigate toxic data. At the core of Refer-Judge is a Jury-and-Judge Chain-of-Thought paradigm, inspired by the deliberative process of the judicial system. This framework targets the root causes of annotation noise: jurors collaboratively assess 3DVG samples from diverse perspectives, providing structured, multi-faceted evaluations. Judges then consolidate these insights …
Marsanywhere: Dataset And Cross-View Diffusion Model For Satellite-To-Ground View Synthesis With Mars Data, Benjamin T. Hinchliff
Marsanywhere: Dataset And Cross-View Diffusion Model For Satellite-To-Ground View Synthesis With Mars Data, Benjamin T. Hinchliff
Master's Theses
Satellite-to-ground view synthesis aims to create a realistic ground view image from a corresponding satellite view image. This is a well-studied problem for street level imagery, with good results being achieved by using modern image synthesis techniques such as diffusion models. However, despite the public availability of satellite and ground level imagery on Mars, these techniques have yet to be applied to the domain due to difficulties in collating and processing the data into a usable form. We address this deficiency by creating a dataset consisting of ground view panorama imagery from the Perseverance rover, along with associated satellite view …
Implementation And Assessment Of The Openbci Platform As An Accessible Brain- Computer Interface, Jewell Norris
Implementation And Assessment Of The Openbci Platform As An Accessible Brain- Computer Interface, Jewell Norris
Honors Theses
OpenBCI is a low-cost, open-source platform for alternative brain-computer interface (BCI) software and hardware. This thesis evaluates OpenBCI’s electroencephalogram (EEG) and electromyography (EMG) capabilities by constructing and testing a 16-channel EEG system using the Ultracortex Mark IV headset and Cyton + Daisy biosensing board. The viability of the system was assessed through real-time BCI control and comparison to a clinical-grade EEG system. Real-time BCI control of an online falling-block game was tested via the use of eye blinks EMG (channels Fp1/Fp2) and head-tilt accelerometer inputs. The BCI game demonstrated reliable control despite minor latency and artifact sensitivity. For clinical comparison, …
Graph Perturbations For Robust Knowledge Discovery And Retrieval, Hanhua Xiao
Graph Perturbations For Robust Knowledge Discovery And Retrieval, Hanhua Xiao
Dissertations and Theses Collection (Open Access)
Graph perturbation, rooted in classical perturbation theory, studies how small topology edits, i.e., adding or deleting edges, affects graph properties (e.g., density, centrality). This fundamental problem underpins applications like bioinformatics, privacy preservation and system defense. While much prior work targets perturbations that influence global graph statistics or model outputs, comparatively little addresses robustness for knowledge discovery and information retrieval. In these settings, graphs are attributed: nodes carry real-world semantics (e.g., locations, people) and edges encode interactions or relationships. This thesis proposes new formulations and algorithms that generate and leverage graph perturbations to make knowledge discovery and retrieval more robust. Specifically, …
Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo
Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergence of new cultural datasets. However, these datasets frequently fall short of providing cultural reasoning while underrepresenting many cultures.In this paper, we introduce the Seeing Culture Benchmark (SCB), focusing on cultural reasoning with a novel approach that requires VLMs to reason on culturally rich images in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual …
Addressing Sparsity For Knowledge Graph Completion: Data And Model Perspectives, Ran Liu
Addressing Sparsity For Knowledge Graph Completion: Data And Model Perspectives, Ran Liu
Dissertations and Theses Collection (Open Access)
Knowledge graphs (KGs) are powerful tools for structuring factual knowledge into relational triples, yet their practical utility is often adversely affected by data sparsity. Many entities and relations are associated with only a few observations, which limits the quality of learned embeddings and weakens generalization in downstream tasks. The problem of sparsity led to two interrelated challenges. Firstly, it restricts the informativeness of training samples: positive examples are scarce, and conventional negative sampling often produces trivial or redundant negatives that resulting in limited guidance. Secondly, in few-shot relation learning scenarios, sparsity worsens distribution shifts between training and test relations, as …
Sketch-Sparsenet: Sparse Convolution Framework For Sketch Recognition, Jingru Yang, Jin Wang, Yang Zhou, Guodong Lu, Yu Sun, Huan Yu, Heming Fang, Zhihui Li, Shengfeng He
Sketch-Sparsenet: Sparse Convolution Framework For Sketch Recognition, Jingru Yang, Jin Wang, Yang Zhou, Guodong Lu, Yu Sun, Huan Yu, Heming Fang, Zhihui Li, Shengfeng He
Research Collection School Of Computing and Information Systems
In free-hand sketch recognition, state-of-the-art methods often struggle to extract spatial features from sketches with sparse distributions, which are characterized by significant blank regions devoid of informative content. To address this challenge, we introduce a novel framework for sketch recognition, termed Sketch-SparseNet. This framework incorporates an advanced convolutional component: the Sketch-Driven Dilated Deformable Block (SD3B). This component excels at extracting spatial features and accurately recognizing free-hand sketches with sparse distributions. The SD3B component innovatively bridges gaps in the blank areas of sketches by establishing spatial relationships among disconnected stroke points through adaptive reshaping of convolution kernels. These kernels are deformable, …
Unresolved Image Simulation For Space Situational Awareness Applications, Fox Coniglario
Unresolved Image Simulation For Space Situational Awareness Applications, Fox Coniglario
Doctoral Dissertations and Master's Theses
The knowledge of what lies in orbit around Earth is at best a guess. Decades of spaceflight, debris buildup, and vehicle collisions have contributed to a large number of objects that are simply not able to be catalogued. Ongoing efforts to catalog debris in orbit have reached limits by conventional measures and as such, research is active in the field of in-orbit space situational awareness. This thesis intends to help fill a hole in the development of such orbital platforms by assisting the development of image processing software pipelines though the simulation of unresolved space imagery. The simulation uses accurate …
Look Before You Decide: Prompting Active Deduction Of Mllms For Assumptive Reasoning, Yian Li, Wentao Tian, Yang Jiao, Jingjing Chen, Tianwen Qian, Bin Zhu, Na Zhao, Yu‑Gang Jiang
Look Before You Decide: Prompting Active Deduction Of Mllms For Assumptive Reasoning, Yian Li, Wentao Tian, Yang Jiao, Jingjing Chen, Tianwen Qian, Bin Zhu, Na Zhao, Yu‑Gang Jiang
Research Collection School Of Computing and Information Systems
Recently, Multimodal Large Language Models (MLLMs) have achieved significant success across multiple disciplines due to their exceptional instruction-following capabilities and extensive world knowledge. However, whether these MLLMs possess human-like compositional reasoning abilities remains an open problem. To unveil their reasoning behaviors, we first curate a Multimodal Assumptive Reasoning Benchmark (MARS-Bench) in this paper. Interestingly, we find that most prevalent MLLMs can be easily fooled by the introduction of a presupposition into the question, whereas such presuppositions appear naive to human reasoning. Besides, we also propose a simple yet effective method, Active Deduction (AD), a novel reinforcement learning paradigm to encourage …
Viewsrd: 3d Visual Grounding Via Structured Multi-View Decomposition, Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Viewsrd: 3d Visual Grounding Via Structured Multi-View Decomposition, Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
3Dvisual grounding aims to identify and localize objects in a 3Dspacebasedontextualdescriptions. However, existing methods struggle with disentangling targets from anchors in complex multi-anchor queries and resolving inconsisten cies in spatial descriptions caused by perspective variations. To tackle these challenges, we propose ViewSRD, a frame work that formulates 3D visual grounding as a structured multi-view decomposition process. First, the Simple Rela tion Decoupling (SRD) module restructures complex multi anchor queries into a set of targeted single-anchor state ments, generating a structured set of perspective-aware de scriptions that clarify positional relationships. These de composed representations serve as the foundation for the Multi-view …
Omnivton: Training-Free Universal Virtual Try-On, Zhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li, Yangyang Xu, Junyu Dong, Yong Du
Omnivton: Training-Free Universal Virtual Try-On, Zhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li, Yangyang Xu, Junyu Dong, Yong Du
Research Collection School Of Computing and Information Systems
Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unified, training-free solution that works across both scenarios remains an open challenge. We propose OmniVTON, the first training-free universal VTON framework that decouples garment and pose conditioning to achieve both texture fidelity and pose consistency across diverse settings. To preserve garment details, we introduce a garment prior generation mechanism that aligns clothing with the body, followed by continuous boundary stitching technique to achieve …
Genwardrobe: A Fully Generative System For Travel Fashion Wardrobe Construction, Peng Jin, Yilin Wen, Mingzhe Yu, Yunshan Ma, Rong Zheng, Jin‑Tu Fan, Chong Wah Ngo
Genwardrobe: A Fully Generative System For Travel Fashion Wardrobe Construction, Peng Jin, Yilin Wen, Mingzhe Yu, Yunshan Ma, Rong Zheng, Jin‑Tu Fan, Chong Wah Ngo
Research Collection School Of Computing and Information Systems
With the increasing demand for outfit planning in real-world travel scenarios, the need for constructing a travel fashion wardrobe, a series of outfits tailored to a user's personalization and destination-specific context over a short travel period, has grown significantly. However, existing systems or works often focus on isolated factors and rely on retrieval-based methods, with insufficient utilization of generative models, limiting their adaptability to real-world travel scenarios. To address this issue, this study introduces GenWardrobe, a fully generative system for travel fashion wardrobe construction. GenWardrobe consists of three key modules: user query analysis, fashion knowledge retrieval via retrieval-augmented generation and …
Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning, Yangyang Xu, Shengfeng He, Wenqi Shao, Yong Du, Kwan-Yee K. Wong, Yu Qiao, Jun Yu, Ping Luo
Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning, Yangyang Xu, Shengfeng He, Wenqi Shao, Yong Du, Kwan-Yee K. Wong, Yu Qiao, Jun Yu, Ping Luo
Research Collection School Of Computing and Information Systems
In this paper, we introduce DiffusionMat, a novel image matting framework that employs a diffusion model for the transition from coarse to refined alpha mattes. Diverging from conventional methods that utilize trimaps merely as loose guidance for alpha matte prediction, our approach treats image matting as a deterministic sequential refinement learning process. This process begins with the addition of noise to trimaps and iteratively denoises them using a pre-trained diffusion model, which incrementally guides the prediction towards a clean alpha matte. The key innovation of our framework is a correction module that adjusts the output at each denoising step, ensuring …
Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan Wang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Yi Tan, Xiang Wang
Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan Wang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Yi Tan, Xiang Wang
Research Collection School Of Computing and Information Systems
Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional inputs like cooking styles, ingredients, and recipes are utilized. However, a yet-unexplored challenge is generating a sequence of procedural images based on cooking steps from a recipe. This could enhance the cooking experience with visual guidance and possibly lead to an intelligent cooking simulation system. To fill this gap, we introduce a novel task called cooking procedural image generation. This task is inherently demanding, as it strives to create photo-realistic images that align with cooking steps while …
From Holistic To Localized: Local Enhanced Adapters For Efficient Visual Instruction Fine-Tuning, Pengkun Jiao, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yugang Jiang
From Holistic To Localized: Local Enhanced Adapters For Efficient Visual Instruction Fine-Tuning, Pengkun Jiao, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yugang Jiang
Research Collection School Of Computing and Information Systems
Efficient Visual Instruction Fine-Tuning (EVIT) seeks to adapt Multimodal Large Language Models (MLLMs) to downstream tasks with minimal computational overhead. However, as task diversity and complexity increase, EVIT faces significant challenges in resolving data conflicts. To address this limitation, we propose the Dual Low-Rank Adaptation (Dual-LoRA), a holistic-to-local framework that enhances the adapter’s capacity to address data conflict through dual structural optimization. Specifically, we utilize two subspaces: a skill space for stable, holistic knowledge retention, and a rank-rectified task space that locally activates the holistic knowledge. Additionally, we introduce Visual Cue Enhancement (VCE), a multi-level local feature aggregation module designed …
Exploring Object Status Recognition For Recipe Progress Tracking In Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington
Exploring Object Status Recognition For Recipe Progress Tracking In Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington
Research Collection School Of Computing and Information Systems
Cooking plays a vital role in everyday independence and well-being, yet remains challenging for people with vision impairments due to limited support for tracking progress and receiving contextual feedback. Object status — the condition or transformation of ingredients and tools — offers a promising but underexplored foundation for context-aware cooking support. In this paper, we present OSCAR (Object Status Context Awareness for Recipes), a technical pipeline that explores the use of object status recognition to enable recipe progress tracking in non-visual cooking. OSCAR integrates recipe parsing, object status extraction, visual alignment with cooking steps, and time-causal modeling to support real-time …
Teaching Diffusion Models To Ground Alpha Matte, Tianyi Xiang, Weiying Zheng, Yutao Jiang, Tingrui Shen, Hewei Yu, Yangyang Xu, Shengfeng He
Teaching Diffusion Models To Ground Alpha Matte, Tianyi Xiang, Weiying Zheng, Yutao Jiang, Tingrui Shen, Hewei Yu, Yangyang Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
The power of visual language models is showcased in visual understanding tasks, where language-guided models achieve impressive flexibility and precision. In this paper, we ex tend this capability to the challenging domain of image matting by framing it as a soft grounding problem, enabling a single diffusion model to handle diverse objects, textures, and transparencies, all directed by descriptive text prompts. Our method teaches the diffusion model to ground alpha mattes by guiding it through a process of instance-level localization and transparency estimation. First, we introduce an intermediate objective that trains the model to accurately localize semantic components of the …
Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng, Xiao Bai
Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng, Xiao Bai
Research Collection School Of Computing and Information Systems
Recent advancements in CLIP-based out-of-distribution (OOD) detection have shown promising results via regularization on prompt tuning, leveraging background features extracted from a few in-distribution (ID) samples as proxies for OOD features.However, these methods suffer from an inherent limitation: a lack of diversity in the extracted OOD features from the few-shot ID data.To address this issue, we propose to leverage external datasets as auxiliary outlier data (i.e., pseudo OOD samples) to extract rich, diverse OOD features, with the features from not only background regions but also foreground object regions, thereby supporting more discriminative prompt tuning for OOD detection. We further introduce …