Marsanywhere: Dataset And Cross-View Diffusion Model For Satellite-To-Ground View Synthesis With Mars Data,
2025
California Polytechnic State University, San Luis Obispo
Marsanywhere: Dataset And Cross-View Diffusion Model For Satellite-To-Ground View Synthesis With Mars Data, Benjamin T. Hinchliff
Master's Theses
Satellite-to-ground view synthesis aims to create a realistic ground view image from a corresponding satellite view image. This is a well-studied problem for street level imagery, with good results being achieved by using modern image synthesis techniques such as diffusion models. However, despite the public availability of satellite and ground level imagery on Mars, these techniques have yet to be applied to the domain due to difficulties in collating and processing the data into a usable form. We address this deficiency by creating a dataset consisting of ground view panorama imagery from the Perseverance rover, along with associated satellite view …
Examining The Roles Of Embodiment And Theory Of Mind In Shaping User Perceptions Of Llm-Driven Conversational Agents,
2025
Clemson University
Examining The Roles Of Embodiment And Theory Of Mind In Shaping User Perceptions Of Llm-Driven Conversational Agents, Elizabeth A. Schlesener
All Dissertations
Large Language Models (LLMs) have advanced conversational agents, enabling natural, human-like interactions in domains such as education, programming, and workplace collaboration. Yet, user distrust persists over privacy, accuracy, and bias. As developers work to mitigate these issues and human-AI collaboration expands, reinforcing trust in LLM-driven systems is essential. To address this problem, this dissertation explores the role of anthropomorphic form in LLM-driven conversational agents and its impact on user perception.
According to the familiarity thesis, humans attribute human-like characteristics to nonhuman entities — a process known as anthropomorphism — to better comprehend unfamiliar phenomena, based on the assumption that they …
Cssa-Fusion: Channel Selective And Spatial Alignment Infrared-Visible Image Fusion,
2025
Singapore Management University
Cssa-Fusion: Channel Selective And Spatial Alignment Infrared-Visible Image Fusion, Zhen Li, Zhi Zeng, Zhongrui Xiao, Ming Wen, Zhiyuan Zhang, Yibin Tian
Research Collection School Of Computing and Information Systems
Infrared-visible image fusion aims to integrate complementary information from two modalities to generate images with enriched semantic content. However, existing methods often neglect two critical aspects: the design of a local–global feature enhancement architecture and spatial alignment. To address these challenges, we propose Channel Selective and Spatial Alignment Fusion (CSSA-Fusion), a novel framework composed of two synergistic modules. The first is a selective channel and redundancy suppression module, which introduces a dual-branch selective channel attention mechanism to jointly capture local saliency and global channel importance for enhanced feature representation, and an informativeness–redundancy separation strategy to suppress redundant information while preserving …
Registration Is A Powerful Rotation-Invariance Learner For 3d Anomaly Detection,
2025
Singapore Management University
Registration Is A Powerful Rotation-Invariance Learner For 3d Anomaly Detection, Yuyang Yu, Zhengwei Chen, Xuemiao Xu, Lei Zhang, Haoxin Yang, Yongwei Nie, Shengfeng He
Research Collection School Of Computing and Information Systems
3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based methods often suffer from inconsistent feature transformations and limited discriminative capacity, particularly in capturing local geometric details and achieving rotation invariance. These limitations become more pronounced when registration fails, leading to unreliable detection results. We argue that point-cloud registration plays an essential role not only in aligning geometric structures but also in guiding feature extraction toward rotation-invariant and locally discriminative representations. To this end, we propose a registration-induced, rotation-invariant feature extraction framework that integrates the objectives …
Jury-And-Judge Chain-Of-Thought For Uncovering Toxic Data In 3d Visual Grounding,
2025
Singapore Management University
Jury-And-Judge Chain-Of-Thought For Uncovering Toxic Data In 3d Visual Grounding, Kaixiang Huang, Qifeng Zhang, Jin Wang, Jingru Yang, Yang Zhou, Huan Yu, Guodong Lu, Shengfeng He
Research Collection School Of Computing and Information Systems
3D Visual Grounding (3DVG) faces persistent challenges due to coarse scene-level observations and logically inconsistent annotations, which introduce ambiguities that compromise data quality and hinder effective model supervision. To address these challenges, we introduce Refer-Judge, a novel framework that harnesses the reasoning capabilities of Multimodal Large Language Models (MLLMs) to identify and mitigate toxic data. At the core of Refer-Judge is a Jury-and-Judge Chain-of-Thought paradigm, inspired by the deliberative process of the judicial system. This framework targets the root causes of annotation noise: jurors collaboratively assess 3DVG samples from diverse perspectives, providing structured, multi-faceted evaluations. Judges then consolidate these insights …
Implementation And Assessment Of The Openbci Platform As An Accessible Brain- Computer Interface,
2025
Mississippi State University
Implementation And Assessment Of The Openbci Platform As An Accessible Brain- Computer Interface, Jewell Norris
Honors Theses
OpenBCI is a low-cost, open-source platform for alternative brain-computer interface (BCI) software and hardware. This thesis evaluates OpenBCI’s electroencephalogram (EEG) and electromyography (EMG) capabilities by constructing and testing a 16-channel EEG system using the Ultracortex Mark IV headset and Cyton + Daisy biosensing board. The viability of the system was assessed through real-time BCI control and comparison to a clinical-grade EEG system. Real-time BCI control of an online falling-block game was tested via the use of eye blinks EMG (channels Fp1/Fp2) and head-tilt accelerometer inputs. The BCI game demonstrated reliable control despite minor latency and artifact sensitivity. For clinical comparison, …
Graph Perturbations For Robust Knowledge Discovery And Retrieval,
2025
Singapore Management University
Graph Perturbations For Robust Knowledge Discovery And Retrieval, Hanhua Xiao
Dissertations and Theses Collection (Open Access)
Graph perturbation, rooted in classical perturbation theory, studies how small topology edits, i.e., adding or deleting edges, affects graph properties (e.g., density, centrality). This fundamental problem underpins applications like bioinformatics, privacy preservation and system defense. While much prior work targets perturbations that influence global graph statistics or model outputs, comparatively little addresses robustness for knowledge discovery and information retrieval. In these settings, graphs are attributed: nodes carry real-world semantics (e.g., locations, people) and edges encode interactions or relationships. This thesis proposes new formulations and algorithms that generate and leverage graph perturbations to make knowledge discovery and retrieval more robust. Specifically, …
Seeing Culture: A Benchmark For Visual Reasoning And Grounding,
2025
Singapore Management University
Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergence of new cultural datasets. However, these datasets frequently fall short of providing cultural reasoning while underrepresenting many cultures.In this paper, we introduce the Seeing Culture Benchmark (SCB), focusing on cultural reasoning with a novel approach that requires VLMs to reason on culturally rich images in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual …
Addressing Sparsity For Knowledge Graph Completion: Data And Model Perspectives,
2025
Singapore Management University
Addressing Sparsity For Knowledge Graph Completion: Data And Model Perspectives, Ran Liu
Dissertations and Theses Collection (Open Access)
Knowledge graphs (KGs) are powerful tools for structuring factual knowledge into relational triples, yet their practical utility is often adversely affected by data sparsity. Many entities and relations are associated with only a few observations, which limits the quality of learned embeddings and weakens generalization in downstream tasks. The problem of sparsity led to two interrelated challenges. Firstly, it restricts the informativeness of training samples: positive examples are scarce, and conventional negative sampling often produces trivial or redundant negatives that resulting in limited guidance. Secondly, in few-shot relation learning scenarios, sparsity worsens distribution shifts between training and test relations, as …
Sketch-Sparsenet: Sparse Convolution Framework For Sketch Recognition,
2025
Singapore Management University
Sketch-Sparsenet: Sparse Convolution Framework For Sketch Recognition, Jingru Yang, Jin Wang, Yang Zhou, Guodong Lu, Yu Sun, Huan Yu, Heming Fang, Zhihui Li, Shengfeng He
Research Collection School Of Computing and Information Systems
In free-hand sketch recognition, state-of-the-art methods often struggle to extract spatial features from sketches with sparse distributions, which are characterized by significant blank regions devoid of informative content. To address this challenge, we introduce a novel framework for sketch recognition, termed Sketch-SparseNet. This framework incorporates an advanced convolutional component: the Sketch-Driven Dilated Deformable Block (SD3B). This component excels at extracting spatial features and accurately recognizing free-hand sketches with sparse distributions. The SD3B component innovatively bridges gaps in the blank areas of sketches by establishing spatial relationships among disconnected stroke points through adaptive reshaping of convolution kernels. These kernels are deformable, …
Unresolved Image Simulation For Space Situational Awareness Applications,
2025
Embry-Riddle Aeronautical University
Unresolved Image Simulation For Space Situational Awareness Applications, Fox Coniglario
Doctoral Dissertations and Master's Theses
The knowledge of what lies in orbit around Earth is at best a guess. Decades of spaceflight, debris buildup, and vehicle collisions have contributed to a large number of objects that are simply not able to be catalogued. Ongoing efforts to catalog debris in orbit have reached limits by conventional measures and as such, research is active in the field of in-orbit space situational awareness. This thesis intends to help fill a hole in the development of such orbital platforms by assisting the development of image processing software pipelines though the simulation of unresolved space imagery. The simulation uses accurate …
Look Before You Decide: Prompting Active Deduction Of Mllms For Assumptive Reasoning,
2025
Singapore Management University
Look Before You Decide: Prompting Active Deduction Of Mllms For Assumptive Reasoning, Yian Li, Wentao Tian, Yang Jiao, Jingjing Chen, Tianwen Qian, Bin Zhu, Na Zhao, Yu‑Gang Jiang
Research Collection School Of Computing and Information Systems
Recently, Multimodal Large Language Models (MLLMs) have achieved significant success across multiple disciplines due to their exceptional instruction-following capabilities and extensive world knowledge. However, whether these MLLMs possess human-like compositional reasoning abilities remains an open problem. To unveil their reasoning behaviors, we first curate a Multimodal Assumptive Reasoning Benchmark (MARS-Bench) in this paper. Interestingly, we find that most prevalent MLLMs can be easily fooled by the introduction of a presupposition into the question, whereas such presuppositions appear naive to human reasoning. Besides, we also propose a simple yet effective method, Active Deduction (AD), a novel reinforcement learning paradigm to encourage …
Viewsrd: 3d Visual Grounding Via Structured Multi-View Decomposition,
2025
Singapore Management University
Viewsrd: 3d Visual Grounding Via Structured Multi-View Decomposition, Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
3Dvisual grounding aims to identify and localize objects in a 3Dspacebasedontextualdescriptions. However, existing methods struggle with disentangling targets from anchors in complex multi-anchor queries and resolving inconsisten cies in spatial descriptions caused by perspective variations. To tackle these challenges, we propose ViewSRD, a frame work that formulates 3D visual grounding as a structured multi-view decomposition process. First, the Simple Rela tion Decoupling (SRD) module restructures complex multi anchor queries into a set of targeted single-anchor state ments, generating a structured set of perspective-aware de scriptions that clarify positional relationships. These de composed representations serve as the foundation for the Multi-view …
Omnivton: Training-Free Universal Virtual Try-On,
2025
Singapore Management University
Omnivton: Training-Free Universal Virtual Try-On, Zhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li, Yangyang Xu, Junyu Dong, Yong Du
Research Collection School Of Computing and Information Systems
Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unified, training-free solution that works across both scenarios remains an open challenge. We propose OmniVTON, the first training-free universal VTON framework that decouples garment and pose conditioning to achieve both texture fidelity and pose consistency across diverse settings. To preserve garment details, we introduce a garment prior generation mechanism that aligns clothing with the body, followed by continuous boundary stitching technique to achieve …
Genwardrobe: A Fully Generative System For Travel Fashion Wardrobe Construction,
2025
Singapore Management University
Genwardrobe: A Fully Generative System For Travel Fashion Wardrobe Construction, Peng Jin, Yilin Wen, Mingzhe Yu, Yunshan Ma, Rong Zheng, Jin‑Tu Fan, Chong Wah Ngo
Research Collection School Of Computing and Information Systems
With the increasing demand for outfit planning in real-world travel scenarios, the need for constructing a travel fashion wardrobe, a series of outfits tailored to a user's personalization and destination-specific context over a short travel period, has grown significantly. However, existing systems or works often focus on isolated factors and rely on retrieval-based methods, with insufficient utilization of generative models, limiting their adaptability to real-world travel scenarios. To address this issue, this study introduces GenWardrobe, a fully generative system for travel fashion wardrobe construction. GenWardrobe consists of three key modules: user query analysis, fashion knowledge retrieval via retrieval-augmented generation and …
Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning,
2025
Singapore Management University
Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning, Yangyang Xu, Shengfeng He, Wenqi Shao, Yong Du, Kwan-Yee K. Wong, Yu Qiao, Jun Yu, Ping Luo
Research Collection School Of Computing and Information Systems
In this paper, we introduce DiffusionMat, a novel image matting framework that employs a diffusion model for the transition from coarse to refined alpha mattes. Diverging from conventional methods that utilize trimaps merely as loose guidance for alpha matte prediction, our approach treats image matting as a deterministic sequential refinement learning process. This process begins with the addition of noise to trimaps and iteratively denoises them using a pre-trained diffusion model, which incrementally guides the prediction towards a clean alpha matte. The key innovation of our framework is a correction module that adjusts the output at each denoising step, ensuring …
Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion,
2025
Singapore Management University
Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan Wang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Yi Tan, Xiang Wang
Research Collection School Of Computing and Information Systems
Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional inputs like cooking styles, ingredients, and recipes are utilized. However, a yet-unexplored challenge is generating a sequence of procedural images based on cooking steps from a recipe. This could enhance the cooking experience with visual guidance and possibly lead to an intelligent cooking simulation system. To fill this gap, we introduce a novel task called cooking procedural image generation. This task is inherently demanding, as it strives to create photo-realistic images that align with cooking steps while …
From Holistic To Localized: Local Enhanced Adapters For Efficient Visual Instruction Fine-Tuning,
2025
Singapore Management University
From Holistic To Localized: Local Enhanced Adapters For Efficient Visual Instruction Fine-Tuning, Pengkun Jiao, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yugang Jiang
Research Collection School Of Computing and Information Systems
Efficient Visual Instruction Fine-Tuning (EVIT) seeks to adapt Multimodal Large Language Models (MLLMs) to downstream tasks with minimal computational overhead. However, as task diversity and complexity increase, EVIT faces significant challenges in resolving data conflicts. To address this limitation, we propose the Dual Low-Rank Adaptation (Dual-LoRA), a holistic-to-local framework that enhances the adapter’s capacity to address data conflict through dual structural optimization. Specifically, we utilize two subspaces: a skill space for stable, holistic knowledge retention, and a rank-rectified task space that locally activates the holistic knowledge. Additionally, we introduce Visual Cue Enhancement (VCE), a multi-level local feature aggregation module designed …
Exploring Object Status Recognition For Recipe Progress Tracking In Non-Visual Cooking,
2025
Singapore Management University
Exploring Object Status Recognition For Recipe Progress Tracking In Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington
Research Collection School Of Computing and Information Systems
Cooking plays a vital role in everyday independence and well-being, yet remains challenging for people with vision impairments due to limited support for tracking progress and receiving contextual feedback. Object status — the condition or transformation of ingredients and tools — offers a promising but underexplored foundation for context-aware cooking support. In this paper, we present OSCAR (Object Status Context Awareness for Recipes), a technical pipeline that explores the use of object status recognition to enable recipe progress tracking in non-visual cooking. OSCAR integrates recipe parsing, object status extraction, visual alignment with cooking steps, and time-causal modeling to support real-time …
Teaching Diffusion Models To Ground Alpha Matte,
2025
Singapore Management University
Teaching Diffusion Models To Ground Alpha Matte, Tianyi Xiang, Weiying Zheng, Yutao Jiang, Tingrui Shen, Hewei Yu, Yangyang Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
The power of visual language models is showcased in visual understanding tasks, where language-guided models achieve impressive flexibility and precision. In this paper, we ex tend this capability to the challenging domain of image matting by framing it as a soft grounding problem, enabling a single diffusion model to handle diverse objects, textures, and transparencies, all directed by descriptive text prompts. Our method teaches the diffusion model to ground alpha mattes by guiding it through a process of instance-level localization and transparency estimation. First, we introduce an intermediate objective that trains the model to accurately localize semantic components of the …
