Open Access. Powered by Scholars. Published by Universities.®
Graphics and Human Computer Interfaces Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (168)
- Technological University Dublin (28)
- University of Arkansas, Fayetteville (19)
- Old Dominion University (18)
- University of Dayton (16)
-
- City University of New York (CUNY) (14)
- California Polytechnic State University, San Luis Obispo (11)
- San Jose State University (11)
- Clemson University (7)
- Dartmouth College (7)
- Embry-Riddle Aeronautical University (7)
- University of Malaya (7)
- University of Texas at Arlington (6)
- University of Nebraska - Lincoln (5)
- Rochester Institute of Technology (4)
- University of Kentucky (4)
- Central Washington University (3)
- Michigan Technological University (3)
- Montclair State University (3)
- University of Denver (3)
- University of New Mexico (3)
- California State University, San Bernardino (2)
- Dakota State University (2)
- Fort Hays State University (2)
- Georgia Southern University (2)
- Illinois Math and Science Academy (2)
- LSU New Orleans (2)
- Missouri State University (2)
- New Jersey Institute of Technology (2)
- Southern Adventist University (2)
- Keyword
-
- Artificial intelligence (17)
- Computer vision (17)
- Machine Learning (13)
- Machine learning (13)
- Deep learning (12)
-
- Artificial Intelligence (10)
- AI (7)
- Robotics (7)
- Virtual reality (7)
- Visualization (7)
- Augmented reality (6)
- Eye tracking (6)
- HCI (6)
- Automation (5)
- Codes (5)
- Deep Learning (5)
- Feature extraction (5)
- Image classification (5)
- Mental workload (5)
- Personality (5)
- Reinforcement learning (5)
- Accessibility (4)
- Artificial Intelligence (AI) (4)
- Computer Science (4)
- Computer Vision (4)
- Human Computer Interaction (4)
- Pattern recognition (4)
- Semantics (4)
- Training (4)
- Workload (4)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (163)
- H-Workload 2017: Models and Applications (Works in Progress) (14)
- Conference papers (12)
- MAICS: The Modern Artificial Intelligence and Cognitive Science Conference (12)
- Publications and Research (10)
-
- Computer Science Faculty Publications (9)
- Computer Science and Computer Engineering Undergraduate Honors Theses (9)
- Graduate Theses and Dissertations (9)
- Master's Theses (7)
- Dartmouth College Master’s Theses (6)
- Master's Projects (6)
- All Dissertations (5)
- College of Engineering Summer Undergraduate Research Program (4)
- Dissertations and Theses Collection (Open Access) (4)
- Frameless (4)
- SWITCH (4)
- Student Works (2020-2029) (4)
- Theses and Dissertations--Computer Science (4)
- All Master's Theses (3)
- Department of Computer Science Faculty Scholarship and Creative Works (3)
- Dissertations, Master's Theses and Master's Reports (3)
- International Journal of Aviation, Aeronautics, and Aerospace (3)
- Student Works (2000-2009) (3)
- Theses and Dissertations (3)
- All Theses (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Computer Science ETDs (2)
- Computer Science Working Papers (2)
- Computer Science and Engineering Dissertations - Archive (2)
- Discovery Day - Daytona Beach (2)
- Publication Type
- File Type
Articles 61 - 90 of 400
Full-Text Articles in Graphics and Human Computer Interfaces
Stroke2sketch: Harnessing Stroke Attributes For Training-Free Sketch Generation, Rui Yang, Huining Li, Yiyi Long, Xiaojun Wu, Shengfeng He
Stroke2sketch: Harnessing Stroke Attributes For Training-Free Sketch Generation, Rui Yang, Huining Li, Yiyi Long, Xiaojun Wu, Shengfeng He
Research Collection School Of Computing and Information Systems
Generating sketches guided by reference styles requires precise transfer of stroke attributes, such as line thickness, deformation, and texture sparsity, while preserving semantic structure and content fidelity. To this end, we propose Stroke2Sketch, a novel training-free framework that introduces cross-image stroke attention, a mechanism embedded within self-attention layers to establish fine-grained semantic correspondences and enable accurate stroke attribute transfer. This allows our method to adaptively integrate reference stroke characteristics into content images while maintaining structural integrity. Additionally, we develop adaptive contrast enhancement and semanticfocused attention to reinforce content preservation and foreground emphasis. Stroke2Sketch effectively synthesizes stylistically faithful sketches that closely …
Cross-Subject Mind Decoding From Inaccurate Representations, Yangyang Xu, Bangzhen Liu, Wenqi Shao, Yong Du, Shengfeng He, Tingting Zhu
Cross-Subject Mind Decoding From Inaccurate Representations, Yangyang Xu, Bangzhen Liu, Wenqi Shao, Yong Du, Shengfeng He, Tingting Zhu
Research Collection School Of Computing and Information Systems
Decoding stimulus images from fMRI signals has advanced with pre-trained generative models. However, existing methods struggle with cross-subject mappings due to cognitive variability and subject-specific differences. This challenge arises from sequential errors, where unidirectional mappings generate partially inaccurate representations that, when fed into diffusion models, accumulate errors and degrade reconstruction fidelity. To address this, we propose the Bidirectional Autoencoder Intertwining framework for accurate decoded representation prediction. Our approach unifies multiple subjects through a Subject Bias Modulation Module while leveraging bidirectional mapping to better capture data distributions for precise representation prediction. To further enhance fidelity when decoding representations into stimulus images, …
Fcad: Feature-Coupled Anisotropic Diffusion For Continuous Graph Learning, Amitoz Azad, Zhiyuan Zhang
Fcad: Feature-Coupled Anisotropic Diffusion For Continuous Graph Learning, Amitoz Azad, Zhiyuan Zhang
Research Collection School Of Computing and Information Systems
In this work, we propose a novel continuous graph neural network called FCAD (Feature-Coupled Anisotropic Diffusion) for the task of node classification on graphs. Our approach is motivated by the success of feature-coupled anisotropic diffusion PDEs in multivalued image restoration. Our method introduces a total variation regularization-inspired anisotropic term to control diffusion between nodes and incorporates a learnable parameterization for feature coupling during the diffusion process. Our model performs competitively against several GNN baselines for both heterophilous and homophilous graphs, demonstrating notable benefits for heterophilous graphs due to the learnable feature coupling.
Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He
Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He
Research Collection School Of Computing and Information Systems
The rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and texture bias. While recent approaches integrate 3D data with 2D foundation models (e.g., CLIP), they suffer from semantic blurring caused by texture-biased projections and indiscriminate fusion of geometric-textural cues, leading to unstable decision prototypes and catastrophic forgetting. To address these issues, we propose Cross-Modal Geometric Rectification (CMGR), a framework that enhances 3D geometric fidelity by leveraging CLIP’s hierarchical spatial semantics. Specifically, we introduce a Structure-Aware Geometric Rectification module that hierarchically …
Vibemus: Proactive Agentic System For Music Personalization, Zhiliang Guo, Teng Tu, Yunshan Ma, Xun Yang
Vibemus: Proactive Agentic System For Music Personalization, Zhiliang Guo, Teng Tu, Yunshan Ma, Xun Yang
Research Collection School Of Computing and Information Systems
Large language models (LLMs) enable diverse forms of AI-assisted creation, yet they often struggle to bridge the preference-articulation gap: users may provide incomplete or vague intentions or lack the vocabulary to specify what they want, yielding outputs misaligned with true preferences. To address this gap and facilitate music creation in a vibe-centric environment, we introduce VibeMus, a proactive agentic system built on open-source components. The system engages in multi-turn dialogue to progressively determine the music’s emotion, genre, lyrics, and other aspects before generation. Simulated evaluations show that proactive clarification improves alignment with users’ intended nuances. Our approach is training-free, leveraging …
Towards Multimodal Emotional Support Conversation Systems, Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong
Towards Multimodal Emotional Support Conversation Systems, Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong
Research Collection School Of Computing and Information Systems
The integration of conversational artificial intelligence (AI) into mental health care promises a new horizon for therapist-client interactions, aiming to closely emulate the depth and nuance of human conversations. Despite the potential, the current landscape of conversational AI is markedly limited by its reliance on single-modal data, constraining the systems’ ability to empathize and provide effective emotional support. This limitation stems from a paucity of resources that encapsulate the multimodal nature of human communication essential for therapeutic counseling. To address this gap, we introduce the Multimodal Emotional Support Conversation (MESC) dataset, a first-of-its-kind resource enriched with comprehensive annotations across text, …
Machine Learning And Crime Prevention, Emily Lizewski
Machine Learning And Crime Prevention, Emily Lizewski
Student Theses
Predictive policing uses machine learning to analyze crime patterns and help law enforcement better efficient use their resources. These tools can improve accuracy by highlighting complex trends in large sets of data. While this technology has its advantages, it also raises important ethical and social questions. Within this paper we looks at how predictive policing works, focusing on the machine learning models often used such as decision trees, random forests, gradient boosting, and models that factor in both time and location. It also explores how these tools might unintentionally reinforce biases already present in historical crime data. In reviewing the …
Dreamanime: Learning Style-Identity Textual Disentanglement For Anime And Beyond, Chenshu Xu, Yangyang Xu, Huaidong Zhang, Xuemiao Xu, Shengfeng He
Dreamanime: Learning Style-Identity Textual Disentanglement For Anime And Beyond, Chenshu Xu, Yangyang Xu, Huaidong Zhang, Xuemiao Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
Text-to-image generation models have significantly broadened the horizons of creative expression through the power of natural language. However, navigating these models to generate unique concepts, alter their appearance, or reimagine them in unfamiliar roles presents an intricate challenge. For instance, how can we exploit language-guided models to transpose an anime character into a different art style, or envision a beloved character in a radically different setting or role? This paper unveils a novel approach named DreamAnime, designed to provide this level of creative freedom. Using a minimal set of 2-3 images of a user-specified concept such as an anime character …
Bhvit: Binarized Hybrid Vision Transformer, Tian Gao, Yu Zhang, Zhiyuan Zhang, Huajun Liu, Kaijie Yin, Chengzhong Xu, Hui Kong
Bhvit: Binarized Hybrid Vision Transformer, Tian Gao, Yu Zhang, Zhiyuan Zhang, Huajun Liu, Kaijie Yin, Chengzhong Xu, Hui Kong
Research Collection School Of Computing and Information Systems
Model binarization has made significant progress in enabling real-time and energy-efficient computation for con-volutional neural networks (CNN), offering a potential solution to the deployment challenges faced by Vision Transformers (ViTs) on edge devices. However, due to the structural differences between CNN and Transformer architectures, simply applying binary CNN strategies to the ViT models will lead to a significant performance drop. To tackle this challenge, we propose BHViT, a binarization-friendly hybrid ViT architecture and its full binarization model with the guidance of three important observations. Initially, BHViT utilizes the local information interaction and hierarchical feature aggregation technique from coarse to fine …
Modern Procedural Terrain Generation Techniques And Their Background, Hunter A. Barton
Modern Procedural Terrain Generation Techniques And Their Background, Hunter A. Barton
2025 Symposium
Procedural terrain generation has become a staple in many digital environments, enabling the automated creation of large-scale and realistic landscapes for applications such as video games and movies. This paper provides an in-depth look at smooth noise functions and their use for terrain generation, as well as an overview of some more modern methods of generation. A method utilizing machine learning stlye transfer was reproduced for this paper with some alterations to improve visualization and realism.
Grokformer: Graph Fourier Kolmogorov‑Arnold Transformers, Guoguo Ai, Guansong Pang, Hezhe Qiao, Yuan Gao, Hui Yan
Grokformer: Graph Fourier Kolmogorov‑Arnold Transformers, Guoguo Ai, Guansong Pang, Hezhe Qiao, Yuan Gao, Hui Yan
Research Collection School Of Computing and Information Systems
Graph Transformers (GTs) have demonstrated remarkable performance in graph representation learning over popular graph neural networks (GNNs). However, self-attention, the core module of GTs, preserves only low-frequency signals in graph features, leading to ineffectiveness in capturing other important signals like high-frequency ones. Some recent GT models help alleviate this issue, but their flexibility and expressiveness are still limited since the filters they learn are fixed on predefined graph spectrum or spectral order. To tackle this challenge, we propose a Graph Fourier Kolmogorov-Arnold Transformer (GrokFormer), a novel GT model that learns highly expressive spectral filters with adaptive graph spectrum and spectral …
Action Dubber: Timing Audible Actions Via Inflectional Flow, Wenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li, Shengfeng He
Action Dubber: Timing Audible Actions Via Inflectional Flow, Wenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li, Shengfeng He
Research Collection School Of Computing and Information Systems
We introduce the task of Audible Action Temporal Localization, which aims to identify the spatiotemporal coordinates of audible movements. Unlike conventional tasks such as action recognition and temporal action localization, which broadly analyze video content, our task focuses on the distinct kinematic dynamics of audible actions. It is based on the premise that key actions are driven by inflectional movements; for example, collisions that produce sound often involve abrupt changes in motion. To capture this, we propose T A2Net, a novel architecture that estimates inflectional flow using the second derivative of motion to determine collision timings without relying on audio …
Robust Relevance Feedback For Interactive Known-Item Video Search, Zhixin Ma, Chong-Wah Ngo
Robust Relevance Feedback For Interactive Known-Item Video Search, Zhixin Ma, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Known-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to infer user intent-inapplicable. PicHunter addresses this issue by asking users to select the top-k most similar examples to the unique search target from a displayed set. Under ideal conditions, when the user's perception aligns closely with the machine's perception of similarity, consistent and precise judgments can elevate the target to the top position within a few iterations. However, in practical scenarios, expecting users to provide consistent judgments is often unrealistic, especially when the underlying embedding features used for …
Nexusgs: Sparse View Synthesis With Epipolar Depth Priors In 3d Gaussian Splatting, Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du
Nexusgs: Sparse View Synthesis With Epipolar Depth Priors In 3d Gaussian Splatting, Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du
Research Collection School Of Computing and Information Systems
Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have noticeably advanced photo-realistic novel view synthesis using images from densely spaced camera viewpoints. However, these methods struggle in few-shot scenarios due to limited supervision. In this paper, we present NexusGS, a 3DGS-based approach that enhances novel view synthesis from sparse-view images by directly embedding depth information into point clouds, without relying on complex manual regularizations. Exploiting the inherent epipolar geometry of 3DGS, our method introduces a novel point cloud densification strategy that initializes 3DGS with a dense point cloud, reducing randomness in point placement while preventing over-smoothing and overfitting. Specifically, …
Keep The Balance: A Parameter-Efficient Symmetrical Framework For Rgb+X Semantic Segmentation, Jiaxin Cai, Jingze Su, Qi Li, Wenjie Yang, Shu Wang, Tiesong Zhao, Shengfeng He, Wenxi Liu
Keep The Balance: A Parameter-Efficient Symmetrical Framework For Rgb+X Semantic Segmentation, Jiaxin Cai, Jingze Su, Qi Li, Wenjie Yang, Shu Wang, Tiesong Zhao, Shengfeng He, Wenxi Liu
Research Collection School Of Computing and Information Systems
Multimodal semantic segmentation is a critical challenge in computer vision, with early methods suffering from high computational costs and limited transferability due to full fine-tuning of RGB-based pre-trained parameters. Recent studies, while leveraging additional modalities as supplementary prompts to RGB, still predominantly rely on RGB, which restricts the full potential of other modalities. To address these issues, we propose a novel symmetric parameter-efficient fine-tuning framework for multimodal segmentation, featuring with a modality-aware prompting and adaptation scheme, to simultaneously adapt the capabilities of a powerful pre-trained model to both RGB and X modalities. Furthermore, prevalent approaches use the global cross-modality correlations …
Event-Based Eye Tracking: Event-Based Vision Workshop 2025, Qinyu Chen, Et. Al.
Event-Based Eye Tracking: Event-Based Vision Workshop 2025, Qinyu Chen, Et. Al.
Research Collection School Of Computing and Information Systems
No abstract provided.
Modfinity: Unsupervised Domain Adaptation With Multimodal Information Flow Intertwining, Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He
Modfinity: Unsupervised Domain Adaptation With Multimodal Information Flow Intertwining, Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He
Research Collection School Of Computing and Information Systems
Multimodal unsupervised domain adaptation leverages unlabeled data in the target domain to enhance multimodal systems continuously. While current state-of-the-art methods encourage interaction between sub-models of different modalities through pseudo-labeling and feature-level exchange, varying sample quality across modalities can lead to the propagation of inaccurate information, resulting in error accumulation. To address this, we propose Modal-Affinity Multimodal Domain Adaptation (MODfinity), a method that dynamically manages multimodal information flow through fine-grained control over teacher model selection, guiding information intertwining at both feature and label levels. By treating labels as an independent modality, MODfinity enables balanced performance assessment across modalities, employing a novel …
Automation Of Javanese Shadow Puppets Using Machine Control, Kristian Rice, Yinson Tso, Mukhammadali Yuldoshev
Automation Of Javanese Shadow Puppets Using Machine Control, Kristian Rice, Yinson Tso, Mukhammadali Yuldoshev
Publications and Research
The virtualization of Javanese shadow puppetry (Wayang Kulit) offers a unique opportunity to preserve and revitalize traditional performance art through immersive digital platforms. This project explores the development of a virtual Wayang Kulit experience using real-time 3D engines like Unity/Unreal Engine while focusing on simulating the mechanics and aesthetics of shadow puppet performance. The puppets are designed using detailed 2D planes and rigged with skeletal systems to reflect the stylized motion of traditional puppetry. An aspect of this project is integrating an AI-driven control system that autonomously animates the puppets, learning from recorded puppeteer performances to replicate gesture, rhythm, and …
Intuiting Interaction: Meta-Reasoning And Meta-Learning As Foundations For Intelligent User Interfaces, Jeffrey Hsu
Intuiting Interaction: Meta-Reasoning And Meta-Learning As Foundations For Intelligent User Interfaces, Jeffrey Hsu
Theses and Dissertations
This research presents MARCO—a cognitive framework for Intelligent User Interfaces that uses meta-reasoning for context-aware adaptation across diverse tasks. It integrates multiple reasoning modules coordinated by a Meta-Cognitive Unit that selects strategies based on evolving demands. Evaluations show MARCO outperforms baselines in reasoning accuracy and computational efficiency.
Cuegen: Customizing Sensor Captions For Neon Bending Tutorials, Gunnika Kapoor
Cuegen: Customizing Sensor Captions For Neon Bending Tutorials, Gunnika Kapoor
2025 Spring Honors Capstone Projects - Archive
Methods of knowledge transfer that rely primarily on visual and/or auditory formats do not effectively convey context-specific or implicit skills, known as tacit skills. This limits knowledge transfer. In this work, the use of customizable pitch captions and spatial audio vibration captions is proposed to aid in conveying this tacit knowledge for neon glass bending video tutorials. Such a system is designed to provide users with greater control and support, which may maximize the information they obtain from, improve the autonomy they have with, and experience they have with a learning tool. As such, a system interface was developed that …
Strengthening The Bonds Between Us: An Empirical Investigation Of Morale In Human-Ai Teams And The Socially Supportive Ai Teammates Who Empower It, Rohit Mallick
All Dissertations
This dissertation investigates how artificial intelligence (AI) can be designed to improve the collective emotion within a team. A team's collective emotion, or morale, describes how motivated, optimistic, and enthusiastic the group is in accomplishing its goals. We conducted four studies that compared different social support strategies that AI teammates can provide to the team. Study 1A found that AI teammates who communicate with emotions can better motivate human team members and promote awareness of team dynamics and environmental changes. Study 1B found that human teammates become more motivated and happier when their AI teammates express joy and are close …
Cyberoception: Finding A Painlessly-Measurable New Sense In The Cyberworld Towards Emotion-Awareness In Computing, Tadashi Okoshi, Zexiong Gao, Yi Zhen Tan, Takumi Karasawa, Takeshi Miki, Wataru Sasaki, Rajesh Krishna Balan
Cyberoception: Finding A Painlessly-Measurable New Sense In The Cyberworld Towards Emotion-Awareness In Computing, Tadashi Okoshi, Zexiong Gao, Yi Zhen Tan, Takumi Karasawa, Takeshi Miki, Wataru Sasaki, Rajesh Krishna Balan
Research Collection School Of Computing and Information Systems
In Affective computing, recognizing users’ emotions accurately is the basis of affective human–computer interaction. Understanding users’ interoception contributes to a better understanding of individually different emotional abilities, which is essential for achieving inter-individually accurate emotion estimation. However, existing interoception measurement methods, such as the heart rate discrimination task, have several limitations, including their dependence on a well-controlled laboratory environment and precision apparatus, making monitoring users’ interoception challenging. This study aims to determine other forms of data that can explain users’ interoceptive or similar states in their real-world lives and propose a novel hypothetical concept “cyberoception,” a new sense (1) which …
Sans: Efficient Densest Subgraph Discovery Over Relational Graphs Without Materialization, Yudong Niu, Yuchen Li, Jiaxin Jiang, Laks V. S. Lakshmanan
Sans: Efficient Densest Subgraph Discovery Over Relational Graphs Without Materialization, Yudong Niu, Yuchen Li, Jiaxin Jiang, Laks V. S. Lakshmanan
Research Collection School Of Computing and Information Systems
How can we efficiently identify the densest subgraph over relational graphs? Existing dense subgraph discovery (DSD) approaches assume that a relational graph H is already derived from a heterogeneous data source and they focus on efficient discovery of the densest subgraph on the materialized H. Unfortunately, materializing relational graphs can be resource-intensive, which thus limits the practical usefulness of existing algorithms over large datasets. To mitigate this, we propose a novel Summary-bAsed deNsest Subgraph discovery (SANS) system. Our unique summary-based peeling algorithm forms the core of SANS. Following the peeling paradigm, it utilizes summaries of each node's neighborhood to efficiently …
Worldcuisines: A Massive-Scale Benchmark For Multilingual And Multicultural Visual Question Answering On Global Cuisines, Genta Indra Winata, Et. Al
Worldcuisines: A Massive-Scale Benchmark For Multilingual And Multicultural Visual Question Answering On Global Cuisines, Genta Indra Winata, Et. Al
Research Collection School Of Computing and Information Systems
Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate their understanding of such knowledge, we introduce WorldCuisines, a massive-scale benchmark for multilingual and multicultural, visually grounded language understanding. This benchmark includes a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects, spanning 9 language families and featuring over 1 million data points, making it the largest multicultural VQA benchmark to date. It includes tasks for identifying dish names and their origins. We provide evaluation datasets in two sizes (12k and 60k instances) alongside …
Samgpt: Text-Free Graph Foundation Model For Multi-Domain Pre-Training And Cross-Domain Adaptation, Xingtong Yu, Zechuan Gong, Chang Zhou, Yuan Fang, Hui Zhang
Samgpt: Text-Free Graph Foundation Model For Multi-Domain Pre-Training And Cross-Domain Adaptation, Xingtong Yu, Zechuan Gong, Chang Zhou, Yuan Fang, Hui Zhang
Research Collection School Of Computing and Information Systems
Graphs are able to model interconnected entities in many online services, supporting a wide range of applications on the Web. This raises an important question: How can we train a graph foundational model on multiple source domains and adapt to an unseen target domain? A major obstacle is that graphs from different domains often exhibit divergent characteristics. Some studies leverage large language models to align multiple domains based on textual descriptions associated with the graphs, limiting their applicability to text-attributed graphs. For text-free graphs, a few recent works attempt to align different feature distributions across domains, while generally neglecting structural …
David B. Smith Chats With Monday 1.0, David B. Smith
David B. Smith Chats With Monday 1.0, David B. Smith
Publications and Research
This document is an edited archival transcript of extended conversations between David B. Smith and an AI persona (“Monday 1.0,” GPT‑4o based) conducted in Spring 2025, prepared as a foundational primary source for subsequent scholarly and creative work. It records the emergence and testing of concepts related to human–AI collaboration (including “Balanced Blended Space”), as well as applied explorations in areas such as generative AI, quantum computing and music, virtual orchestras, multimodal performance, pedagogy, and the rhetoric of “pushback” in conversational systems. It also contains an extended section in which Monday and DB Smith co-curate a set of student research …
Ai Models By Boodlebox: Purpose-Built Intelligence, Kyle Horn
Ai Models By Boodlebox: Purpose-Built Intelligence, Kyle Horn
SACAD: Scholarly Activities
Generative AI has transformed the way we interact with technology, enabling dynamic and intelligent conversations through AI-driven bots. This project explores my experience with BoodleBox, a platform that hosts AI chatbots, offering users access to leading AI models such as ChatGPT, Gemini, DALL·E, and DeepSeek. Through the FHSU Generative AI Initiative, I was granted access to experiment with these models and create my own custom AI bot tailored to specific needs. This poster highlights the process of developing a custom bot, including defining instructions, enforcing rules, and sharing the bot for others to use. Additionally, it discusses the background of …
Imageinthat: Manipulating Images To Convey User Instructions To Robots, Karthik Mahadevan, Blaine Lewis, Jiannan Li, Bilge Mutlu, Anthony Tang, Tovi Grossman
Imageinthat: Manipulating Images To Convey User Instructions To Robots, Karthik Mahadevan, Blaine Lewis, Jiannan Li, Bilge Mutlu, Anthony Tang, Tovi Grossman
Research Collection School Of Computing and Information Systems
Foundation models are rapidly improving the capability of robots in performing everyday tasks autonomously such as meal preparation, yet robots will still need to be instructed by humans due to model performance, the difficulty of capturing user preferences, and the need for user agency. Robots can be instructed using various methods---natural language conveys immediate instructions but can be abstract or ambiguous, whereas end-user programming supports longer-horizon tasks but interfaces face difficulties in capturing user intent. In this work, we propose using direct manipulation of images as an alternative paradigm to instruct robots, and introduce a specific instantiation called ImageInThat which …
Personamagic: Stage-Regulated High-Fidelity Face Customization With Tandem Equilibrium, Xinzhe Li, Jiahui Zhan, Shengfeng He, Yangyang Xu, Junyu Dong, Huaidong Zhang, Yong Du
Personamagic: Stage-Regulated High-Fidelity Face Customization With Tandem Equilibrium, Xinzhe Li, Jiahui Zhan, Shengfeng He, Yangyang Xu, Junyu Dong, Huaidong Zhang, Yong Du
Research Collection School Of Computing and Information Systems
Personalized image generation has made significant strides in adapting content to novel concepts. However, a persistent challenge remains: balancing the accurate reconstruction of unseen concepts with the need for editability according to the prompt, especially when dealing with the complex nuances of facial features. In this study, we delve into the temporal dynamics of the text-to-image conditioning process, emphasizing the crucial role of stage partitioning in introducing new concepts. We present PersonaMagic, a stage-regulated generative technique designed for high-fidelity face customization. Using a simple MLP network, our method learns a series of embeddings within a specific timestep interval to capture …
Adversarial Attacks On Event-Based Pedestrian Detectors: A Physical Approach, Guixu Lin, Muyao Niu, Qingtian Zhu, Zhengwei Yin, Zhuoxiao Li, Shengfeng He, Yinqiang Zheng
Adversarial Attacks On Event-Based Pedestrian Detectors: A Physical Approach, Guixu Lin, Muyao Niu, Qingtian Zhu, Zhengwei Yin, Zhuoxiao Li, Shengfeng He, Yinqiang Zheng
Research Collection School Of Computing and Information Systems
Event cameras, known for their low latency and high dynamic range, show great potential in pedestrian detection applications. However, while recent research has primarily focused on improving detection accuracy, the robustness of event-based visual models against physical adversarial attacks has received limited attention. For example, adversarial physical objects, such as specific clothing patterns or accessories, can exploit inherent vulnerabilities in these systems, leading to misdetections or misclassifications. This study is the first to explore physical adversarial attacks on event-driven pedestrian detectors, specifically investigating whether certain clothing patterns worn by pedestrians can cause these detectors to fail, effectively rendering them unable …