Open Access. Powered by Scholars. Published by Universities.®

Artificial Intelligence and Robotics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1 - 30 of 400

Full-Text Articles in Graphics and Human Computer Interfaces

From Data To Decision-Making: The Role Of Local Digital Twins In Cross-Domain Management Within Municipalities – A Research-In-Progress Study In Veenendaal, Diana M.E. Boekman, Koen Smit, Guido Ongena, Rob Peters Sep 2026

From Data To Decision-Making: The Role Of Local Digital Twins In Cross-Domain Management Within Municipalities – A Research-In-Progress Study In Veenendaal, Diana M.E. Boekman, Koen Smit, Guido Ongena, Rob Peters

Communications of the IIMA

Municipalities are facing increasingly complex, interconnected challenges in areas like housing, climate adaptation, mobility, and social policy. Local Digital Twins (LDTs) are seen as a promising tool to make this complexity more understandable and support decision-making. At the same time, both literature and practice show that few initiatives get past the pilot phase, even though getting through that phase is essential for successful long-term adoption.

This paper presents a research-in-progress study on the development and application of an implementation method for LDT technology within the municipality of Veenendaal, based on human values rather than driven by technological possibilities. Based on …


Restoring Linguistic Grounding In Vla Models Via Train-Free Attention Recalibration, Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen Sep 2026

Restoring Linguistic Grounding In Vla Models Via Train-Free Attention Recalibration, Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen

Research Collection School Of Computing and Information Systems

Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generalist robotic policies. However, their reliability under Out-Of-Distribution (OOD) instructions remains underexplored. In this paper, we reveal a critical failure mode in which VLA policies continue executing visually plausible actions even when the language instruction contradicts the scene. We refer to this phenomenon as linguistic blindness, where VLA policies prioritize visual priors over instruction semantics during action generation. To systematically analyze this issue, we introduce ICBench, a diagnostic benchmark constructed from the LIBERO dataset that probes language–action coupling …


Ai And The Environment: Solutions For Advancing Technology Safely, Kaitlynn Baker Aug 2026

Ai And The Environment: Solutions For Advancing Technology Safely, Kaitlynn Baker

Discovery Day - Daytona Beach

Since 2022, the world of Artificial Intelligence (AI) has boomed. AI went from a special and rare entity to a commonly used resource available to all through web sites, and phone apps. AI has benefitted everyday activities by making office, class, and personal tasks easier through grammar help, informational citations, and as someone to bounce ideas off of. Additionally, many companies have begun utilizing AI to improve customer service and experience, and train workers more efficiently, therefore, saving thousands of dollars. Despite the benefits humans reap from its use, AI has been harming our environment at growing rates. Data centers …


How Ai Influences The Design Process Of Unmanned Underwater Vehicles’ (Uuvs) 3d Sonar System, Eden Tsouklaris, Abriella Smith, Brianna Broderick, Carissa Aumack, Victoria Cornaro Aug 2026

How Ai Influences The Design Process Of Unmanned Underwater Vehicles’ (Uuvs) 3d Sonar System, Eden Tsouklaris, Abriella Smith, Brianna Broderick, Carissa Aumack, Victoria Cornaro

Discovery Day - Daytona Beach

With the exponential growth of Artificial Intelligence (AI), user interface (UI) designers have explored using AI to shorten design time. This study assessed the effectiveness of UIs designed with AI programs versus manual methods for an Unmanned Underwater Vehicle (UUV) control system. Participants were tasked with designing an interface that would allow submarine operators to monitor and coordinate three UUVs repairing a severed underwater communication cable at a depth of 2,000 meters. The scenario presented several operational challenges (zero visibility, sonar-only perception, data latency, and potential system degradation), requiring participants' designs to maintain spatial awareness and support remote repair tasks. …


Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo Jul 2026

Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo

Research Collection School Of Computing and Information Systems

Rapid advances in AI-generated image (AIGI) technology enable highly realistic synthesis, threatening public information integrity and security. Recent studies have demonstrated that incorporating texture-level artifact features alongside semantic features into multimodal large language models (MLLMs) can enhance their AIGI detection capability. However, our preliminary analyses reveal that artifact features exhibit high intra-feature similarity, leading to an almost uniform attention map after the softmax operation. This phenomenon causes attention dilution, thereby hindering effective fusion between semantic and artifact features. To overcome this limitation, we propose a lightweight fusion adapter, TranX-Adapter, which integrates a Task-aware Optimal-Transport Fusion that leverages the Jensen-Shannon divergence …


Chi Meta-Project Ecosystem Overview - Spring 2026, David B. Smith Jul 2026

Chi Meta-Project Ecosystem Overview - Spring 2026, David B. Smith

Publications and Research

This paper offers a high-level account of the Center for Holistic Integration’s (CHI) meta-project ecosystem as visualized in the included system map. CHI provides an organizational structure framed around persistent meta-projects that support and extend individual initiatives across curriculum, scholarly and applied research, infrastructure, artistic production, AI development, cultural inquiry, and external partnerships. Rather than presenting the map as a static inventory of projects, the paper examines how its core domains function as living systems through which knowledge, tools, documentation, participants, and collaborations can accumulate over time. It also considers how CHI-mediated connectivity, institutional integration, and external funding allow the …


Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen Jul 2026

Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen

Research Collection School Of Computing and Information Systems

Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object’s state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action–object interactions into …


Enhancing Pointing Gestures Of Non-Hmd Users In Asymmetric Collocated Mixed Reality Collaboration, Nam-Dang Vo, Van-Vinh Thai, Anthony Tang, Khanh-Duy Le Jun 2026

Enhancing Pointing Gestures Of Non-Hmd Users In Asymmetric Collocated Mixed Reality Collaboration, Nam-Dang Vo, Van-Vinh Thai, Anthony Tang, Khanh-Duy Le

Research Collection School Of Computing and Information Systems

A common collocated group setting in mixed-reality (MR) collaboration is a person wearing a MR headset (HMD user) and presenting MR contents to audiences who are not provided with such specialized devices (Non-HMD users). In this setting, while Non-HMD users can view the MR environment shown on a large physical display, it still remains challenging for the HMD user to interpret their pointing gesture when they spatially refer to objects in the MR environment. To address this, we designed and evaluated two pointing techniques—SCREEN and SCREEN+SPACE—that support Non-HMD users in referring to MR content. Screen pointing allows users to refer …


Videocreator: An Agentic System For Multi-Turn Video Production, Zhengyang Liang, Yan Shu, Cathal Gurrin, Nicu Sebe, Lizi Liao Jun 2026

Videocreator: An Agentic System For Multi-Turn Video Production, Zhengyang Liang, Yan Shu, Cathal Gurrin, Nicu Sebe, Lizi Liao

Research Collection School Of Computing and Information Systems

Recent advances in video generation models enable visually compelling single clips. However, real-world video creation is inherently continuous and iterative: creators refine content over multiple rounds while maintaining narrative, style, and entity consistency. Existing standalone generators are largely stateless and lack memory of previously generated segments, making it difficult to produce a coherent and consistent video project. To address this gap, we present VideoCreator, a unified video agent that integrates generation and understanding with a project-level memory system. VideoCreator leverages understanding capabilities to perform fine-grained analysis of newly produced content and uses persistent memory to retain and reuse prior context …


Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang Jun 2026

Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang

Research Collection School Of Computing and Information Systems

Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we …


History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu Jun 2026

History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu

Research Collection School Of Computing and Information Systems

Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution in addressing the drawbacks of poor generalization in the training-based paradigm. However, existing LLM-based works primarily perform naive reasoning for decision-making and lack feedback, e.g., reviewing historical errors and predicting future potentials. Consequently, it may suffer from continuous failure for those initial error tasks. In this paper, we rethink LLM-based zero-shot VLN-CE and propose a new paradigm, named EvoNav, to improve …


Ai Interview Helper: A Tool For Assisting Search And Rescue Long-Profile Interviews, Dylan P. Starink Jun 2026

Ai Interview Helper: A Tool For Assisting Search And Rescue Long-Profile Interviews, Dylan P. Starink

Master's Theses

In Search and Rescue (SAR) operations, time pressure and limited interviewer experience can lead to missed opportunities when interviewing a missing person’s friends and family. This thesis presents a real-time, end-to-end system that provides context-aware follow-up question suggestions as interviews unfold. Leveraging large language models (LLMs) and agentic design patterns, the system is intended to support interviewers by helping them identify relevant follow-up questions and pursue potentially overlooked lines of inquiry.

The system was evaluated through three mock interviews with two SAR interviewer participants across two events. Given the limited sample size, the results provide early insights into the feasibility …


Toward Learning-Based Reconstruction And Part Decomposition Of Man-Made 3d Geometry: Neural Implicit Representations And Scalable Supervision, Shen Fan May 2026

Toward Learning-Based Reconstruction And Part Decomposition Of Man-Made 3d Geometry: Neural Implicit Representations And Scalable Supervision, Shen Fan

Dissertations

Digital three-dimensional (3D) models are central to engineering design, analysis, and manufacturing, but learning pipelines for man-made geometry often operate on sampled carriers that do not preserve all of the structure present in exact CAD representations. This dissertation studies learning-based reconstruction and part decomposition for structured man-made 3D geometry, from general object benchmarks to CAD-derived datasets, with a focus on neural implicit representations trained from signed-distance samples, point clouds, and tessellated meshes. The goal is to make these models more accurate, more part-aware, and more consistently supervised.

First, signed distance function (SDF) reconstruction with implicit neural representations is improved through …


Du Undergraduate Showcase Abstracts: Research, Scholarship, And Creative Works, Sophia Wismar, Henry Staats, Allison Metzler, Chloe Puckett, Rachel Levine, Christa Kilpatrick, Scott Wolf, Joe Walsh, Grace Doolittle, John Engebreston, Zoe Lopez, Christopher Aaby, Audrey Duff, Timothy Sisk, Katelyn Lamberton, Angela Narayan, Gilkah Argueta, Habiba Samir, Girena Tesfazghi, Genet Kenore, Sinit Tesfamariam, Effley Brooks, Abi Newell, Megan Doherty, Natalie Baer, Lexi Blood, Talya Riciputi, Jessica Jimenez, Devin Hernandez, Lynn Clark, Taj Kumar, Sunil Kumar, Allen Rutman, Mira Pronobis, Tess Carson, Anna Sher, Frankie Stroud, Tamra Pearson D'Estree, Alyssa Wilson, Emily Melnick, Jenalee Doom, Yihang Gao, Gwendolyn Geiger, Noah Gettle, Scott Nichols, Clare Ayoub, Cara Dienno, Sunny Walker, Zoe Hansen, Maya Wheeler, Addison Rice, Patrick Martin, Sanjana Acharya, Daniel Mcintosh, Amanda Mckellips, Calli Cain, Justin Blake, Peter Sokol-Hessner, Natalie Miller, Max Weisbuch, Sophia Dellota, John Macikas, Charlotte Snow, Mark Siemens, Zoe Lynch, Alex Huffman, Prachi Shah, Jason Roney, Halcyon Levi, Nicole Herzog, Andrea Koly, Daniel Linseman, Annie London, Xi Yang, Avery Zwisler, Jane Smith, Chaz Contag, Michael Kerwin, Lucy Rand, Grace Schroeder, Michelle Rozenman, Nissa Tapper, Guiming Zhang, Mateo Mazariego-Halpern, Keith Meyer, Julie Do, Dakota Park-Ozee, Travis Herink, Kara Neu, Jonathan Plomin, Eve-Odine Duchaufour, Debbie Gale Mitchell, Tennyson Anderson-Stricklin, Lily Treitz, Samantha Rosenberger, Sierra Griffith, Finley Joseph, Daniel Sampson, Emmy Davis, Skyler Kasnoff, Evon Lopez, Vivian Nguyen, Cassy Young, Franklin Sellner, Martin Tobon, Ila Graham, Zach Billings, Holden Hedit, Decatur Boland, Paul Kosempel, Cory Chandler, Jay Mahoney, Sam Dragan, Susan Dagget, Yarrow Ator, Heidi Vuletich, Owen Weber, Andrew Kloeppel, Petersen Gray, Mandi Schaeffer-Fry, Razleen Bassra, Bryanna Rodriguez, Christina Blue, Taubie Sanders, Rachel Epstein, Luke Milburn, Camryn Evans, Ezra Martinez, Mary Westwood, Gabri Notov, Robin Tinghitella, Lilou Cabrol, Eli Barbour, Juliet Mendik, Selma Myers, Zac Wise, Noah Fahlin, Michelle Knowles, Abigail Hopper, Michael Greenberger, Romi Laclair, Sarah Watamura, Sabrina Efroymson, Casey Barker, Sydney Seltzer, Bryn Yehle, Jennifer Hoffman, Sara Garcia, Ryuka Nagamine, Trevor Briggs, Remy Le Boeuf, Elena Krone, Eileen Farrell, Regan O'Rourke, Elena Roel, Greg Mortimer, Ali Ayoub, Stefani Langehennig, Caitlin Turk, Logan Scmid, Stefan Chavez-Norgaard, Karen Kim, Tatiana Peccedi, Courtney Cassidy, John Sebesta, Rhianna Lewis, Janice Bening-Lacek, Vivian Lawless, Mckenna Hanson, Jeffrey Amidon, Riya Joshi, Ram Ambre, Brady Worrell, Perrin Schneider, Ali Azadani, Brooke Agulnek, Lyndsie Salvagio, Elise Siemanowki, Yan Qin, Andre Allen, Melodie Nguyen, Megan Livengood, Abby Reams, Saffron Hartreeve, Bri Wylie, Sarah Brookman, Mariah Loiacono, Green Russo, Abhia Lodhi, Gabrielle Welsh, Nika Spehar, Shahked Levin, Evrim Baykal, Kimberly Chiew, Jocelyn Torres, Kailey Hicks, Mykaela Tanino-Springsteen, Audrey Bellows, Akam Chahal, Madeline Tepper, Shannon Murphy, Alexa Fonseca, Deborah Han, Cassandra Perez, Oluwatoyin Alaba, Julia Roncoroni, Vy Nguyen, Nana Burn, Sarah Sasse, Rubin Tuder, Anthony Gerber, Nancy Lorenzon, Christine Vohwinkel, Camryn Gunter, Tristan Weber, Sam Rommel, Brian Michel, Muskan Fatima, Alannah Oleson, Kira Frey, Edward Garrido, Beckett Morris, Kerstin Haring, Drew Middleton, Abigail Walpert, Liam Dee, Gabby Ishaw, Cole Carnes, Maddie Weiser, Claire Fox, Valeriia Vlasenko, Kateri Mcrae, Riley Smith, Abigail Templin, Kushani Rajapaksha May 2026

Du Undergraduate Showcase Abstracts: Research, Scholarship, And Creative Works, Sophia Wismar, Henry Staats, Allison Metzler, Chloe Puckett, Rachel Levine, Christa Kilpatrick, Scott Wolf, Joe Walsh, Grace Doolittle, John Engebreston, Zoe Lopez, Christopher Aaby, Audrey Duff, Timothy Sisk, Katelyn Lamberton, Angela Narayan, Gilkah Argueta, Habiba Samir, Girena Tesfazghi, Genet Kenore, Sinit Tesfamariam, Effley Brooks, Abi Newell, Megan Doherty, Natalie Baer, Lexi Blood, Talya Riciputi, Jessica Jimenez, Devin Hernandez, Lynn Clark, Taj Kumar, Sunil Kumar, Allen Rutman, Mira Pronobis, Tess Carson, Anna Sher, Frankie Stroud, Tamra Pearson D'Estree, Alyssa Wilson, Emily Melnick, Jenalee Doom, Yihang Gao, Gwendolyn Geiger, Noah Gettle, Scott Nichols, Clare Ayoub, Cara Dienno, Sunny Walker, Zoe Hansen, Maya Wheeler, Addison Rice, Patrick Martin, Sanjana Acharya, Daniel Mcintosh, Amanda Mckellips, Calli Cain, Justin Blake, Peter Sokol-Hessner, Natalie Miller, Max Weisbuch, Sophia Dellota, John Macikas, Charlotte Snow, Mark Siemens, Zoe Lynch, Alex Huffman, Prachi Shah, Jason Roney, Halcyon Levi, Nicole Herzog, Andrea Koly, Daniel Linseman, Annie London, Xi Yang, Avery Zwisler, Jane Smith, Chaz Contag, Michael Kerwin, Lucy Rand, Grace Schroeder, Michelle Rozenman, Nissa Tapper, Guiming Zhang, Mateo Mazariego-Halpern, Keith Meyer, Julie Do, Dakota Park-Ozee, Travis Herink, Kara Neu, Jonathan Plomin, Eve-Odine Duchaufour, Debbie Gale Mitchell, Tennyson Anderson-Stricklin, Lily Treitz, Samantha Rosenberger, Sierra Griffith, Finley Joseph, Daniel Sampson, Emmy Davis, Skyler Kasnoff, Evon Lopez, Vivian Nguyen, Cassy Young, Franklin Sellner, Martin Tobon, Ila Graham, Zach Billings, Holden Hedit, Decatur Boland, Paul Kosempel, Cory Chandler, Jay Mahoney, Sam Dragan, Susan Dagget, Yarrow Ator, Heidi Vuletich, Owen Weber, Andrew Kloeppel, Petersen Gray, Mandi Schaeffer-Fry, Razleen Bassra, Bryanna Rodriguez, Christina Blue, Taubie Sanders, Rachel Epstein, Luke Milburn, Camryn Evans, Ezra Martinez, Mary Westwood, Gabri Notov, Robin Tinghitella, Lilou Cabrol, Eli Barbour, Juliet Mendik, Selma Myers, Zac Wise, Noah Fahlin, Michelle Knowles, Abigail Hopper, Michael Greenberger, Romi Laclair, Sarah Watamura, Sabrina Efroymson, Casey Barker, Sydney Seltzer, Bryn Yehle, Jennifer Hoffman, Sara Garcia, Ryuka Nagamine, Trevor Briggs, Remy Le Boeuf, Elena Krone, Eileen Farrell, Regan O'Rourke, Elena Roel, Greg Mortimer, Ali Ayoub, Stefani Langehennig, Caitlin Turk, Logan Scmid, Stefan Chavez-Norgaard, Karen Kim, Tatiana Peccedi, Courtney Cassidy, John Sebesta, Rhianna Lewis, Janice Bening-Lacek, Vivian Lawless, Mckenna Hanson, Jeffrey Amidon, Riya Joshi, Ram Ambre, Brady Worrell, Perrin Schneider, Ali Azadani, Brooke Agulnek, Lyndsie Salvagio, Elise Siemanowki, Yan Qin, Andre Allen, Melodie Nguyen, Megan Livengood, Abby Reams, Saffron Hartreeve, Bri Wylie, Sarah Brookman, Mariah Loiacono, Green Russo, Abhia Lodhi, Gabrielle Welsh, Nika Spehar, Shahked Levin, Evrim Baykal, Kimberly Chiew, Jocelyn Torres, Kailey Hicks, Mykaela Tanino-Springsteen, Audrey Bellows, Akam Chahal, Madeline Tepper, Shannon Murphy, Alexa Fonseca, Deborah Han, Cassandra Perez, Oluwatoyin Alaba, Julia Roncoroni, Vy Nguyen, Nana Burn, Sarah Sasse, Rubin Tuder, Anthony Gerber, Nancy Lorenzon, Christine Vohwinkel, Camryn Gunter, Tristan Weber, Sam Rommel, Brian Michel, Muskan Fatima, Alannah Oleson, Kira Frey, Edward Garrido, Beckett Morris, Kerstin Haring, Drew Middleton, Abigail Walpert, Liam Dee, Gabby Ishaw, Cole Carnes, Maddie Weiser, Claire Fox, Valeriia Vlasenko, Kateri Mcrae, Riley Smith, Abigail Templin, Kushani Rajapaksha

DU Undergraduate Research Journal Archive

Abstracts from the DU Undergraduate Research Showcase.


Improving Semantic Precision In Text-To-Image Diffusion Models Via Latent-Space Optimization And Semantically-Parsed Evaluation, Mohammad Rouie Miab May 2026

Improving Semantic Precision In Text-To-Image Diffusion Models Via Latent-Space Optimization And Semantically-Parsed Evaluation, Mohammad Rouie Miab

McKelvey School of Engineering Graduate Student Theses & Dissertations

Text-to-image diffusion models can produce visually impressive images from natural-language prompts, but they often fail to satisfy the detailed semantic constraints expressed in compositional prompts. Typical failure modes include omitted objects, merged entities, incorrect quantities, incorrect attribute binding, and leakage of one entity's attributes onto another. This thesis studies the problem of semantic precision in text-to-image generation: how faithfully a generated image satisfies the structured meaning of its prompt.   The thesis makes two linked contributions. First, it presents a training-free inference-time refinement method for diffusion-based image generation. The method operates directly in latent space during denoising and uses noun-phrase-aware cross-attention …


Teacher-Student Diffusion Model For Text-Driven 3d Hand Motion Generation, Ching Lam Cheng, Bin Zhu, Shengfeng He May 2026

Teacher-Student Diffusion Model For Text-Driven 3d Hand Motion Generation, Ching Lam Cheng, Bin Zhu, Shengfeng He

PhD Student’s Publications Collection

Generating realistic 3D hand motion from natural language is vital for VR, robotics, and human-computer interaction. Existing methods either focus on full-body motion, overlooking detailed hand gestures, or require explicit 3D object meshes, limiting generality. We propose TSHaMo, a model-agnostic teacher-student diffusion framework for text-driven hand motion generation. The student model learns to synthesize motions from text alone, while the teacher leverages auxiliary signals (e.g., MANO parameters) to provide structured guidance during training. A co-training strategy enables the student to benefit from the teacher’s intermediate predictions while remaining text-only at inference. Evaluated using two diffusion backbones on GRAB and H2O, …


Beyond The Interface: Human Perceptions Of Generative-Ai Chatbots As Conversational Partners, Browning W.E. Blair May 2026

Beyond The Interface: Human Perceptions Of Generative-Ai Chatbots As Conversational Partners, Browning W.E. Blair

All Theses

Generative AI (gen-AI) chatbots are becoming embedded in everyday communicative life, yet it remains unclear whether users perceive these systems as socially reciprocative conversational partners. Therefore, this study examines how young adults understand and interact with gen-AI chatbots, focusing on perceptions of conversational partnership, anthropomorphism, politeness, discomfort, and technical understanding. Guided by the CASA framework, Media Equation Theory, and the uncanny valley hypothesis, this study employed four semi-structured, online focus groups with 15 undergraduate students and recent college graduates in the United States. Findings indicate that participants did not broadly perceive gen-AI chatbots as conversational partners in the interpersonal sense. …


Portrait Shadow Removal Via Self-Exemplar Illumination Equalization, Qian Huang, Cheng Xu, Guiqing Li, Ziheng Wu, Shengxin Liu, Shengfeng He Apr 2026

Portrait Shadow Removal Via Self-Exemplar Illumination Equalization, Qian Huang, Cheng Xu, Guiqing Li, Ziheng Wu, Shengxin Liu, Shengfeng He

Research Collection School Of Computing and Information Systems

We introduce the Self-Exemplar Illumination Equalization Network, designed specifically for effective portrait shadow removal. The core idea of our method is that partially shadowed portraits can find ideal exemplars within their non-shadowed facial regions. Rather than directly fusing two distinct classes of facial features, our approach utilizes non-shadowed regions as an illumination indicator to equalize the shadowed regions, generating deshadowed results without boundary-merging artifacts. Our network comprises cascaded Self-Exemplar Illumination Equalization Blocks (SExmBlock), each containing two modules: a self-exemplar feature matching module and a feature-level illumination rectification module. The former identifies and applies internal illumination exemplars to shadowed areas, producing …


Match-A-Fit, Brianna Mendoza, Adan Diaz De Leon, Pedro Jacobo, Juan Marco Saca Dada Apr 2026

Match-A-Fit, Brianna Mendoza, Adan Diaz De Leon, Pedro Jacobo, Juan Marco Saca Dada

Posters - 2026

Welcome to Match-a-Fit! Match-a-Fit is an iOS application that allows the user to create a digital closet by uploading images of their clothing items. With AI, the program can generate outfits based on the digital closet, the time, and the occasion. Match-a-Fit’s purpose is designed to help users who struggle to get ready, run out of time, or can’t decide on an outfit, by easily generating outfit options based on the occasion.


Match-A-Fit, Adan Diaz De Leon, Juan Marco Saca Dada, Brianna Mendoza, Arsalan Kataneh, Theophile Nsabimana, Pedro Jacobo Apr 2026

Match-A-Fit, Adan Diaz De Leon, Juan Marco Saca Dada, Brianna Mendoza, Arsalan Kataneh, Theophile Nsabimana, Pedro Jacobo

Presentations - 2026

Welcome to Match-a-Fit! Match-a-Fit is an iOS application that allows the user to create a digital closet by uploading images of their clothing items. With AI, the program can generate outfits based on the digital closet, the time, and the occasion. Match-a-Fit’s purpose is designed to help users who struggle to get ready, run out of time, or can’t decide on an outfit, by easily generating outfit options based on the occasion.


Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong Apr 2026

Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong

Research Collection School Of Computing and Information Systems

Conventionally, human intuition defines vision as a modality of passive optical sensing, relying on ambient light to perceive the environment. However, active optical sensing, which involves emitting and receiving signals, offers unique advantages by capturing both radiometric and geometric properties of the environment, independent of external illumination conditions. This work focuses on advancing active optical sensing using Light Detection and Ranging (LiDAR), which captures intensity data, enabling the estimation of surface reflectance that remains invariant under varying illumination. Such properties are crucial for robotic perception tasks, including detection, recognition, segmentation, and Simultaneous Localization and Mapping (SLAM). A key challenge with …


Weakly Supervised Video Anomaly Detection And Localization With Spatio-Temporal Prompts, Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang, Qingsen Yan, Peng Wang, Yanning Zhang Apr 2026

Weakly Supervised Video Anomaly Detection And Localization With Spatio-Temporal Prompts, Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang, Qingsen Yan, Peng Wang, Yanning Zhang

Research Collection School Of Computing and Information Systems

Current weakly supervised video anomaly detection (WSVAD) task aims to achieve frame-level anomalous event detection with only coarse video-level annotations available. Existing works typically involve extracting global features from full-resolution video frames and training frame-level classifiers to detect anomalies in the temporal dimension. However, most anomalous events tend to occur in localized spatial regions rather than the entire video frames, which implies existing frame-level feature based works may be misled by the dominant background information and lack the interpretation of the detected anomalies. To address this dilemma, this paper introduces a novel method called STPrompt that learns spatio-temporal prompt embeddings …


Semat: Semantic Enhanced Natural Image Interactive Matting, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Qianru Sun, Yang Tang, Bo Li, Pan Zhou Apr 2026

Semat: Semantic Enhanced Natural Image Interactive Matting, Ruihao Xia, Yu Liang, Peng-Tao Jiang, Hao Zhang, Qianru Sun, Yang Tang, Bo Li, Pan Zhou

Research Collection School Of Computing and Information Systems

Recent approaches attempt to adapt powerful interactive segmentation models, such as SAM, to interactive matting and fine-tune the models based on synthetic matting datasets. However, models trained on synthetic data fail to generalize to complex and occlusion scenes. We address this challenge by proposing a new matting dataset based on the COCO dataset, namely COCO-Matting. It selects real-world complex images from COCO and converts semantic segmentation masks to matting labels. The built COCO-Matting comprises an extensive collection of 36,980 human instance-level alpha mattes in complex natural scenarios. Furthermore, existing SAM-based matting methods extract intermediate features and masks from a frozen …


Dragging With Geometry: From Pixels To Geometry-Guided Image Editing, Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou Apr 2026

Dragging With Geometry: From Pixels To Geometry-Guided Image Editing, Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou

Research Collection School Of Computing and Information Systems

Interactive point-based image editing serves as a controllable editor, enabling precise and flexible manipulation of image content. However, most drag-based methods operate primarily on the 2D pixel plane with limited use of 3D cues. As a result, they often produce imprecise and inconsistent edits, particularly in geometry-intensive scenarios such as rotations and perspective transformations. To address these limitations, we propose a novel geometry-guided drag-based image editing method—GeoDrag, which addresses three key challenges: 1) incorporating 3D geometric cues into pixel-level editing, 2) mitigating discontinuities caused by geometry-only guidance, and 3) resolving conflicts arising from multi-point dragging. Built upon a unified displacement …


Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou Apr 2026

Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou

Research Collection School Of Computing and Information Systems

While text-to-3D generation has attracted growing interest, existing methods often struggle to produce 3D assets that align well with human preferences. Current preference alignment techniques for 3D content typically rely on hardly-collected preference-paired multi-view 2D images to train 2D reward models, when then guide 3D generation — leading to geometric artifacts, such as the Janus face problem and geometric incompleteness, due to their inherent 2D bias. To address these limitations, we construct 3D-MeshPref, the first large-scale unpaired 3D preference dataset, featuring diverse 3D meshes annotated by a large language model and refined by human evaluators. We then develop RewardCS, the …


From Spatial To Actions: Grounding Vision-Language-Action Model In Spatial Foundation Priors, Zhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu, Lei Zhou, Chenchen Liu, Dong Wang, Francis E. H. Tay, Sijin Chen, Ziwei Liu, Yuxiao Liu, Xinghang Li, Pan Zhou Apr 2026

From Spatial To Actions: Grounding Vision-Language-Action Model In Spatial Foundation Priors, Zhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu, Lei Zhou, Chenchen Liu, Dong Wang, Francis E. H. Tay, Sijin Chen, Ziwei Liu, Yuxiao Liu, Xinghang Li, Pan Zhou

Research Collection School Of Computing and Information Systems

Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require specialized sensors and transfer poorly across modalities, or inject weak cues that lack geometry and degrade vision-language alignment. In this work, we introduce FALCON (From Spatial to Action), a novel paradigm that injects rich 3D spatial tokens into the action head. FALCON leverages spatial foundation models to deliver strong geometric priors from RGB alone, and includes an Embodied Spatial Model that can optionally fuse depth, or pose …


Who You Explain To Matters: Learning By Explaining To Conversational Agents With Different Pedagogical Roles, Zhengtao Xu, Junti Zhang, Anthony Tang, Yi-Chieh Lee Apr 2026

Who You Explain To Matters: Learning By Explaining To Conversational Agents With Different Pedagogical Roles, Zhengtao Xu, Junti Zhang, Anthony Tang, Yi-Chieh Lee

Research Collection School Of Computing and Information Systems

Conversational agents are increasingly used in education for learning support. An application is “learning by explaining”, where learners explain their understanding to an agent. However, existing research focuses on single roles, leaving it unclear how different pedagogical roles influence learners’ interaction patterns, learning outcomes and experiences. We conducted a between-subjects study (N=96) comparing agents with three pedagogical roles (Tutee, Peer, Challenger) and a control condition while learning an economics concept. We found that different pedagogical roles shaped learning dynamics, including interaction patterns and experiences. Specifically, the Tutee agent elicited the most cognitive investment but led to high pressure. The Peer …


Lagrangian Motion Fields For Long-Term Motion Generation, Yifei Yang, Zikai Huang, Chenshu Xu, Shengfeng He Feb 2026

Lagrangian Motion Fields For Long-Term Motion Generation, Yifei Yang, Zikai Huang, Chenshu Xu, Shengfeng He

Research Collection School Of Computing and Information Systems

Long-term motion generation is a challenging task that requires producing coherent and realistic sequences over extended durations. Current methods primarily rely on framewise motion representations, which capture only static spatial details and overlook temporal dynamics. This approach leads to significant redundancy across the temporal dimension, complicating the generation of effective long-term motion. To overcome these limitations, we introduce the novel concept of Lagrangian Motion Fields, specifically designed for long-term motion generation. By treating each joint as a Lagrangian particle with uniform velocity over short intervals, our approach condenses motion representations into a series of "supermotions" (analogous to superpixels). This method …


Modeling Joint Visual Attention In Naturalistic Dyadic Interactions, Kuushini Thennakoon, Yasasi Abeysinghe, Bhanuka Mahanama, Vikas Ashok, Sampath Jayarathna Jan 2026

Modeling Joint Visual Attention In Naturalistic Dyadic Interactions, Kuushini Thennakoon, Yasasi Abeysinghe, Bhanuka Mahanama, Vikas Ashok, Sampath Jayarathna

Computer Science Faculty Publications

Joint visual attention (JVA) provides important insight into how individuals coordinate attention during social interaction. Egocentric eye tracking enables the study of JVA in natural, multi-user settings. This work presents a multi-stage framework to identify and analyze JVA using egocentric video and gaze data. The approach consists of three steps: spatiotemporal tube-based visual similarity, gaze-guided object detection, and attention pattern analysis using the ambient–focal coefficient K. Results show that object-focused collaborative activities exhibit high JVA, with object detection capturing higher joint attention than visual similarity, whereas conversation-based or independent activities show lower and more fragmented joint attention. Analysis of K …


Lost In Instructions: Study Of Blind Users' Experiences With Diy Manuals And Ai-Rewritten Instructions For Assembly, Operation, And Troubleshooting Of Tangible Products, Monalika Padma Reddy, Aruna Balasubramanian, Jiawei Zhou, Xiaojun Bi, Iv Ramakrishnan, Vikas Ashok Jan 2026

Lost In Instructions: Study Of Blind Users' Experiences With Diy Manuals And Ai-Rewritten Instructions For Assembly, Operation, And Troubleshooting Of Tangible Products, Monalika Padma Reddy, Aruna Balasubramanian, Jiawei Zhou, Xiaojun Bi, Iv Ramakrishnan, Vikas Ashok

Computer Science Faculty Publications

AI tools like ChatGPT and Be-My-AI are increasingly being used by blind individuals. Although prior work has explored their use in some Do-It-Yourself (DIY) tasks by blind individuals, little is known about how they use these tools and the available product-manual resources to assemble, operate, and troubleshoot physical/tangible products – tasks requiring spatial reasoning, structural understanding, and precise execution. We address this knowledge gap via an interview study and a usability study with blind participants, investigating how they leverage AI tools and product manuals for DIY tasks with physical products. Findings show that manuals are essential resources, but product-manual instructions …