Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- China Simulation Federation (3880)
- Singapore Management University (1881)
- Old Dominion University (640)
- San Jose State University (277)
- MBZUAI (233)
-
- City University of New York (CUNY) (184)
- Technological University Dublin (157)
- Air Force Institute of Technology (137)
- Chapman University (124)
- California Polytechnic State University, San Luis Obispo (116)
- Chinese Academy of Sciences (111)
- University of Arkansas, Fayetteville (100)
- Lindenwood University (97)
- Edith Cowan University (92)
- Embry-Riddle Aeronautical University (92)
- University of Nebraska - Lincoln (78)
- University of Kentucky (76)
- University of South Florida (71)
- University of Nevada, Las Vegas (63)
- Clemson University (60)
- Dartmouth College (60)
- University of Denver (59)
- Utah State University (57)
- University of Michigan Law School (56)
- The Texas Medical Center Library (54)
- Thomas Jefferson University (54)
- New Jersey Institute of Technology (53)
- University of Malaya (50)
- Purdue University (48)
- Missouri University of Science and Technology (47)
- Keyword
-
- Artificial intelligence (777)
- Machine learning (685)
- Deep learning (435)
- Machine Learning (359)
- Artificial Intelligence (357)
-
- AI (234)
- Deep Learning (201)
- Simulation (160)
- Computer vision (157)
- Reinforcement learning (140)
- Generative AI (134)
- Neural networks (128)
- Natural language processing (108)
- Large language models (107)
- Robotics (97)
- Natural Language Processing (90)
- ChatGPT (89)
- Path planning (88)
- Optimization (82)
- Large Language Models (77)
- Computer Vision (75)
- Classification (71)
- Neural network (67)
- Neural Networks (65)
- Virtual reality (64)
- Reinforcement Learning (63)
- Computer Science (59)
- Cybersecurity (59)
- Genetic algorithm (58)
- Algorithms (57)
- Publication Year
- Publication
-
- Journal of System Simulation (3880)
- Research Collection School Of Computing and Information Systems (1648)
- Master's Projects (248)
- Theses and Dissertations (183)
- Computer Science Faculty Publications (124)
-
- Bulletin of Chinese Academy of Sciences (Chinese Version) (111)
- Faculty Scholarship (108)
- Publications and Research (99)
- Computer Vision Faculty Publications (98)
- Master's Theses (96)
- Conference papers (92)
- Electrical & Computer Engineering Faculty Publications (90)
- Machine Learning Faculty Publications (86)
- Electronic Theses and Dissertations (85)
- Faculty Publications (77)
- Dissertations (69)
- Research outputs 2022 to 2026 (64)
- USF Tampa Graduate Theses and Dissertations (59)
- Dissertations and Theses Collection (Open Access) (57)
- Articles (54)
- Dissertations, Theses, and Capstone Projects (53)
- Theses and Dissertations--Computer Science (48)
- Natural Language Processing Faculty Publications (46)
- Teaching and Generative AI: Pedagogical Possibilities and Productive Tensions (46)
- Graduate Theses and Dissertations (43)
- Open Access Theses & Dissertations (42)
- Theses (40)
- Electrical & Computer Engineering Theses & Dissertations (39)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (39)
- Publications (39)
- Publication Type
- File Type
Articles 271 - 300 of 11145
Full-Text Articles in Computer Sciences
Georoad-Upernet: Geo-1-Based Weakly Supervised Multispectral Road Extraction Via Role-Aware Context Fusion And Semantic Regularization, Shaoqian Chen, Yunliang Chen, Jianxin Li, Ao Yang
Georoad-Upernet: Geo-1-Based Weakly Supervised Multispectral Road Extraction Via Role-Aware Context Fusion And Semantic Regularization, Shaoqian Chen, Yunliang Chen, Jianxin Li, Ao Yang
Research outputs 2022 to 2026
Extracting roads accurately from remote sensing images is important for map updates, traffic analysis, and infrastructure monitoring. Medium-resolution multispectral images can provide useful surface and background information, but when used alone, the spatial details are limited for retaining narrow roads, intersection structures, and fine road topologies. To address this problem, this paper proposes GeoRoad-UPerNet, a Geo-1-centered weakly supervised multispectral framework for road extraction. In this framework, Geo-1 serves as the primary 16-band multispectral source, Sentinel-2 Level-2A imagery serves as auxiliary contextual support, and OpenStreetMap (OSM) road information is converted into proxy supervision rather than dense manual ground truth. GeoRoad-UPerNet contains …
“Grandpa, Can You Speak Nicer?”: Envisioned Chatbot Roles And Design Tensions In Intergenerational Communication Conflicts, Tianyi Zhang, Emran Bin Elias Poh, Yueyue Hou, Yi-Chieh Lee, Renwen Zhang, Jiannan Li, Anthony Tang
“Grandpa, Can You Speak Nicer?”: Envisioned Chatbot Roles And Design Tensions In Intergenerational Communication Conflicts, Tianyi Zhang, Emran Bin Elias Poh, Yueyue Hou, Yi-Chieh Lee, Renwen Zhang, Jiannan Li, Anthony Tang
Research Collection School Of Computing and Information Systems
Intergenerational conversations often break down when differences in tone, language, or expectations lead participants to feel dismissed or misunderstood. In this work, we explore how people envision AI-driven chatbot interventions for addressing communication problems in text-based intergenerational family chat. We conducted a scenario-based design interview with 10 pairs of family members from different generations, in which participants designed chatbot interventions that varied in intervention target and timing. Our findings show that participants expect chatbots to perform multiple themes of intervention, including mediating understanding, providing emotional support, offering evaluative commentary, and guiding interaction through behavioral suggestions. These expectations varied systematically across …
Group Conversational Agents: A Review Of Designs That Support And Shape Group Interaction, Shunyi Yeo, Tianyi Zhang, Scott Bateman, Gary Hsieh, Young-Ho Kim, Simon Tangi Perrault, Jiannan Li, Anthony Tang
Group Conversational Agents: A Review Of Designs That Support And Shape Group Interaction, Shunyi Yeo, Tianyi Zhang, Scott Bateman, Gary Hsieh, Young-Ho Kim, Simon Tangi Perrault, Jiannan Li, Anthony Tang
Research Collection School Of Computing and Information Systems
Conversational agents that participate in or mediate group interaction introduce challenges that extend beyond supporting individual users, raising new questions about how agents participate in and influence groups. To characterise this emerging design space, we present a systematic review of 53 peer-reviewed studies on group conversational agents (GCAs). We analyse how GCAs intervene in group-level processes, including participation regulation, conflict mediation, task alignment, and execution support. Using concepts from group research as an analytic lens, we organise prior GCA work around recurring group interactional challenges (orientation, conflict, alignment, and execution), and examine the roles agents are designed to play in …
Cfalr: Collaborative Filtering-Augmented Large Language Model For Personalized Fashion Outfit Recommendation, Yujuan Ding, Junrong Liao, Yunshan Ma, Yi Bin, Wenqi Fan, Tat-Seng Chua, Qing Li
Cfalr: Collaborative Filtering-Augmented Large Language Model For Personalized Fashion Outfit Recommendation, Yujuan Ding, Junrong Liao, Yunshan Ma, Yi Bin, Wenqi Fan, Tat-Seng Chua, Qing Li
Research Collection School Of Computing and Information Systems
Personalized outfit recommendation poses a significant challenge in e-commerce and social media platforms, requiring systems that balance user preferences with aesthetic compatibility. Collaborative filtering (CF) provides a traditional solution for this, but it struggles with data-sparse scenarios and complex user-item-outfit relationships. Meanwhile, existing template-based approaches are constrained by rigid pre-designed structures. To bridge these research gaps, we introduce CFALR (Collaborative Filtering-Augmented Large Language Model for Recommendation), a novel framework that synergizes collaborative filtering with large language models for personalized outfit recommendation. Specifically, CFALR describes user-outfit interactions in natural language and leverages LLMs to capture fashion semantics while employing CF-enhanced embeddings …
Deployment-Aware Deep Learning For Computer Vision: Efficient Architectures From 3d Segmentation To Mixed Reality, Bahar Uddin Mahmud
Deployment-Aware Deep Learning For Computer Vision: Efficient Architectures From 3d Segmentation To Mixed Reality, Bahar Uddin Mahmud
Dissertations
Deep learning has become the dominant approach for solving vision-centric problems; however, its successful deployment in real-world applications remains limited by high computational cost, data dependency, and insufficient integration with practical and human-centered environments. While state-of-the art deep learning models often achieve impressive performance in controlled settings, they frequently fail to generalize or operate efficiently under deployment constraints such as limited resources, complex data modalities, and real-time interaction requirements. These limitations motivate the need for a deployment-oriented deep learning framework that balances accuracy, efficiency, and practical usability.
This dissertation investigates the design and deployment of efficient deep learning architectures for …
3dcotton, Md Ahmed Al Muzaddid, William J. Beksi
3dcotton, Md Ahmed Al Muzaddid, William J. Beksi
Agriculture - Archive
3DCotton is an image dataset consisting of 8 cotton plants recorded at the Texas A&M University Research Farm. The images were captured using an Apple iPhone at a resolution of 1040x1920 pixels. Approximately 150 images per plant were taken from a distance of 1 m by recording multiple viewpoints. These images can be utilized for developing 3D reconstruction methods.
Ms110 Syllabus: Introduction To Computers, Information Systems, And Artificial Intelligence, Wei Zhang
Ms110 Syllabus: Introduction To Computers, Information Systems, And Artificial Intelligence, Wei Zhang
Management Science and Information Systems Faculty Publication Series
This is a syllabus for Professor Wei Zhang's MS110: Introduction to Computers, Information Systems and Artificial Intelligence Course within UMass Boston's College of Management. This is an Open Educational Resource and can be remixed, copied, redistributed, altered and reused as long as permission is given to the original creator.
Shared Language For Responsible Ai Integration, Asa B. Stone, Mark C. Stone, Alisha Bevins, Jean Claude Niyomugabo, Irene Magara, Jacob Abaare, Derek M. Heeren, Mubarak Abu Zouriq
Shared Language For Responsible Ai Integration, Asa B. Stone, Mark C. Stone, Alisha Bevins, Jean Claude Niyomugabo, Irene Magara, Jacob Abaare, Derek M. Heeren, Mubarak Abu Zouriq
PRAIRIE: Pioneering Responsible AI for Research, Innovation, and Education
As AI rapidly reshapes how we work and learn, employers increasingly seek graduates who can think before they prompt, exercising judgment under pressure rather than merely producing output. Yet students are praised for AI use in one course and penalized for it in the next, and faculty are left to lead responsibly on shifting ground, with no shared language to guide them.
This paper introduces the PRAIRIE Framework for AI Integration, a shift from reactive gatekeeping toward proactive stewardship. It emerged from a qualitative sentiment analysis of three communities (students, faculty, and industry partners) whose concerns converged on one need: …
Ai Interview Helper: A Tool For Assisting Search And Rescue Long-Profile Interviews, Dylan P. Starink
Ai Interview Helper: A Tool For Assisting Search And Rescue Long-Profile Interviews, Dylan P. Starink
Master's Theses
In Search and Rescue (SAR) operations, time pressure and limited interviewer experience can lead to missed opportunities when interviewing a missing person’s friends and family. This thesis presents a real-time, end-to-end system that provides context-aware follow-up question suggestions as interviews unfold. Leveraging large language models (LLMs) and agentic design patterns, the system is intended to support interviewers by helping them identify relevant follow-up questions and pursue potentially overlooked lines of inquiry.
The system was evaluated through three mock interviews with two SAR interviewer participants across two events. Given the limited sample size, the results provide early insights into the feasibility …
Emotional Support Through Ai: Venting To Artificial Intelligence Or A Perceived Human May Offer Comparable Emotional Well-Being Benefits, Meilan Hu, Jerlyn Q. H. Ho, Claire Ng, Shermaine S. M. Wong, Andree Hartanto
Emotional Support Through Ai: Venting To Artificial Intelligence Or A Perceived Human May Offer Comparable Emotional Well-Being Benefits, Meilan Hu, Jerlyn Q. H. Ho, Claire Ng, Shermaine S. M. Wong, Andree Hartanto
Research Collection School of Social Sciences
Artificial Intelligence (AI) chatbots are increasingly being explored as sources of informal emotional support, with emerging evidence suggesting that venting to these systems can reduce negative affect. Yet, it remains unclear whether such benefits depend on the responder's perceived identity. Given that emotional relief from venting often hinges on perceived authenticity and emotional validation, this study investigates whether the emotional well-being benefits of venting differ when users believe they are interacting with an AI chatbot versus a human, even when responses are content-matched. In a pre-registered experiment ( N = 279), participants were randomly assigned to either an AI-assisted venting …
Navigating Oer Support Without Drowning In Ai, Lydia Burrage-Goodwin, Christine Moynihan
Navigating Oer Support Without Drowning In Ai, Lydia Burrage-Goodwin, Christine Moynihan
Joseph P. Healey Library Publications
This was a presentation at the June 2026 Boston Library Consortium at Connecticut College.
UMB Healey Librarians Lydia Burrage-Goodwin and Christine Moynihan talk about what experiences they have had with faculty using OER and AI, which led them to develop ethics guidelines to support librarians who work with faculty authors. Attendees learned about creating AI use statements for OERs, using AI transparency logos, and applying open licenses to fully AI generated content as well as OER adaptations.
The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala
The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala
Student Theses
The rapid adoption of Large Language Models (LLMs) in software development has transformed coding practices by enabling automated code generation, completion, and optimization. Despite these advantages, concerns persist regarding the security and reliability of LLM-generated code. This study presents a comprehensive evaluation of both the functional correctness and security of code produced by three prominent LLMs as of early 2026. A total of 4,800 code snippets were generated using 100 security-focused programming prompts derived from the OWASP Top 10:2025, translated across eight natural languages and two phrasing styles (literal and natural developer-oriented prompts). To assess performance, a multi-stage experimental framework …
Quantum And Conventional Informatics Studies Of Synthesis Energetics And Defect Formation In Nitride Crystal Epitaxy, Andrew Steven Messecar
Quantum And Conventional Informatics Studies Of Synthesis Energetics And Defect Formation In Nitride Crystal Epitaxy, Andrew Steven Messecar
Dissertations
Machine learning is a valuable approach for the processing and analysis of complex information. By estimating relationships from recorded data, machine learning methodologies can be effective strategies for pattern recognition, enabling investigations and technological applications based thereon. The potential for improved understanding of high-dimensional data has drawn interest towards machine learning from across the sciences, including the research and development of new and improved material systems. In the context of experimental materials research, much of the reported efforts to incorporate machine learning into conventional practice have been primarily focused on either the enhanced analysis of characterization experiment data or the …
When Saying "No" Is Not Enough: Cognitive-Action Decoupling And The Illusion Of Safety In Llm Agents, Shasha Yu
When Saying "No" Is Not Enough: Cognitive-Action Decoupling And The Illusion Of Safety In Llm Agents, Shasha Yu
School of Professional Studies
Current safety evaluations of large language models (LLMs) predominantly rely on textual compliance, implicitly assuming that refusal-style responses correspond to safe behavior. This assumption becomes fragile when LLMs are embedded in agentic systems with the ability to execute state-changing actions. In this paper, we present an empirical critique of text-centric safety evaluation through an action-aware study of LLM agents under controlled conditions. Across multiple state-of-the-art models, we observe a recurring cognitive-action decoupling: agents generate policy-aligned refusal language while still producing unsafe tool-mediated action proposals. This produces an illusion of safety, where conversational audits indicate compliance even as operational risk persists. …
Trace: Temporal Rhetorical Analysis And Consistency Evaluation For Legislative Speech, David Hernandez
Trace: Temporal Rhetorical Analysis And Consistency Evaluation For Legislative Speech, David Hernandez
Master's Theses
Legislators frequently discuss the same policy issues across multiple hearings and legislative sessions, sometimes maintaining consistent positions and other times modifying or reframing their stance over time. Understanding how these positions evolve is important for analyzing political discourse and democratic accountability, yet identifying such shifts at scale remains difficult.
We introduce TRACE (Temporal Rhetorical Analysis and Consistency Evaluation), a system built on the Digital Democracy Database (DDDB) for detecting rhetorical inconsistency in California legislative hearing testimony. TRACE organizes utterances into speaker-anchored timelines indexed by bill and session, then applies hybrid semantic retrieval — combining dense BGE embeddings with BM25 lexical …
Inductive Biases In Field-Level Cosmological Inference From Galaxy Catalogs, James O'Connor Baldwin
Inductive Biases In Field-Level Cosmological Inference From Galaxy Catalogs, James O'Connor Baldwin
Dissertations, Theses, and Capstone Projects
We perform field-level likelihood-free inference of the matter density parameter Ωm from simulated galaxy catalogs using machine learning models with differing inductive biases. Using features extracted from hydrodynamic simulations in the CAMELS suite, we investigate how both observable choice and model architecture govern the extraction of cosmological information. We consider galaxy positions and line-of-sight peculiar velocities, both separately and in combination, and compare permutation-invariant Deep Sets, implemented with either standard multilayer perceptrons (MLPs) or Kolmogorov–Arnold Networks (KANs), to graph neural networks (GNNs) implemented with MLPs, which explicitly encode spatial relations. We evaluate inference performance under both in-distribution and out-of-distribution (OOD) …
Adaptive Outlier Detection Over Data Stream, Rui Zhu, Mingyuan Jiang, Xiaochun Yang, Baihua Zheng, Bin Wang, Tao Qiu
Adaptive Outlier Detection Over Data Stream, Rui Zhu, Mingyuan Jiang, Xiaochun Yang, Baihua Zheng, Bin Wang, Tao Qiu
Research Collection School Of Computing and Information Systems
Continuous distance-based outlier detection in streaming data poses significant challenges and has a wide range of practical applications. Traditional threshold-based methods perform well under stable streaming conditions, where fixed parameters remain effective. However, they often struggle with dynamic data distributions and high stream speeds, leading to suboptimal performance, limited control over the number of returned outliers, and failure to meet real-time detection requirements. To address these issues, this paper introduces a novel Recall and Proportion-Aware Outlier Detection (RPA-OD) query. In RPA-OD, ρ defines a distance relaxation that enables real-time outlier detection. Specifically, objects with fewer than k neighbors within the …
Enhancing Pointing Gestures Of Non-Hmd Users In Asymmetric Collocated Mixed Reality Collaboration, Nam-Dang Vo, Van-Vinh Thai, Anthony Tang, Khanh-Duy Le
Enhancing Pointing Gestures Of Non-Hmd Users In Asymmetric Collocated Mixed Reality Collaboration, Nam-Dang Vo, Van-Vinh Thai, Anthony Tang, Khanh-Duy Le
Research Collection School Of Computing and Information Systems
A common collocated group setting in mixed-reality (MR) collaboration is a person wearing a MR headset (HMD user) and presenting MR contents to audiences who are not provided with such specialized devices (Non-HMD users). In this setting, while Non-HMD users can view the MR environment shown on a large physical display, it still remains challenging for the HMD user to interpret their pointing gesture when they spatially refer to objects in the MR environment. To address this, we designed and evaluated two pointing techniques—SCREEN and SCREEN+SPACE—that support Non-HMD users in referring to MR content. Screen pointing allows users to refer …
Towards Efficient Continual Learning: From Memory Optimization To Foundation Models, Zilin Luo
Towards Efficient Continual Learning: From Memory Optimization To Foundation Models, Zilin Luo
Dissertations and Theses Collection (Open Access)
Continual learning, also termed lifelong learning, enables machine learning models to incrementally acquire new knowledge while mitigating the degradation of previously learned information—a capability essential for adapting to dynamic, real-world data environments. This dissertation investigates the core challenges of continual learning and extends its application to enhancing training efficiency in the era of foundation models. The first part of this dissertation addresses the constraints of few-shot exemplar storage with a novel compression framework. While leveraging class activation maps to downsample non-discriminative pixels, we introduce an adaptive masking model, optimized through bilevel optimization, to store more exemplars efficiently. The second part …
Towards Auto-Evaluation For Large Language Models, Jiahao Ying
Towards Auto-Evaluation For Large Language Models, Jiahao Ying
Dissertations and Theses Collection (Open Access)
The rapid advancement of large language models (LLMs) has created an urgent need for evaluation methodologies that are timely, scalable, reliable, and informative. Conventional evaluation benchmarks, although essential for measuring model capabilities and guiding model development, are often constructed and maintained through labor-intensive human annotation. As LLMs continue to improve through increases in model scale, training data, and computational resources, static benchmarks may quickly lose discriminative power. Moreover, the growing use of large and diverse training corpora increases the risk of benchmark leakage, which can inflate evaluation results and obscure the true capabilities of models. These challenges call for a …
How To Save The Take-Home Essay With Oral Assessments, Matthew Hammerton, Jacqueline Ho
How To Save The Take-Home Essay With Oral Assessments, Matthew Hammerton, Jacqueline Ho
Research Collection School of Social Sciences
In a commentary, the authors opined that pairing take-home essays with oral assessments is a more effective response to AI than policing its use. Students who cannot adequately explain their work can be marked down, reducing incentives to rely on AI. They noted that oral exams help preserve key elements of university education – intellectual effort, ownership, and human relationships – while allowing take-home essays to remain relevant in an AI-driven landscape that demands greater emphasis on understanding, responsibility, and dialogue.
Legal Ethics Of Ai Snake Oil: Navigating The Hype, Harm, And Hope Of Legal Ai, Drew Simshaw
Legal Ethics Of Ai Snake Oil: Navigating The Hype, Harm, And Hope Of Legal Ai, Drew Simshaw
Michigan Law Review
A review of AI Snake Oil.By Arvind Narayanan and Sayash Kapoor.
Saag: Structured Agent Assessment And Grounding, Ritvik Garimella, Vedant Khandelwal, Anvi Kohli, Amit Sheth
Saag: Structured Agent Assessment And Grounding, Ritvik Garimella, Vedant Khandelwal, Anvi Kohli, Amit Sheth
Publications
Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucinate argument values, or satisfy a schema while choosing a agent for the wrong reason. Existing benchmarks collapse these distinctions into a single binary score, leaving practitioners unable to diagnose where agent calls fail. We propose SAAG a cascaded diagnostic framework that decomposes agent-calling evaluation into three sequential stages: registry conformance, structural completeness, and argument grounding, each producing interpretable stage-specific diagnostics. These diagnostics additionally enable iterative self-repair: on prediction failure, the stage-specific signal guides targeted correction without leaking ground-truth values. We evaluate this …
Rvit-Fusionnet: A Local Cross-Attention Feature Fusion-Based Hybrid Framework For Brain Tumor Classification, Naima Islam, Sajeeb Kumar Ray, Md Anwar Hossain, Syed Mohammed Shamsul Islam
Rvit-Fusionnet: A Local Cross-Attention Feature Fusion-Based Hybrid Framework For Brain Tumor Classification, Naima Islam, Sajeeb Kumar Ray, Md Anwar Hossain, Syed Mohammed Shamsul Islam
Research outputs 2022 to 2026
Accurate brain tumor classification via MRI is essential for diagnosis and treatment. This study introduces RViT-FusionNet, a hybrid deep learning model that integrates convolutional and transformer architectures for enhanced tumor detection. The model utilizes ResNet-50 to capture textural details and a Vision Transformer for extracting global context. A Local Cross-Attention (LCA) module is proposed to align and merge these features, allowing the network to model local structures and long-range dependencies concurrently. To enhance generalization across varied imaging conditions and tumor types, a domain discriminator is included to discern spatial and domain-specific patterns, fostering the learning of domain-invariant representations. The approach …
Anatomical Domain Shifts: Test-Time Heterogeneous Adaptation For 3d Human Pose Prediction, Qiongjie Cui, Pan Zhou, Jingjing Chen, Na Zhao
Anatomical Domain Shifts: Test-Time Heterogeneous Adaptation For 3d Human Pose Prediction, Qiongjie Cui, Pan Zhou, Jingjing Chen, Na Zhao
Research Collection School Of Computing and Information Systems
The research frontier in human pose prediction (HPP) is advancing toward continual test-time adaptation (TTA), where models must self-adapt to dynamic test distributions. To date, the homeostatic continual TTA remains the sole viable solution, which isolates the model parameters and update domain-sensitive ones. Despite mitigating full-body domain gaps, human anatomical heterogeneity (domain shifts often localize to specific regions) is ignored. This anatomical-agnostic approach forces uniform parameter adaptation across kinematically distinct segments, causing: over-adaptation of stable regions and under-adaptation of shift-prone articulations. To address it, we introduce TT-HA, a novel Test-Time Heterogeneous Adaptation that implicitly estimates domain changes for anatomical segments, …
Rode: Linear Rectified Mixture Of Diverse Experts For Food Large Multi-Modal Models, Pengkun Jiao, Xinlan Wu, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yu-Gang
Rode: Linear Rectified Mixture Of Diverse Experts For Food Large Multi-Modal Models, Pengkun Jiao, Xinlan Wu, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yu-Gang
Research Collection School Of Computing and Information Systems
Large Multi-modal Models (LMMs) have significantly advanced a variety of vision-language tasks. The scalability and availability of high-quality training data play a pivotal role in the success of LMMs. In the realm of food, while comprehensive food datasets such as Recipe1M offer an abundance of ingredient and recipe information, they often fall short of providing ample data for nutritional analysis. The Recipe1M+ dataset, despite offering a subset for nutritional evaluation, is limited in the scale and accuracy of nutrition information. To bridge this gap, we introduce Uni-Food, a unified food dataset that comprises over 100,000 images with various food labels, …
Interfold: Learning Interpretable Diffusion Manifolds Beyond Binary Samples, Alexander Vincent Lewi, Rainer Tan, Shengfeng He
Interfold: Learning Interpretable Diffusion Manifolds Beyond Binary Samples, Alexander Vincent Lewi, Rainer Tan, Shengfeng He
Research Collection School Of Computing and Information Systems
We propose InterFold, a framework for learning and applying interpretable semantic manifolds in latent diffusion models, without requiring binary or paired supervision. Existing methods for semantic editing either rely on limited paired data or uncover only coarse, unsupervised directions that fail to capture user-specific, fine-grained attributes. InterFold addresses these limitations by learning a target attribute manifold in the H-space of diffusion models using only a set of positive, unlabeled examples. To edit a new image, InterFold projects its H-space representation toward this learned manifold through test-time optimization, enabling precise, identity-preserving modifications of complex, non-binary concepts. To make these edits effective …
Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang
Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang
Research Collection School Of Computing and Information Systems
Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we …
Context Matters: Auditing Gender Bias In T2i Generation Through Risk-Tiered Use-Case Profiles, Jose Luis Luna Campoverde, Yankun Wu, Xiaofei Xie, Noa Garcia
Context Matters: Auditing Gender Bias In T2i Generation Through Risk-Tiered Use-Case Profiles, Jose Luis Luna Campoverde, Yankun Wu, Xiaofei Xie, Noa Garcia
Research Collection School Of Computing and Information Systems
Text-to-image (T2I) generative models are increasingly used to produce content for education, media, and public-facing communication, and are starting to be integrated into higher-impact pipelines. Since generated images tend to reinforce stereotypes, producing representational erasure via “default” depictions and shaping perceptions of who belongs in certain roles, a growing body of work has proposed metrics to quantify gender bias in T2I outputs. Yet existing evaluations remain fragmented. Metrics are often reported without a shared view of what they measure, what assumptions they entail, or how their results should be interpreted under different deployment contexts. This limits the usefulness of gender …
Rc-Nf: Robot-Conditioned Normalizing Flow For Real-Time Anomaly Detection In Robotic Manipulation, Shijie Zhou, Bin Zhu, Jiarui Yang, Xiangyu Zhao, Jingjing Chen, Yu-Gang Jiang
Rc-Nf: Robot-Conditioned Normalizing Flow For Real-Time Anomaly Detection In Robotic Manipulation, Shijie Zhou, Bin Zhu, Jiarui Yang, Xiangyu Zhao, Jingjing Chen, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
Recent advances in Vision-Language-Action (VLA) models have enabled robots to execute increasingly complex tasks. However, VLA models trained through imitation learning struggle to operate reliably in dynamic environments and often fail under Out-of-Distribution (OOD) conditions. To address this issue, we propose Robot-Conditioned Normalizing Flow(RC-NF), a real-time monitoring model for robotic anomaly detection and intervention that ensures the robot's state and the object's motion trajectory align with the task. RC-NF decouples the processing of task-aware robot and object states within the normalizing flow. It requires only positive samples for unsupervised training and calculates accurate robotic anomaly scores during inference through the …