Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Artificial Intelligence and Robotics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 511 - 540 of 11148

Full-Text Articles in Computer Sciences

Architecture-Agnostic Test-Time Adaptation Via Backprop-Free Embedding Alignment, Xiao Ma, Young D. Kwon, Pan Zhou, Dong Ma Apr 2026

Architecture-Agnostic Test-Time Adaptation Via Backprop-Free Embedding Alignment, Xiao Ma, Young D. Kwon, Pan Zhou, Dong Ma

PhD Student’s Publications Collection

Test-Time Adaptation (TTA) adapts a deployed model during online inference to mitigate the impact of domain shift. While achieving strong accuracy, most existing methods rely on backpropagation, which is memory and computation intensive, making them unsuitable for resource-constrained devices. Recent attempts to reduce this overhead often suffer from high latency or are tied to specific architectures such as ViT-only or CNN-only. In this work, we revisit domain shift from an embedding perspective. Our analysis reveals that domain shift induces three distinct structural changes in the embedding space: translation (mean shift), scaling (variance shift), and rotation (covariance shift). Based on this …


Scalable Multi-Task Low-Rank Model Adaptation, Zichen Tian, Antoine Ledent, Qianru Sun Apr 2026

Scalable Multi-Task Low-Rank Model Adaptation, Zichen Tian, Antoine Ledent, Qianru Sun

PhD Student’s Publications Collection

Scaling multi-task low-rank adaptation (LoRA) to a large number of tasks induces catastrophic performance degradation, such as an accuracy drop from 88.2% to 2.0% on DOTA when scaling from 5 to 15 tasks. This failure is due to parameter and representation misalignment. We find that existing solutions, like regularization and dynamic routing, fail at scale because they are constrained by a fundamental trade-off: strengthening regularization to reduce inter-task conflict inadvertently suppresses the essential feature discrimination required for effective routing. In this work, we identify two root causes for this trade-off. First, uniform regularization disrupts inter-task knowledge sharing: shared underlying knowledge …


Semantic Entanglement In Vector-Based Retrieval: A Formal Framework And Context-Conditioned Disentanglement Pipeline For Agentic Rag Systems, Nick Loghmani Apr 2026

Semantic Entanglement In Vector-Based Retrieval: A Formal Framework And Context-Conditioned Disentanglement Pipeline For Agentic Rag Systems, Nick Loghmani

iSchool - All Scholarship

Retrieval-Augmented Generation (RAG) systems deployed in agentic environments depend on the geometric properties of vector representations to retrieve contextually appropriate evidence for autonomous reasoning. When source documents conflate multiple topics within contiguous text regions, standard vectorization pipelines produce embedding spaces in which semantically distinct content occupies overlapping geometric neighborhoods — a condition we term semantic entanglement. This paper formalizes semantic entanglement as a model-relative measure of cross-topic overlap, defines an Entanglement Index (EI) as a quantitative proxy, and argues that higher EI is associated with reduced attainable Top-K retrieval precision under cosine similarity retrieval. We introduce the Semantic Disentanglement …


Columnas: The Honors Program Newsletter At Bentley University, Amanda Li, Wilson Jan, Michael Raphael, Alexandra Rieckehoff, Karina Wu, Michael Shehata, Nilufar Noorian, Eloise Weintraub Apr 2026

Columnas: The Honors Program Newsletter At Bentley University, Amanda Li, Wilson Jan, Michael Raphael, Alexandra Rieckehoff, Karina Wu, Michael Shehata, Nilufar Noorian, Eloise Weintraub

Honors Program

INSIDE THE MODERN WORLD

Page 2: Stepping Out by Amanda Li

Page 3: Inside the Corporate Slop Bowl by Wilson Jan

Page 4: The Silencing: An Evaluation of the Global Attacks on the Right to Protest by Michael Raphael

THE SOUND OF CHANGE

Page 5: The Social, Cultural, and Economic Impact of Bad Bunny by Alexandra Rieckehoff

Page 6: Streaming Changed Music, But Is It Fair to Artists? by Karina Wu

Page 7: Feeling the Music: How Haptic Wearables Are Changing the Way We Experience Sound by Michael Shehata

SHIFTING SYSTEMS

Page 8: The Story Behind Davos, One of the …


Storycomposerai: Supporting Human-Ai Story Co-Creation Through Decomposition And Linking, Shuo Niu, Dylan Clements, Marina Margalit Nemanov, Hyungsin Kim Apr 2026

Storycomposerai: Supporting Human-Ai Story Co-Creation Through Decomposition And Linking, Shuo Niu, Dylan Clements, Marina Margalit Nemanov, Hyungsin Kim

Computer Science

GenAI's ability to produce text and images is increasingly incorporated into human-AI co-creation tasks such as storytelling and video editing. However, integrating GenAI into these tasks requires enabling users to retain control over editing individual story elements while ensuring that generated visuals remain coherent with the storyline and consistent across multiple AI-generated outputs. This work examines a paradigm of creative decomposition and linking, which allows creators to clearly communicate creative intent by prompting GenAI to tailor specific story elements, such as storylines, personas, locations, and scenes, while maintaining coherence among them. We implement and evaluate StoryComposerAI, a system that exemplifies …


Scenarioxp: A Complete Scenario-Based Testing Framework For The Exploration And Exploitation Of Autonomous Vehicle Validation Scenarios, Quentin Goss Apr 2026

Scenarioxp: A Complete Scenario-Based Testing Framework For The Exploration And Exploitation Of Autonomous Vehicle Validation Scenarios, Quentin Goss

Doctoral Dissertations and Master's Theses

Today is an age of exciting emerging technology where cutting-edge research in autonomous vehicles (AVs) reduces the active human participation in driving and extends awareness beyond human limitations of perception and reaction, improving driving safety and quality of the user experience as a result. The ever-increasing complexity of these autonomous systems poses many challenges towards the validation and verification (V\&V) of these complex systems under time and resource constraints, as the use of artificial intelligence and also the intricacy of the operating environment means that these systems are also black-box and non-deterministic. Scenario-based V\&V testing of such systems, which involves …


Study Of Output And Behavior Of Llms Using Confidence Framing In Prompt Engineering, Micah Parrilla Apr 2026

Study Of Output And Behavior Of Llms Using Confidence Framing In Prompt Engineering, Micah Parrilla

Doctoral Dissertations and Master's Theses

While prompt engineering is pivotal for shaping Large Language Model (LLM) outputs, the impact of confidence framing on behavioral calibration remains underexplored. This study investigates the ways in which psychological framing, utilizing techniques such as capability praise, role amplification, and doubt induction, affects linguistic tone, objective accuracy, and internal calibration. A 1,080-trial experimental matrix evaluated six diverse models across factual, logical, coding, and cyber security domains. Analysis using the Kruskal-Wallis H-test revealed highly significant behavioral shifts across all measured dimensions, providing conclusive evidence that the applied frames exert a substantial influence on model performance.

The findings identify a distinct cognitive …


Knowledge Distillation From A Large Vision-Language Model To Compact Students For Architectural Floor Plan Understanding, Kiran Silwal Apr 2026

Knowledge Distillation From A Large Vision-Language Model To Compact Students For Architectural Floor Plan Understanding, Kiran Silwal

Honors Theses

In this research, the use of a large vision-language model to train smaller, deployable models for architectural floor plan question answering is investigated. Reading a floor plan today requires either a human expert or a paid query to a proprietary model, and neither option is practical for real-estate platforms that must process thousands of units at scale. To address this problem, a knowledge distillation approach is employed in which a large teacher model (GPT-4.1-mini) generates labeled question-answer pairs from floor plan images, and smaller student models learn from those labels. The teacher produced 37,027 labeled pairs from 12,343 floor plan …


Rethinking News Classification Through A Multi-Dimensional Framework, Luana De Jesus Ferreira Apr 2026

Rethinking News Classification Through A Multi-Dimensional Framework, Luana De Jesus Ferreira

Honors Theses

This thesis proposes a multi-dimensional framework for news classification that evaluates articles across three independent dimensions: headline accuracy, language neutrality, and content reliability. These dimensions produce both a continuous reliability score and a five-tier interpretive scale, while additionally classifying articles by genre and topic. To operationalize this framework, a structured annotation protocol was developed and applied to a dataset of 373 news articles drawn from 79 outlets spanning a wide range of contemporary media ecosystem. A binary Logistic Regression classifier trained on the ISOT Fake News Dataset was then evaluated against this dataset to examine how a model trained on …


Deep Learning-Based Automated Pneumonia Detection From Chest X-Rays: A Comparative Study Of Custom Cnn And Transfer Learning Architectures, Ahmed Sajim Apr 2026

Deep Learning-Based Automated Pneumonia Detection From Chest X-Rays: A Comparative Study Of Custom Cnn And Transfer Learning Architectures, Ahmed Sajim

Honors Theses

Pneumonia is a leading global cause of mortality, claiming approximately 2.5 million lives an-nually and placing exceptional diagnostic pressure on radiologists in resource-limited settings. Manual interpretation of chest X-ray (CXR) images is time-consuming, subject to inter-observer variability, and limited by radiologist availability. This thesis presents a systematic investiga-tion into deep learning-based automated pneumonia detection comparing five convolutional neural network (CNN) architectures: a custom-designed 2D CNN and four pretrained transfer learning models—ResNet, DenseNet, MobileNet, and VGG19.

A targeted data augmentation pipeline addresses the severe class imbalance in the Kag-gle Chest X-Ray Pneumonia dataset, expanding the Normal class from 1,583 to 9,495 …


Sentra, David Aguilar Apr 2026

Sentra, David Aguilar

Posters - 2026

In the fast-paced space of event organization, fostering continuous collaboration among participants is essential. However, organizers often lose valuable time monitoring multiple, disconnected systems once an event is underway. Enter Sentra: an all-in-one Discord bot tailored specifically for weekend events like hackathons. Sentra bridges the gap between participants and organizers by consolidating seamless team matchmaking, robust support ticketing, and automated AI moderation into a single, unified interface.


Synapse, Nicolas Diaz, Alexander Murphy, Sonia Cerrillo, Naomi Ramirez, Jesse Kemmer Apr 2026

Synapse, Nicolas Diaz, Alexander Murphy, Sonia Cerrillo, Naomi Ramirez, Jesse Kemmer

Posters - 2026

People tend to accumulate a great deal of notes throughout their lives with no coherent way to organize them. Even with the built-in notes app, the notes eventually accumulate until it becomes borderline impossible to find what is needed. Our proposed solution is Synapse, an LLM powered notes app with a tagging system that allows notes to be sorted by topic. The LLM will be able to read the user's notes and recommend tags


Ai Dependence And Its Impact On Human Decision-Making Quality And Supply Chain Efficiency, Jesus Salazar, Leonardo Fabbri, Axel Villegas Apr 2026

Ai Dependence And Its Impact On Human Decision-Making Quality And Supply Chain Efficiency, Jesus Salazar, Leonardo Fabbri, Axel Villegas

Posters - 2026

  • Artificial Intelligence (AI) is transforming supply chain management by enabling:
    • Data-driven decision-making
    • Improved forecasting accuracy
    • Enhanced operational efficiency (Choudhary et al., 2023; Ivanov & Dolgui, 2021)
  • AI applications such as predictive analytics support:
    • Inventory optimization, Logistics planning
    • Procurement decisions in real time
  • However, increasing reliance on AI introduces risks:
    • Automation bias (over-trusting AI outputs)
    • Reduced human critical thinking
    • Overdependence on algorithmic recommendations (Raisch & Krakowski, 2021)
  • This study examines the dual impact of AI dependence on:
    • Decision-making quality
    • Supply chain efficiency
  • Objective:
    • Identify whether AI improves performance or reduces human effectiveness
    • Determine the optimal balance between AI support and human …


From Attention To Reasoning: Beyond Accuracy In Multimodal Ai, Wayner Barrios Apr 2026

From Attention To Reasoning: Beyond Accuracy In Multimodal Ai, Wayner Barrios

Dartmouth College Ph.D Dissertations

Multimodal large language models have achieved impressive performance on vision-language benchmarks by integrating visual encoders with large language models. Yet a critical gap persists between benchmark accuracy and genuine multimodal understanding: current evaluation frameworks assess performance by final answers alone, rewarding confident predictions while leaving systematic reasoning failures undetected.

This thesis addresses this gap through a unified framework that progresses from understanding to reasoning, using video as the most comprehensive multimodal testbed. Video inherently combines vision, audio, and language with temporal dynamics and massive token redundancy; techniques developed for video's comprehensive challenges transfer naturally to simpler multimodal tasks.

On understanding …


Discrete Diffusion For Bundle Construction, Teng Tu, Ai Li, Yunshan Ma, Shuo Xu, Xiaohao Liu, Haokai Ma, Liang Pang, Tat-Seng Chua Apr 2026

Discrete Diffusion For Bundle Construction, Teng Tu, Ai Li, Yunshan Ma, Shuo Xu, Xiaohao Liu, Haokai Ma, Liang Pang, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

As a central task in product bundling, bundle construction aims to select a subset of items from large item catalogs to build an entire bundle or, more practically, complete a partial bundle. Existing methods often rely on the sequential construction paradigm that predicts items one at a time, nevertheless, this paradigm is fundamentally unsuitable for the essentially unordered bundles. In contrast, non-sequential methods model a bundle as a set, but still face two dimensionality curses: the combinatorial space grows exponentially with both bundle length and catalog size. Accordingly, we identify two technical challenges: 1) how to effectively and efficiently model …


Enhancing Financial Audit Operations Through Ai Anomaly Detection, Nadya Cousin, Rebekah Garza, Talisa Gomez Apr 2026

Enhancing Financial Audit Operations Through Ai Anomaly Detection, Nadya Cousin, Rebekah Garza, Talisa Gomez

Posters - 2026

❖ Financial auditing plays a critical role in ensuring accuracy, regulatory compliance, and fraud detection in financial reporting

❖ Traditional audit approaches rely heavily on sampling and manual review processes, limiting their ability to scale with increasing data complexity

❖ The rapid growth of high-volume, high-velocity financial data (big data) has exposed significant limitations in traditional auditing, including:

  • Incomplete data coverage
  • Delayed anomaly detection
  • Increased risk of material misstatements

❖ These limitations create a need for scalable, automated, and data-driven audit solutions

❖ Artificial Intelligence (AI), particularly anomaly detection models, enables:

  •  Full-population testing
  •  Real-time pattern recognition
  •  Proactive risk identification


The Psychology Behind Ai-Generated Phishing And Social Engineering Attacks, A’Shya Reynolds Apr 2026

The Psychology Behind Ai-Generated Phishing And Social Engineering Attacks, A’Shya Reynolds

School of Cybersecurity Master's Level Projects and Papers

Cybercrime has evolved significantly with the integration of artificial intelligence (AI), transforming traditional phishing and social engineering attacks into highly sophisticated and personalized threats. While early phishing attempts relied on generic messaging and low success rates, modern AI-driven attacks leverage advanced data analytics, natural language processing, and behavioral prediction to manipulate victims more effectively.

This research examines how cybercriminals utilize AI to enhance psychological manipulation techniques in phishing and social engineering attacks, increasing victim susceptibility. Drawing from interdisciplinary literature in cybersecurity and psychology, this study explores key psychological mechanisms, including cognitive biases, emotional triggers, and decision-making processes that influence victim …


Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng Apr 2026

Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng

Research Collection Lee Kong Chian School Of Business

Existing research on AI cultural biases predominantly focuses on Western models, overlooking critical gaps in non-Western models. We conduct a comparative analysis of AI models – ChatGPT (U.S. developed) and ErnieBot (China developed) – from different cultures to investigate how corresponding cultural biases manifest in their outputs. Additionally, we examine how cultural alignment between human users and AI models impacts their collaborative creative performance and the underlying psychological mechanisms. In Study 1, multi-choice prompt with zero-shot technique was used to evaluate cultural biases in four widely used AI models – ChatGPT-3.5/4, ErnieBot-3.5/4 – comparing their responses to established cultural psychometric …


Deep Learning Based Approaches For Low Cost Defense Detection, Adele J. Noel-Rickert Apr 2026

Deep Learning Based Approaches For Low Cost Defense Detection, Adele J. Noel-Rickert

All NMU Master's Theses

Pulmonary fibrosis is a progressive interstitial lung disease characterized by the accumulation of fibrotic tissue within the lungs, leading to impaired respiratory function and reduced quality of life. Early detection is important for disease management; however, accurate diagnosis often relies on high-resolution computed tomography (CT), which may not be accessible in all clinical settings. Chest radiography provides a lower-cost and widely available imaging modality, but interpretation of chest X-rays for fibrotic disease can be challenging due to subtle radiographic patterns and overlapping anatomical structures. This thesis investigates the use of multimodal deep learning techniques to assist in pul- monary fibrosis …


Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy Apr 2026

Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy

Publications and Research

Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools—training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures—encodes a set of assumptions that systematically disadvantages speakers of underrepresented languages before a single model is trained. This paper examines those assumptions through the lens of Bengali, one of the world’s most widely spoken languages with roughly 285 million speakers (Ethnologue, 2025; International Communication and Leadership School, 2026), and the structural barriers that emerge when attempting to build AI-assisted educational tools for Bengali-speaking learners in low-connectivity …


Portrait Shadow Removal Via Self-Exemplar Illumination Equalization, Qian Huang, Cheng Xu, Guiqing Li, Ziheng Wu, Shengxin Liu, Shengfeng He Apr 2026

Portrait Shadow Removal Via Self-Exemplar Illumination Equalization, Qian Huang, Cheng Xu, Guiqing Li, Ziheng Wu, Shengxin Liu, Shengfeng He

Research Collection School Of Computing and Information Systems

We introduce the Self-Exemplar Illumination Equalization Network, designed specifically for effective portrait shadow removal. The core idea of our method is that partially shadowed portraits can find ideal exemplars within their non-shadowed facial regions. Rather than directly fusing two distinct classes of facial features, our approach utilizes non-shadowed regions as an illumination indicator to equalize the shadowed regions, generating deshadowed results without boundary-merging artifacts. Our network comprises cascaded Self-Exemplar Illumination Equalization Blocks (SExmBlock), each containing two modules: a self-exemplar feature matching module and a feature-level illumination rectification module. The former identifies and applies internal illumination exemplars to shadowed areas, producing …


Weakly Supervised Video Anomaly Detection And Localization With Spatio-Temporal Prompts, Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang, Qingsen Yan, Peng Wang, Yanning Zhang Apr 2026

Weakly Supervised Video Anomaly Detection And Localization With Spatio-Temporal Prompts, Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang, Qingsen Yan, Peng Wang, Yanning Zhang

Research Collection School Of Computing and Information Systems

Current weakly supervised video anomaly detection (WSVAD) task aims to achieve frame-level anomalous event detection with only coarse video-level annotations available. Existing works typically involve extracting global features from full-resolution video frames and training frame-level classifiers to detect anomalies in the temporal dimension. However, most anomalous events tend to occur in localized spatial regions rather than the entire video frames, which implies existing frame-level feature based works may be misled by the dominant background information and lack the interpretation of the detected anomalies. To address this dilemma, this paper introduces a novel method called STPrompt that learns spatio-temporal prompt embeddings …


Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong Apr 2026

Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong

Research Collection School Of Computing and Information Systems

Conventionally, human intuition defines vision as a modality of passive optical sensing, relying on ambient light to perceive the environment. However, active optical sensing, which involves emitting and receiving signals, offers unique advantages by capturing both radiometric and geometric properties of the environment, independent of external illumination conditions. This work focuses on advancing active optical sensing using Light Detection and Ranging (LiDAR), which captures intensity data, enabling the estimation of surface reflectance that remains invariant under varying illumination. Such properties are crucial for robotic perception tasks, including detection, recognition, segmentation, and Simultaneous Localization and Mapping (SLAM). A key challenge with …


Generative Ai In Enterprises: Optimizing Applications With Large Language Models, Donghao Huang Apr 2026

Generative Ai In Enterprises: Optimizing Applications With Large Language Models, Donghao Huang

Dissertations and Theses Collection (Open Access)

This dissertation investigates how to deploy Large Language Models (LLMs) effectively in enterprise settings, where accuracy, reliability, cost, privacy, and operational constraints often matter more than benchmark performance alone. Drawing on seventeen peer-reviewed publications (eleven published and six accepted for publication), the work develops and validates optimization strategies across three connected themes: retrieval-augmented generation (RAG), agentic AI for workflow automation, and deployment guidelines for real-world enterprise environments.

First, we study RAG optimization through systematic evaluation of open and proprietary models, highlighting conditions under which efficient open-weight models can match or exceed proprietary alternatives. To address a pervasive failure mode in …


Low-Complexity Structured Neural Networks And Their Usage In Image And Signal Processing, Adam Kuzmicki Apr 2026

Low-Complexity Structured Neural Networks And Their Usage In Image And Signal Processing, Adam Kuzmicki

Doctoral Dissertations and Master's Theses

Conventional neural networks face significant challenges due to high computational costs, large parameter counts, and reliance on backpropagation, which restricts their application in resource-constrained and real-time settings. To address these challenges, this thesis proposes three structured neural network (NN) architectures grounded in the theories of sparse and self-contained factorizations of transforms, with applications to image compression, reconstruction, classification, encryption, and also adaptive wideband multi-beam beamforming. The first neural network architecture, named DCTrix-Net, replaces conventional spatial con- volution with highly sparse factorization of the discrete Cosine transform (DCT) complemented by Toeplitz-structured weight initialization, achieving at least 97% FLOP reduction over CNNs, …


Ai Interpretability In Healthcare Communication, Ananya Jeyappragash Apr 2026

Ai Interpretability In Healthcare Communication, Ananya Jeyappragash

Dartmouth College Master’s Theses

Artificial intelligence has increasingly been adopted in healthcare, largely for specialized tasks and under significant human oversight. The use of large black-box systems raises important concerns about transparency in high-stakes environments such as clinical decision-making. Clinical communication is fundamentally human-centered, and failures in judgment can have serious consequences for patient care. Overestimating the reasoning abilities of large language models may lead to undue trust in fabricated or “hallucinated” outputs, while rejecting AI-assisted tools altogether may preserve inefficient workflows and contribute to missed or delayed diagnoses. These concerns reflect a broader tradeoff between accuracy and interpretability: although more complex models may …


Reducing Class-Wise Performance Disparity Via Margin Regularization, Beier Zhu, Kesen Zhao, Jiequan Cui, Qianru Sun, Yuan Zhou, Xun Yang, Hanwang Zhang Apr 2026

Reducing Class-Wise Performance Disparity Via Margin Regularization, Beier Zhu, Kesen Zhao, Jiequan Cui, Qianru Sun, Yuan Zhou, Xun Yang, Hanwang Zhang

Research Collection School Of Computing and Information Systems

Deep neural networks often exhibit substantial disparities in class-wise accuracy, even when trained on class-balanced data—posing concerns for reliable deployment. While prior efforts have explored empirical remedies, a theoretical understanding of such performance disparities in classification remains limited. In this work, we present Margin Regularization for performance disparity Reduction (MR2 ), a theoretically principled regularization for classification by dynamically adjusting margins in both the logit and representation spaces. Our analysis establishes a margin-based, class-sensitive generalization bound that reveals how per-class feature variability contributes to error, motivating the use of larger margins for “hard” classes. Guided by this insight, MR2 optimizes …


Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou Apr 2026

Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou

Research Collection School Of Computing and Information Systems

While text-to-3D generation has attracted growing interest, existing methods often struggle to produce 3D assets that align well with human preferences. Current preference alignment techniques for 3D content typically rely on hardly-collected preference-paired multi-view 2D images to train 2D reward models, when then guide 3D generation — leading to geometric artifacts, such as the Janus face problem and geometric incompleteness, due to their inherent 2D bias. To address these limitations, we construct 3D-MeshPref, the first large-scale unpaired 3D preference dataset, featuring diverse 3D meshes annotated by a large language model and refined by human evaluators. We then develop RewardCS, the …


Real-Time Motion-Controllable Autoregressive Video Diffusion, Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou, Xiaolong Shen, Yuan Zhou, Qianru Sun, Hanwang Zhang Apr 2026

Real-Time Motion-Controllable Autoregressive Video Diffusion, Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou, Xiaolong Shen, Yuan Zhou, Qianru Sun, Hanwang Zhang

Research Collection School Of Computing and Information Systems

Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to simple control signals or text-to-video generation, and often suffer from quality degradation and motion artifacts in few-step generation. To address these challenges, we propose AR-Drag, the first RL-enhanced few-step AR video diffusion model for real-time image-to-video generation with diverse motion control. We first fine-tune a base I2V model to support basic motion control, then further improve it via reinforcement learning with a trajectory-based reward model. Our design preserves the …


Be Responsible In Your Answers! Monitoring Out-Of-Domain Behaviors In Domain-Specific Llms, Boquan Li, Chenzhe Lou, Zhe Ren, Peixin Zhang, Zirui Fu, Jun Sun, Yaowen Zheng Apr 2026

Be Responsible In Your Answers! Monitoring Out-Of-Domain Behaviors In Domain-Specific Llms, Boquan Li, Chenzhe Lou, Zhe Ren, Peixin Zhang, Zirui Fu, Jun Sun, Yaowen Zheng

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have accelerated the rapid development of chatbot web applications in various domains, such as coding, biomedicine and psychology. Compared to general LLMs like ChatGPT, domain-specific LLMs require a greater sense of responsibility. For instance, if a programming LLM casually answers medical or psychological questions, it not only misleads the public but also poses legal risks. This highlights new demands for monitoring and preventing such irresponsible behaviors. Existing efforts attempt to monitor LLMs from multiple aspects, such as lying, jailbreaks, and toxic content, while overlooking out-of-domain behaviors. In this work, we propose an innovative LLM domain monitoring …