Open Access. Powered by Scholars. Published by Universities.®

Digital Commons Network™

Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1261 - 1290 of 63035

Full-Text Articles in Entire DC Network

Aiw26s: Machine Learning Of Structured Data, Moumita Saha Apr 2026

Aiw26s: Machine Learning Of Structured Data, Moumita Saha

Paul English Applied Artificial Intelligence (AI) Institute Publications

This workshop introduces the fundamentals of machine learning for structured data, focusing on tabular datasets and real-world applications. Participants explore key concepts such as data types, data preprocessing, feature engineering, and supervised learning methods. The session covers commonly used models, including linear regression, logistic regression, decision trees, and neural networks, along with evaluation metrics such as RMSE, accuracy, and confusion matrices. By the end of the workshop, participants will have gained a practical understanding of how to build, interpret, and evaluate machine learning models for structured data.


From Attention To Reasoning: Beyond Accuracy In Multimodal Ai, Wayner Barrios Apr 2026

From Attention To Reasoning: Beyond Accuracy In Multimodal Ai, Wayner Barrios

Dartmouth College Ph.D Dissertations

Multimodal large language models have achieved impressive performance on vision-language benchmarks by integrating visual encoders with large language models. Yet a critical gap persists between benchmark accuracy and genuine multimodal understanding: current evaluation frameworks assess performance by final answers alone, rewarding confident predictions while leaving systematic reasoning failures undetected.

This thesis addresses this gap through a unified framework that progresses from understanding to reasoning, using video as the most comprehensive multimodal testbed. Video inherently combines vision, audio, and language with temporal dynamics and massive token redundancy; techniques developed for video's comprehensive challenges transfer naturally to simpler multimodal tasks.

On understanding …


Discrete Diffusion For Bundle Construction, Teng Tu, Ai Li, Yunshan Ma, Shuo Xu, Xiaohao Liu, Haokai Ma, Liang Pang, Tat-Seng Chua Apr 2026

Discrete Diffusion For Bundle Construction, Teng Tu, Ai Li, Yunshan Ma, Shuo Xu, Xiaohao Liu, Haokai Ma, Liang Pang, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

As a central task in product bundling, bundle construction aims to select a subset of items from large item catalogs to build an entire bundle or, more practically, complete a partial bundle. Existing methods often rely on the sequential construction paradigm that predicts items one at a time, nevertheless, this paradigm is fundamentally unsuitable for the essentially unordered bundles. In contrast, non-sequential methods model a bundle as a set, but still face two dimensionality curses: the combinatorial space grows exponentially with both bundle length and catalog size. Accordingly, we identify two technical challenges: 1) how to effectively and efficiently model …


Chopchop: The Digital Cookbook, Dominc Mcdevitt, Shane Misley, Adolfo Duran, Katie Cerda, Kobie Henson Apr 2026

Chopchop: The Digital Cookbook, Dominc Mcdevitt, Shane Misley, Adolfo Duran, Katie Cerda, Kobie Henson

Posters - 2026

Many of today's home chefs still use the same limited methods of saving recipes that have been used for decades, i.e. handwritten notes, disorganized pdfs, screenshots, saved text messages, etc. Not only are these formats hard to keep track of and easily lost, but they also suffer the risk of becoming irrevocably damaged or stained in the cooking process. They are also notoriously hard to edit, which limits a chef's ability to tailor recipes to their taste, their available ingredients, or even just a different serving size. Another major issue with these approaches is the lack of easy sharing. Giving …


Fallen Light, Joshua Do Apr 2026

Fallen Light, Joshua Do

Posters - 2026

Lucifer is often portrayed in popular media as purely evil; however, this interpretation overlooks the more complex idea of gradual moral corruption through deception and pride. The purpose of Fallen Light is to explore how doubt, pride, and subtle deception can lead even a highly exalted being away from God over time.. The game emphasizes the internal struggle between obedience and self-exaltation rather than immediate rebellion, and the problems with Being overly prideful.


Enhancing Financial Audit Operations Through Ai Anomaly Detection, Nadya Cousin, Rebekah Garza, Talisa Gomez Apr 2026

Enhancing Financial Audit Operations Through Ai Anomaly Detection, Nadya Cousin, Rebekah Garza, Talisa Gomez

Posters - 2026

❖ Financial auditing plays a critical role in ensuring accuracy, regulatory compliance, and fraud detection in financial reporting

❖ Traditional audit approaches rely heavily on sampling and manual review processes, limiting their ability to scale with increasing data complexity

❖ The rapid growth of high-volume, high-velocity financial data (big data) has exposed significant limitations in traditional auditing, including:

  • Incomplete data coverage
  • Delayed anomaly detection
  • Increased risk of material misstatements

❖ These limitations create a need for scalable, automated, and data-driven audit solutions

❖ Artificial Intelligence (AI), particularly anomaly detection models, enables:

  •  Full-population testing
  •  Real-time pattern recognition
  •  Proactive risk identification


The Psychology Behind Ai-Generated Phishing And Social Engineering Attacks, A’Shya Reynolds Apr 2026

The Psychology Behind Ai-Generated Phishing And Social Engineering Attacks, A’Shya Reynolds

School of Cybersecurity Master's Level Projects and Papers

Cybercrime has evolved significantly with the integration of artificial intelligence (AI), transforming traditional phishing and social engineering attacks into highly sophisticated and personalized threats. While early phishing attempts relied on generic messaging and low success rates, modern AI-driven attacks leverage advanced data analytics, natural language processing, and behavioral prediction to manipulate victims more effectively.

This research examines how cybercriminals utilize AI to enhance psychological manipulation techniques in phishing and social engineering attacks, increasing victim susceptibility. Drawing from interdisciplinary literature in cybersecurity and psychology, this study explores key psychological mechanisms, including cognitive biases, emotional triggers, and decision-making processes that influence victim …


Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng Apr 2026

Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng

Research Collection Lee Kong Chian School Of Business

Existing research on AI cultural biases predominantly focuses on Western models, overlooking critical gaps in non-Western models. We conduct a comparative analysis of AI models – ChatGPT (U.S. developed) and ErnieBot (China developed) – from different cultures to investigate how corresponding cultural biases manifest in their outputs. Additionally, we examine how cultural alignment between human users and AI models impacts their collaborative creative performance and the underlying psychological mechanisms. In Study 1, multi-choice prompt with zero-shot technique was used to evaluate cultural biases in four widely used AI models – ChatGPT-3.5/4, ErnieBot-3.5/4 – comparing their responses to established cultural psychometric …


Deep Learning Based Approaches For Low Cost Defense Detection, Adele J. Noel-Rickert Apr 2026

Deep Learning Based Approaches For Low Cost Defense Detection, Adele J. Noel-Rickert

All NMU Master's Theses

Pulmonary fibrosis is a progressive interstitial lung disease characterized by the accumulation of fibrotic tissue within the lungs, leading to impaired respiratory function and reduced quality of life. Early detection is important for disease management; however, accurate diagnosis often relies on high-resolution computed tomography (CT), which may not be accessible in all clinical settings. Chest radiography provides a lower-cost and widely available imaging modality, but interpretation of chest X-rays for fibrotic disease can be challenging due to subtle radiographic patterns and overlapping anatomical structures. This thesis investigates the use of multimodal deep learning techniques to assist in pul- monary fibrosis …


Explainable Artificial Intelligence In The Image Domain And Its Applications To The Medical Field, Mirtha Lucas Apr 2026

Explainable Artificial Intelligence In The Image Domain And Its Applications To The Medical Field, Mirtha Lucas

Theses and Dissertations from DePaul University

This dissertation investigates the development of Explainable Artificial Intelligence (XAI) methods for deep learning models in the image domain, with a particular focus on medical imaging applications. Although neural networks achieve high predictive performance, their lack of interpretability limits their adoption in critical domains such as healthcare, where transparency and trust are essential. This work addresses this challenge by proposing novel approaches that improve the interpretability and reliability of model predictions.   A primary contribution is the introduction of Riemann–Stieltjes Integrated Grad-CAM (RSI Grad-CAM), a gradient-based attribution method that generates more relevant and spatially localized saliency maps. The method is evaluated …


Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy Apr 2026

Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy

Publications and Research

Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools—training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures—encodes a set of assumptions that systematically disadvantages speakers of underrepresented languages before a single model is trained. This paper examines those assumptions through the lens of Bengali, one of the world’s most widely spoken languages with roughly 285 million speakers (Ethnologue, 2025; International Communication and Leadership School, 2026), and the structural barriers that emerge when attempting to build AI-assisted educational tools for Bengali-speaking learners in low-connectivity …


Llm-Driven Mission Control And Autonomous Planning For Search-And-Rescue Uavs: A Simulation-Based Evaluation, Naser Bader Alsaedi Apr 2026

Llm-Driven Mission Control And Autonomous Planning For Search-And-Rescue Uavs: A Simulation-Based Evaluation, Naser Bader Alsaedi

Theses

Unmanned aerial vehicles (UAVs) are increasingly used in search‑and‑rescue (SAR) missions, yet many systems still rely on fragmented software where mission design, perception, and flight control are configured separately. This thesis examines whether a unified AI‑driven framework can reduce configuration effort and operator workload in UAV‑based SAR operations. The proposed system integrates natural‑language mission specification using a large language model (LLM) (LLaMA 3.1), autonomous coverage planning, YOLOv8‑based victim detection, and PX4/MAVSDK control within a single architecture. Operators describe missions through free‑form text or a graphical interface; the model converts these descriptions into structured mission parameters that are automatically planned and …


Portrait Shadow Removal Via Self-Exemplar Illumination Equalization, Qian Huang, Cheng Xu, Guiqing Li, Ziheng Wu, Shengxin Liu, Shengfeng He Apr 2026

Portrait Shadow Removal Via Self-Exemplar Illumination Equalization, Qian Huang, Cheng Xu, Guiqing Li, Ziheng Wu, Shengxin Liu, Shengfeng He

Research Collection School Of Computing and Information Systems

We introduce the Self-Exemplar Illumination Equalization Network, designed specifically for effective portrait shadow removal. The core idea of our method is that partially shadowed portraits can find ideal exemplars within their non-shadowed facial regions. Rather than directly fusing two distinct classes of facial features, our approach utilizes non-shadowed regions as an illumination indicator to equalize the shadowed regions, generating deshadowed results without boundary-merging artifacts. Our network comprises cascaded Self-Exemplar Illumination Equalization Blocks (SExmBlock), each containing two modules: a self-exemplar feature matching module and a feature-level illumination rectification module. The former identifies and applies internal illumination exemplars to shadowed areas, producing …


Weakly Supervised Video Anomaly Detection And Localization With Spatio-Temporal Prompts, Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang, Qingsen Yan, Peng Wang, Yanning Zhang Apr 2026

Weakly Supervised Video Anomaly Detection And Localization With Spatio-Temporal Prompts, Peng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang, Qingsen Yan, Peng Wang, Yanning Zhang

Research Collection School Of Computing and Information Systems

Current weakly supervised video anomaly detection (WSVAD) task aims to achieve frame-level anomalous event detection with only coarse video-level annotations available. Existing works typically involve extracting global features from full-resolution video frames and training frame-level classifiers to detect anomalies in the temporal dimension. However, most anomalous events tend to occur in localized spatial regions rather than the entire video frames, which implies existing frame-level feature based works may be misled by the dominant background information and lack the interpretation of the detected anomalies. To address this dilemma, this paper introduces a novel method called STPrompt that learns spatio-temporal prompt embeddings …


Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong Apr 2026

Super Lidar Intensity For Robotic Perception, Wei Gao, Jie Zhang, Mingle Zhao, Zhiyuan Zhang, Shu Kong, Maani Ghaffari, Dezhen Song, Chengzhong Xu, Hui Kong

Research Collection School Of Computing and Information Systems

Conventionally, human intuition defines vision as a modality of passive optical sensing, relying on ambient light to perceive the environment. However, active optical sensing, which involves emitting and receiving signals, offers unique advantages by capturing both radiometric and geometric properties of the environment, independent of external illumination conditions. This work focuses on advancing active optical sensing using Light Detection and Ranging (LiDAR), which captures intensity data, enabling the estimation of surface reflectance that remains invariant under varying illumination. Such properties are crucial for robotic perception tasks, including detection, recognition, segmentation, and Simultaneous Localization and Mapping (SLAM). A key challenge with …


Managing Reproducibility Debt In Scientific Software: A Practical Framework, Zara Hassan, Christoph Treude, Graham Williams, Michael Norrish, Alex Potanin Apr 2026

Managing Reproducibility Debt In Scientific Software: A Practical Framework, Zara Hassan, Christoph Treude, Graham Williams, Michael Norrish, Alex Potanin

Research Collection School Of Computing and Information Systems

Scientific software includes end-user applications, modelling tools, research software for publications, and production systems for real users. It plays a key role across various scientific disciplines by enabling large-scale computation, simulation, and data analysis. Unlike commercial software, scientific software is often developed in dynamic research environments with limited engineering practices, documentation, or testing. This makes it fragile and difficult to reproduce results, even when code and data are available, conditions in which Reproducibility Debt (RpD) accumulates. This paper presents the Reproducibility Debt Management Framework (RpD-MF), which is grounded in evidence from a systematic literature review, practitioner interviews, and a global …


Generative Ai In Enterprises: Optimizing Applications With Large Language Models, Donghao Huang Apr 2026

Generative Ai In Enterprises: Optimizing Applications With Large Language Models, Donghao Huang

Dissertations and Theses Collection (Open Access)

This dissertation investigates how to deploy Large Language Models (LLMs) effectively in enterprise settings, where accuracy, reliability, cost, privacy, and operational constraints often matter more than benchmark performance alone. Drawing on seventeen peer-reviewed publications (eleven published and six accepted for publication), the work develops and validates optimization strategies across three connected themes: retrieval-augmented generation (RAG), agentic AI for workflow automation, and deployment guidelines for real-world enterprise environments.

First, we study RAG optimization through systematic evaluation of open and proprietary models, highlighting conditions under which efficient open-weight models can match or exceed proprietary alternatives. To address a pervasive failure mode in …


Fingerprinting Voice Commands Of Vpn-Protected Smart Speakers, Xiaoguang Guo, Keyang Yu, Qi Li, Dong Chen Apr 2026

Fingerprinting Voice Commands Of Vpn-Protected Smart Speakers, Xiaoguang Guo, Keyang Yu, Qi Li, Dong Chen

Computer Science Faculty Research and Publications

Extensive recent research has shown that it is surprisingly easy to infer Amazon Alexa voice commands over their network traffic data. To prevent these traffic analytics (TA)-based inference attacks, smart home owners are considering deploying virtual private networks (VPNs) to safeguard their smart speakers. In this work, we design a new machine learning-powered attack framework—VoiceAttack that could still accurately fingerprint voice commands on VPN-encrypted voice speaker network traffic. We evaluate VoiceAttack under 5 different real-world settings using Amazon Alexa and Google Home. Our results show that VoiceAttack could correctly infer voice command sentences with a Matthews Correlation Coefficient (MCC) of …


Low-Complexity Structured Neural Networks And Their Usage In Image And Signal Processing, Adam Kuzmicki Apr 2026

Low-Complexity Structured Neural Networks And Their Usage In Image And Signal Processing, Adam Kuzmicki

Doctoral Dissertations and Master's Theses

Conventional neural networks face significant challenges due to high computational costs, large parameter counts, and reliance on backpropagation, which restricts their application in resource-constrained and real-time settings. To address these challenges, this thesis proposes three structured neural network (NN) architectures grounded in the theories of sparse and self-contained factorizations of transforms, with applications to image compression, reconstruction, classification, encryption, and also adaptive wideband multi-beam beamforming. The first neural network architecture, named DCTrix-Net, replaces conventional spatial con- volution with highly sparse factorization of the discrete Cosine transform (DCT) complemented by Toeplitz-structured weight initialization, achieving at least 97% FLOP reduction over CNNs, …


Ai Interpretability In Healthcare Communication, Ananya Jeyappragash Apr 2026

Ai Interpretability In Healthcare Communication, Ananya Jeyappragash

Dartmouth College Master’s Theses

Artificial intelligence has increasingly been adopted in healthcare, largely for specialized tasks and under significant human oversight. The use of large black-box systems raises important concerns about transparency in high-stakes environments such as clinical decision-making. Clinical communication is fundamentally human-centered, and failures in judgment can have serious consequences for patient care. Overestimating the reasoning abilities of large language models may lead to undue trust in fabricated or “hallucinated” outputs, while rejecting AI-assisted tools altogether may preserve inefficient workflows and contribute to missed or delayed diagnoses. These concerns reflect a broader tradeoff between accuracy and interpretability: although more complex models may …


Short-Term Electricity Price Forecasting With Constrained Regressors, Mucun Sun, Li Zhang, Yifeng Gao Apr 2026

Short-Term Electricity Price Forecasting With Constrained Regressors, Mucun Sun, Li Zhang, Yifeng Gao

Computer Science Faculty Publications

The volatility of electricity price presents a challenge to market participants as their decision-making process are highly depend on the accuracy of price forecasts. However, there is growing empirical evidence of increasing price volatility and price spikes in electricity markets as a result of variable renewable energy generation, extreme weather events, and other factors. The distribution shift caused by spikes in electricity price data differentiates the forecasting tasks from other renewable energy sources. Moreover, the observations may be compromised by cyberattacks and thus not available in the testing phase. To this end, we propose a Similarity-Enhanced Electricity Decomposition Forecasting model …


Reducing Class-Wise Performance Disparity Via Margin Regularization, Beier Zhu, Kesen Zhao, Jiequan Cui, Qianru Sun, Yuan Zhou, Xun Yang, Hanwang Zhang Apr 2026

Reducing Class-Wise Performance Disparity Via Margin Regularization, Beier Zhu, Kesen Zhao, Jiequan Cui, Qianru Sun, Yuan Zhou, Xun Yang, Hanwang Zhang

Research Collection School Of Computing and Information Systems

Deep neural networks often exhibit substantial disparities in class-wise accuracy, even when trained on class-balanced data—posing concerns for reliable deployment. While prior efforts have explored empirical remedies, a theoretical understanding of such performance disparities in classification remains limited. In this work, we present Margin Regularization for performance disparity Reduction (MR2 ), a theoretically principled regularization for classification by dynamically adjusting margins in both the logit and representation spaces. Our analysis establishes a margin-based, class-sensitive generalization bound that reveals how per-class feature variability contributes to error, motivating the use of larger margins for “hard” classes. Guided by this insight, MR2 optimizes …


Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou Apr 2026

Dreamcs: Geometry-Aware Text-To-3d Generation With Unpaired 3d Reward Supervision, Xiandong Zou, Ruihao Xia, Hongsong Wang, Pan Zhou

Research Collection School Of Computing and Information Systems

While text-to-3D generation has attracted growing interest, existing methods often struggle to produce 3D assets that align well with human preferences. Current preference alignment techniques for 3D content typically rely on hardly-collected preference-paired multi-view 2D images to train 2D reward models, when then guide 3D generation — leading to geometric artifacts, such as the Janus face problem and geometric incompleteness, due to their inherent 2D bias. To address these limitations, we construct 3D-MeshPref, the first large-scale unpaired 3D preference dataset, featuring diverse 3D meshes annotated by a large language model and refined by human evaluators. We then develop RewardCS, the …


Real-Time Motion-Controllable Autoregressive Video Diffusion, Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou, Xiaolong Shen, Yuan Zhou, Qianru Sun, Hanwang Zhang Apr 2026

Real-Time Motion-Controllable Autoregressive Video Diffusion, Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou, Xiaolong Shen, Yuan Zhou, Qianru Sun, Hanwang Zhang

Research Collection School Of Computing and Information Systems

Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to simple control signals or text-to-video generation, and often suffer from quality degradation and motion artifacts in few-step generation. To address these challenges, we propose AR-Drag, the first RL-enhanced few-step AR video diffusion model for real-time image-to-video generation with diverse motion control. We first fine-tune a base I2V model to support basic motion control, then further improve it via reinforcement learning with a trajectory-based reward model. Our design preserves the …


Be Responsible In Your Answers! Monitoring Out-Of-Domain Behaviors In Domain-Specific Llms, Boquan Li, Chenzhe Lou, Zhe Ren, Peixin Zhang, Zirui Fu, Jun Sun, Yaowen Zheng Apr 2026

Be Responsible In Your Answers! Monitoring Out-Of-Domain Behaviors In Domain-Specific Llms, Boquan Li, Chenzhe Lou, Zhe Ren, Peixin Zhang, Zirui Fu, Jun Sun, Yaowen Zheng

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have accelerated the rapid development of chatbot web applications in various domains, such as coding, biomedicine and psychology. Compared to general LLMs like ChatGPT, domain-specific LLMs require a greater sense of responsibility. For instance, if a programming LLM casually answers medical or psychological questions, it not only misleads the public but also poses legal risks. This highlights new demands for monitoring and preventing such irresponsible behaviors. Existing efforts attempt to monitor LLMs from multiple aspects, such as lying, jailbreaks, and toxic content, while overlooking out-of-domain behaviors. In this work, we propose an innovative LLM domain monitoring …


Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo Apr 2026

Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo

Research Collection School Of Computing and Information Systems

Penetration testing is essential for identifying vulnerabilities in web applications before real adversaries can exploit them. Recent work has explored automating this process with Large Language Model (LLM)-powered agents, but existing approaches either rely on a single generic agent that struggles in complex scenarios or narrowly specialized agents that cannot adapt to diverse vulnerability types. We therefore introduce PenForge, a framework that dynamically constructs expert agents during testing rather than relying on those prepared beforehand. By integrating automated reconnaissance of potential attack surfaces with agents instantiated on the fly for context-aware exploitation, PenForge achieves a 30.0% exploit success rate (12/40) …


Stacked From One: Multi-Scale Self-Injection For Context Window Extension, Wei Han, Pan Zhou, Shuicheng Yan Apr 2026

Stacked From One: Multi-Scale Self-Injection For Context Window Extension, Wei Han, Pan Zhou, Shuicheng Yan

Research Collection School Of Computing and Information Systems

The limited context window of contemporary large language models (LLMs) remains a primary bottleneck for their broader application across diverse domains. Although continual pre-training on long-context data offers a straightforward solution, it incurs prohibitive data acquisition and computational costs. To address this challenge, we propose SHAREDLLM, a novel framework based on multi-grained context compression and query-aware information acquisition. SHAREDLLM comprises two stacked short-context LLMs: a lower model serving as a compressor and an upper model acting as a decoder. The lower model compresses long inputs into compact, multi-grained representations, which are then forwarded to the upper model for context-aware processing. …


Thinktank-Me: A Multi-Expert Framework For Middle East Event Forecasting, Haoxuan Li, He Chang, Yunshan Ma, Yi Bin, Yang Yang, See-Kiong Ng, Tat-Seng Chua Apr 2026

Thinktank-Me: A Multi-Expert Framework For Middle East Event Forecasting, Haoxuan Li, He Chang, Yunshan Ma, Yi Bin, Yang Yang, See-Kiong Ng, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Event forecasting is inherently influenced by multifaceted considerations, including international relations, regional historical dynamics, and cultural contexts. However, existing LLM-based approaches employ single-model architectures that generate predictions along a singular explicit trajectory, constraining their ability to capture diverse geopolitical nuances across complex regional contexts. To address this limitation, we introduce ThinkTank-ME, a novel Think Tank framework for Middle East event forecasting that emulates collaborative expert analysis in real-world strategic decision-making. To facilitate expert specialization and rigorous evaluation, we construct POLECAT-FOR-ME, a Middle East–focused event forecasting benchmark. Experimental results demonstrate the superiority of multi-expert collaboration in handling complex temporal geopolitical forecasting …


Causality-Aware Safety Testing For Autonomous Driving Systems, Wenbing Tang, Mingfei Cheng, Renzhi Wang, Yuan Zhou, Chengwei Liu, Yang Liu, Zuohua Ding Apr 2026

Causality-Aware Safety Testing For Autonomous Driving Systems, Wenbing Tang, Mingfei Cheng, Renzhi Wang, Yuan Zhou, Chengwei Liu, Yang Liu, Zuohua Ding

Research Collection School Of Computing and Information Systems

Simulation-based testing is essential for evaluating the safety of Autonomous Driving Systems (ADSs). Comprehensive evaluation requires testing across diverse scenarios that can trigger various types of violations under different conditions. While existing methods typically focus on individual diversity metrics, such as input scenarios, ADS-generated motion commands, and system violations, they often fail to capture the complex interrelationships among these elements. For instance, identical motion commands can produce different collision risks in varying scenes, and the same collision may result from different commands under different scenarios. This oversight leads to gaps in testing coverage, potentially missing critical issues in the ADS …


From Spatial To Actions: Grounding Vision-Language-Action Model In Spatial Foundation Priors, Zhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu, Lei Zhou, Chenchen Liu, Dong Wang, Francis E. H. Tay, Sijin Chen, Ziwei Liu, Yuxiao Liu, Xinghang Li, Pan Zhou Apr 2026

From Spatial To Actions: Grounding Vision-Language-Action Model In Spatial Foundation Priors, Zhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu, Lei Zhou, Chenchen Liu, Dong Wang, Francis E. H. Tay, Sijin Chen, Ziwei Liu, Yuxiao Liu, Xinghang Li, Pan Zhou

Research Collection School Of Computing and Information Systems

Existing vision-language-action (VLA) models act in 3D real-world but are typically built on 2D encoders, leaving a spatial reasoning gap that limits generalization and adaptability. Recent 3D integration techniques for VLAs either require specialized sensors and transfer poorly across modalities, or inject weak cues that lack geometry and degrade vision-language alignment. In this work, we introduce FALCON (From Spatial to Action), a novel paradigm that injects rich 3D spatial tokens into the action head. FALCON leverages spatial foundation models to deliver strong geometric priors from RGB alone, and includes an Embodied Spatial Model that can optionally fuse depth, or pose …