Open Access. Powered by Scholars. Published by Universities.®

Digital Commons Network™

Open Access. Powered by Scholars. Published by Universities.®

Singapore Management University

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 31 - 60 of 8668

Full-Text Articles in Entire DC Network

A Framework For Top-K Queries With Constrained Preferences, Kyriakos Mouratidis, Nikolaos Chaloulakos, Bo Tang Jul 2026

A Framework For Top-K Queries With Constrained Preferences, Kyriakos Mouratidis, Nikolaos Chaloulakos, Bo Tang

Research Collection School Of Computing and Information Systems

Traditional rank-aware processing assumes a dataset that contains available options to cover a specific need (e.g., restaurants, hotels, etc) and users who browse that dataset via top-k queries with linear scoring functions, i.e., by ranking the options according to the weighted sum of their attributes, for a set of given weights. In practice, however, user preferences (weights) may only be estimated with bounded accuracy, or may be inherently imprecise due to the inability of a human user to specify exact weight values with absolute accuracy. Motivated by this, we define the constrained-preference top-k (CT) query. Given an approximate description of …


Income Tax Over-Withholding And Household Investment Decisions, Xi Novia Chen, Jungbae Kim, Ben Lourie, Chenqi Zhu Jul 2026

Income Tax Over-Withholding And Household Investment Decisions, Xi Novia Chen, Jungbae Kim, Ben Lourie, Chenqi Zhu

Research Collection School Of Accountancy

Over three-quarters of U.S. taxpayers over-withhold taxes, leading to tax refunds. This study explores how over-withholding impacts stock investments by comparing individuals' investment patterns from wages and tax refunds. We find that while some portion of tax refunds are promptly invested, the investment rate is lower than that of wages. This suggests that over-withholding, which alters the label and timing of wages into refunds, influences investment behavior. Our cross-sectional analysis indicates that the differential propensity to invest wages versus tax refunds are more pronounced for individuals with automatic investment setups and lower financial sophistication. These findings underscore the importance of …


Social Mobility Beliefs Predict Competence Stereotype Gaps Between The Poor And Rich, Gregory Tee Hng Tan Jul 2026

Social Mobility Beliefs Predict Competence Stereotype Gaps Between The Poor And Rich, Gregory Tee Hng Tan

Dissertations and Theses Collection (Open Access)

The poor tend to be viewed as less competent than the rich. This dissertation examined whether social mobility beliefs shape perceived rich-poor competence gaps. When society is seen as highly mobile, people may attribute economic outcomes to internal abilities than external constraints, thereby widening the rich-poor competence gap. Conversely, perceiving low mobility may shift explanations away from internal abilities toward external constraints and narrow the gap. Four studies tested this theory. Study 1 examined the relationships between self-reported mobility beliefs, attributions, and competence of rich and poor targets among Singapore and UK participants. While higher mobility beliefs predicted stronger internal …


Accountable Agents In Software Engineering: An Analysis Of Terms Of Service And A Research Roadmap, Christoph Treude Jul 2026

Accountable Agents In Software Engineering: An Analysis Of Terms Of Service And A Research Roadmap, Christoph Treude

Research Collection School Of Computing and Information Systems

AI coding assistants and autonomous agents are becoming integral to software development workflows, reshaping how code is produced, reviewed, and maintained. While recent research has focused mainly on the capabilities and impacts of productivity of these systems, much less attention has been paid to accountability: who is responsible when agents generate, modify, or recommend code? In practice, accountability is defined through the Terms of Service (ToS) and related policy documents that govern the use of AI-powered development tools.In this vision paper, we present a comparative analysis of the Terms of Service for widely used AI coding assistants and agent-enabled development …


Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang Jul 2026

Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely underexplored. In this paper, we identify spatiotemporal sycophancy, a failure mode in which Vid-LLMs retract initially correct, visually grounded judgments and conform to misleading user feedback under negation-based gaslighting. Rather than merely changing their answers, the models often fabricate unsupported temporal or spatial explanations to justify incorrect revisions. To systematically investigate this phenomenon, we propose a negation-based gaslighting evaluation framework and introduce GasVideo-1000, a curated benchmark designed to probe spatiotemporal sycophancy with clear visual grounding and temporal reasoning requirements. …


Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen Jul 2026

Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen

Research Collection School Of Computing and Information Systems

Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object’s state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action–object interactions into …


Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun Jul 2026

Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun

Research Collection School Of Computing and Information Systems

Large language models (LLMs) are increasingly trained on massive, heterogeneous text corpora, raising serious concerns about the unauthorised use of proprietary or personal data during model training. In this work, we address the problem of data protection against unwanted model learning in a realistic blackbox setting. We propose Disclaimer Injection, a novel data-level defence that renders text unlearnable to LLMs. Rather than relying on model-side controls or explicit data removal, our approach exploits the models’ own alignment mechanisms: injecting carefully designed alignment-triggers to prevent effective learning. Through layer-wise analysis, we find that finetuning on such protected data induces persistent activation …


Operational Agency: A Permeable Legal Fiction For Tracing Culpability In Ai Systems, Anirban Mukherjee, Hannah H. Chang Jul 2026

Operational Agency: A Permeable Legal Fiction For Tracing Culpability In Ai Systems, Anirban Mukherjee, Hannah H. Chang

Research Collection Lee Kong Chian School Of Business

Modern artificial intelligence (AI) systems act with a high degree of independence yet lack legal personhood—a paradox that fractures doctrines grounded in human-centric notions of mens rea and actus reus. This Article introduces Operational Agency (OA)—a permeable legal fiction structured as an ex post evidentiary framework—and Operational Agency Graph (OAG)—a tool for mapping causal interactions among human actors, organizations, and AI systems. OA evaluates an AI’s observable operational characteristics: its goal-directedness (as a proxy for intent), predictive processing (as a proxy for foresight), and safety architecture (as a proxy for standard of care). OAG operationalizes that analysis by embedding these …


Modeling And Forecasting Intraday Spot Volatility, Adam Clements, Daniel P. A. Preve Jul 2026

Modeling And Forecasting Intraday Spot Volatility, Adam Clements, Daniel P. A. Preve

Research Collection School Of Economics

We propose a multiple-equation regression-based method for modeling and forecasting intraday spot volatility. In this approach, intraday intervals are treated as individual time series, deviating from the common practice of treating the data as one continuous sample. Our empirical study, which spans more than two decades and encompasses six US blue-chip stocks, employs the recent OK volatility estimator developed by Li, Wang, and Zhang (2024) to expose the dynamics of latent intraday spot volatility over time. We demonstrate that the proposed method effectively captures the intricate dynamics of intraday spot volatility and find strong evidence that it outperforms a competing …


How Do Sectoral Shocks Shape Future Gdp?, Paul Ho, Danial Lashkari, Pierre-Daniel Sarte Jul 2026

How Do Sectoral Shocks Shape Future Gdp?, Paul Ho, Danial Lashkari, Pierre-Daniel Sarte

Research Collection School Of Economics

A production sector’s size, as measured by its Domar weight, captures the contemporaneous aggregate effect of its productivity shocks. Any future aggregate effects, however, depend on that sector’s participation in the investment network. We derive a dynamic generalization of Hulten’s theorem in an environment with intermediate-input and investment networks. This generalization, implied by production efficiency alone, decomposes each sector’s Domar weight into an impact and a propagation component. The relative size of these components then determines how persistent the aggregate effects of sectoral shocks are, but cannot be known absent information on the economy’s production structure. We show in a …


Integrated Optimization Of Farmland Cultivation And Fertilizer Application: Implications For Farm Management And Crop Production, Onur Boyabatli, Lusheng Shao, Yangfang (Helen) Zhou Jul 2026

Integrated Optimization Of Farmland Cultivation And Fertilizer Application: Implications For Farm Management And Crop Production, Onur Boyabatli, Lusheng Shao, Yangfang (Helen) Zhou

Research Collection Lee Kong Chian School Of Business

Motivated by the fresh produce industry, this paper studies a farmer’s joint cultivation and fertilizer (a representative farm input) application decisions facing uncertainties in yield, crop price, and harvesting cost where the latter two are yield dependent and yield is stochastically increasing in the fertilizer application rate. We develop a two-stage stochastic model of a farmer growing a commodity crop in a single season to maximize the expected profit. We then use the model to evaluate the optimal expected harvest volume (a measure of crop production). Our analytical analysis is complemented with numerical experiments calibrated to data. We characterize how …


The Power And Peril Of Awe In Leadership: Transforming Follower Identity And Behavior, Jack Mcguire, Daniel Mcallister, Jochen Menges, David De Cremer Jul 2026

The Power And Peril Of Awe In Leadership: Transforming Follower Identity And Behavior, Jack Mcguire, Daniel Mcallister, Jochen Menges, David De Cremer

Research Collection Lee Kong Chian School Of Business

Awe is a profound emotion that has captured significant attention within psychological research. While the potential for leaders to inspire awe in followers has received some recognition, systematic research on the nature and effects of awe in leadership—and within organizational contexts more broadly—remains limited. In this article, we offer a conceptual framework that explains the multifaceted and transformative nature of leadership through the power of awe. Specifically, we identify four leader behaviors—charismatic leadership tactics, exceptional performance, problem reframing, and self-sacrificial behavior—that elicit awe among followers. We further propose three variants of awe-inspiring leaders, describing how variation in a leader’s self-construal …


Knowledge-State Generative Agents For Pre-Assessment Question Evaluation, Ping Fan Ke, Yi Meng Lau, Siaw Ling Lo Jul 2026

Knowledge-State Generative Agents For Pre-Assessment Question Evaluation, Ping Fan Ke, Yi Meng Lau, Siaw Ling Lo

Research Collection School Of Computing and Information Systems

This paper introduces a Knowledge‑State Generative Agent framework for evaluating the quality of pre‑assessment questions. The framework employs large language model (LLM)–based agents prompted to adopt a teacher persona to simulate the responses of students with and without mastery of targeted knowledge components. A preliminary empirical study using archival data from 424 students enrolled in an Information Systems Management course indicates that the proposed approach yields interpretable metrics under Classical Test Theory. Results further show that agents instantiated with the relevant mastered knowledge components exhibit systematically higher performance than agents lacking such mastery. In addition, the study suggests that teacher-persona …


Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma Jul 2026

Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma

Research Collection School Of Computing and Information Systems

Personalized generative recommender systems have emerged as a promising solution for fashion recommendation. However, existing methods primarily rely on implicit visual embeddings from historical interactions, which often contain preference-irrelevant information and result in insufficient user behavior modeling. Moreover, these models typically generate only item images, providing limited interpretability. To address these limitations, we propose DualFashion, a Dual-Diffusional Generative Fashion Recommendation Architecture that jointly models image and text modalities for personalized and explainable recommendation. DualFashion adopts a dual-diffusion Transformer with image and text branches, where structured attribute-level captions and visual outfit information are jointly used as conditioning signals to model user …


Activity Transition Graph Generation: How Far Are We?, Jiakun Liu, Peixin Zhang, Han Hu, Yonghui Liu, Wei Minn, Ferdian Thung, Shahar Maoz, Eran Toch, Debin Gao, David Lo Jul 2026

Activity Transition Graph Generation: How Far Are We?, Jiakun Liu, Peixin Zhang, Han Hu, Yonghui Liu, Wei Minn, Ferdian Thung, Shahar Maoz, Eran Toch, Debin Gao, David Lo

Research Collection School Of Computing and Information Systems

Android applications (i.e., apps) are indispensable nowadays and are getting bigger and bigger with an increasing number offunctionalities. To understand how to access functionalities in an app, prior studies proposed tools to model the transitionsbetween functionalities with the activity transition graph (ATG). ATG is an important data structure and has been used forvarious Android app analyses, including app design, understanding, and testing. However, there is no benchmarking work onATG generation. It is still unclear whether the transitions identified by tools are correct and how many transitions are missed.To fill this gap, we manually identified all transitions in 98 applications to …


Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo Jul 2026

Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo

Research Collection School Of Computing and Information Systems

Learning-based dynamic fault localization techniques play a crucial role in the field of software engineering. These techniques dynamically execute test cases to meticulously extract useful knowledge from the execution information in the program, with the aim of identifying fault locations by leveraging machine learning, deep learning, and large language models. Currently, there is already a flourishing body of research that is intensely focused on learning-based dynamic fault localization. Research literature can be categorized into two main aspects for learning-based dynamic fault localization: data-based enhancements (i.e., the datasets) and model-based enhancements (i.e., the suspiciousness algorithms). Thus, we conduct an extensive literature …


Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning, Brahmanage Janaka Chathuranga Thilakarathna, Akshat Kumar Jul 2026

Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning, Brahmanage Janaka Chathuranga Thilakarathna, Akshat Kumar

Research Collection School Of Computing and Information Systems

Sequential decision-making using Markov Decision Process underpins many real-world applications. Both model-based and model-free methods have achieved strong results in these settings. However, real-world tasks must balance reward maximization with safety constraints, often conflicting objectives, that can lead to unstable min–max, adversarial optimization. A promising alternative is safety reachability analysis, which precomputes a forward-invariant safe state–action set, ensuring that an agent starting inside this set remains safe indefinitely. Yet, most reachability-based methods address only hard safety constraints, and little work extends reachability to cumulative cost constraints. To address this, first, we define a safety-conditioned reachability set that decouples reward maximization …


Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma Jul 2026

Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma

Research Collection School Of Computing and Information Systems

Despite the importance of open-ended event forecasting for risk management, current LLM-based methods predominantly target only the most probable outcomes, neglecting the intrinsic uncertainty of real-world events. To bridge this gap, we advance open-ended event forecasting from pinpoint forecasting to scatter forecasting by introducing the proxy task of hypothesis generation. This paradigm aims to generate an inclusive and diverse set of hypotheses that broadly cover the space of plausible future events. To this end, we propose SCATTER, a reinforcement learning framework that jointly optimizes inclusiveness and diversity of the hypothesis. Specifically, we design a novel hybrid reward that consists of …


Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang Jul 2026

Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang

Research Collection School Of Computing and Information Systems

Document Question Answering (DQA) involves generating answers from a document based on a user’s query, representing a key task in document understanding. This task requires interpreting visual layouts, which has prompted recent studies to adopt multimodal Retrieval-Augmented Generation (RAG) that processes page images for answer generation. However, in multimodal RAG, visual DQA struggles to utilize a large number of images effectively, as the retrieval stage often retains only a few candidate pages (e.g., Top-4), causing informative but less visually salient content to be overlooked in favor of common yet low-information pages. To address this issue, we propose a Multi-Armed Bandit–based …


Learning By Writing: Exploring Authentic Legal Learning Through Case Summaries, Ee-Ing Ong, Wei Yang Quek, Duan Ning, Magdeleine Lew Jul 2026

Learning By Writing: Exploring Authentic Legal Learning Through Case Summaries, Ee-Ing Ong, Wei Yang Quek, Duan Ning, Magdeleine Lew

Research Collection Yong Pung How School Of Law

We use authentic learning as a pedagogical framework in a collaboration between our law school and the national Supreme Court of a Southeast Asian country, which facilitates law students’ development of their legal analytical and writing skills, and helps them better bridge the gap between existing legal curricula and the needs of legal practice. Akin to a writing apprenticeship, students write summaries on selected Supreme Court judgments, with their output reviewed by faculty as well as judicial law clerks from the court. The results are published on the court’s website and circulated to other stakeholders. In the post-exercise survey, participating …


Online Planning Of Power Flows For Power Systems Against Bushfires Using Spatial Context, Jianyu Xu, Qiuzhuang Sun, Yang Yang, Huadong Mo, Daoyi Dong Jul 2026

Online Planning Of Power Flows For Power Systems Against Bushfires Using Spatial Context, Jianyu Xu, Qiuzhuang Sun, Yang Yang, Huadong Mo, Daoyi Dong

Research Collection College of Integrative Studies

A power station or transmission line can be affected due to bushfires, increasing operation costs. We study a fundamental but challenging problem of planning the optimal power flow (OPF) for power systems under bushfires. We develop a model to capture the stochastic nature of bushfire spread based on Moore’s neighborhood model and propose an online optimization modeling framework to sequentially plan power flows in the electricity network. Our framework assumes that bushfire spread is non-stationary over time and that the spread and containment probabilities are unknown. To address these challenges, we develop a contextual online learning algorithm that treats the …


Purifai: Detecting And Fixing Search-Induced Distortions In Web-Augmented Llms, Guoqing Wang, Zhao Zhang, Zeyu Sun, Xiaofei Xie, Yizhou Chen, Yanchao Tan, Dan Hao Jul 2026

Purifai: Detecting And Fixing Search-Induced Distortions In Web-Augmented Llms, Guoqing Wang, Zhao Zhang, Zeyu Sun, Xiaofei Xie, Yizhou Chen, Yanchao Tan, Dan Hao

Research Collection School Of Computing and Information Systems

As Large Language Models (LLMs) increasingly serve as interfaces for proprietary data (e.g., enterprise knowledge bases, legal statutes), ensuring their fidelity to trusted internal information is paramount. While integrating real-time web search can enhance model utility, it introduces a critical vulnerability: the ingestion of conflicting, misleading, or hallucinated content from the open web can override the model's adherence to its verified internal knowledge. We define this failure mode as search-induced distortion, a significant risk in high-stakes domains where the internal knowledge base serves as the absolute ground truth.To address this challenge, we present PurifAI, a proactive, model-agnostic, cache-level purification system …


Verbalizing Lightgcn: Direct Learning Of Textual Representations From User-Item Interaction Graph Via Llms, Manh-Khanh Ngo Huu, Hady Wirawan Lauw Jul 2026

Verbalizing Lightgcn: Direct Learning Of Textual Representations From User-Item Interaction Graph Via Llms, Manh-Khanh Ngo Huu, Hady Wirawan Lauw

Research Collection School Of Computing and Information Systems

In this work, we propose VerbaLightGCN, a novel LLM-based recommendation framework that integrates the semantic understanding of LLMs with user-item interaction modeling. Traditional collaborative filtering (CF) models typically embed user and item IDs into a latent space to capture interaction signals. However, pretrained LLMs cannot natively interpret these learned embeddings. To bridge this gap, VerbaLightGCN adopts a CF-as-text paradigm, in which collaborative signals are encoded in textual form and directly learned from the user–item interaction graph, and are then combined with semantic information to construct user and item profiles that function as latent embeddings. Inspired by LightGCN, our method retains …


Generation-Augmented Video Corpus Moment Retrieval, Mingjin Kuai, Qianyin Xiao, Juncheng Li, Jin Peng, Lizi Liao, Wei Ji Jul 2026

Generation-Augmented Video Corpus Moment Retrieval, Mingjin Kuai, Qianyin Xiao, Juncheng Li, Jin Peng, Lizi Liao, Wei Ji

Research Collection School Of Computing and Information Systems

Video Corpus Moment Retrieval (VCMR) requires models to efficiently retrieve and precisely locate specific moments relevant to natural language queries within a massive, untrimmed video corpus. However, existing discriminative approaches typically rely on shallow visual-textual feature matching mechanisms, which often struggle to capture fine-grained semantic differences. To address this limitation, we propose Video-GAR, a novel framework that reframes the conventional retrieval task from superficial matching to generative understanding, positing that the capability for query reconstruction evidences deep semantic comprehension. Specifically, Video-GAR orchestrates three synergistic components: To overcome the computational efficiency bottleneck, we construct a Bi-Mamba backbone that leverages the linear …


Larger Is Not Always Better: Exploring Small Open-Source Language Models In Logging Statement Generation, Renyi Zhong, Yichen Li, Guangba Yu, Wenwei Gu, Jinxi Kuang, Yintong Huo, Michael R. Lyu Jul 2026

Larger Is Not Always Better: Exploring Small Open-Source Language Models In Logging Statement Generation, Renyi Zhong, Yichen Li, Guangba Yu, Wenwei Gu, Jinxi Kuang, Yintong Huo, Michael R. Lyu

Research Collection School Of Computing and Information Systems

Developers use logging statements to create logs that document system behavior and aid in software maintenance. As such, high-quality logging is essential for effective maintenance; however, manual logging often leads to errors and inconsistency. Recent methods emphasize using large language models (LLMs) for automated logging statement generation, but these present privacy and resource issues, hindering their suitability for enterprise use. This paper presents the first large-scale empirical study evaluating small open-source language models (SOLMs) for automated logging statement generation. We evaluate four prominent SOLMs using various prompt strategies and parameter-efficient fine-tuning techniques, such as Low-Rank Adaptation (LoRA) and Retrieval-Augmented Generation …


Anomaly Management In Unmanned Aerial Vehicles: A Systematic Literature Review, Ivan Tan Wei Han, Christopher M. Poskitt, Lingxiao Jiang, Lwin Khin Shar Jul 2026

Anomaly Management In Unmanned Aerial Vehicles: A Systematic Literature Review, Ivan Tan Wei Han, Christopher M. Poskitt, Lingxiao Jiang, Lwin Khin Shar

Research Collection School Of Computing and Information Systems

Unmanned Aerial Vehicles (UAVs) are increasingly deployed in safety-critical applications such as logistics, surveillance, disaster response, and urban air mobility. While their autonomy enables powerful capabilities, it also introduces vulnerabilities due to hardware faults, software defects, communication failures, and adversarial interference. This survey presents a comprehensive review of research studies closely related to UAV anomalies published between 2015 and 2025, covering 111 papers from academic and industrial sources. We introduce a unified five-pillar taxonomy—anomaly generation, prevention, detection, recovery, and analysis—that organizes existing work across the full anomaly management lifecycle. In contrast to prior surveys that focus primarily on detection algorithms, …


Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou Jul 2026

Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou

Research Collection School Of Computing and Information Systems

Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize single greedy trajectories, decoding involves verifying and ranking multiple sampled draft paths. We propose Variational Speculative Decoding (VSD), formulating draft training as variational inference over latent proposals (draft paths). VSD maximizes the marginal probability of target-model acceptance, yielding an ELBO that promotes high-quality latent proposals while minimizing divergence from the target distribution. To enhance quality and reduce variance, we incorporate a path-level utility and optimize via an Expectation-Maximization procedure. The E-step draws MCMC samples from an oracle-filtered posterior, while the M-step maximizes weighted likelihood using …


Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen Jul 2026

Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen

Research Collection School Of Computing and Information Systems

The widespread availability of large-scale code datasets has accelerated the development of code large language models (CodeLLMs), raising concerns about unauthorized dataset usage. Dataset poisoning offers a proactive defense by reducing the utility of such unauthorized training. However, existing poisoning methods often require full-dataset poisoning and introduce transformations that break code compilability. In this paper, we introduce FunPoison, a functionality-preserving poisoning approach that injects short, compilable weak-use fragments into executed code paths. FunPoison leverages reusable statement-level templates with automatic repair and conservative safety checking to ensure side-effect freedom, while a type-aware synthesis module preserves type correctness, suppresses static-analysis warnings, and …


How Systems Use The Research Organization Registry (Ror) For Research Integrity, Maria Gould, Adam Buttrick Jun 2026

How Systems Use The Research Organization Registry (Ror) For Research Integrity, Maria Gould, Adam Buttrick

FORCE 2026

As scholarly publishers strive to ensure the integrity of the scholarly record, verified author affiliations have become a key metric to assist with assessment of manuscript submissions. For instance, a recent article in Nature used author affiliations to identify problematic research practices at particular institutions, and a recent STM report on Trusted Identity in Academic Publishing points out that a verified author affiliation valuable as it links an individual to an organisation, providing a route for accountability where there would otherwise be none.

Knowing that there is a confirmed link between an author and a reputable research organization is a …


Promoting Crisis-Ready Data Policies In Science Publishing: Implementing Unesco’S Open Science Guidance For Resilient Preservation, Response, And Recovery, Francis P. Crawley Jun 2026

Promoting Crisis-Ready Data Policies In Science Publishing: Implementing Unesco’S Open Science Guidance For Resilient Preservation, Response, And Recovery, Francis P. Crawley

FORCE 2026

Crises disrupt not only societies but also the continuity of scientific data and publications essential for evidence-based decision-making. Preserving the integrity, accessibility, and long-term usability of these resources during disruptions has become a global priority. In response, UNESCO, with the support of CODATA, has developed a suite of Open Science Toolkit instruments (Factsheet, Guidance, and Checklist) on ‘Developing Data Policies for Times of Crisis Facilitated by Open Science’. This presentation examines how these instruments can strengthen data preservation, response, and recovery within scholarly communication. By promoting crisis-ready data policies supported by AI-enabled systems in science publishing, the UNESCO approach connects …