Open Access. Powered by Scholars. Published by Universities.®

Digital Commons Network™

Open Access. Powered by Scholars. Published by Universities.®

Singapore Management University

Discipline
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 121 - 150 of 1690

Full-Text Articles in Entire DC Network

Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He Oct 2025

Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He

Research Collection School Of Computing and Information Systems

In an NBA game scenario, consider the challenge of locating and analyzing the 3D poses of players performing a user-specified action, such as attempting a shot. Traditional 3D human pose estimation (3DHPE) methods often fall short in such complex, multi-person scenes due to their lack of semantic integration and reliance on isolated pose data. To address these limitations, we introduce Language-Driven 3D Human Pose Estimation (L3DHPE), a novel approach that extends 3DHPE to general multi-person contexts by incorporating detailed language descriptions. We present Panoptic-L3D, the first dataset designed for L3DHPE, featuring 3,838 linguistic annotations for 1,476 individuals across 588 videos, …


Unsupervised Visual Chain-Of-Thought Reasoning Via Preference Optimization, Kesen Zhao, Beier Zhu, Qianru Sun, Hanwang Zhang Oct 2025

Unsupervised Visual Chain-Of-Thought Reasoning Via Preference Optimization, Kesen Zhao, Beier Zhu, Qianru Sun, Hanwang Zhang

Research Collection School Of Computing and Information Systems

Chain-of-thought (CoT) reasoning greatly improves the interpretability and problem-solving abilities of multimodal large language models (MLLMs). However, existing ap proaches focus on text CoT, limiting their ability to lever age visual cues. Visual CoT remains underexplored, and the only work [35] is based on supervised fine-tuning that relies on extensive labeled bounding-box data and is hard to generalize to unseen cases. In this paper, we introduce Unsupervised Visual CoT (UV-CoT), a novel framework for image-level CoT reasoning via preference optimization. UV-CoTperforms preference comparisons between model generated bounding boxes (one is preferred and the other is dis-preferred), eliminating the need for …


Everyday Norms Have Become More Permissive Over Time And Vary Across Cultures, K. Eriksson, P. Strimling, I. Vartanova, Andree Hartanto Oct 2025

Everyday Norms Have Become More Permissive Over Time And Vary Across Cultures, K. Eriksson, P. Strimling, I. Vartanova, Andree Hartanto

Research Collection School of Social Sciences

Every social situation that people encounter in their daily lives comes with a set of unwritten rules about what behavior is considered appropriate or inappropriate. These everyday norms can vary across societies: some societies may have more permissive norms in general or for certain behaviors, or for certain behaviors in specific situations. In a preregistered survey of 25,422 participants across 90 societies, we map societal differences in 150 everyday norms and show that they can be explained by how societies prioritize individualizing moral foundations such as care and liberty versus binding moral foundations such as purity. Specifically, societies with more …


Teaching Diffusion Models To Ground Alpha Matte, Tianyi Xiang, Weiying Zheng, Yutao Jiang, Tingrui Shen, Hewei Yu, Yangyang Xu, Shengfeng He Oct 2025

Teaching Diffusion Models To Ground Alpha Matte, Tianyi Xiang, Weiying Zheng, Yutao Jiang, Tingrui Shen, Hewei Yu, Yangyang Xu, Shengfeng He

Research Collection School Of Computing and Information Systems

The power of visual language models is showcased in visual understanding tasks, where language-guided models achieve impressive flexibility and precision. In this paper, we ex tend this capability to the challenging domain of image matting by framing it as a soft grounding problem, enabling a single diffusion model to handle diverse objects, textures, and transparencies, all directed by descriptive text prompts. Our method teaches the diffusion model to ground alpha mattes by guiding it through a process of instance-level localization and transparency estimation. First, we introduce an intermediate objective that trains the model to accurately localize semantic components of the …


Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan Zhang, Yixin Cao, Lizi Liao Oct 2025

Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan Zhang, Yixin Cao, Lizi Liao

Research Collection School Of Computing and Information Systems

Translating chart images into executable plotting scripts-referred to as the chart-to-code generation task-requires Multimodal Large Language Models (MLLMs) to perform fine-grained visual parsing, precise code synthesis, and robust cross-modal reasoning. However, this task is inherently under-constrained: multiple valid code implementations can produce the same visual chart, and evaluation must consider both code correctness and visual fidelity across diverse dimensions. This makes it difficult to learn accurate and generalizable mappings through standard supervised fine-tuning. To address these challenges, we propose a dual preference-guided refinement framework that combines a feedback-driven, dual-modality reward mechanism with iterative preference learning. Our approach introduces a structured …


Merit Transference And The Paradox Of Merit Inflation, Matthew Hammerton Sep 2025

Merit Transference And The Paradox Of Merit Inflation, Matthew Hammerton

Research Collection School of Social Sciences

Many religious traditions and ethical systems hold that individuals accrue merit through their good intentions, acts, and character, and demerit through their bad intentions, acts, and character. This merit and demerit, accumulated by individuals throughout their lives, gives each person a kind of ethical “score” that can determine what they deserve, and influence whether good or bad things happen to them (e.g., divine punishments and rewards, a favourable or unfavourable rebirth, etc.). In some traditions (most notably Buddhism, but also to a limited extent in Hinduism, Islam, and Christianity), “merit transference” is a feature of these merit-based ethical systems. This …


Large Lithium-Ion Battery Model For Secure Shared E-Bike Battery In Smart Cities, Donghui Ding, Zhao Li, Linhao Luo, Ming Jin, Bin Zhu, Yichen Zhong, Junhao Hu, Peng Cai, Huiqi Hu Sep 2025

Large Lithium-Ion Battery Model For Secure Shared E-Bike Battery In Smart Cities, Donghui Ding, Zhao Li, Linhao Luo, Ming Jin, Bin Zhu, Yichen Zhong, Junhao Hu, Peng Cai, Huiqi Hu

Research Collection School Of Computing and Information Systems

Electric bikes powered by lithium-ion batteries are increasingly used in smart cities to promote sustainable mobility and efficient delivery services. However, limited battery range and slow plug-in charging remain key challenges. Shared electric bike battery systems, facilitated by battery swapping stations, offer a promising solution by enabling quick and efficient battery replacements. However, their success hinges on accurate anomaly detection, battery health estimation and remain range prediction. These tasks remain challenging due to data scarcity, battery diversity and environmental variability. Here we show that a large-scale lithium-ion battery model trained on over ten million battery time series data enables robust …


Vibemus: Proactive Agentic System For Music Personalization, Zhiliang Guo, Teng Tu, Yunshan Ma, Xun Yang Sep 2025

Vibemus: Proactive Agentic System For Music Personalization, Zhiliang Guo, Teng Tu, Yunshan Ma, Xun Yang

Research Collection School Of Computing and Information Systems

Large language models (LLMs) enable diverse forms of AI-assisted creation, yet they often struggle to bridge the preference-articulation gap: users may provide incomplete or vague intentions or lack the vocabulary to specify what they want, yielding outputs misaligned with true preferences. To address this gap and facilitate music creation in a vibe-centric environment, we introduce VibeMus, a proactive agentic system built on open-source components. The system engages in multi-turn dialogue to progressively determine the music’s emotion, genre, lyrics, and other aspects before generation. Simulated evaluations show that proactive clarification improves alignment with users’ intended nuances. Our approach is training-free, leveraging …


Learning Camp, Language Of Instruction, And Education Outcomes: Evidence From A Field Experiment In Malawi, Hyuncheol Bryant. Kim, Kim Sep 2025

Learning Camp, Language Of Instruction, And Education Outcomes: Evidence From A Field Experiment In Malawi, Hyuncheol Bryant. Kim, Kim

Research Collection School Of Economics

We conducted a randomized experiment to study the impacts of a summer learning camp and the language of instruction on education outcomes. The program, covering social studies and mathematics, provided additional high-quality learning time for 4th and 5th graders in Malawi. The program significantly increased test scores in social studies and mathematics (0.24–0.36 standard deviations). We find suggestive evidence that the impact on test scores in social studies was greater for students instructed in Chichewa, the local language, than for those instructed in English. However, we find no such evidence of differential impacts on test scores in mathematics by the …


Lighttransfer: Your Long-Context Llm Is Secretly A Hybrid Model With Effortless Adaptation, Xuan Zhang, Fengzhuo Zhang, Cunxiao Du, Chao Du, Tianyu Pang, Wei Gao, Min Lin Sep 2025

Lighttransfer: Your Long-Context Llm Is Secretly A Hybrid Model With Effortless Adaptation, Xuan Zhang, Fengzhuo Zhang, Cunxiao Du, Chao Du, Tianyu Pang, Wei Gao, Min Lin

Research Collection School Of Computing and Information Systems

Scaling language models to handle longer contexts introduces substantial memory challenges due to the growing cost of key-value (KV) caches. Motivated by the efficiency gains of hybrid models and the broad availability of pretrained large transformer backbones, we explore transitioning transformer models into hybrid architectures for a more efficient generation. In this work, we propose LightTransfer, a lightweight method that transforms models such as LLaMA into hybrid variants. Our approach identifies lazy layers -- those focusing on recent or initial tokens -- and replaces their full attention with streaming attention. This transformation can be performed without any training for long-context …


Guiding Multiple Remote Users In Physical Tasks With Language-Driven Robotic Telepresence, Ruyi Li, Jingfei Guo, Xinyi Zhang, Xuji Zhang, Zeqing Li, Jiannan Li, Jiangtao Gong Sep 2025

Guiding Multiple Remote Users In Physical Tasks With Language-Driven Robotic Telepresence, Ruyi Li, Jingfei Guo, Xinyi Zhang, Xuji Zhang, Zeqing Li, Jiannan Li, Jiangtao Gong

Research Collection School Of Computing and Information Systems

Remote assistance through robotic telepresence could involve both control and memory challenges, particularly in one expert to multiple workers situation. In this work, we proposed a novelty language-driven interface to facilitate remote collaboration through telepresence robots. Through operations and maintenance expert interviews and a scenario simulation study, we identified key pain points in executing one-expert-multiple-workers remote guidance using the telepresence robot and proposed two design goals, which together consist of five sub-design goals with corresponding features. These features were integrated into a standard telepresence robot, resulting in the development of a Collaborative LLM-based Embodied Assistant Robot, named CLEAR Robot. A …


In Crypto We Trust: A Cross-National Comparison Of Factors Affecting Trust In Cryptocurrencies, Arif Perdana, W. Eric Lee, Chu Yeong Lim, Gary Pan, Poh Sun Seow Sep 2025

In Crypto We Trust: A Cross-National Comparison Of Factors Affecting Trust In Cryptocurrencies, Arif Perdana, W. Eric Lee, Chu Yeong Lim, Gary Pan, Poh Sun Seow

Research Collection School Of Accountancy

Background: Trust in cryptocurrencies is a complex and evolving construct influenced by technological, economic, and social factors. Unlike traditional information systems that rely on institutional trust, cryptocurrency users often depend on underlying technology and community insights. Despite growing interest, existing research has primarily focused on narrow aspects of trust, with limited cross-national perspectives and comprehensive models.

Method: This study investigates how geographic differences impact user trust in cryptocurrencies by surveying 644 respondents from China, Germany, and the USA. Drawing on established trust formation theories, it utilizes a comprehensive model encompassing technological, economic, and social factors.

Results: Our analysis reveals multiple …


Towards Multimodal Emotional Support Conversation Systems, Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong Sep 2025

Towards Multimodal Emotional Support Conversation Systems, Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong

Research Collection School Of Computing and Information Systems

The integration of conversational artificial intelligence (AI) into mental health care promises a new horizon for therapist-client interactions, aiming to closely emulate the depth and nuance of human conversations. Despite the potential, the current landscape of conversational AI is markedly limited by its reliance on single-modal data, constraining the systems’ ability to empathize and provide effective emotional support. This limitation stems from a paucity of resources that encapsulate the multimodal nature of human communication essential for therapeutic counseling. To address this gap, we introduce the Multimodal Emotional Support Conversation (MESC) dataset, a first-of-its-kind resource enriched with comprehensive annotations across text, …


Mpo: Multilingual Safety Alignment Via Reward Gap Optimization, Weixiang Zhao, Yulin Hu, Yang Deng, Tongtong Wu, Wenxuan Zhang, Jiahe Guo, An Zhang, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu Aug 2025

Mpo: Multilingual Safety Alignment Via Reward Gap Optimization, Weixiang Zhao, Yulin Hu, Yang Deng, Tongtong Wu, Wenxuan Zhang, Jiahe Guo, An Zhang, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu

Research Collection School Of Computing and Information Systems

Large language models (LLMs) have become increasingly central to AI applications worldwide, necessitating robust multilingual safety alignment to ensure secure deployment across diverse linguistic contexts. Existing preference learning methods for safety alignment, such as RLHF and DPO, are primarily monolingual and struggle with noisy multilingual data. To address these limitations, we introduce Multilingual reward gaP Optimization (MPO), a novel approach that leverages the well-aligned safety capabilities of the dominant language (e.g., English) to improve safety alignment across multiple languages. MPO directly minimizes the reward gap difference between the dominant language and target languages, effectively transferring safety capabilities while preserving the …


Think Both Ways: Teacher-Student Bidirectional Reasoning Enhances Mcq Generation And Distractor Quality, Yimiao Qiu, Yang Deng, Quanming Yao, Zhimeng Zhang, Zhiang Dong, Chang Yao, Jingyuan Chen Aug 2025

Think Both Ways: Teacher-Student Bidirectional Reasoning Enhances Mcq Generation And Distractor Quality, Yimiao Qiu, Yang Deng, Quanming Yao, Zhimeng Zhang, Zhiang Dong, Chang Yao, Jingyuan Chen

Research Collection School Of Computing and Information Systems

Generating high-quality Multiple Choice Questions (MCQs) remains challenging for educational tools due to the need for contextual relevance and plausible distractors. Existing methods still struggle with these dual requirements, leading to questions that lack depth and distractors that are either too obvious or irrelevant. In this paper, we propose BiFlow, a novel framework that integrates bidirectional reasoning perspectives: teacher reasoning generates contextually relevant questions and plausible distractors, while student reasoning evaluates question clarity and the misleading nature of the distractors. To further enhance reasoning, we introduce PathFinder, a mechanism that employs breadth-first search and Chainof-Thought (CoT) strategies to explore diverse …


Causalabstain: Enhancing Multilingual Llms With Causal Reasoning For Trustworthy Abstention, Yuxi Sun, Aoqi Zuo, Wei Gao, Jing Ma Aug 2025

Causalabstain: Enhancing Multilingual Llms With Causal Reasoning For Trustworthy Abstention, Yuxi Sun, Aoqi Zuo, Wei Gao, Jing Ma

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) often exhibit knowledge disparities across languages. Encouraging LLMs to abstain when faced with knowledge gaps is a promising strategy to reduce hallucinations in multilingual settings. Current abstention strategies for multilingual scenarios primarily rely on generating feedback in various languages using LLMs and performing self-reflection. However, these methods can be adversely impacted by inaccuracies and biases in the generated feedback. To address this, from a causal perspective, we introduce CausalAbstain, a method that helps LLMs determine whether to utilize multiple generated feedback responses and how to identify the most useful ones. Extensive experiments demonstrate that CausalAbstain effectively …


Trust In Healthcare Ai Can’T Just Be Designed – It Must Be Felt By Clinicians And Patients, Adriana Banozic-Tang, Heng Wang Aug 2025

Trust In Healthcare Ai Can’T Just Be Designed – It Must Be Felt By Clinicians And Patients, Adriana Banozic-Tang, Heng Wang

Research Collection Yong Pung How School Of Law

Trust in healthcare AI currently over-relies on system design, not lived medical realities.Continuous feedback loops are necessary to embed trust in healthcare AI that is responsive to clinician and patient needs.Initiatives in South-East Asia show how trust in technology can be extended from policy to practice.


Computational Fact-Checking With Limited Resources, Fengzhu Zeng Aug 2025

Computational Fact-Checking With Limited Resources, Fengzhu Zeng

Dissertations and Theses Collection (Open Access)

The rapid dissemination of information through online platforms has sparked widespread concern about the propagation of misinformation. Manual fact-checking by pro- fessional fact-checkers is time-consuming and lacks scalability to address the vast volume of daily information. Consequently, computational fact-checking, driven by automated techniques in natural language processing (NLP), has garnered interest as
a potential solution. However, computational fact-checking faces critical challenges limited resources, particularly due to the issues of data scarcity and computing resource constraints. One key challenge is data scarcity, which arises from the constant generation of new information and emerging events on social media. This scarcity manifests in …


Socioeconomic Status Shapes Dyadic Interactions: Examining Behavioral And Physiologic Responses, Jacinth J. X. Tan, Tessa V. West, Wendy B. Mendes Aug 2025

Socioeconomic Status Shapes Dyadic Interactions: Examining Behavioral And Physiologic Responses, Jacinth J. X. Tan, Tessa V. West, Wendy B. Mendes

Research Collection School of Social Sciences

With more opportunities for diverse interactions, little is known about how social interactions involving people of different socioeconomic status (SES) may unfold. We investigated social attunement patterns in dyadic interactions involving SES. Unacquainted individuals recruited from the community interacted with similar-or-different-SES partners in the lab (Ndyads = 130). Attunement was assessed throughout the interaction by examining physiological linkage—how much a person’s physiological change is predicted by another’s physiological change, over time. Overall, low-SES participants showed stronger physiological linkage—indicating greater attunement—to partners across SES. Participants also appeared more comfortable when interacting with low-SES partners. There were no SES differences in dominance …


Language Brokering Conditions The Indirect Association Between Mexican‐Origin Adolescents’ Academic Discrimination And Educational Expectations, Su Yeong Kim, Yayu Du, Chantal Alvarado, Wei Xiang Sim, Wen Wen, Tianlu Zhang, Jingyi Shen Aug 2025

Language Brokering Conditions The Indirect Association Between Mexican‐Origin Adolescents’ Academic Discrimination And Educational Expectations, Su Yeong Kim, Yayu Du, Chantal Alvarado, Wei Xiang Sim, Wen Wen, Tianlu Zhang, Jingyi Shen

Research Collection School of Social Sciences

Mexican‐origin adolescents, a significant portion of the US Latino population, often experience a decline in educational expectations from early to late adolescence. Contextual factors such as academic discrimination and language brokering for parents may contribute to this decline. This study investigates the indirect effect of academic discrimination experienced in middle school on educational expectations in young adulthood through high school grades and engagement, and the moderating role of language brokering experiences in these relations. Data were collected from 604 Mexican‐origin adolescents across four waves from 2012 to 2023. Academic discrimination experiences in middle school were negatively associated with school grades …


Dreamanime: Learning Style-Identity Textual Disentanglement For Anime And Beyond, Chenshu Xu, Yangyang Xu, Huaidong Zhang, Xuemiao Xu, Shengfeng He Aug 2025

Dreamanime: Learning Style-Identity Textual Disentanglement For Anime And Beyond, Chenshu Xu, Yangyang Xu, Huaidong Zhang, Xuemiao Xu, Shengfeng He

Research Collection School Of Computing and Information Systems

Text-to-image generation models have significantly broadened the horizons of creative expression through the power of natural language. However, navigating these models to generate unique concepts, alter their appearance, or reimagine them in unfamiliar roles presents an intricate challenge. For instance, how can we exploit language-guided models to transpose an anime character into a different art style, or envision a beloved character in a radically different setting or role? This paper unveils a novel approach named DreamAnime, designed to provide this level of creative freedom. Using a minimal set of 2-3 images of a user-specified concept such as an anime character …


Focus: Evaluating Pre-Trained Vision-Language Models On Underspecification Reasoning, Kankan Zhou, Yibin Lai, Kyriakos Mouratidis, Jing Jiang Aug 2025

Focus: Evaluating Pre-Trained Vision-Language Models On Underspecification Reasoning, Kankan Zhou, Yibin Lai, Kyriakos Mouratidis, Jing Jiang

Research Collection School Of Computing and Information Systems

Humans possess a remarkable ability to interpret underspecified ambiguous statements by inferring their meanings from contexts such as visual inputs. This ability, however, may not be as developed in recent pre-trained visionlanguage models (VLMs). In this paper, we introduce a novel probing dataset called FOCUS to evaluate whether state-of-the-art VLMs have this ability. FOCUS consists of underspecified sentences paired with image contexts and carefully designed probing questions. Our experiments reveal that VLMs still fall short in handling underspecification even when visual inputs that can help resolve the ambiguities are available. To further support research in underspecification, FOCUS will be released …


Cami: A Counselor Agent Supporting Motivational Interviewing Through State Inference And Topic Exploration, Yizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang, Phey Ling Kit, Nicholas Gabriel Lim, Cameron Shi Ern Tan, Ee-Peng Lim Aug 2025

Cami: A Counselor Agent Supporting Motivational Interviewing Through State Inference And Topic Exploration, Yizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang, Phey Ling Kit, Nicholas Gabriel Lim, Cameron Shi Ern Tan, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Conversational counselor agents have become essential tools for addressing the rising demand for scalable and accessible mental health support. This paper introduces CAMI, a novel automated counselor agent grounded in Motivational Interviewing (MI) – a client-centered counseling approach designed to address ambivalence and facilitate behavior change. CAMI employs a novel STAR framework, consisting of client’s state inference, motivation topic exploration, and response generation modules, leveraging large language models (LLMs). These components work together to evoke change talk, aligning with MI principles and improving counseling outcomes for diverse clients. We evaluate CAMI’s performance through both automated and expert evaluations, utilizing simulated …


How To Enable Effective Cooperation Between Humans And Nlp Models: A Survey Of Principles, Formalizations, And Beyond, Chen Huang, Yang Deng, Wenqiang Lei, Jiancheng Lv, Tat-Seng Chua, Jimmy Huang Aug 2025

How To Enable Effective Cooperation Between Humans And Nlp Models: A Survey Of Principles, Formalizations, And Beyond, Chen Huang, Yang Deng, Wenqiang Lei, Jiancheng Lv, Tat-Seng Chua, Jimmy Huang

Research Collection School Of Computing and Information Systems

With the advancement of large language models (LLMs), intelligent models have evolved from mere tools to autonomous agents with their own goals and strategies for cooperating with humans. This evolution has birthed a novel paradigm in NLP, i.e., human-model cooperation, that has yielded remarkable progress in numerous NLP tasks in recent years. In this paper, we take the first step to present a thorough review of human-model cooperation, exploring its principles, formalizations, and open challenges. In particular, we introduce a new taxonomy that provides a unified perspective to summarize existing approaches. Also, we discuss potential frontier areas and their corresponding …


Evowiki: Evaluating Llms On Evolving Knowledge, Wei Tang, Yixin Cao, Yang Deng, Jiahao Ying, Bo Wang, Yizhe Yang, Yuyue Zhao, Qi Zhang, Xuanjing Huang, Yu-Gang Jiang, Yong Liao Aug 2025

Evowiki: Evaluating Llms On Evolving Knowledge, Wei Tang, Yixin Cao, Yang Deng, Jiahao Ying, Bo Wang, Yizhe Yang, Yuyue Zhao, Qi Zhang, Xuanjing Huang, Yu-Gang Jiang, Yong Liao

Research Collection School Of Computing and Information Systems

Knowledge utilization is a critical aspect of LLMs, and understanding how they adapt to evolving knowledge is essential for their effective deployment. However, existing benchmarks are predominantly static, failing to capture the evolving nature of LLMs and knowledge, leading to inaccuracies and vulnerabilities such as contamination. In this paper, we introduce EvoWiki, an evolving dataset designed to reflect knowledge evolution by categorizing information into stable, evolved, and uncharted states. EvoWiki is fully auto-updated, enabling precise evaluation of continuously changing knowledge and newly released LLMs. Through experiments with Retrieval-Augmented Generation (RAG) and Continual Learning (CL), we evaluate how effectively LLMs adapt …


Fact-Audit: An Adaptive Multi-Agent Framework For Dynamic Fact-Checking Evaluation Of Large Language Models, Hongzhan Lin, Yang Deng, Yuxuan Gu, Wenxuan Zhang, Jing Ma, See-Kiong Ng, Tat-Seng Chua Aug 2025

Fact-Audit: An Adaptive Multi-Agent Framework For Dynamic Fact-Checking Evaluation Of Large Language Models, Hongzhan Lin, Yang Deng, Yuxuan Gu, Wenxuan Zhang, Jing Ma, See-Kiong Ng, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have significantly advanced the fact-checking studies. However, existing automated fact-checking evaluation methods rely on static datasets and classification metrics, which fail to automatically evaluate the justification production and uncover the nuanced limitations of LLMs in fact-checking. In this work, we introduce FACT-AUDIT, an agent-driven framework that adaptively and dynamically assesses LLMs’ fact-checking capabilities. Leveraging importance sampling principles and multi-agent collaboration, FACT-AUDIT generates adaptive and scalable datasets, performs iterative model-centric evaluations, and updates assessments based on model-specific responses. By incorporating justification production alongside verdict prediction, this framework provides a comprehensive and evolving audit of LLMs’ factual reasoning …


Xfinbench: Benchmarking Llms In Complex Financial Problem Solving And Reasoning, Zhihan Zhang, Yixin Cao, Lizi Liao Aug 2025

Xfinbench: Benchmarking Llms In Complex Financial Problem Solving And Reasoning, Zhihan Zhang, Yixin Cao, Lizi Liao

Research Collection School Of Computing and Information Systems

Solving financial problems demands complex reasoning, multimodal data processing, and a broad technical understanding, presenting unique challenges for current large language models (LLMs). We introduce **XFinBench**, a novel benchmark with 4,235 examples designed to evaluate LLM’s ability in solving comple**X**, knowledge-intensive **Fin**ancial problems across diverse graduate-level finance topics with multi-modal context. We identify five core capabilities of LLMs using XFinBench, i.e., _terminology understanding_, _temporal reasoning_, _future forecasting_, _scenario planning_, and _numerical modelling_. Upon XFinBench, we conduct extensive experiments on 18 leading models. The result shows that o1 is the best-performing text-only model with an overall accuracy of 67.3%, but still …


Taclr: A Scalable And Efficient Retrieval-Based Method For Industrial Product Attribute Value Identification, Yindu Su, Huike Zou, Lin Sun, Ting Zhang, Haiyang Yang, Chen Li Yu, David Lo, Qingheng Zhang, Shuguang Han, Jufeng Chen Aug 2025

Taclr: A Scalable And Efficient Retrieval-Based Method For Industrial Product Attribute Value Identification, Yindu Su, Huike Zou, Lin Sun, Ting Zhang, Haiyang Yang, Chen Li Yu, David Lo, Qingheng Zhang, Shuguang Han, Jufeng Chen

Research Collection School Of Computing and Information Systems

Product Attribute Value Identification (PAVI) involves identifying attribute values from product profiles, a key task for improving product search, recommendation, and business analytics on e-commerce platforms. However, existing PAVI methods face critical challenges, such as inferring implicit values, handling outof-distribution (OOD) values, and producing normalized outputs. To address these limitations, we introduce Taxonomy-Aware Contrastive Learning Retrieval (TACLR), the first retrieval-based method for PAVI. TACLR formulates PAVI as an information retrieval task by encoding product profiles and candidate values into embeddings and retrieving values based on their similarity. It leverages contrastive training with taxonomy-aware hard negative sampling and employs adaptive inference …


I See Sick People: Beliefs About Sensory Detection Of Infectious Disease Are Largely Consistent Across Cultures, Joshua M. Ackerman, Theodore Samore, Daniel M. T. Fessler, Norman P. Li, Et Al Aug 2025

I See Sick People: Beliefs About Sensory Detection Of Infectious Disease Are Largely Consistent Across Cultures, Joshua M. Ackerman, Theodore Samore, Daniel M. T. Fessler, Norman P. Li, Et Al

Research Collection School of Social Sciences

Identifying cues to contagious disease is critical for effectively tracking and defending against interpersonal infection threats. People hold lay beliefs about the types of sensory information most relevant for identifying whether others are sick with transmissible illnesses. Are these beliefs universal, or do they vary along cultural and ecological dimensions? Participants in 58 countries (N = 19,217) judged how effective, and how likely they were to use, cues involving each of the five major sensory modalities in an imagined social interaction during a flu outbreak. Belief patterns were strongly consistent across countries (sight > audition > touch > smell > taste), suggesting a largely …


Faithfulrag: Fact-Level Conflict Modeling For Context-Faithful Retrieval-Augmented Generation, Qinggang Zhang, Zhishang Xiang, Yilin Xiao, Le Wang, Junhui Li, Xinrun Wang, Jinsong Su Aug 2025

Faithfulrag: Fact-Level Conflict Modeling For Context-Faithful Retrieval-Augmented Generation, Qinggang Zhang, Zhishang Xiang, Yilin Xiao, Le Wang, Junhui Li, Xinrun Wang, Jinsong Su

Research Collection School Of Computing and Information Systems

Large language models (LLMs) augmented with retrieval systems have demonstrated significant potential in handling knowledge-intensive tasks. However, these models often struggle with unfaithfulness issues, generating outputs that either ignore the retrieved context or inconsistently blend it with the LLM’s parametric knowledge. This issue is particularly severe in cases of knowledge conflict, where the retrieved context conflicts with the model’s parametric knowledge. While existing faithful RAG approaches enforce strict context adherence through well-designed prompts or modified decoding strategies, our analysis reveals a critical limitation: they achieve faithfulness by forcibly suppressing the model’s parametric knowledge, which undermines the model’s internal knowledge structure …