Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2641 - 2670 of 63009

Full-Text Articles in Computer Sciences

Advancing Security Safeguards In Large Language Models Through Multi-Agent Systems, Mohammed Rashed Alnuaimi Nov 2025

Advancing Security Safeguards In Large Language Models Through Multi-Agent Systems, Mohammed Rashed Alnuaimi

Theses

This thesis focused on enhancing the safe use of Large Language Model (LLM) through the innovative use of a Multi-Agent System (MAS). As LLMs like ChatGPT became essential to our everyday interactions, the need to maintain the safe use of these systems increased. This research thoroughly assessed the current security measures in place for LLM, pointed out their limitations and developed new and more effective security strategies. The core of the proposed solution was a MAS designed to ensure that all data processed by LLM met guidelines including Privacy, Confidentiality, and Ethical standards before reaching the user. The system involved …


Evaluating Large Language Models For Automated Cv Ranking: A Hybrid Embedding Approach For Enhanced Recruitment, Sarah Mohamed Alhindaassi Nov 2025

Evaluating Large Language Models For Automated Cv Ranking: A Hybrid Embedding Approach For Enhanced Recruitment, Sarah Mohamed Alhindaassi

Theses

Increasing numbers of applications have revealed limitations in legacy keyword-filtering-based Applicant Tracking Systems (ATS), which commonly overlook candidate potential and ignore contextual or transferable skills. Advances in Natural Language Processing (NLP) and Large Language Models (LLMs) offer an exhilarating alternative, supporting context-sensitive and human-crafted reasoning in candidate evaluation. This thesis systematically evaluates four classes of approaches, lexical models, embedding-based methods, Large Language Models (LLMs), and hybrid ensembles, for automation of Curriculum Vitae (CV) to Job Description (JD) matching without exploiting prior annotations or annotations at match time. Using a combination of publicly available datasets and real-world sample data covering three …


The Impact Of Sanctions On Github Developers And Activities, Youmei Fan, Ani Hovhannisyan, Hideaki Hata, Christoph Treude, Raula G. Kula Nov 2025

The Impact Of Sanctions On Github Developers And Activities, Youmei Fan, Ani Hovhannisyan, Hideaki Hata, Christoph Treude, Raula G. Kula

Research Collection School Of Computing and Information Systems

The GitHub platform has fueled the creation of truly global software, enabling contributions from developers across various geographical regions of the world. As software becomes more entwined with global politics and social regulations, it becomes similarly subject to government sanctions. In 2019, GitHub restricted access to certain services for users in specific locations but rolled back these restrictions for some communities (e.g., the Iranian community) in 2021. We conducted a largescale empirical study, collecting approximately 156 thousand user profiles and their 41 million activity points from 2008 to 2022, to understand the response of developers. Our results indicate that many …


Explainable Sentiment Analysis With Deepseek-R1: Performance, Efficiency, And Few-Shot Learning, Donghao Huang, Zhaoxia Wang Nov 2025

Explainable Sentiment Analysis With Deepseek-R1: Performance, Efficiency, And Few-Shot Learning, Donghao Huang, Zhaoxia Wang

Research Collection School Of Computing and Information Systems

Large language models (LLMs) have transformed sentiment analysis, yet balancing accuracy, efficiency, and explainability remains a critical challenge. This study presents the first comprehensive evaluation of DeepSeek-R1—an open-source reasoning model—against OpenAI’s GPT-4o and GPT-4o-mini. We test the full 671B model and its distilled variants, systematically documenting few-shot learning curves. Our experiments show DeepSeek-R1 achieves a 91.39% F1 score on 5-class sentiment and 99.31% accuracy on binary tasks with just 5 shots, an eightfold improvement in few-shot efficiency over GPT-4o. Architecture-specific distillation effects emerge, where a 32B Qwen2.5-based model outperforms the 70B Llama-based variant by 6.69 percentage points. While its reasoning …


Generative Ai And Empirical Software Engineering: A Paradigm Shift, Christoph Treude, Margaret-Anne Storey Nov 2025

Generative Ai And Empirical Software Engineering: A Paradigm Shift, Christoph Treude, Margaret-Anne Storey

Research Collection School Of Computing and Information Systems

The widespread adoption of generative AI in software engineering marks a paradigm shift, offering new opportunities to design and utilize software engineering tools while influencing both developers and the artifacts they create. Traditional empirical methods in software engineering, including quantitative, qualitative, and mixed-method approaches, are well established. However, this paradigm shift introduces novel data types and redefines many concepts in the software engineering process. The roles of developers, users, agents, and researchers increasingly overlap, blurring the distinctions between these social and technical actors within the field. This paper examines how integrating AI into software engineering challenges traditional research paradigms. It …


Adasteer: Your Aligned Llm Is Inherently An Adaptive Jailbreak Defender, Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu Nov 2025

Adasteer: Your Aligned Llm Is Inherently An Adaptive Jailbreak Defender, Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu

Research Collection School Of Computing and Information Systems

Despite extensive efforts in safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks. Activation steering offers a training-free defense method but relies on fixed steering coefficients, resulting in suboptimal protection and increased false rejections of benign inputs. To address this, we propose AdaSteer, an adaptive activation steering method that dynamically adjusts model behavior based on input characteristics. We identify two key properties: Rejection Law (R-Law), which shows that stronger steering is needed for jailbreak inputs opposing the rejection direction, and Harmfulness Law (H-Law), which differentiates adversarial and benign inputs. AdaSteer steers input representations along both the Rejection Direction …


One Planner To Guide Them All! Learning Adaptive Conversational Planners For Goal-Oriented Dialogues, Huy Dao, Lizi Liao Nov 2025

One Planner To Guide Them All! Learning Adaptive Conversational Planners For Goal-Oriented Dialogues, Huy Dao, Lizi Liao

Research Collection School Of Computing and Information Systems

Goal-oriented dialogues, such as recommendation and negotiation, often require balancing multiple, conflicting objectives. Existing methods typically involve training separate models for specific combinations of objectives, leading to computational and scalability issues. In this work, we aim to develop a new dialogue policy method that can adapt to varying objective preferences at inference time without retraining. This raises several challenges in terms of both (1) optimization strategy and (2) knowledge utilization. To address these, we propose a novel learning framework, Preference Adaptive Dialogue Policy Planner (PADPP), for multi-objective goal-oriented dialogues. Specifically, to tackle the former, we introduce a novel policy optimization …


Enhancing Spatial Understanding In Mixed-Reality Presentations, Nam-Dang Vo, Van-Vinh Thai, Nam-Hoi Do, Viet-Tham Huynh, Anthony Tang, Khan-Duy Le Nov 2025

Enhancing Spatial Understanding In Mixed-Reality Presentations, Nam-Dang Vo, Van-Vinh Thai, Nam-Hoi Do, Viet-Tham Huynh, Anthony Tang, Khan-Duy Le

Research Collection School Of Computing and Information Systems

Mixed reality (MR) presentations often involve a presenter wearing a head-mounted display (HMD) and an audience watching via a large display, making it difficult for audiences to perceive spatial relationships between the presenter and virtual objects. We report two experiments testing three design variations: (1) scene camera placement (audience-aligned vs. opposite), (2) overlaying the presenter’s first-person view, and (3) highlighting objects in the presenter’s view. Results show that audience-aligned cameras and object highlighting improve spatial understanding, while combining third- and first-person views can further aid perception. We derive design guidelines for configuring MR presentations to better support audience comprehension.


Image Captioning Through The Lens Of The Gricean Maxims: Generating Meaningful And Relevant Image Descriptions, Annika Lindh Nov 2025

Image Captioning Through The Lens Of The Gricean Maxims: Generating Meaningful And Relevant Image Descriptions, Annika Lindh

Doctoral

Image captioning models enable us to automatically generate natural language image descriptions for previously unseen images. It combines the two fields of computer vision and natural language generation, allowing models to interpret the con tent of an image and communicate that knowledge through natural language text.

Research into image captioning has the potential benefit of reducing the gap in digital information availability between fully sighted individuals and those who are visually impaired. However, automatically generated captions often fail to provide the required level of detail and specificity to achieve this goal. Furthermore, current standard evaluation methods are insufficient at measuring …


Symbolic Execution Engine For Dynamic Analysis Of System Software, Pansilu Madhura Bhashana Pitigala Arachchillage Nov 2025

Symbolic Execution Engine For Dynamic Analysis Of System Software, Pansilu Madhura Bhashana Pitigala Arachchillage

Dissertations and Theses Collection (Open Access)

System software, like any regular software, is prone to errors. It plays a specific role in a computer system by managing the underlying hardware and providing a platform to execute the application software. Defective or vulnerable system software can be exploited by attackers to compromise the entire system. Therefore, the system software must be studied and thoroughly analyzed to evaluate its security. However, due to the inherent complexity and its close interactions with the hardware, analyzing system software is a challenging task. As a result, there is a lack of tools and techniques capable of effectively analyzing system software.

This …


Graph Perturbations For Robust Knowledge Discovery And Retrieval, Hanhua Xiao Nov 2025

Graph Perturbations For Robust Knowledge Discovery And Retrieval, Hanhua Xiao

Dissertations and Theses Collection (Open Access)

Graph perturbation, rooted in classical perturbation theory, studies how small topology edits, i.e., adding or deleting edges, affects graph properties (e.g., density, centrality). This fundamental problem underpins applications like bioinformatics, privacy preservation and system defense. While much prior work targets perturbations that influence global graph statistics or model outputs, comparatively little addresses robustness for knowledge discovery and information retrieval. In these settings, graphs are attributed: nodes carry real-world semantics (e.g., locations, people) and edges encode interactions or relationships. This thesis proposes new formulations and algorithms that generate and leverage graph perturbations to make knowledge discovery and retrieval more robust. Specifically, …


Enhancing Multi-View, Multi-Modal Sensing, Perception And Actuation For Edge Intelligence, Dhanuja Tharith Wanniarachchige Nov 2025

Enhancing Multi-View, Multi-Modal Sensing, Perception And Actuation For Edge Intelligence, Dhanuja Tharith Wanniarachchige

Dissertations and Theses Collection (Open Access)

Artificial Intelligence of Things (AIoT) technologies have ushered in exciting new advances in intelligent sensing, perception, and actuation for many real-world cyberphysical systems (CPS) applications. These technologies have had a formidable impact in domains such as large-scale video surveillance, autonomous transportation and robotics, precision healthcare, and industrial automation. In these applications, sensors and actuators are often collocated with processing nodes, and such nodes are typically interconnected via wireless networks. Vision-based machine intelligence, exemplified by tasks such as object detection, object tracking, and activity analysis, is a very common enabler of such CPS applications. Efficient execution of Deep Neural Network (DNN) …


Scaling Up Cooperative Multi-Agent Reinforcement Learning, Minghong Geng Nov 2025

Scaling Up Cooperative Multi-Agent Reinforcement Learning, Minghong Geng

Dissertations and Theses Collection (Open Access)

Multi-agent systems (MAS) involve multiple autonomous agents that coordinate their actions to achieve shared or competing objectives in dynamic environments. Over the past decade, multi-agent reinforcement learning (MARL) has emerged as a powerful paradigm for enabling collaborative behaviors among autonomous agents within MAS to solve complex tasks. This dissertation discusses a critical scalability gap that exists between current MARL capabilities and real-world deployment requirements. Most existing MARL research focuses on small-scale laboratory problems, often struggling to coordinate large agent populations and facing challenges with extended decision-making horizons. In contrast, many real-world applications demand coordination among hundreds or thousands of agents …


Attorneys And Ai: How Lawyers Use Artificial Intelligence And Analyze Its Impacts, Matthew I. Hall, Christian Turner, Eddie A. Gomez Schieber, Nathaniel Kite, Ari Schlesinger Nov 2025

Attorneys And Ai: How Lawyers Use Artificial Intelligence And Analyze Its Impacts, Matthew I. Hall, Christian Turner, Eddie A. Gomez Schieber, Nathaniel Kite, Ari Schlesinger

Scholarly Works

AI systems are testing lawyers' professional ethics obligations of competence, confidentiality, and candor. In the legal profession, the widespread availability of AI systems presents opportunities, like improving the review of documents during the discovery stage of a lawsuit, and challenges, illustrated by the handful of high-profile incidents where lawyers submitted legal briefs in court citing and describing fictitious cases based on AI-generated output. We conducted interviews with 44 legal professionals in the U.S. to understand how attorneys are making sense of AI technology and the impacts these technologies are having on their profession, legal ethics, and legal institutions. We describe …


Addressing Sparsity For Knowledge Graph Completion: Data And Model Perspectives, Ran Liu Nov 2025

Addressing Sparsity For Knowledge Graph Completion: Data And Model Perspectives, Ran Liu

Dissertations and Theses Collection (Open Access)

Knowledge graphs (KGs) are powerful tools for structuring factual knowledge into relational triples, yet their practical utility is often adversely affected by data sparsity. Many entities and relations are associated with only a few observations, which limits the quality of learned embeddings and weakens generalization in downstream tasks. The problem of sparsity led to two interrelated challenges. Firstly, it restricts the informativeness of training samples: positive examples are scarce, and conventional negative sampling often produces trivial or redundant negatives that resulting in limited guidance. Secondly, in few-shot relation learning scenarios, sparsity worsens distribution shifts between training and test relations, as …


The Ai-Powered Learning Loop In Higher Education, Oualid Abidi, Vladimir Dzenopoljac, Aleksandra Dzenopoljac Nov 2025

The Ai-Powered Learning Loop In Higher Education, Oualid Abidi, Vladimir Dzenopoljac, Aleksandra Dzenopoljac

All Works

Purpose – This study examines how generative AI tools affect business students’ academic performance by investigating whether flexible AI policies promote deeper learning, enhance self-efficacy and facilitate tacit knowledge acquisition in a Middle Eastern context, while ensuring efficiency and academic integrity. Design/methodology/approach – A qualitative, exploratory study observed 20 final-year business students in Kuwait during five in-class activities using generative AI tools. Semi-structured interviews complemented the researcher’s observations. Thematic analysis revealed patterns in benefits, challenges and learning processes, leading to the development of the AI-powered learning loop framework to explain academic performance outcomes. Findings – The study indicates that generative …


Designing For Novice Debuggers: A Pilot Study On An Ai-Assisted Debugging Tool, Oka Kurniawan, Erick Chandra, Christopher M. Poskitt, Yannic Noller, Kenny T.W. Choo, Cyrille Jegourel Nov 2025

Designing For Novice Debuggers: A Pilot Study On An Ai-Assisted Debugging Tool, Oka Kurniawan, Erick Chandra, Christopher M. Poskitt, Yannic Noller, Kenny T.W. Choo, Cyrille Jegourel

Research Collection School Of Computing and Information Systems

Debugging is a fundamental skill that novice programmers must develop. Numerous tools have been created to assist novice programmers in this process. Recently, large language models (LLMs) have been integrated with automated program repair techniques to generate fixes for students' buggy code. However, many of these tools foster an over-reliance on AI and do not actively engage students in the debugging process. In this work, we aim to design an intuitive debugging assistant, CodeHinter, that combines traditional debugging tools with LLM-based techniques to help novice debuggers fix semantic errors while promoting active engagement in the debugging process. We present findings …


Simulated Interactive Debugging, Yannic Noller, Erick Chandra, Srinidhi Chandrashekar, Kenny Choo, Cyrille Jegourel, Oka Kurniawan, Christopher M. Poskitt Nov 2025

Simulated Interactive Debugging, Yannic Noller, Erick Chandra, Srinidhi Chandrashekar, Kenny Choo, Cyrille Jegourel, Oka Kurniawan, Christopher M. Poskitt

Research Collection School Of Computing and Information Systems

Debugging software, i.e., the localization of faults and their repair, is a key activity in software engineering. Therefore, effective and efficient debugging is one of the core skills a software engineer must develop. However, the teaching of debugging techniques is usually very limited or only taught in indirect ways, e.g., during software projects. As a result, most Computer Science (CS) students learn debugging only in an ad-hoc and unstructured way. In this work, we present our approach called Simulated Interactive Debugging that interactively guides students along the debugging process. The guidance aims to empower the students to repair their solutions …


Seeing Is Fixing: Cross-Modal Reasoning With Multimodal Llms For Visual Software Issue Fixing, Kai Huang, Jian Zhang, Xiaofei Xie, Chunyang Chen Nov 2025

Seeing Is Fixing: Cross-Modal Reasoning With Multimodal Llms For Visual Software Issue Fixing, Kai Huang, Jian Zhang, Xiaofei Xie, Chunyang Chen

Research Collection School Of Computing and Information Systems

Large language model (LLM)-based automated program repair (APR) techniques have shown promising results in resolving real-world github issue tasks. Existing APR systems are primarily evaluated in unimodal settings (e.g., SWE-bench), relying solely on textual issue descriptions and source code. However, these autonomous systems struggle to resolve multimodal problem scenarios (e.g., SWE-bench M) due to limitations in interpreting and leveraging visual information. In multimodal scenarios, LLMs need to rely on visual information in the graphical user interface (GUI) to understand bugs and generate fixes. To bridge this gap, we propose GUIRepair, a cross-modal reasoning approach for resolving multimodal issue scenarios by …


Deep Reinforcement Learning For Solving The Stochastic E-Waste Collection Problem, Dang Viet Anh Nguyen, Aldy Gunawan, Mustafa Misir, Kwan Hui Lim, Pieter Vansteenwegen Nov 2025

Deep Reinforcement Learning For Solving The Stochastic E-Waste Collection Problem, Dang Viet Anh Nguyen, Aldy Gunawan, Mustafa Misir, Kwan Hui Lim, Pieter Vansteenwegen

Research Collection School Of Computing and Information Systems

With the growing influence of the internet and information technology, Electrical and Electronic Equipment (EEE) has become a gateway to technological innovations. However, discarded devices, also called e-waste, pose a significant threat to the environment and human health if not properly treated, disposed of, or recycled. In this study, we extend a novel model for the e-waste collection in an urban context: the Heterogeneous VRP with Multiple Time Windows and Stochastic Travel Times (HVRP-MTWSTT). We propose a solution method that employs deep reinforcement learning to guide local search heuristics (DRL-LSH). The contributions of this paper are as follows: (1) HVRP-MTWSTT …


How Behavioral Science Can Improve The Return On Ai Investments, David De Cremer, Shane Schweitzer, Jack Mcguire, Devesh Narayanan Nov 2025

How Behavioral Science Can Improve The Return On Ai Investments, David De Cremer, Shane Schweitzer, Jack Mcguire, Devesh Narayanan

Research Collection Lee Kong Chian School Of Business

Many AI projects fail because leaders treat adoption as a tech purchase instead of a behavioral change problem. People resist tools that disrupt routines, overreact to visible AI errors, and prefer familiar human judgment. As a result, even good systems fail to gain purchase. Leaders can address this problem by applying “Behavioral Human-Centered AI” across the AI adoption cycle. In the design phrase, companies should co-design with diverse users, add purposeful friction where it improves scrutiny, require beta tests with subgroup results and behavioral input. During adoption, they should frame AI as an augmenter, disclose limits and safeguards, use explainability …


Building Confidence For Class Participation, Tamas Makany, Ivy Seow Nov 2025

Building Confidence For Class Participation, Tamas Makany, Ivy Seow

Research Collection Lee Kong Chian School Of Business

What happens to students’ critical thinking when half the class filters their thoughts through AI? During a recent debate on AI policy in education, one student mentioned they routinely run their ideas through ChatGPT before speaking up. When I asked who else did the same, more than half the class raised their hands.


Interaction2code: Benchmarking Mllm-Based Interactive Webpage Code Generation From Interactive Prototyping, Jingyu Xiao, Yuxuan Wan, Yintong Huo, Zixin Wang, Xinyi Xu, Wenxuan Wang, Zhiyao Xu, Yuhang Wang, Michael R. Lyu Nov 2025

Interaction2code: Benchmarking Mllm-Based Interactive Webpage Code Generation From Interactive Prototyping, Jingyu Xiao, Yuxuan Wan, Yintong Huo, Zixin Wang, Xinyi Xu, Wenxuan Wang, Zhiyao Xu, Yuhang Wang, Michael R. Lyu

Research Collection School Of Computing and Information Systems

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance on the design-to-code task, i.e., generating UI code from UI mock-ups. However, existing benchmarks only contain static web pages for evaluation and ignore the dynamic interaction, limiting the practicality, usability and user engagement of the generated webpages. To bridge these gaps, we present the first systematic investigation of MLLMs in generating interactive webpages. Specifically, we formulate the Interaction-to-Code task and establish the Interaction2Code benchmark, encompassing 127 unique webpages and 374 distinct interactions across 15 webpage types and 31 interaction categories. Through comprehensive experiments utilizing state-of-theart (SOTA) MLLMs, evaluated via both automatic …


Sustainable Llm Inference For Edge Ai: Evaluating Quantized Llms For Energy Efficiency, Output Accuracy, And Inference Latency, Erik Johanne Husom, Arda Goknil, Merve Astekin, Lwin Khin Shar, Andre Kasen, Sagar Sen, Benedikt Andreas Mithassel, Ahmet Soylu Nov 2025

Sustainable Llm Inference For Edge Ai: Evaluating Quantized Llms For Energy Efficiency, Output Accuracy, And Inference Latency, Erik Johanne Husom, Arda Goknil, Merve Astekin, Lwin Khin Shar, Andre Kasen, Sagar Sen, Benedikt Andreas Mithassel, Ahmet Soylu

Research Collection School Of Computing and Information Systems

Deploying Large Language Models (LLMs) on edge devices presents significant challenges due to computational constraints, memory limitations, inference speed, and energy consumption. Model quantization has emerged as a key technique to enable efficient LLM inference by reducing model size and computational overhead. In this study, we conduct a comprehensive analysis of 28 quantized LLMs from the Ollama library, which applies by default Post-Training Quantization (PTQ) and weight-only quantization techniques, deployed on an edge device (Raspberry Pi 4 with 4GB RAM). We evaluate energy efficiency, inference performance, and output accuracy across multiple quantization levels and task types. Models are benchmarked on …


Do Code Semantics Help? A Comprehensive Study On Execution Trace-Based Information For Code Large Language Models, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Yi Li Nov 2025

Do Code Semantics Help? A Comprehensive Study On Execution Trace-Based Information For Code Large Language Models, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Yi Li

Research Collection School Of Computing and Information Systems

Code Large Language Models (Code LLMs) have opened a new era in programming with their impressive capabilities. However, recent research has revealed critical limitations in their ability to reason about runtime behavior and understand the actual functionality of programs, which poses significant challenges for their post-training and practical deployment. Specifically, Code LLMs encounter two principal issues: (1) a lack of proficiency in reasoning about program execution behavior, as they struggle to interpret what programs actually do during runtime, and (2) inconsistent and fragmented representation of semantic information, such as execution traces, across existing methods, which hinders their ability to generalize …


Efficient Integration Of External Knowledge To Llm-Based World Models Via Retrieval-Augmented Generation And Reinforcement Learning, Chang Yang, Xinrun Wang, Qinggang Zhang, Qi Jiang, Xiao Huang Nov 2025

Efficient Integration Of External Knowledge To Llm-Based World Models Via Retrieval-Augmented Generation And Reinforcement Learning, Chang Yang, Xinrun Wang, Qinggang Zhang, Qi Jiang, Xiao Huang

Research Collection School Of Computing and Information Systems

World models achieve remarkable success in predicting future states and planning in complex environments and Large Language Models (LLMs) serve as promising foundation to build general world models. However, their performances are usually constrained by the limited external knowledge to specific environments. Existing research attempts to enhance LLM-based world models through prompting or fine-tuning approaches, which are either requiring human knowledge or computationally extensive. Therefore, we introduce Retrieval-Augmented World Models (RAWM), a novel framework that leverages retrieval-augmented generation to efficiently integrate the external knowledge to LLM-based world models. Our main contributions are threefold: (i) We introduce a memory system and …


Mmlu-Prox: A Multilingual Benchmark For Advanced Large Language Model Evaluation, Weihao Xuan, Et. Al. Nov 2025

Mmlu-Prox: A Multilingual Benchmark For Advanced Large Language Model Evaluation, Weihao Xuan, Et. Al.

Research Collection School Of Computing and Information Systems

Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting comprehensively. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-lingual comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves …


From Personas To Talks: Revisiting The Impact Of Personas On Llm-Synthesized Emotional Support Conversations, Shenghan Wu, Yimo Zhu, Wynne Hsu, Mong-Li Lee, Yang Deng Nov 2025

From Personas To Talks: Revisiting The Impact Of Personas On Llm-Synthesized Emotional Support Conversations, Shenghan Wu, Yimo Zhu, Wynne Hsu, Mong-Li Lee, Yang Deng

Research Collection School Of Computing and Information Systems

The rapid advancement of Large Language Models (LLMs) has revolutionized the generation of emotional support conversations (ESC), offering scalable solutions with reduced costs and enhanced data privacy. This paper explores the role of personas in the creation of ESC by LLMs. Our research utilizes established psychological frameworks to measure and infuse persona traits into LLMs, which then generate dialogues in the emotional support scenario. We conduct extensive evaluations to understand the stability of persona traits in dialogues, examining shifts in traits post-generation and their impact on dialogue quality and strategy distribution. Experimental results reveal several notable findings: 1) LLMs can …


Chain Of Strategy Optimization Makes Large Language Models Better Emotional Supporter, Weixiang Zhao, Xingyu Sui, Xinyang Han, Yang Deng, Yulin Hu, Jiahe Guo, Libo Qin, Qianyun Du, Shijin Wang, Yanyan Zhao, Bing Qin, Ting Liu Nov 2025

Chain Of Strategy Optimization Makes Large Language Models Better Emotional Supporter, Weixiang Zhao, Xingyu Sui, Xinyang Han, Yang Deng, Yulin Hu, Jiahe Guo, Libo Qin, Qianyun Du, Shijin Wang, Yanyan Zhao, Bing Qin, Ting Liu

Research Collection School Of Computing and Information Systems

The growing emotional stress in modern society has increased the demand for Emotional Support Conversations (ESC). While Large Language Models (LLMs) show promise for ESC, they face two key challenges: (1) low strategy selection accuracy, and (2) preference bias, limiting their adaptability to users’ emotional needs. Existing supervised fine-tuning (SFT) struggles to address these issues, as it rigidly trains models on single gold-standard responses without modeling nuanced strategy trade-offs. To overcome these limitations, we propose a novel two-stage framework that optimizes strategy selection preferences at each dialogue turn. We first leverage Monte Carlo Tree Search to construct ESC-Pro, a high-quality …


Envisioning Future Interactive Web Development: Editing Webpage With Natural Language, Truong Hai Dang, Jingyu Xiao, Yintong Huo Nov 2025

Envisioning Future Interactive Web Development: Editing Webpage With Natural Language, Truong Hai Dang, Jingyu Xiao, Yintong Huo

Research Collection School Of Computing and Information Systems

The evolution of web applications relies on iterative code modifications, a process that is traditionally manual and time-consuming. While Large Language Models (LLMs) can generate UI code, their ability to edit existing code from new design requirements (e.g., ”center the logo”) remains a challenge. This is largely due to the absence of large-scale, high-quality tuning data to align model performance with human expectations. In this paper, we introduce a novel, automated data generation pipeline that uses LLMs to synthesize a high-quality fine-tuning dataset for web editing, named Instruct4Edit. Our approach generates diverse instructions, applies the corresponding code modifications, and performs …